Reddit Scraper — Posts, Comments, Search, Subreddit Monitoring
Pricing
from $1.00 / 1,000 posts
Reddit Scraper — Posts, Comments, Search, Subreddit Monitoring
Subreddit listings, keyword search and full comment trees, no login and no proxy to configure. Posts with full text, score, flair and media; comments with depth and parent. Schedule it with onlyNew to get just what appeared since the last run. Pay per row, no start fee.
Pricing
from $1.00 / 1,000 posts
Rating
0.0
(0)
Developer
Moxlade
Maintained by CommunityActor stats
0
Bookmarked
25
Total users
20
Monthly active users
7 hours ago
Last modified
Categories
Share
Subreddit listings, keyword search and full comment trees — no Reddit login, no proxy to set up, no start fee. Give it subreddits, search queries or post URLs and get posts with their full text, score, upvote ratio, flair and media, and comments with author, score, depth and parent. Turn on onlyNew and put it on a schedule: every run brings just the posts that are new since the run before, and a quiet day costs nothing. Pay per delivered row — $2.00 per 1,000 posts, $1.00 per 1,000 comments.
Newest posts in a subreddit
Give it this:
{"subreddits": ["Upwork"],"sort": "new","maxItems": 50}
Back comes one row per post or comment (kind says which) — the fields below are the ones this real row carried; every row has the same 27, and a field the source did not show arrives as null, not as a zero:
{"kind": "post","id": "t3_1wir0ob","subreddit": "Upwork","title": "results of competing in the 50+ proposals jobs?","author": "acryptostewie","created_utc": "2026-09-17T11:10:13.675000+00:00","score": 1,"upvote_ratio": 1,"comment_count": 1,"post_type": "text","url": "https://www.reddit.com/r/Upwork/comments/1wir0ob/results_of_competing_in_the_50_proposals_jobs/","domain": "self.Upwork","permalink": "https://www.reddit.com/r/Upwork/comments/1wir0ob/results_of_competing_in_the_50_proposals_jobs/","body_text": "as a Top-Rated dev, 100% JSS i see plenty of high paying jobs but always 50+ proposals with people bidding a lot of conn …","body_html": "<p>\n as a Top-Rated dev, 100% JSS i see plenty of high paying jobs but always 50+ proposals with people bidding a l …","nsfw": false,"stickied": false,"locked": false,"language": "en","award_count": 0,"found_by": "r/Upwork/new","fetched_at": "2026-09-17T11:31:02Z"}
The same actor, asked differently:
- Top posts of the month with their comments —
{"subreddits": ["ClaudeCode"], "sort": "top", "time": "month", "includeComments": true, "maxCommentsPerPost": 50, "maxItems": 25} - Search a topic across Reddit since a date —
{"searches": ["claude code"], "searchSort": "new", "postedSince": "7d", "maxItems": 100} - Search inside chosen subreddits —
{"subreddits": ["ClaudeCode", "ChatGPTCoding"], "searches": ["hooks", "mcp server"], "maxItems": 100} - Specific posts with every comment —
{"postUrls": ["https://www.reddit.com/r/Upwork/comments/1wd745i/here_are_my_tips_for_upwork_newcomers/"], "includeComments": true, "maxCommentsPerPost": 500} - Monitor: only what is new since the last run —
{"subreddits": ["Upwork", "freelance"], "sort": "new", "onlyNew": true, "maxItems": 200}
What a row carries
Three ways in, one row shape. subreddits walks a listing by new, hot, top (with a time range) or rising; searches runs Reddit's own post search, sitewide or inside each subreddit you name; postUrls fetches specific posts. Every row is the same 27 typed fields, kind = post or comment, found_by says which input produced it.
Full text on every post, not a preview. Listing posts carry the complete body as Reddit renders it (body_text with paragraph breaks, body_html with links), the score, upvote ratio, comment count, flair, type, media URLs, language, stickied/locked flags. Search hits are filled with author and body (fetchBodies, on by default).
Comment trees with structure. includeComments delivers each post's comments right after it — depth, parent_id, post_id, author, score, permalink — in Reddit's Best/Top/New/… order, deep branches expanded up to maxCommentsPerPost.
Monitoring that remembers. onlyNew remembers the ids it delivered and, on a chronological walk, the newest post it delivered per source; the next run stops at that point. Schedule a task and each run is the delta — the same lever as postedSince (12h, 7d, 2w, a date) for a one-off window.
Nothing to configure, nothing to babysit. No cookies, no account, no proxy field. Reddit is read from a maintained residential exit that rotates on any wall and waits out the rate limit; a page that fails is retried before the run is failed, and the failure reason is in the status message.
Measured, not promised. The acceptance suite runs the deployed build before every release. On the release runs: 25 listing rows with score, ratio, counts, type, flags and language on 25 of 25; 20 search rows with author and body on 20 of 20; a thread of 40 comments with a resolved parent on 22 of 22 nested comments; the second monitoring run at 0 rows and $0.00.
All 27 fields of a row
Every row carries all of them. A field the source did not show arrives as null, never as a zero or an empty string, and the note says when that happens.
The post or the comment
| field | type | what it holds |
|---|---|---|
kind | string | post or comment. |
id | string | Reddit's own id: t3_… for a post, t1_… for a comment. Stable; deduplicate and join on it. |
subreddit | string | Subreddit name without the r/. |
title | string | Post title. Null on comment rows. |
author | string | Username without the u/. Null when the account is deleted. |
created_utc | string | When it was posted, ISO 8601 UTC. |
permalink | string | The post's or comment's own URL on reddit.com. |
url | string | What the post links to: the external URL, the media, or the crossposted post. For a text post this is its own permalink. |
domain | string | Domain of url as Reddit labels it (self.<sub> for text posts, i.redd.it, youtube.com, …). |
What it says
| field | type | what it holds |
|---|---|---|
body_text | string | Full text of the post or comment, paragraphs separated by blank lines. Empty for a link post. |
body_html | string | The same body as Reddit's rendered HTML, links kept. |
media | array | Media URLs the post carries (images, galleries, video). |
flair | string | The post's flair text, when it has one. |
post_type | string | text, image, gallery, video, multi_media, crosspost, link or poll. Subreddit listings only. |
language | string | Reddit's language tag for the post (en, de, …). Subreddit listings only. |
How it is doing
| field | type | what it holds |
|---|---|---|
score | integer | Net upvotes as Reddit showed them at fetch time. Null on a post reached by URL alone — Reddit's post page carries no score for a visitor; a post reached through a subreddit listing or a search has it. |
upvote_ratio | number | Share of upvotes, 0–1. Subreddit listings only. |
comment_count | integer | Comments on the post as Reddit counts them (a comment row carries the count of its post). |
award_count | integer | Awards on the post. Subreddit listings only. |
nsfw | boolean | Marked NSFW. Known on search hits; null where Reddit did not say. |
stickied | boolean | Pinned by the moderators. Subreddit listings only. |
locked | boolean | Comments locked. Subreddit listings only. |
Where a comment sits in the thread
Filled on comment rows, so a tree can be rebuilt without a second pass.
| field | type | what it holds |
|---|---|---|
post_id | string | Comment rows: the t3_… id of the post the comment belongs to. |
parent_id | string | Comment rows: the id of the parent comment (t1_…), or null for a top-level comment. |
depth | integer | Comment rows: 0 for a top-level comment, 1 for a reply to it, and so on. |
How the row was reached
| field | type | what it holds |
|---|---|---|
found_by | string | How the row was reached: r/<sub>/<sort> for a subreddit listing, search:<query> (with @r/<sub> when scoped), url for a post given by URL, and comments:<post id> for a comment. |
fetched_at | string | When this row was read from Reddit, ISO 8601 UTC. |
Ways to run it
List a subreddit. Set subreddits and a sort. new is chronological — pair it with postedSince for a window or onlyNew for a schedule. top with time gives the period's best. maxPostsPerSource says how far down each listing to look (Reddit itself stops at about 1,000). Ready-made: Newest posts in a subreddit and Top posts of the month, with their comments.
Search for posts. Set searches; each query runs across Reddit, or inside each of subreddits when both are given. searchSort is Reddit's own ordering (new, relevance, hot, top, comments). Hits come with author and full text unless you turn fetchBodies off to save the extra request per post. Ready-made: Search Reddit for a topic, last 7 days.
Fetch posts by URL, with their comments. Paste post URLs or ids into postUrls, turn on includeComments, set maxCommentsPerPost and commentSort. Comments follow their post in the dataset, each with depth and parent_id. A post reached by URL alone carries no score (Reddit's post page shows a visitor none); posts reached through a listing or a search do.
Monitor on a schedule. Turn on onlyNew, save the run as a task, add a schedule. The first run seeds the memory; every later run delivers only what appeared since, or nothing at all — and nothing is charged for nothing. Name the memory with stateStoreName to share it between tasks or to reset it. Ready-made, schedule it: Monitor subreddits: only new posts each run.
Everything you can ask for
Where to read
Three ways in, one row shape: a subreddit listing, Reddit's own post search, or specific posts by URL. Mix them freely — found_by on each row says which input produced it.
| input | accepts | default | what it changes |
|---|---|---|---|
subreddits | array | — | Subreddit names or URLs (Upwork, r/Upwork, https://www.reddit.com/r/Upwork/). Each is listed by sort and time. With searches set as well, every search runs inside each of these subreddits instead. |
sort | new / hot / top / rising | "new" | How each subreddit is listed. new is chronological, which is what postedSince and onlyNew walk most cheaply. |
time | hour / day / week / month / year / all | "week" | Applies to sort: top only. |
searches | array | — | Keyword searches for posts, across all of Reddit — or inside each of subreddits when both are given. One row per matching post. |
searchSort | new / relevance / hot / top / comments | "new" | Order of search results, as Reddit's own search offers it. |
postUrls | array | — | Specific posts to fetch: URLs (https://www.reddit.com/r/Upwork/comments/1w8w2hb/…, old.reddit.com or redd.it links) or ids (1w8w2hb, t3_1w8w2hb). Each comes back with its full body and, with includeComments, its comment tree. A subreddit URL here is read as that subreddit. |
maxPostsPerSource | integer | 100 | How far down each listing or search to look. Reddit itself stops a listing at about 1,000 posts. |
Comments
Each post's comments are delivered right after it, in Reddit's own order, with the tree structure on every row.
| input | accepts | default | what it changes |
|---|---|---|---|
includeComments | boolean | false | Fetch the comment tree of every delivered post — one comment row per comment, with author, score, depth and parent, right after its post. Billed as comment per row. |
maxCommentsPerPost | integer | 100 | Cap on comment rows per post. Deep threads are expanded breadth-first until the cap. |
commentSort | confidence / top / new / controversial / old / qa | "confidence" | Reddit's comment ordering for the tree. |
Bodies and the time window
What each row carries, and how far back to go.
| input | accepts | default | what it changes |
|---|---|---|---|
fetchBodies | boolean | true | A search hit alone carries title, score, comment count and date. Leave this on to read each hit's author and full text as well (one extra request per post). Subreddit listings and post URLs always carry the body. |
postedSince | string | — | Keep only posts created after this point: a date (2026-09-01), a timestamp, or a window before now — 12h, 7d, 2w, 1m (or 7 days, 2 weeks). With sort: new the listing walk stops at the cutoff, so nothing older is fetched. |
Monitoring
For a scheduled task: the run remembers what it delivered and the next one returns only what is new.
| input | accepts | default | what it changes |
|---|---|---|---|
onlyNew | boolean | false | Remember every post delivered and skip it next time. On a scheduled task, every run brings just the posts and comments that are new since the run before; with nothing new it delivers no rows and charges nothing. The memory lives with us, keyed to your Apify account, under stateStoreName, or under a name built from the targets when that is empty. |
stateStoreName | string | — | Which memory onlyNew reads and writes. It lives with us, keyed to your Apify account; your own account storage stays untouched. Empty means a name built from the subreddits, searches and post URLs. Give two tasks the same name to share one memory; pick a new name to start again from zero. |
Run size
The spending cap: nothing is charged for a row that is not in your dataset.
| input | accepts | default | what it changes |
|---|---|---|---|
maxItems | integer | 100 | Stop after this many post rows across the whole run. Comment rows are bounded separately by maxCommentsPerPost. With no start fee, this is also your spending cap: posts × $0.002 (+ comments × $0.001). |
If an input can't be used
The run still ends Succeeded and charges nothing. Its status message begins INPUT REJECTED and says what to change. The dataset holds a single row, {"error": true, "code": …, "message": …}, with no results, and the key-value record ERROR repeats it, so an agent should check error on the first row. A search that finds nothing is not an error, only an empty dataset. A run marked Failed is a fault on our side, not your input; it costs nothing and reaches us.
Before you run it
How do I choose between this and the other Reddit actors? There are a lot of them — 569 actors on this store carry Reddit in their name or title, and they ran 1.1M times in the last 30 days (our own store census, read 2026-09-23); the ones that charge per delivered row sit between about $0.001 and $0.004. This one is $2.00 per 1,000 posts and $1.00 per 1,000 comments, with no start fee and a zero-row run charging nothing. What it gives: the complete post body as Reddit renders it, the comment tree with depth and parent_id so replies keep their shape, and onlyNew state so a scheduled run returns the delta instead of the same page again. What it does not do: user timelines, subreddit metadata, media re-hosted on our side, or anything behind a Reddit login.
Does it need a Reddit account or a proxy? No. Reddit is read logged out, through an exit we maintain, and nothing is stored on your side. There is no cookie or proxy field because there is nothing for you to supply.
Why does a post fetched by URL have no score? Reddit's post page renders nothing but a script shell to a visitor; the post's body and its comment tree come from two other endpoints that do not carry the post's score. A post reached through a subreddit listing or a search has its score, upvote ratio and comment count, because those pages do.
How many comments does a post return? Up to maxCommentsPerPost (default 100, up to 2,000). The first page of a thread holds about 25 top-level comments with their visible replies; deeper branches are expanded breadth-first until the cap. A thread's comment_count is Reddit's total; the delivered count can be lower when comments are deleted or collapsed.
Is this legal, and what about personal data? Everything is read from what Reddit serves a logged-out visitor: no account, no cookie, nothing behind a login. A row carries what a public post carries — the subreddit, the title, the body, the score and the account name that wrote it, which is the pseudonymous handle Reddit shows and the only identifier it exposes. No e-mail, no IP, no profile history. Every row is read at run time, so a post or comment Reddit has removed cannot be returned by a later run. What you may do with public posts depends on your purpose and your jurisdiction, and that part is yours, not ours.
How do I get the rows out? They land in the run's dataset: download them as JSON, CSV, Excel or XML from the run page, or read the same rows over the Apify API (the API tab has the call for this Actor). Save a run as a task, schedule it, and point the store's integrations at it — webhooks, Zapier, Make, Slack, Google Drive — or let an agent call the Actor over MCP. Every row carries a stable id and permalink, so repeated runs deduplicate against your own table.
Who makes this, and what else is there? Moxlade — corpora your agent can ask. Its other actors are the Upwork freelancer census (search by rate, JSS, badge, country; exact earnings; every contract) and the Vinted seller census (who carries a brand, how far under the market they price, what each closet holds), and its MCP endpoint for Upwork buyer intelligence is buyer.moxlade.com. Working code for the same client calls, in Python, JavaScript and curl: github.com/getmoxlade/upwork-freelancers-examples — this actor takes the same calls with its own input.
Is the data current? Every row is read from Reddit at run time — fetched_at says when — so scores and comment counts are what Reddit showed at that moment.
What it costs
Charged per row, and only once the row is in your dataset, so maxItems and maxCommentsPerPost together are your spending cap. No start fee, no minimum, and an empty run is free: on the release runs the charged event count equalled the row count on every run, and the zero-row run charged $0.00. Speed on the same runs: 25 listing rows in a 6 s run, 20 search rows with bodies in 26 s, one post with 40 comments in 17 s.
| event | what one event is | per 1,000 |
|---|---|---|
post | One post row — from a subreddit listing, a search, or a post URL — with its full text where Reddit shows one. Charged only after the row is in your dataset. | $2.00 |
comment | One comment row (includeComments), with author, score, depth and parent. | $1.00 |
Who asks this
Community and brand monitoring — watch the subreddits and keywords that matter, on a schedule, and get only the new posts — with the text, not a link to click.
Research and content teams — pull a subreddit's top posts of the month with their comment trees, or every post matching a query since a date, into one typed dataset.
Product and lead-gen builders — a stable id (id), permalink, created_utc and found_by on every row, so deduplication, joins and incremental pulls are one line each.
Data pipelines and AI agents — one call, JSON rows with full text and structure, no browser and no login to keep alive.
Questions
Every figure on this page was measured on our own runs or read from our own tables, on the date given next to it.