Reddit Scraper - Posts, Comments, Search, Users (no login)
Pricing
from $2.50 / 1,000 post fetcheds
Reddit Scraper - Posts, Comments, Search, Users (no login)
Scrape Reddit posts, comments, search results and user history without a login. Paste keywords, r/subreddits, u/users or post links in one list. Only-new-posts monitoring per line, date filter, and a plain-English reason for every line that returns nothing.
Pricing
from $2.50 / 1,000 post fetcheds
Rating
5.0
(1)
Developer
Omar Eldeeb
Maintained by CommunityActor stats
2
Bookmarked
15
Total users
5
Monthly active users
3 days ago
Last modified
Categories
Share
Reddit Scraper — Posts, Comments, Search & Users (no login)
Paste what you want from Reddit — keywords, r/subreddits, u/users, or post links — one per line, and press Start. Each line is detected automatically. You get clean rows (posts, comments) you can export to JSON, CSV or Excel, send to an AI/RAG pipeline, or schedule as a monitor. No Reddit account, app registration or API key needed.
What does Reddit Scraper do?
- Keyword search across all of Reddit, or inside one subreddit (
r/rust async runtime). - Subreddit feeds — hot, new, top, rising, controversial, with a time window.
- Full posts with complete comment threads — paste a post link (redd.it and app share links work too), or turn on comments for every post a search or feed finds. Every collapsed "load more comments" and "continue this thread" branch is expanded, up to your per-post limit.
- User history — a user's posts and comments, with their karma and account age.
- Monitoring — "only new posts since my last run", tracked separately for every keyword and subreddit line.
- Date filter — only posts newer than a date or a relative time ("7 days").
- Rich post data on every row — upvote ratio, author flair, every image of a gallery, poll options with their votes, and the MP4 / HLS links of Reddit-hosted videos, read from Reddit's own JSON for each page.
- Never an unexplained empty run — every line that returns nothing gets a plain-English reason (does not exist, private, banned, quarantined, no matches, filtered out, no new posts, blocked).
How to use it
- Open the actor's Input tab.
- In What to scrape, paste one item per line:
| You type | What happens |
|---|---|
python or "agentic payments" OR x402 | Search all of Reddit (plain text is always a search; quotes, OR, subreddit:, author: work as on Reddit) |
r/rust async runtime | Search inside r/rust |
r/technology or https://www.reddit.com/r/technology/top/?t=week | The subreddit's feed (a sort in the link wins for that line) |
https://www.reddit.com/r/.../comments/abc123/... · https://redd.it/abc123 · reddit.com/gallery/abc123 · app share link | That post + its comments. A link to a single comment returns the whole thread (same as the post link, so the two are not scraped twice). |
u/spez or https://www.reddit.com/user/spez/submitted/ | The user's posts & comments |
https://www.reddit.com/search/?q=x402&sort=new | That exact search |
- Set Max posts per line and Max total rows (spending cap).
- Click Start. Export the dataset, or read it via the Apify API.
Tip: plain text is always searched. To scrape a community's feed, start the line with
r/. If a one-word line looks like a community name (e.g.sekai_multiverse) and finds nothing, the result tells you to writer/sekai_multiverse.
Recipes (copy the JSON into the Input tab)
Prices are the current pay-per-event prices; every post and comment row is charged, from the first one.
1. Brand / keyword monitoring — schedule daily with Apify Schedules
{ "targets": ["\"agentic payments\" OR x402"], "sort": "new", "onlyNew": true }
First run saves a baseline (up to 25 posts ≈ $0.08); later runs return only posts published since the previous run — a quiet keyword costs cents per day.
2. Lead generation / pain-point mining
{ "targets": ["r/smallbusiness looking for a tool", "r/SaaS recommend crm"], "postedAfter": "30 days", "commentsMode": "all", "maxCommentsPerPost": 20 }
Returns every matching post from the last 30 days up to 25 per line — how many that is depends on the community (a busy one fills all 25; r/SaaS "recommend crm" had 1 in September 2026). With comments, capped at 200 rows: at most ≈ $0.27.
3. Research / sentiment dataset
{ "targets": ["r/ChatGPT"], "sort": "top", "timeFilter": "year", "maxPostsPerTarget": 500, "commentsMode": "all", "maxItems": 20000 }
500 posts ($1.25) + up to 19,500 comments ($15.60) — set a lower maxItems to cap spend.
4. Community analytics
{ "targets": ["r/rust", "r/golang"], "sort": "new", "postedAfter": "7 days" }
Up to 25 recent posts per community with subscriber counts ≈ $0.13.
5. Thread deep-dive
{ "targets": ["https://www.reddit.com/r/AskProgramming/comments/1woumpx/i_find_learning_python_hard_after_learning_java/"] }
The post + up to 100 comments ≈ $0.08.
6. Person research
{ "targets": ["u/spez"] }
25 profile items with karma and account age ≈ $0.10.
Input reference
| Field | Default | What it does |
|---|---|---|
targets | — | The lines to scrape (see the table above). If it is empty, the field names other Reddit scrapers use — startUrls (strings or { "url": … } objects), urls, url, searches, searchTerms, keywords, searchQueries — are read as lines. With nothing usable, the run succeeds with 0 rows and names any fields it did not recognise. |
maxPostsPerTarget | 25 | Posts per keyword / subreddit / user line (for users: posts + comments). |
sort | auto | auto = relevance for searches, hot for subreddits. Also new, top, hot, rising, relevance, comments, controversial. |
timeFilter | all | hour … all. Limits every search; for subreddits/users only Top and Controversial use it. |
postedAfter | — | 2026-09-01, 7 days, 3 weeks, 1 month (API: also 24 hours). Older posts are skipped and not charged. With Sort = "Best match", lines are read newest-first so the cut-off can be reached quickly; pick another sort to override. A future date is rejected. |
onlyNew | false | Monitoring: return only posts newer than the previous run, per subreddit and keyword line. New posts are delivered oldest-first up to "Max posts per line"; any extra wait for the next run (none are skipped). If Reddit fails mid-check, or the run's time limit cuts the check short, the line delivers nothing yet, keeps its position and saves where it stopped — the next run continues from there (status stopped, not a failure), so a busy community is never skipped. If that saved spot can no longer be read (its post was removed, or Reddit refuses it), the run checks again from the newest post with its full reach instead; a spot that fails two runs in a row is dropped the same way. More than ~1,000 new posts between runs is more than Reddit lists; the oldest are then out of reach (status partial) — schedule busy lines more often. Keyword lines re-check the last 30 minutes (search indexes some posts late) without repeating posts. State lives in your account's key-value store reddit-monitor-state, separately for each saved task. If two runs check the same line at the same time, only one delivers it; the other skips it (status skipped, nothing charged). |
commentsMode | postUrls | postUrls = comments for pasted post links only; all = also for every post a search/feed/user line finds; none = posts only. |
maxCommentsPerPost | 100 | Up to this many comments per post, in Reddit's default (best) order — the whole thread is read, collapsed branches included (max 50,000). Big threads have thousands of comments and each one is charged ($0.0008): a 5,000-comment thread ≈ $4. maxItems still caps the run. |
maxCommentDepth | 10 | 1 = top-level only, 0 = post only. |
maxItems | 200 | Hard cap on posts + comments for the whole run (status rows are not charged and not counted). |
includeNsfw | true | Off = skip NSFW posts (not charged) — in feeds and searches, and for pasted post links too (the line then gets a filtered_out status row saying why). |
statusRows | true | Add one recordType: "status" row (never charged) per line that returned nothing or failed. |
proxyConfiguration | Apify Proxy | Datacenter works. With Apify Proxy, a US residential IP is used only after network errors or regional age gates (residentialFallback); if you picked a country or group, every such request is counted in OUTPUT.residentialFetches and mentioned in the run's status message. Your own proxy URLs are used as they are — never replaced — and a run with no proxy stays without one. If the proxy you chose cannot reach Reddit, the message says so (and what to change) instead of asking you to re-run. |
Output
Every row has recordType (post, comment or status), input (the line that produced it), inputType, source (which page type it was read from: reddit-listing, reddit-search, reddit-user-page, reddit-post-page), a canonical url on www.reddit.com, and scrapedAt. Unknown values are null; an empty text is "".
How pages are read. Every page — subreddit feeds, searches, user pages and full comment threads — is read from old.reddit's JSON for that page (the same page, as data). If Reddit refuses the JSON for one page, that page is read from its HTML version instead: every row is still delivered and charged once, only the JSON-only details (see below) are null on those rows, and the line's message and OUTPUT.targets[].htmlFallbackPages say how many pages that happened to.
Post row
{"recordType": "post","input": "r/rust async runtime","inputType": "search","id": "1u4k55x","fullId": "t3_1u4k55x","url": "https://www.reddit.com/r/rust/comments/1u4k55x/why_are_async_runtimes_so_big_and_complex/","title": "Why are async runtimes so big and complex?","subreddit": "rust","author": "lelelesdx","authorId": "t2_g8ulx","body": "I've been trying to learn how async and futures work. …","score": 123,"upvoteRatio": 0.97,"numComments": 53,"createdAt": "2026-06-13T07:09:42.000Z","postType": "self","externalUrl": null,"domain": "self.rust","flair": "🙋 seeking help & advice","isSelf": true,"isNsfw": false,"source": "reddit-search","scrapedAt": "2026-09-28T08:12:16.073Z","authorFlair": null,"galleryImages": null,"poll": null,"video": null}
| Post field | Notes |
|---|---|
title, subreddit, author, authorId (t2_…) | |
body (alias selfText) | The post's text on every page (subreddit and user feeds, search results, post links), including the text some link/video posts carry. null (not "") when the post was removed or deleted (see isRemoved), or when only old.reddit's "content not supported on old Reddit" stub exists; "" only for a self post that has a title and no text. On rows from a page read through the HTML fallback, subreddit/user feed rows have null (old.reddit's HTML loads it lazily). |
score, numComments, createdAt | score is null when Reddit hides it. |
upvoteRatio | 0–1, on every post. (HTML-fallback rows: only from the post's own page.) |
postType | self, link, image, video, gallery; imageUrl for image posts. |
externalUrl, domain, thumbnail | Link posts. externalUrl is the link exactly as submitted (a link to another Reddit post uses www.reddit.com), and domain is Reddit's own value (e.g. self.<name> for a post on a user profile; null when Reddit gives none, as for some crossposts) on every page — before 0.3 search rows derived both from old.reddit's search page. |
flair, isNsfw, isSpoiler, isOriginalContent, isStickied | isStickied is true when the post is pinned in its subreddit — also when it shows up on r/popular, a multireddit, a search or a user page (before 0.3 only a subreddit's own feed marked it). |
isLocked, isArchived | Comments closed by a moderator / post older than Reddit's archive age (no new votes or comments). Known on every page. HTML-fallback rows: isLocked is null on search results and on archived posts (old.reddit hides the lock marker on archived posts). |
isRemoved, removedBy | The post was removed or deleted — its body is then null (never ""). removedBy is Reddit's category when known: deleted (by the author), moderator, reddit, automod_filtered… isRemoved is null only where the page cannot tell (HTML-fallback rows of search results and feeds). |
numCrossposts, awards, crosspostParent, crosspostParentId | crosspostParent = {title, subreddit, author, score, numComments, body} of the original and crosspostParentId its t3_… id. The crosspost row's own body stays null (the text belongs to the original author); the shared post's text is crosspostParent.body (null when it has none, and on HTML-fallback rows). numCrossposts is the post's own count (on a crosspost: how often the crosspost itself was crossposted; null there on HTML-fallback feed rows, where old.reddit only shows the original's count). |
subredditType, subredditSubscribers | Subscribers of the post's subreddit, on every post (also on r/all, searches and user pages). HTML-fallback rows: subreddit lines only. |
authorPostKarma, authorCommentKarma, authorCreatedAt | On user lines (the profile's post karma, comment karma and account creation date). |
authorFlair | The author's flair text in that subreddit (text parts only; null when none). |
galleryImages | Gallery posts: every image in gallery order, [{url, width, height, caption}] (full-size source image; caption null when none). A crosspost of a gallery lists the original's images. null for other posts. |
poll | Poll posts: {options: [{text, votes}], totalVotes, votingEndsAt}. The post's body is the author's own text without Reddit's "View Poll" link (null for a poll with no text). Reddit hides per-option votes until voting ends (null until then); totalVotes is always shown. null for other posts. |
video | Reddit-hosted videos (v.redd.it): {url, hlsUrl, durationSec, width, height, isGif} — url is the progressive MP4 fallback (Reddit usually serves its audio as a separate track) and is stable; hlsUrl is the adaptive stream, but it is a signed link that Reddit issues per request and that expires (about 30 days) — use it soon after the run and don't store it as a permanent link (store url, or re-scrape the post). null for other posts, including YouTube and other embeds (see externalUrl). |
| Comment field | Notes |
|---|---|
body, author, authorId, score, createdAt, editedAt | An image in a comment appears as its link (https://preview.redd.it/…), not as old.reddit's <image> placeholder. |
postId, postTitle | The post the comment belongs to. |
parentId, depth | depth is 1-based: a top-level comment is 1 (Reddit's own JSON counts from 0). parentId is t3_… for top-level comments, else the parent t1_…. If the parent comment was deleted, its replies are still returned and parentId still names the deleted comment (which itself is not a row). Comments listed on a user page have depth null (Reddit does not list where they sit in the thread) and their real parentId (null on HTML-fallback rows). |
numReplies, isSubmitter (written by the post's author), distinguished (admin/moderator) | numReplies counts all replies below the comment (every level), as old.reddit shows it. null for comments listed on a user page: Reddit does not list replies there (before 0.3 this was 0 — old.reddit prints "0 children" there even for a comment with 38 replies). |
isStickied | Pinned by a moderator at the top of the thread. |
authorFlair | The author's flair text in that subreddit (null when none). |
Status row (never charged — one per line that returned nothing or failed)
{"recordType": "status","input": "r/CenturyClub","inputType": "subreddit","status": "private","message": "r/CenturyClub is private (invite-only). Only public communities can be scraped.","rowsReturned": 0,"url": "https://www.reddit.com/r/CenturyClub/"}
status is one of partial (rows delivered, but part of the line is missing — e.g. Reddit stopped answering on page 3, Reddit refused to page past a post it no longer shows (the message says so — not a block), some posts came without their comments or with only their first comment page, a comment thread was cut by the time limit, or the line used up its fair share of a short run so the next lines could run), skipped (another run was already checking this "only new" line), no_results (Reddit returned no posts at all — never used when posts were found but filtered or cut short; a zero-result search is re-checked twice before it is reported), not_found, banned, private, premium_only, suspended, quarantined, filtered_out, no_new_posts, invalid_input, blocked, rate_limited, timed_out, stopped. Filter them out with recordType != "status" or turn off Explain empty results in the dataset.
Dataset views. The Console's Posts, Comments, Media & polls and Line status views choose which columns are shown; Apify views cannot filter rows, so every view lists all rows (a comment row in the Posts view simply has the post columns empty). To keep one kind of row, filter on recordType in the table, your export or your code.
Run summary (OUTPUT record)
Every run writes an OUTPUT record to its key-value store: complete, stoppedEarly, stopReason, failed, a message, and targets[] with each line's status, message, rows returned, pages read and htmlFallbackPages (pages read from Reddit's HTML because its JSON was refused). Use it to detect partial runs from the API.
When does a run fail? When it delivered no rows at all and at least one line failed on our side (Reddit or the network blocked it, or the run timed out while fetching it) — you get a FAILED run with a clear message instead of a silent empty dataset, and no rows are charged. A run where every line is explained by your input (does not exist, private, no matches…) succeeds and explains each one — and so does a run with no lines at all ("No targets given…", nothing charged). When the failure comes from the proxy you chose (your own proxy URLs, no proxy, or an Apify country/group without the fallback), the message says so and what to change. If the platform moves the run to another server mid-way, it continues where it stopped without pushing or charging any row twice.
How much does it cost to scrape Reddit?
Pay per row. Every post, comment, search result and user-history row is charged at the price below, starting with the first row of the run:
| Event | Price |
|---|---|
| Post fetched (subreddit feeds, post links) | $0.0025 |
| Search result fetched | $0.003 |
| Comment fetched | $0.0008 |
| User history item fetched | $0.004 |
New post found by monitoring (onlyNew on subreddit lines) | $0.0025 |
Status rows are never charged. Posts removed by your filters (date, NSFW, duplicates) are never charged.
Existing API callers (0.1.x input)
Integrations and schedules built on 0.1.x keep working unchanged: legacy lines read pages of 25 posts, so maxPagesPerTarget still means what it did and a legacy line never returns (or charges) more than maxPagesPerTarget × 25 rows. When targets is empty, the actor accepts the old mode input (subreddit_posts, search, comments_thread, user_profile, monitor with subreddits, searchQuery, restrictToSubreddit, postUrls, usernames, includeComments, maxPagesPerTarget) and behaves as before: same pages, same sort, same events, and the monitor keeps its saved state in reddit-monitor-state. New-only options are ignored on that path. If you send both, targets wins and the old fields are ignored with a warning. comments_thread still returns about one comment page per post (up to 200 comments) so existing bills do not change; to get more or fewer, send maxCommentsPerPost (any value, e.g. 5000 for whole threads or 100) — the thread is then expanded up to that many comments per post, bounded by maxItems. Since 0.2.6 a mode whose field is empty succeeds with 0 rows and a "No targets given" message instead of failing. One visible difference: url permalinks now use https://www.reddit.com/… (0.1.x emitted https://old.reddit.com/…) — if you de-duplicate on url across old and new runs, compare the path or use id.
Limits
- Comments that were deleted or removed are not rows (Reddit no longer shows them); their replies are, with
parentIdnaming the removed comment. Reddit'snumCommentsalso counts removed and spam-filtered comments, so a full thread usually has 1–5% fewer rows than that number. - A full thread takes time: about one request per 30–80 comments, a few requests per second — a 10,000-comment thread takes roughly 5–20 minutes depending on how fast Reddit answers. If the run's time limit cuts a thread, what was read is delivered (and charged) and the line is marked
partial. - If Reddit's full-thread endpoint is unavailable, the post's first comment page is returned instead and the line is marked
partial("Full thread unavailable, returned the first page"). If thread requests keep failing for 3 posts in a row on one line, that line's later posts come with their first comment page only (linepartial, message says so) instead of each waiting for retries. - A post whose thread was cut by the time limit or a server move / restart is finished when the line continues (its second pass, or after the restart) — comments already delivered are not pushed or charged again, and the per-post limit counts them however many restarts happen. Recovered rows are never charged again; a row delivered in the last instant before a hard restart may stay uncharged (in your favour).
- The actor needs at least 512 MB of memory.
- Flair of posts listed on
/user/pages is not shown by Reddit there, so it isnullon user lines. - Quarantined, private, banned and premium-only communities need a logged-in account and are reported, not scraped.
- Monitoring (
onlyNew) applies to subreddit and keyword lines; user and post lines are scraped normally.
Legal & responsible use
This Actor collects publicly available content from Reddit. You are responsible for complying with Reddit's User Agreement and applicable laws (including GDPR/CCPA) when collecting and processing data, and for respecting the rights and privacy of the people whose content you scrape. Do not use it to collect personal data unlawfully, to harass, or to redistribute content in violation of Reddit's terms.
FAQ
- Do I need a Reddit account or API key? No.
- A line returned nothing — why? Open the dataset's Line status view (or the run's
OUTPUTrecord): each empty line has a reason, e.g. "r/SekaiAI does not exist on Reddit" or "No new posts since your last run". - Are scores exact? Reddit fuzzes scores slightly and hides them on very new posts/comments (
null). - Issues or requests? Open an issue from the Actor's Issues tab.