Reddit Scraper - Posts, Comments & Search
Pricing
from $3.00 / 1,000 results
Reddit Scraper - Posts, Comments & Search
Scrape Reddit posts, comments, subreddits, user profiles and keyword search results. Filter by flair, score and date; export reply trees, text and media to JSON, CSV or Excel. No Reddit API key required. Run through the Apify API or connect n8n, Make and Zapier.
Pricing
from $3.00 / 1,000 results
Rating
5.0
(2)
Developer
Dev
Maintained by CommunityActor stats
0
Bookmarked
9
Total users
8
Monthly active users
3 days ago
Last modified
Categories
Share
Reddit Scraper — Posts, Comments, Search and Flairs
Collect public Reddit posts, comments, subreddit feeds, keyword search results, user profiles and flair metadata. Export text, scores, dates, media and reply relationships to JSON, CSV or Excel, or connect the Actor to n8n, Make, Zapier and the Apify API.
You do not need a Reddit account, cookies or a proxy. Access is managed for you. The Actor uses Reddit's public web routes, not the official Reddit developer API, and is not affiliated with Reddit.
Quick start
- Put one or more community names in Subreddits, such as
technologyorn8n. The box starts with an example — replace it. - Optionally add a Keyword to search inside those communities.
- Set Maximum results, then press Start.
{"subreddits": ["n8n"],"sort": "new","maxItems": 10}
Prefer links? Paste them into Reddit URLs instead and clear the Subreddits box.
What it costs
$3.00 per 1,000 saved results, plus a small Actor start fee ($0.0005 at the default 512 MB). The Pricing tab on this Actor is always authoritative.
Only saved dataset rows are charged. Duplicates, filtered-out candidates, scanned pages and unavailable comments are not. A post with nested comments is one row — its replies are not charged separately. Runs that return nothing still pay the start fee.
If you set a spending limit on a run, the Actor stops cleanly before the next result would exceed it and records stopReason: "budget-limit" in the run's SUMMARY. If a run returns fewer rows than you asked for, check there first.
What you can collect
| Source | Example |
|---|---|
| Subreddit feeds | technology, or /r/n8n/top/?t=week |
| Keyword search | a subreddit plus a keyword, or a keyword alone for all of Reddit |
| A single post | a canonical post URL, or a t3_ id |
| User profiles | /user/name/submitted/, /comments/, /overview/ |
| Comments | see Comments and reply trees below |
| Community details | /r/n8n/about/ |
| Flair templates | Flair templates mode, or /r/n8n/flairs/ |
Use an explicit public subreddit such as r/all rather than a personalized home feed, and canonical post URLs rather than share links. Private, quarantined and unverified-visibility content is never returned.
Search, ranking and filters
{"subreddits": ["n8n"],"search": "workflow","sort": "top","time": "month","minScore": 5,"maxItems": 25}
Searches can be ordered by relevance, hot, top, new or most comments, over an hour, day, week, month, year or all time. A plain feed URL keeps its own order; adding a keyword turns that feed into a search and uses Result order instead.
Keyword search uses Reddit's own search index, so it is not a guaranteed literal substring match.
Multi-word keywords are tokenized. Searching artificial intelligence asks Reddit for posts matching those words, not for that exact phrase, so some results contain only one of them. In a 249-row sample, about 69% of rows contained the full phrase; a single-word keyword scoped to subreddits scored 97%. This is how Reddit's index answers, not a filtering failure — if you need the exact phrase, filter the output on the title and body fields. Comment keyword search is different: it verifies every row literally contains your keyword before saving it, and measured 100% in the same test.
Optional filters: dates, post flair, minimum score and minimum comment count, plus NSFW selection. Filters are applied before results count against your limit.
{"startUrls": ["https://www.reddit.com/r/n8n/new/"],"flair": "Help","postedAfter": "2026-01-01","minComments": 2,"maxItems": 20}
Comments and reply trees
{"startUrls": ["https://www.reddit.com/r/SUBREDDIT/comments/POST_ID/TITLE/"],"mode": "posts-and-comments","maxItems": 1,"maxComments": 100,"maxCommentsPerPost": 100}
Two ways to get comments: set Maximum comments per post, which turns comment collection on and nests replies inside each post row, or choose a mode — Posts and their comments for nested replies, Comments from post URLs for one row per comment. Plain Auto with no comment limit returns posts only.
Comments can be ordered by confidence, top, new, controversial, old or Q&A. Reddit's displayed comment count often exceeds what any tool can retrieve, because it includes deleted, removed and unavailable replies.
Comment keyword search
Search comment text directly across subreddits, without searching parent posts first:
{"mode": "comment-search","search": "workflow","subreddits": ["n8n", "automation"],"sort": "new","maxItems": 100}
Every row is a comment that still exists and genuinely contains your keyword, filtered by the comment's creation date rather than the parent post's — old threads often hold recent comments. Only emitted matches count against your limit and your bill.
Advanced JSON options: keywordMatch (substring default, whole-word, phrase), caseSensitive, maxScannedComments (default 10000) and endPage (default 20 pages per subreddit).
This mode does not expand reply trees, and post/flair/score/NSFW filters do not apply to it. Reddit's index decides which comments are discoverable, so this is not an exhaustive scan of a subreddit's history.
Output
Posts include IDs, title, text and HTML, author, subreddit, permalink, created/edited times, score, upvote ratio, comment count, flair, media links, flags and source context. Comments include comment/post/parent IDs, author, text, score, timestamps, depth and post context. Profiles, communities and flair templates have their own fields.
Fields Reddit does not provide stay null rather than being invented, and removed content is not recovered.
Each run also writes a SUMMARY record with per-source progress, warnings, stop reasons and timings — the first place to look if a result count surprises you.
Limits
maxItems caps output across the whole run; comment settings cap comment expansion separately. Setting 0 means "no cap from me", not a promise of complete history — Reddit's own listing and search ceilings still apply.
Any single search returns at most ~250 results, whatever maxItems says. Measured: a query certain to have thousands of matches returned exactly 250 and reached back only two months. This is Reddit's ceiling, so a search asking for 1,000 rows legitimately finishes with fewer. To go deeper, split the work — separate runs per subreddit, per result order, or per date window — and de-duplicate on redditId.
Maximum comments is a total for the whole run, not per post. Leave it empty and each post keeps its own Maximum comments per post allowance; set it only when you want a hard ceiling. Replies nest inside the post row, so they are not billed as extra results.
Managed access is intended for low-volume workloads. It paces requests within a run and stops on rate limits or access failures rather than retrying around them.
Small cloud tests returned single rows in roughly 3–7 seconds and 25 rows in about 3.4 seconds. These are observations on small workloads, not guaranteed response times; visibility checks, comment expansion, filters and Reddit throttling all add time.
Notes for existing users
Customer JavaScript hooks (customMapFunction, extendOutputFunction), customer cookies, proxy overrides and browser overrides are disabled in this managed release and are rejected at runtime, not merely hidden. Remove them from older tasks before running.
Legacy JSON field names still work (subreddit, profile, postUrls, maxPosts, postSort, timeRange, includeComments, pageSize and others). A keyword combined with subreddit feeds now searches those communities; add "searchScope": "independent" to restore the older separate-jobs behaviour.
Operator credentials are stored as secret runtime configuration — never in inputs, source files, datasets or build arguments. The Actor only reads public Reddit content: it never votes, posts, messages or changes accounts.
For programmatic use, start the Actor with your Apify token and the same JSON input, wait for completion or use a webhook, then read its default dataset. See Apify run and build options.