Reddit Scraper: Posts & All Comments, No API Key
Pricing
Pay per event
Reddit Scraper: Posts & All Comments, No API Key
Scrape Reddit posts, complete comment threads, search results and user profiles without an API key. Expands every "load more comments", 4x more comments than first-page scrapers. Export to CSV, Excel or JSON. For data scientists, engineers and researchers.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Kalil Fagundes
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 hours ago
Last modified
Categories
Share
Reddit Scraper: Posts, Full Comment Threads & Search (No API Key)
Reddit Scraper extracts Reddit posts, complete comment threads, search results and user profiles without a Reddit account, API key or OAuth app. It expands every "load more comments" and "continue this thread" link, so you get the whole conversation, not the first ~500 comments. Export Reddit data to CSV, Excel, JSON or Google Sheets, or call it from Python, JavaScript, n8n, Make, Zapier and AI agents over MCP.
$1.55 per 1,000 posts, $1.00 per 1,000 comments. Built for software engineers, data scientists, ML and NLP engineers, academic researchers and market analysts.
Reddit Scraper in 30 seconds
- What it scrapes: subreddit posts (hot, new, top, rising, controversial), Reddit search results, user profiles, and full comment threads from any post URL.
- What makes it different: it opens every hidden comment. On a 2,873-comment r/AskReddit thread it returned 1,970 comments, 4x the 499 a first-page scraper gets.
- No Reddit API needed: no account, no API key, no OAuth app, no rate-limit approval.
- Output: one clean JSON item per post or comment, with the reply tree (
parentId,depth), media URLs, and analysis flags such as deleted, removed, moderator and controversial. - Formats: JSON, CSV, Excel, XML, HTML, JSONL, plus API, webhooks and integrations.
- Speed: about 2,000 comments a minute from a single thread.
- Price: $1.55 per 1,000 posts and $1.00 per 1,000 comments. No start fee, no compute or proxy charges.
Why most Reddit scrapers miss most of the comments
When Reddit serves a thread it does not send every comment. It sends the first ~500 and hides the rest behind two kinds of links:
- "load more comments": whole branches of replies, folded away
- "continue this thread": deep reply chains, cut off after a few levels
A scraper that only reads the page Reddit sends gets that first slice and stops. Nothing in the output warns you: the dataset looks complete, but on any popular thread the biggest and deepest part of the discussion is missing. Sentiment scores, topic models, training sets and research findings built on that data rest on a biased sample, skewed toward the top-voted, top-level comments.
How this scraper gets every Reddit comment
Reddit Scraper follows every "load more comments" and every "continue this thread" link, in batches of up to 100 comments per request, until the thread is exhausted or your limit is reached. Each comment comes with parentId and depth, so you can rebuild the exact reply tree.
Measured on a real r/AskReddit thread in October 2026 (Reddit counter: 2,873 comments):
| Method | Comments returned |
|---|---|
| First page only (what most scrapers return) | 499 |
| Reddit Scraper, all hidden comments expanded | 1,970 |
That is 4x more data from the same thread. The remaining difference to Reddit's counter is comments that were removed or deleted, which Reddit no longer serves to anyone. Every returned comment linked back to a parent in the dataset, with no orphans and correct depths.
Reddit Scraper vs. the Reddit API and other options
| Reddit Scraper | Official Reddit API | First-page scrapers | Copy and paste | |
|---|---|---|---|---|
| Account, API key or app approval | Not needed | Required | Usually not needed | Not needed |
| Comments behind "load more comments" | All | Extra calls you write yourself | Missing | Click by hand |
Reply tree (parentId, depth) | Yes | Yes | Often missing | No |
| Export to CSV, Excel, JSON | Built in | Code it yourself | Varies | Manual |
| Cost | $1.00β$1.55 per 1,000 items | Paid for commercial use since 2023 | Varies | Your time |
What can you do with Reddit data?
- Brand monitoring and social listening: track mentions of your brand, product or competitors across all of Reddit, with exact-phrase search and a
searchQuerytag on every match. Schedule a daily run to follow sentiment over time. - Market research and customer pain points: find what people complain about, ask for and compare in their own words. The long-tail replies where users describe workarounds rarely sit in the first 500 comments.
- Sentiment and opinion analysis: analyse the full discussion, not a sample biased toward the most upvoted takes.
controversialityflags polarizing comments;isDeletedandisRemovedlet you drop placeholders. - AI training data, fine-tuning and RAG: complete conversation trees for dialogue datasets, reply-pair extraction, classifier training and retrieval corpora for LLMs.
postTitleon every comment keeps the context. - Academic research: reproducible collection of whole threads for discourse, polarization, misinformation and online community studies.
- Content and trend research: see which posts, formats and topics take off in a subreddit, with scores, upvote ratios, comment counts and community size.
Why it fits your work
- Data scientists and analysts: sentiment and opinion analysis on the whole thread, with stable, typed fields ready for pandas, R or BigQuery.
- ML and NLP engineers: full reply trees for fine-tuning and RAG, with
parentIdanddepthfor reply-pair extraction. - Software engineers: clean JSON with stable field names, an API you can call from any language, and a schema that does not change between runs.
- Academic researchers: whole threads collected the same way every time, with the time of collection in
scrapedAt. - Market and product researchers: the real problems, workarounds and competitor complaints buried deep in threads.
What does Reddit Scraper do?
| Mode | What you give it | What you get |
|---|---|---|
| π Subreddit Posts | One or more subreddits | Posts sorted by hot, new, top, rising or controversial, optionally with their full comment threads |
| π Search Reddit | One keyword or a list of them | Matching posts from all of Reddit or one subreddit, deduplicated across queries, optionally with full comment threads |
| π€ User Profile | One or more usernames | A user's posts, comments, or both |
| π¬ Post Comments | One or more post URLs | The post and its complete comment thread, every hidden comment included |
What data can I extract from Reddit?
| Posts | Comments |
|---|---|
| Title, body text, author | Comment text, author |
| Score, upvote ratio, comment count | Score, controversial flag |
| Subreddit, subreddit member count | Subreddit, title of the post it belongs to |
| Created and edited dates | Created and edited dates |
| Permalink, external link, domain | Permalink |
| Image, gallery and video URLs (full size, every gallery image) | postId, parentId, depth to rebuild the reply tree |
| Flair, author flair, post type, thumbnail | Whether the author is the post's OP |
| NSFW, spoiler, pinned, locked, archived, OC, crosspost parent | Moderator or admin comment |
| Deleted or removed by moderators | Deleted or removed by moderators |
| The search query that found it (Search mode) | Time it was scraped |
Every item also carries scrapedAt, the time it was collected, so repeated runs of the same search can be compared over time.
Fields built for analysis:
isDeleted/isRemovedflag the[deleted]and[removed]placeholders, so you can drop them before sentiment analysis or model training.distinguishedmarks official moderator and admin posts, such as the AutoModerator notice pinned on most threads, so they do not skew your results.controversialityis Reddit's own flag for comments with many up and down votes: a ready-made signal for polarizing opinions.postTitleon every comment keeps the context when you export comments to a spreadsheet or feed them to an LLM.searchQuerytells you which of your queries found each post when you search several phrases in one run.
How to scrape Reddit, step by step
- Click Try for free (or Start) and pick a Scraping Mode.
- Fill in the field for that mode: subreddits, search terms, usernames or post URLs.
- Choose whether to Extract Comments and set Max Results.
- Click Start and wait for the run to finish.
- Download the results from the Output tab as JSON, CSV, Excel, XML or HTML.
How to scrape a subreddit
Choose Subreddit Posts, enter one or more subreddit names (python or r/python both work), pick a sort and, for top, a time range such as year. Leave Extract Comments on to get each post's comments too.
How to get all comments from a Reddit post
Choose Post Comments, paste the post URL (www, old or plain reddit.com links all work), set Max Comments per Post to 0 and Max Results above the size of the thread. Every hidden comment is included.
How to search Reddit by keyword
Choose Search Reddit and enter a keyword, or several in Search Queries List. Wrap phrases in double quotes, such as "looking for an alternative to", for exact matching. Use Limit Search to Subreddit to search one community only, and sort by new, or top with a time range, for recent posts.
How to scrape a Reddit user's posts and comments
Choose User Profile, enter usernames (with or without u/) and pick posts, comments or both.
How to export Reddit data to CSV, Excel or Google Sheets
After the run, open the Output tab and pick a format: CSV and Excel open directly in spreadsheets. For Google Sheets, use the Apify Google Sheets integration or the n8n workflow included with the source. Filter on the type column to keep only posts or only comments.
How to download Reddit images and videos
Every post has mediaUrls: full-size image links, every image of a gallery in order, and Reddit-hosted video files. Links to YouTube and other sites are in externalUrl.
How to monitor Reddit automatically
Save your input as a task and add an Apify Schedule (for example daily), sorting by new. Connect a webhook, Slack, email, n8n or Make to get each run's results, and deduplicate on id across runs.
How to use Reddit data in Python, pandas or an LLM
Call the Actor with the Apify Python client (example below), load the items into a pandas DataFrame, or connect the Actor to Claude, ChatGPT-compatible clients, Cursor or VS Code through Apify's MCP server so an AI agent can search Reddit on its own.
How much does it cost to scrape Reddit?
| Event | Price |
|---|---|
| Post saved to the dataset | $1.55 per 1,000 ($0.00155 each) |
| Comment saved to the dataset | $1.00 per 1,000 ($0.001 each) |
Nothing else is billed: no start fee, no charge for compute time or residential proxies. Comments are priced lower on purpose: full threads are what this scraper is for, and they are where the volume is.
| You scrape | You pay |
|---|---|
| 100 posts, no comments | $0.16 |
| 10 posts with 1,990 comments | $2.01 |
| One 2,000-comment thread in full | $2.00 |
| 1,000 posts with 9,000 comments | $10.55 |
Set a maximum cost per run in the run options and the scraper stops at that amount: it never scrapes results you would not be charged for.
To spend less:
- Uncheck Extract Comments when you only need posts.
- Keep Max Results close to what you need. Posts and comments both count toward it.
- Lower Max Comments per Post if you only need the top of each thread.
Input
Only mode plus the field for that mode is required. Everything else has sensible defaults.
| Field | Default | What it does |
|---|---|---|
includeComments | true | Also fetch each post's comments in Subreddit and Search mode |
mode | subreddit_posts | subreddit_posts, search, user_profile or post_comments |
subreddits | Subreddit names, with or without r/ | |
sort | hot | hot, new, top, rising, controversial |
timeFilter | week | hour, day, week, month, year, all. Only used when the sort is top |
searchQuery | One search term. Wrap a phrase in double quotes for an exact match | |
searchQueriesList | Several search terms in one run. Overrides searchQuery | |
searchSubreddit | Limit the search to one subreddit | |
searchSort | relevance | relevance, hot, top, new, comments |
usernames | Usernames, with or without u/ | |
userContentType | overview | overview (posts and comments), submitted, comments |
postUrls | Full post URLs (www, old or plain reddit.com) | |
maxCommentsPerPost | 100 | Comments per post. 0 means no limit |
maxResults | 100 | Total items for the whole run, posts and comments together (1 to 10,000) |
includeNsfw | false | Include adult posts and their comments |
proxyConfiguration | Residential | Apify Proxy settings. Residential IPs are strongly recommended |
Input examples
Top posts of the year from a subreddit, with comments
{"mode": "subreddit_posts","subreddits": ["enem", "vestibular"],"sort": "top","timeFilter": "year","includeComments": true,"maxCommentsPerPost": 0,"maxResults": 2000}
Brand monitoring with several exact phrases
{"mode": "search","searchQueriesList": ["\"notion alternative\"", "\"switched from notion\""],"searchSort": "new","includeComments": false,"maxResults": 300}
Every comment of a few threads
{"mode": "post_comments","postUrls": ["https://www.reddit.com/r/AskReddit/comments/1wrj0s2/"],"maxCommentsPerPost": 0,"maxResults": 5000}
A user's recent comments
{"mode": "user_profile","usernames": ["spez"],"userContentType": "comments","maxResults": 200}
Output
Each item is either a post or a comment, so you can filter on type. Comments follow their post in the dataset.
Post example
searchQuery only appears on posts found in Search mode.
{"type": "post","id": "1w6o67w","subreddit": "enem","title": "Um pessoal desse sub nos ΓΊltimos dias","author": "example_user","selftext": "","url": "https://www.reddit.com/r/enem/comments/1w6o67w/um_pessoal_desse_sub_nos_ultimos_dias/","externalUrl": "https://i.redd.it/5znjedyv6enh1.jpeg","mediaUrls": ["https://i.redd.it/5znjedyv6enh1.jpeg"],"score": 2071,"upvoteRatio": 0.99,"numComments": 37,"subredditSubscribers": 412000,"created": "2026-09-03T23:46:38+00:00","editedAt": "","isNSFW": false,"isSpoiler": false,"isPinned": false,"isLocked": false,"isArchived": false,"isDeleted": false,"isRemoved": false,"distinguished": "","flair": "Humor","domain": "i.redd.it","isVideo": false,"thumbnail": "","postHint": "image","isOriginalContent": false,"authorFlair": "","crosspostParent": "","mediaOnly": false,"isGallery": false,"scrapedAt": "2026-10-02T19:42:58+00:00","searchQuery": "\"enem 2026\""}
Comment example
{"type": "comment","id": "p7ok2xq","postId": "1w6o67w","postTitle": "Um pessoal desse sub nos ΓΊltimos dias","parentId": "p7oj9bd","subreddit": "enem","author": "another_user","body": "Literalmente eu na semana da prova.","score": 57,"controversiality": 0,"created": "2026-09-04T00:12:51+00:00","editedAt": "","depth": 1,"isSubmitter": false,"isDeleted": false,"isRemoved": false,"distinguished": "","url": "https://www.reddit.com/r/enem/comments/1w6o67w/comment/p7ok2xq/","scrapedAt": "2026-10-02T19:42:58+00:00"}
parentId is the comment being replied to, or the post id for a top-level comment (depth 0). Comments Reddit shows on first load come in thread order; comments loaded from "load more comments" links come after them, so use parentId to rebuild the exact tree.
Limits and tips
- About 1,000 posts per listing. Reddit stops paginating any listing (a subreddit sort, a search, a user profile) at roughly 1,000 items. That is a Reddit limit, not one of this Actor. To go further, run the same subreddit with several sorts (
new,topwith differenttimeFiltervalues,controversial) or several search terms, then deduplicate onid. - Comments are not affected by that limit. A single thread can return thousands of comments.
- Comment counts differ from Reddit's.
numCommentsincludes removed and deleted comments, which Reddit no longer serves, so a thread usually returns fewer comments than that number. - One long thread cannot use up a run. In Subreddit and Search mode, each post may use at most a fifth of
maxResultsfor its comments, so a run always reaches several posts. The log says when this lowersmaxCommentsPerPost. Post Comments mode has no such cap. maxResultsis shared across queries insearchQueriesListand is used in order, so later queries get nothing once it runs out. Allow at least 25 per query.- Quote your search phrases. Unquoted, Reddit matches loose words and returns mostly unrelated posts.
relevancefavours old, popular posts. For recent results usenew, ortopwith atimeFilter.- Only public content. Private, quarantined-by-invite and banned subreddits, and deleted accounts, return nothing.
FAQ
Is it legal to scrape Reddit?
Reddit Scraper only reads public pages that any logged-out visitor can see. Scraping public data is generally allowed, but you are responsible for how you use it: follow Reddit's terms, data protection laws such as GDPR and LGPD, and do not collect personal data without a legitimate reason. If unsure, ask a lawyer.
Can I scrape Reddit without an API key?
Yes. Reddit Scraper needs no Reddit account, API key or OAuth app. It reads Reddit's public pages through a real headless browser.
What is the best alternative to the Reddit API?
For collecting public posts and comments, a scraper like this one avoids app approval, rate-limit negotiations and Reddit's paid commercial API access, which started in 2023. You pay only per result, and you get full comment threads that would otherwise take many extra API calls.
Is there an alternative to Pushshift?
Pushshift is no longer publicly available. Reddit Scraper covers the common Pushshift use cases for current data: subreddit posts, keyword search, user histories and full comment threads. It cannot fetch posts beyond Reddit's ~1,000-item listing limit, so it is not a full historical archive.
How do I get all comments from a Reddit thread?
Use Post Comments mode with maxCommentsPerPost set to 0 and a maxResults larger than the thread. The Actor follows every "load more comments" and "continue this thread" link, so you get the whole thread, not the ~500 comments Reddit shows on first load.
Why do other Reddit scrapers return fewer comments for the same thread?
Reddit only sends the first ~500 comments of a thread and hides the rest behind "load more comments" and "continue this thread" links. Scrapers that do not open those links stop there without telling you. This one opens all of them.
Why is the comment count lower than the number Reddit shows?
Reddit's counter includes removed and deleted comments, which it no longer serves to anyone. Every comment Reddit still serves is returned.
How many Reddit posts can I scrape?
Up to 10,000 results per run, posts and comments together. Each Reddit listing (one subreddit sort, one search, one profile) stops at about 1,000 posts, a limit set by Reddit. Combine sorts, time ranges and search terms to collect more, then deduplicate on id.
Can I scrape Reddit posts from a specific date range?
Not by exact dates: Reddit does not offer a date filter. Use top with a time range (hour, day, week, month, year, all), or new and filter on the created field.
How fast is it?
About 2,000 comments a minute from a single thread, and up to 25 posts per request when comments are off. Requests are paced to avoid blocks.
Can I scrape private or NSFW subreddits?
Private subreddits: no, only public content. NSFW posts are skipped by default; turn on Include NSFW Content to include them and their comments.
Does it download images and videos?
It returns the links: mediaUrls has full-size images, every image of a gallery and Reddit-hosted videos. Download them with any tool or script.
Can AI agents like Claude or ChatGPT use this scraper?
Yes. Add it to any MCP client (Claude, Cursor, VS Code and others) through Apify's MCP server, as shown below, and the agent can search Reddit and read full threads by itself.
Do I need proxies?
Yes. Reddit blocks most datacenter IPs. Keep the default Residential Apify Proxy group; its cost is already included in the price.
What happens with a subreddit, user or URL that doesn't exist?
The Actor logs a warning, skips it and carries on with the rest of the run. A run that finds nothing at all ends with a status message explaining the likely causes.
Can I schedule runs or get results through the API?
Yes. Use Apify Schedules for recurring runs, and the API, webhooks or integrations (n8n, Make, Zapier, Google Sheets) to collect the dataset.
Use the Reddit Scraper API from code
JavaScript
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: '<APIFY_TOKEN>' });const run = await client.actor('kalilfagundes.w/reddit-scraper').call({mode: 'subreddit_posts',subreddits: ['enem'],sort: 'top',timeFilter: 'month',maxResults: 500,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(items.length);
Python
from apify_client import ApifyClientclient = ApifyClient("<APIFY_TOKEN>")run = client.actor("kalilfagundes.w/reddit-scraper").call(run_input={"mode": "search","searchQuery": "\"redaΓ§Γ£o nota 1000\"","searchSort": "new","maxResults": 200,})items = client.dataset(run["defaultDatasetId"]).list_items().itemsprint(len(items))
MCP (Claude, Cursor, VS Code and other AI clients)
{"mcpServers": {"reddit-scraper": {"url": "https://mcp.apify.com?tools=kalilfagundes.w/reddit-scraper","headers": { "Authorization": "Bearer <APIFY_TOKEN>" }}}}
An n8n workflow that searches Reddit for complaint phrases and logs them to Google Sheets is included in the source under examples/n8n.
Changelog
- 1.5: richer output. Posts get
mediaUrls(full-size images, every gallery image, Reddit videos),subredditSubscribers,isLocked,isArchivedand, in Search mode,searchQuery. Comments getpostTitleandcontroversiality. Both getisDeleted,isRemoved,distinguished,editedAt(a date, replacing the oldeditednumber) andscrapedAt. RemovedawardsandisPromoted, which Reddit no longer fills (awards were discontinued in 2023). - 1.4: follows "load more comments" and "continue this thread" links, so threads are complete instead of stopping at ~500 comments. New
parentIdfield on comments. A run migrated to another server mid-run no longer repeats items. Posts and comments are priced separately, and runs stop at the user's maximum cost. - 1.3: "Extract Comments" moved to the top of the form and enabled by default.
Feedback
Found a bug or want a feature? Open an issue on the Issues tab.