Reddit Scraper: Posts & All Comments, No API Key avatar

Reddit Scraper: Posts & All Comments, No API Key

Pricing

Pay per event

Go to Apify Store
Reddit Scraper: Posts & All Comments, No API Key

Reddit Scraper: Posts & All Comments, No API Key

Scrape Reddit posts, complete comment threads, search results and user profiles without an API key. Expands every "load more comments", 4x more comments than first-page scrapers. Export to CSV, Excel or JSON. For data scientists, engineers and researchers.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Kalil Fagundes

Kalil Fagundes

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 hours ago

Last modified

Categories

Share

Reddit Scraper: Posts, Full Comment Threads & Search (No API Key)

Reddit Scraper extracts Reddit posts, complete comment threads, search results and user profiles without a Reddit account, API key or OAuth app. It expands every "load more comments" and "continue this thread" link, so you get the whole conversation, not the first ~500 comments. Export Reddit data to CSV, Excel, JSON or Google Sheets, or call it from Python, JavaScript, n8n, Make, Zapier and AI agents over MCP.

$1.55 per 1,000 posts, $1.00 per 1,000 comments. Built for software engineers, data scientists, ML and NLP engineers, academic researchers and market analysts.

Reddit Scraper in 30 seconds

  • What it scrapes: subreddit posts (hot, new, top, rising, controversial), Reddit search results, user profiles, and full comment threads from any post URL.
  • What makes it different: it opens every hidden comment. On a 2,873-comment r/AskReddit thread it returned 1,970 comments, 4x the 499 a first-page scraper gets.
  • No Reddit API needed: no account, no API key, no OAuth app, no rate-limit approval.
  • Output: one clean JSON item per post or comment, with the reply tree (parentId, depth), media URLs, and analysis flags such as deleted, removed, moderator and controversial.
  • Formats: JSON, CSV, Excel, XML, HTML, JSONL, plus API, webhooks and integrations.
  • Speed: about 2,000 comments a minute from a single thread.
  • Price: $1.55 per 1,000 posts and $1.00 per 1,000 comments. No start fee, no compute or proxy charges.

Why most Reddit scrapers miss most of the comments

When Reddit serves a thread it does not send every comment. It sends the first ~500 and hides the rest behind two kinds of links:

  • "load more comments": whole branches of replies, folded away
  • "continue this thread": deep reply chains, cut off after a few levels

A scraper that only reads the page Reddit sends gets that first slice and stops. Nothing in the output warns you: the dataset looks complete, but on any popular thread the biggest and deepest part of the discussion is missing. Sentiment scores, topic models, training sets and research findings built on that data rest on a biased sample, skewed toward the top-voted, top-level comments.

How this scraper gets every Reddit comment

Reddit Scraper follows every "load more comments" and every "continue this thread" link, in batches of up to 100 comments per request, until the thread is exhausted or your limit is reached. Each comment comes with parentId and depth, so you can rebuild the exact reply tree.

Measured on a real r/AskReddit thread in October 2026 (Reddit counter: 2,873 comments):

MethodComments returned
First page only (what most scrapers return)499
Reddit Scraper, all hidden comments expanded1,970

That is 4x more data from the same thread. The remaining difference to Reddit's counter is comments that were removed or deleted, which Reddit no longer serves to anyone. Every returned comment linked back to a parent in the dataset, with no orphans and correct depths.

Reddit Scraper vs. the Reddit API and other options

Reddit ScraperOfficial Reddit APIFirst-page scrapersCopy and paste
Account, API key or app approvalNot neededRequiredUsually not neededNot needed
Comments behind "load more comments"AllExtra calls you write yourselfMissingClick by hand
Reply tree (parentId, depth)YesYesOften missingNo
Export to CSV, Excel, JSONBuilt inCode it yourselfVariesManual
Cost$1.00–$1.55 per 1,000 itemsPaid for commercial use since 2023VariesYour time

What can you do with Reddit data?

  • Brand monitoring and social listening: track mentions of your brand, product or competitors across all of Reddit, with exact-phrase search and a searchQuery tag on every match. Schedule a daily run to follow sentiment over time.
  • Market research and customer pain points: find what people complain about, ask for and compare in their own words. The long-tail replies where users describe workarounds rarely sit in the first 500 comments.
  • Sentiment and opinion analysis: analyse the full discussion, not a sample biased toward the most upvoted takes. controversiality flags polarizing comments; isDeleted and isRemoved let you drop placeholders.
  • AI training data, fine-tuning and RAG: complete conversation trees for dialogue datasets, reply-pair extraction, classifier training and retrieval corpora for LLMs. postTitle on every comment keeps the context.
  • Academic research: reproducible collection of whole threads for discourse, polarization, misinformation and online community studies.
  • Content and trend research: see which posts, formats and topics take off in a subreddit, with scores, upvote ratios, comment counts and community size.

Why it fits your work

  • Data scientists and analysts: sentiment and opinion analysis on the whole thread, with stable, typed fields ready for pandas, R or BigQuery.
  • ML and NLP engineers: full reply trees for fine-tuning and RAG, with parentId and depth for reply-pair extraction.
  • Software engineers: clean JSON with stable field names, an API you can call from any language, and a schema that does not change between runs.
  • Academic researchers: whole threads collected the same way every time, with the time of collection in scrapedAt.
  • Market and product researchers: the real problems, workarounds and competitor complaints buried deep in threads.

What does Reddit Scraper do?

ModeWhat you give itWhat you get
πŸ“‹ Subreddit PostsOne or more subredditsPosts sorted by hot, new, top, rising or controversial, optionally with their full comment threads
πŸ” Search RedditOne keyword or a list of themMatching posts from all of Reddit or one subreddit, deduplicated across queries, optionally with full comment threads
πŸ‘€ User ProfileOne or more usernamesA user's posts, comments, or both
πŸ’¬ Post CommentsOne or more post URLsThe post and its complete comment thread, every hidden comment included

What data can I extract from Reddit?

PostsComments
Title, body text, authorComment text, author
Score, upvote ratio, comment countScore, controversial flag
Subreddit, subreddit member countSubreddit, title of the post it belongs to
Created and edited datesCreated and edited dates
Permalink, external link, domainPermalink
Image, gallery and video URLs (full size, every gallery image)postId, parentId, depth to rebuild the reply tree
Flair, author flair, post type, thumbnailWhether the author is the post's OP
NSFW, spoiler, pinned, locked, archived, OC, crosspost parentModerator or admin comment
Deleted or removed by moderatorsDeleted or removed by moderators
The search query that found it (Search mode)Time it was scraped

Every item also carries scrapedAt, the time it was collected, so repeated runs of the same search can be compared over time.

Fields built for analysis:

  • isDeleted / isRemoved flag the [deleted] and [removed] placeholders, so you can drop them before sentiment analysis or model training.
  • distinguished marks official moderator and admin posts, such as the AutoModerator notice pinned on most threads, so they do not skew your results.
  • controversiality is Reddit's own flag for comments with many up and down votes: a ready-made signal for polarizing opinions.
  • postTitle on every comment keeps the context when you export comments to a spreadsheet or feed them to an LLM.
  • searchQuery tells you which of your queries found each post when you search several phrases in one run.

How to scrape Reddit, step by step

  1. Click Try for free (or Start) and pick a Scraping Mode.
  2. Fill in the field for that mode: subreddits, search terms, usernames or post URLs.
  3. Choose whether to Extract Comments and set Max Results.
  4. Click Start and wait for the run to finish.
  5. Download the results from the Output tab as JSON, CSV, Excel, XML or HTML.

How to scrape a subreddit

Choose Subreddit Posts, enter one or more subreddit names (python or r/python both work), pick a sort and, for top, a time range such as year. Leave Extract Comments on to get each post's comments too.

How to get all comments from a Reddit post

Choose Post Comments, paste the post URL (www, old or plain reddit.com links all work), set Max Comments per Post to 0 and Max Results above the size of the thread. Every hidden comment is included.

How to search Reddit by keyword

Choose Search Reddit and enter a keyword, or several in Search Queries List. Wrap phrases in double quotes, such as "looking for an alternative to", for exact matching. Use Limit Search to Subreddit to search one community only, and sort by new, or top with a time range, for recent posts.

How to scrape a Reddit user's posts and comments

Choose User Profile, enter usernames (with or without u/) and pick posts, comments or both.

How to export Reddit data to CSV, Excel or Google Sheets

After the run, open the Output tab and pick a format: CSV and Excel open directly in spreadsheets. For Google Sheets, use the Apify Google Sheets integration or the n8n workflow included with the source. Filter on the type column to keep only posts or only comments.

How to download Reddit images and videos

Every post has mediaUrls: full-size image links, every image of a gallery in order, and Reddit-hosted video files. Links to YouTube and other sites are in externalUrl.

How to monitor Reddit automatically

Save your input as a task and add an Apify Schedule (for example daily), sorting by new. Connect a webhook, Slack, email, n8n or Make to get each run's results, and deduplicate on id across runs.

How to use Reddit data in Python, pandas or an LLM

Call the Actor with the Apify Python client (example below), load the items into a pandas DataFrame, or connect the Actor to Claude, ChatGPT-compatible clients, Cursor or VS Code through Apify's MCP server so an AI agent can search Reddit on its own.

How much does it cost to scrape Reddit?

EventPrice
Post saved to the dataset$1.55 per 1,000 ($0.00155 each)
Comment saved to the dataset$1.00 per 1,000 ($0.001 each)

Nothing else is billed: no start fee, no charge for compute time or residential proxies. Comments are priced lower on purpose: full threads are what this scraper is for, and they are where the volume is.

You scrapeYou pay
100 posts, no comments$0.16
10 posts with 1,990 comments$2.01
One 2,000-comment thread in full$2.00
1,000 posts with 9,000 comments$10.55

Set a maximum cost per run in the run options and the scraper stops at that amount: it never scrapes results you would not be charged for.

To spend less:

  • Uncheck Extract Comments when you only need posts.
  • Keep Max Results close to what you need. Posts and comments both count toward it.
  • Lower Max Comments per Post if you only need the top of each thread.

Input

Only mode plus the field for that mode is required. Everything else has sensible defaults.

FieldDefaultWhat it does
includeCommentstrueAlso fetch each post's comments in Subreddit and Search mode
modesubreddit_postssubreddit_posts, search, user_profile or post_comments
subredditsSubreddit names, with or without r/
sorthothot, new, top, rising, controversial
timeFilterweekhour, day, week, month, year, all. Only used when the sort is top
searchQueryOne search term. Wrap a phrase in double quotes for an exact match
searchQueriesListSeveral search terms in one run. Overrides searchQuery
searchSubredditLimit the search to one subreddit
searchSortrelevancerelevance, hot, top, new, comments
usernamesUsernames, with or without u/
userContentTypeoverviewoverview (posts and comments), submitted, comments
postUrlsFull post URLs (www, old or plain reddit.com)
maxCommentsPerPost100Comments per post. 0 means no limit
maxResults100Total items for the whole run, posts and comments together (1 to 10,000)
includeNsfwfalseInclude adult posts and their comments
proxyConfigurationResidentialApify Proxy settings. Residential IPs are strongly recommended

Input examples

Top posts of the year from a subreddit, with comments

{
"mode": "subreddit_posts",
"subreddits": ["enem", "vestibular"],
"sort": "top",
"timeFilter": "year",
"includeComments": true,
"maxCommentsPerPost": 0,
"maxResults": 2000
}

Brand monitoring with several exact phrases

{
"mode": "search",
"searchQueriesList": ["\"notion alternative\"", "\"switched from notion\""],
"searchSort": "new",
"includeComments": false,
"maxResults": 300
}

Every comment of a few threads

{
"mode": "post_comments",
"postUrls": ["https://www.reddit.com/r/AskReddit/comments/1wrj0s2/"],
"maxCommentsPerPost": 0,
"maxResults": 5000
}

A user's recent comments

{
"mode": "user_profile",
"usernames": ["spez"],
"userContentType": "comments",
"maxResults": 200
}

Output

Each item is either a post or a comment, so you can filter on type. Comments follow their post in the dataset.

Post example

searchQuery only appears on posts found in Search mode.

{
"type": "post",
"id": "1w6o67w",
"subreddit": "enem",
"title": "Um pessoal desse sub nos ΓΊltimos dias",
"author": "example_user",
"selftext": "",
"url": "https://www.reddit.com/r/enem/comments/1w6o67w/um_pessoal_desse_sub_nos_ultimos_dias/",
"externalUrl": "https://i.redd.it/5znjedyv6enh1.jpeg",
"mediaUrls": [
"https://i.redd.it/5znjedyv6enh1.jpeg"
],
"score": 2071,
"upvoteRatio": 0.99,
"numComments": 37,
"subredditSubscribers": 412000,
"created": "2026-09-03T23:46:38+00:00",
"editedAt": "",
"isNSFW": false,
"isSpoiler": false,
"isPinned": false,
"isLocked": false,
"isArchived": false,
"isDeleted": false,
"isRemoved": false,
"distinguished": "",
"flair": "Humor",
"domain": "i.redd.it",
"isVideo": false,
"thumbnail": "",
"postHint": "image",
"isOriginalContent": false,
"authorFlair": "",
"crosspostParent": "",
"mediaOnly": false,
"isGallery": false,
"scrapedAt": "2026-10-02T19:42:58+00:00",
"searchQuery": "\"enem 2026\""
}

Comment example

{
"type": "comment",
"id": "p7ok2xq",
"postId": "1w6o67w",
"postTitle": "Um pessoal desse sub nos ΓΊltimos dias",
"parentId": "p7oj9bd",
"subreddit": "enem",
"author": "another_user",
"body": "Literalmente eu na semana da prova.",
"score": 57,
"controversiality": 0,
"created": "2026-09-04T00:12:51+00:00",
"editedAt": "",
"depth": 1,
"isSubmitter": false,
"isDeleted": false,
"isRemoved": false,
"distinguished": "",
"url": "https://www.reddit.com/r/enem/comments/1w6o67w/comment/p7ok2xq/",
"scrapedAt": "2026-10-02T19:42:58+00:00"
}

parentId is the comment being replied to, or the post id for a top-level comment (depth 0). Comments Reddit shows on first load come in thread order; comments loaded from "load more comments" links come after them, so use parentId to rebuild the exact tree.

Limits and tips

  • About 1,000 posts per listing. Reddit stops paginating any listing (a subreddit sort, a search, a user profile) at roughly 1,000 items. That is a Reddit limit, not one of this Actor. To go further, run the same subreddit with several sorts (new, top with different timeFilter values, controversial) or several search terms, then deduplicate on id.
  • Comments are not affected by that limit. A single thread can return thousands of comments.
  • Comment counts differ from Reddit's. numComments includes removed and deleted comments, which Reddit no longer serves, so a thread usually returns fewer comments than that number.
  • One long thread cannot use up a run. In Subreddit and Search mode, each post may use at most a fifth of maxResults for its comments, so a run always reaches several posts. The log says when this lowers maxCommentsPerPost. Post Comments mode has no such cap.
  • maxResults is shared across queries in searchQueriesList and is used in order, so later queries get nothing once it runs out. Allow at least 25 per query.
  • Quote your search phrases. Unquoted, Reddit matches loose words and returns mostly unrelated posts.
  • relevance favours old, popular posts. For recent results use new, or top with a timeFilter.
  • Only public content. Private, quarantined-by-invite and banned subreddits, and deleted accounts, return nothing.

FAQ

Reddit Scraper only reads public pages that any logged-out visitor can see. Scraping public data is generally allowed, but you are responsible for how you use it: follow Reddit's terms, data protection laws such as GDPR and LGPD, and do not collect personal data without a legitimate reason. If unsure, ask a lawyer.

Can I scrape Reddit without an API key?

Yes. Reddit Scraper needs no Reddit account, API key or OAuth app. It reads Reddit's public pages through a real headless browser.

What is the best alternative to the Reddit API?

For collecting public posts and comments, a scraper like this one avoids app approval, rate-limit negotiations and Reddit's paid commercial API access, which started in 2023. You pay only per result, and you get full comment threads that would otherwise take many extra API calls.

Is there an alternative to Pushshift?

Pushshift is no longer publicly available. Reddit Scraper covers the common Pushshift use cases for current data: subreddit posts, keyword search, user histories and full comment threads. It cannot fetch posts beyond Reddit's ~1,000-item listing limit, so it is not a full historical archive.

How do I get all comments from a Reddit thread?

Use Post Comments mode with maxCommentsPerPost set to 0 and a maxResults larger than the thread. The Actor follows every "load more comments" and "continue this thread" link, so you get the whole thread, not the ~500 comments Reddit shows on first load.

Why do other Reddit scrapers return fewer comments for the same thread?

Reddit only sends the first ~500 comments of a thread and hides the rest behind "load more comments" and "continue this thread" links. Scrapers that do not open those links stop there without telling you. This one opens all of them.

Why is the comment count lower than the number Reddit shows?

Reddit's counter includes removed and deleted comments, which it no longer serves to anyone. Every comment Reddit still serves is returned.

How many Reddit posts can I scrape?

Up to 10,000 results per run, posts and comments together. Each Reddit listing (one subreddit sort, one search, one profile) stops at about 1,000 posts, a limit set by Reddit. Combine sorts, time ranges and search terms to collect more, then deduplicate on id.

Can I scrape Reddit posts from a specific date range?

Not by exact dates: Reddit does not offer a date filter. Use top with a time range (hour, day, week, month, year, all), or new and filter on the created field.

How fast is it?

About 2,000 comments a minute from a single thread, and up to 25 posts per request when comments are off. Requests are paced to avoid blocks.

Can I scrape private or NSFW subreddits?

Private subreddits: no, only public content. NSFW posts are skipped by default; turn on Include NSFW Content to include them and their comments.

Does it download images and videos?

It returns the links: mediaUrls has full-size images, every image of a gallery and Reddit-hosted videos. Download them with any tool or script.

Can AI agents like Claude or ChatGPT use this scraper?

Yes. Add it to any MCP client (Claude, Cursor, VS Code and others) through Apify's MCP server, as shown below, and the agent can search Reddit and read full threads by itself.

Do I need proxies?

Yes. Reddit blocks most datacenter IPs. Keep the default Residential Apify Proxy group; its cost is already included in the price.

What happens with a subreddit, user or URL that doesn't exist?

The Actor logs a warning, skips it and carries on with the rest of the run. A run that finds nothing at all ends with a status message explaining the likely causes.

Can I schedule runs or get results through the API?

Yes. Use Apify Schedules for recurring runs, and the API, webhooks or integrations (n8n, Make, Zapier, Google Sheets) to collect the dataset.

Use the Reddit Scraper API from code

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<APIFY_TOKEN>' });
const run = await client.actor('kalilfagundes.w/reddit-scraper').call({
mode: 'subreddit_posts',
subreddits: ['enem'],
sort: 'top',
timeFilter: 'month',
maxResults: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length);

Python

from apify_client import ApifyClient
client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("kalilfagundes.w/reddit-scraper").call(run_input={
"mode": "search",
"searchQuery": "\"redaΓ§Γ£o nota 1000\"",
"searchSort": "new",
"maxResults": 200,
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items))

MCP (Claude, Cursor, VS Code and other AI clients)

{
"mcpServers": {
"reddit-scraper": {
"url": "https://mcp.apify.com?tools=kalilfagundes.w/reddit-scraper",
"headers": { "Authorization": "Bearer <APIFY_TOKEN>" }
}
}
}

An n8n workflow that searches Reddit for complaint phrases and logs them to Google Sheets is included in the source under examples/n8n.

Changelog

  • 1.5: richer output. Posts get mediaUrls (full-size images, every gallery image, Reddit videos), subredditSubscribers, isLocked, isArchived and, in Search mode, searchQuery. Comments get postTitle and controversiality. Both get isDeleted, isRemoved, distinguished, editedAt (a date, replacing the old edited number) and scrapedAt. Removed awards and isPromoted, which Reddit no longer fills (awards were discontinued in 2023).
  • 1.4: follows "load more comments" and "continue this thread" links, so threads are complete instead of stopping at ~500 comments. New parentId field on comments. A run migrated to another server mid-run no longer repeats items. Posts and comments are priced separately, and runs stop at the user's maximum cost.
  • 1.3: "Extract Comments" moved to the top of the form and enabled by default.

Feedback

Found a bug or want a feature? Open an issue on the Issues tab.