Reddit Scraper API — Posts, Comments, Search, Users avatar

Reddit Scraper API — Posts, Comments, Search, Users

Pricing

from $3.00 / 1,000 post, community or user rows

Go to Apify Store
Reddit Scraper API — Posts, Comments, Search, Users

Reddit Scraper API — Posts, Comments, Search, Users

Used by brand-monitoring teams, market researchers and AI pipelines that need Reddit at volume. Reddit scraper for posts, comments, subreddits, users and search in one flat table, with score and upvote ratio on every row. No API key, no login. Filters run before billing, so discarded rows are free.

Pricing

from $3.00 / 1,000 post, community or user rows

Rating

0.0

(0)

Developer

Blackcube

Blackcube

Maintained by Community

Actor stats

0

Bookmarked

7

Total users

3

Monthly active users

12 hours ago

Last modified

Share

More from this account: YouTube Transcript Suite · Website Contact & Email Suite · Career Site & ATS Jobs Suite · Google News Suite · Keyword Research Suite · Shopify Store Intelligence Suite · eBay Data Suite · Amazon Reviews Suite · Meta Ad Library · Vinted · Trustpilot Review Intelligence Suite · App Store & Google Play Reviews Suite · Etsy Research Suite · YouTube Comments Intelligence Suite · Telegram Channel Intelligence Suite · Business Reviews Suite · Google Trends Suite · Amazon Product Data Suite · Polymarket Order Book · Nordic Marketplaces Suite · Google Sheets Suite · TikTok Suite · Snapchat Suite · LinkedIn Public Data Suite · Instagram Suite · Contact Validation Suite · Google Maps Suite · X / Twitter

Name a community. Get back a flat table of its posts with the score, the upvote ratio, the text and the flair already in it, and, if you ask, the first page of every comment thread.

A community name, a community address in any sort and time window, keywords to look for inside a community, or a list of five hundred of them. Every one comes back in the same flat schema. No API key. No approval queue. No login. No developer token.

What this Actor cannot read. Single post and comment links, redd.it short links, user profiles, About pages and site-wide search. This Actor reads them through Reddit's JSON API, which since September 2026 answers only visitors who pass Reddit's verification page, and this Actor never gets past that page. Each such input comes back as a free blocked row that says so. It is never billed, and it is never passed off as "no posts".

You pay for rows you actually receive. Starting a run costs a fraction of a cent: $0.003, falling to $0.001 by volume (to $0.0015 from 6 October 2026). That is the only tiered start fee in this category, and even the $0.003 Free-plan fee is 6.7× to 72× below what the largest Actors here charge. After that, everything else is free: a community that turns out to be private, a keyword with no matches, an input this Actor cannot read, every row a filter discarded, every duplicate, and every row a previous run already gave you. They come back as rows explaining themselves and they never reach your invoice.

Unofficial. Not affiliated with, endorsed by, or connected to Reddit. It reads only what Reddit publishes publicly to logged-out visitors.


Why this one

1. The score and the upvote ratio are on every post row, by default

This is the field buyers in this category most often do not get. There is no flag to switch on, no "detailed mode" that halves your throughput, and no best-effort caveat.

On the last 1,500-post benchmark run — five communities, new sort, 34 seconds — score, upvoteRatio, createdAt and title were populated on 1,500 of 1,500 rows. author was populated on 1,499; the one gap was a deleted account, which is returned as null rather than as the literal text [deleted].

A missing value is always null, never 0. A post with no score and a post scored zero are different facts, and a filter on score > 0 must not silently swallow the second kind.

2. Paste the URL you already have

You pasteWhat you get
rust or r/rustThe community's posts
https://www.reddit.com/r/rust/top/?t=weekThat exact feed, sort and window already applied
https://www.reddit.com/r/rust/search/?q=asyncr/rust's feed, filtered to the posts containing every word (see below)
https://www.reddit.com/r/rust/comments/1vq17rg/…, https://redd.it/1vq17rg or a comment linkA free blocked row: this Actor reads single posts through Reddit's JSON API, which answers only verified visitors
https://www.reddit.com/user/spez/ or u/spezA free blocked row, for the same reason
https://www.reddit.com/search/?q=rust+asyncA free blocked row, for the same reason
/r/rust, old.reddit.com/…, sh.reddit.com/…All understood

Anything it cannot read comes back as a free row naming the input and the reason. It is never guessed at. A bare word like marketing is not silently treated as a community — that is far more likely to be a search term pasted into the wrong box, and resolving it to a community would bill you for data you never asked for.

3. A spending cap that actually binds

Total row limit is a hard ceiling on billable rows for the entire run, counting posts and comments together, and it is enforced at the moment of delivery rather than checked hopefully in a loop. Set it to 1,000 and you get 1,000 — not 1,004 because four workers were in flight, and not 15,000 because the limit was per-source and you supplied twenty sources.

It survives interruption, too: if the platform migrates the container mid-run, the count continues from what was already delivered instead of restarting.

4. Repeat runs only charge for what is new

Turn on Only rows I have not had before and a scheduled run returns — and bills for — only what has appeared since last time. The first run returns everything; every run after it returns the delta.

Run 1Run 2Run 30
Normal500 rows500 rows500 rows
Only new500 rows~20 rows~20 rows

Each schedule keeps its own memory, and editing the input list does not reset it — adding one more community to a daily job does not make the next run re-deliver and re-charge for everything it already gave you.

5. Date windows and keyword filters, and you are not billed for what they discard

Filtering is the most-requested and least-available feature in this category, and the reason is always the same: when an Actor bills for every row it touches, throwing rows away costs the developer money. Discarding is cheap here, so it is free for you.

  • Posted after / Posted before — an absolute date (2026-01-31) or an age (7 days, 12h, 3 months). Use an age in a scheduled run; a fixed date silently goes stale.
  • Must contain / Must not contain — keyword filters over title, body and flair.
  • Minimum score.

Every row these remove is discarded before billing, and the run log tells you how many and why.

6. Comment threads, and an honest count of what is missing

Each post's thread comes from the page Reddit renders for it: the first page, about 25 comments, nested, with full text, score, depth and author. Reddit's deeper "load more" pages are not read. When Reddit counts more comments on a post than that page gave, and your Comments per post limit was not what stopped it, a free row with errorReason: "truncated" says how many Reddit counts and how many came back. Replies carry depth, parentId and isSubmitter, so the tree can be rebuilt exactly.

7. Deep coverage, past the per-community ceiling

Reddit stops any single community feed at roughly 1,000 posts and then simply stops paginating. Switch on Deep coverage and the same community is read through its other sorts and time windows and the results are de-duplicated.

Measured on r/rust: the best single feed returned 998 unique posts; four feeds combined and de-duplicated returned 1,488.

What this is not: it is not a full historical archive, and no live scraper can be one. If a community has 200,000 posts going back ten years, this reaches a few thousand of them, not all of them. That limit is Reddit's and it applies to Reddit's own official API as well. Where a run stops short, the log says which of the two reasons it was — a limit you set, or the ceiling itself.


What you get

One flat row per post or comment, plus free rows for anything that produced nothing. Every row carries type, sourceRef (which of your inputs produced it) and scrapedAt.

Posts — id · fullId · url · permalink · subreddit · subredditId · title · body · bodyHtml · author · authorId · createdAt · score · upvoteRatio · numComments · ageHours · scorePerHour · commentsPerHour · commentToScoreRatio · wordCount · flair · postType · language · linkUrl · domain · thumbnail · images[] · videoUrl · galleryImages[] · isNsfw · isSpoiler · isOc · isSelf · isVideo · isPinned · isStickied · isLocked · isArchived · totalAwards · editedAt · crosspostParentId

Since September 2026 Reddit's JSON endpoints answer only visitors who pass its verification page, which this Actor never gets past. So community listing rows are read from the feed Reddit renders itself. On those rows title, body, bodyHtml, author, score, upvoteRatio, numComments, createdAt, permalink, linkUrl, domain, flair, postType, language, isNsfw, isSpoiler, isSelf and isVideo are filled, and so is crosspostParentId on a crosspost. postType is Reddit's own name for the post: text, link, image, multi_media (a text post with media inline), crosspost, or another type as Reddit names it. language is Reddit's language tag for the post, such as en. thumbnail is the small preview image the post's card shows, as on a link post with a preview, and null when the card shows none. images carries only a direct image post's own file. videoUrl, isOc, isPinned, isStickied, isLocked, isArchived and editedAt stay null, and galleryImages stays empty. Comment rows come from the first page of each thread (see 6). Community and user rows came from the JSON API's About, user and search endpoints (see above), so runs do not produce them.

Comments — id · fullId · postId · postTitle · parentId · permalink · subreddit · author · authorId · body · bodyHtml · score · createdAt · ageHours · scorePerHour · wordCount · depth · isSubmitter · isStickied · editedAt · totalAwards

Engagement metrics are worked out from the row's own fields at the moment it was read (scrapedAt), so you can rank by them without a spreadsheet. ageHours is the hours since createdAt. scorePerHour is score ÷ ageHours, an average over the post's or comment's life so far, not its speed right now. commentsPerHour is numComments ÷ ageHours, and commentToScoreRatio is numComments ÷ score; both are on post rows only, and the ratio is null when the score is zero or negative. wordCount is the number of words in body. Each metric is null, never 0, when a value it needs is null, as with a post that has no body. They are rounded to four decimal places.

Communities — name · title · description · publicDescription · subscribers · activeUsers · createdAt · isNsfw · lang · iconUrl · bannerUrl

Users — username · createdAt · linkKarma · commentKarma · totalKarma · isGold · isMod · isEmployee · isVerified · iconUrl · bio

Free rows — sourceRef · errorReason · message · statusCode. errorReason is one of notFound, private, banned, quarantined, empty, blocked, badInput, maxItemsReached, chargeLimitReached, requestFailed, truncated — a stable machine-readable value, so a pipeline can branch on it.

Text arrives unescaped. A title containing & arrives as &, not &.


How to scrape a subreddit

Put the community name in Communities — rust, or r/rust, either works — and pick a Sort. top and controversial also use the Time window.

Set Results per source for how many posts you want from each. Beyond roughly 1,000, turn on Deep coverage.

How to scrape Reddit comments

Switch on Include comments and set Comments per post. Every post the run returns then also returns the first page of its thread (about 25 comments), with replies followed to Reply depth. Comments are billed at their own, lower rate.

A single thread's URL cannot be read on its own: this Actor reads single posts through Reddit's JSON API, which answers only visitors who pass Reddit's verification page. It comes back as a free blocked row.

How to search a community for keywords

Put the community in Communities and your terms in Keywords. Reddit's JSON search answers only verified visitors, so the community's feed is read instead, about 200 posts per keyword in the sort you chose. Only the posts whose title or text contains every word come back, marked matchedBy: "feed-keyword-filter". A "quoted phrase" counts as one word. Operators such as title: or author: are not applied, and the order is the feed's, not Reddit's search ranking. One free row per search says how many posts were read and how many matched.

Keywords without a community, which would be a search across all of Reddit, come back as a free blocked row.

Reddit users

This Actor reads user profiles and a user's posts or comments through Reddit's JSON API, which answers only verified visitors, so Users and profile URLs come back as a free blocked row each. Nothing is charged for them.

How to monitor a subreddit for new posts

Sort by new, switch on Only rows I have not had before, set Posted after to something like 2 days, and put the Actor on a schedule. Each run returns only what appeared since the last one, and you are billed only for those rows.


Pricing

Two units, because a thread returns an order of magnitude more comments than posts and charging both at one rate makes a comment-heavy pull absurd.

FREEBRONZESILVERGOLDPLATINUMDIAMOND
Per 1,000 posts, until 6 October 2026$6.00$5.40$3.90$3.00$2.80$2.60
Per 1,000 posts, from 6 October 2026$5.00$4.50$3.25$2.50$2.50$2.50
Per 1,000 comments, until 6 October 2026$0.90$0.80$0.70$0.60$0.55$0.50
Per 1,000 comments, from 6 October 2026$0.90$0.80$0.70$0.60$0.60$0.60
Per run started, until 6 October 2026$0.003$0.0025$0.002$0.0015$0.0012$0.001
Per run started, from 6 October 2026$0.003$0.0025$0.002$0.0015$0.0015$0.0015

The new prices take effect on 6 October 2026 at 09:40 UTC. Communities and user profiles are billed at the post rate. Platform usage is included — what is above is what you pay, with nothing added afterwards.

Everything here is tiered, including the run fee. Every other Actor in this category charges a flat start fee that its largest customers pay at the same rate as its smallest — and charges between $0.02 and $0.216 for it.

Never billed: rows a filter discarded, rows an earlier run already delivered when Only new is on, duplicate rows, and every free row explaining a gap.

A ten-post evaluation run costs about $0.063 on the Free plan ($0.003 to start plus 10 × $0.006), and $0.053 from 6 October 2026.


Honest limits

  • A residential proxy is required, and it is the default (the code falls back to it when the field is left empty). Reddit refuses datacenter addresses outright, so a run configured without one will return free blocked rows rather than data. Leave the proxy setting alone unless you know why you are changing it.
  • Single posts, users, About pages and site-wide search are not available. This Actor reads them through Reddit's JSON API, which answers only visitors who pass Reddit's verification page, and this Actor never gets past that page. They come back as free blocked rows.
  • The first page of each comment thread, about 25 comments. A free truncated row says when Reddit counts more.
  • Roughly 1,000 posts per community feed, ~500 per top window, and about 200 posts read per keyword. Deep coverage combines feeds to beat the feed ceiling; nothing gets you the complete history of a large community.
  • Public content only. A community this Actor cannot read returns a free row saying so. Nothing behind a login is read.
  • Deleted content is gone. An author who deleted their account is null, not a name.

Frequently asked

Do I need a Reddit account, an app, or an API key? No. Nothing to register, nothing to wait for approval on.

Why is there a run fee at all, when the rest of the billing is per row? Because reaching Reddit at all costs something before a single row exists, and a fee that covers it is what lets everything else — the filters, the duplicates, the failures — genuinely stay free. It is a fraction of a cent and it is tiered; the alternative most of this category picked was a flat fee 6.7× to 72× larger than the $0.003 Free-plan fee here.

What if I type a date it cannot read? You get a free row saying so, and the run does not pretend the filter ran. It will never quietly bill you for rows a filter you thought was active should have removed.

What happens if a community in my list is private, banned or misspelled? You get a free row naming it, and the run carries on. A banned community and one that does not exist both come back as an empty row: measured on 1 October 2026, Reddit's feed answers both with a page of no posts and no reason, so the row says the community may be empty, private, banned or misspelled. A private community has not been measured. Depending on how Reddit answers, it comes back as that same empty row, as a notFound or private row, or, once the Actor's retries run out, as a blocked row. Every input you supply produces either data or an explanation — never silence.

Can I get more than 1,000 posts from one community? Yes, with Deep coverage — measured at 1,488 unique on a community where a single feed gave 998. Not unlimited, and not the full archive.

Why did my run return fewer rows than I asked for? The log says which of the two reasons applied: a limit you set, or Reddit's own ceiling. They are reported separately on purpose.

Can I use this from Python / n8n / Make / an LLM agent? Yes — it is a standard Actor with a standard dataset, callable from the Apify API and client libraries.


You are responsible for how you use the data you collect, including compliance with Reddit's terms and with privacy law in your jurisdiction. Scrape public content, and honour deletion requests.

Use it from n8n, MCP, the API or a schedule

Built to be called by a workflow, not only from the Store form. The Actor is vonsensey/reddit-scraper-posts-comments-api; every snippet below sends {}, which runs the defaults shown on the form — replace it with your own input.

n8n

Install the Apify community node (@apify/n8n-nodes-apify under Settings → Community Nodes, or search "Apify" on n8n Cloud). Add Apify → Run Actor with Actor vonsensey/reddit-scraper-posts-comments-api and your input JSON, then Apify → Get Dataset Items on the run's defaultDatasetId and pipe the rows anywhere. For scheduled runs, the On new Apify Event trigger fires when a run of this Actor finishes.

MCP (Claude, Cursor, VS Code, any MCP client)

{
"mcpServers": {
"apify": {
"url": "https://mcp.apify.com?tools=vonsensey/reddit-scraper-posts-comments-api",
"headers": {
"Authorization": "Bearer <YOUR_APIFY_TOKEN>"
}
}
}
}

Your agent then calls vonsensey/reddit-scraper-posts-comments-api as a tool with the same input the form takes and reads the dataset back.

REST API (one call, rows in the response)

curl -X POST "https://api.apify.com/v2/acts/vonsensey~reddit-scraper-posts-comments-api/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" -d '{}'

Python

from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("vonsensey/reddit-scraper-posts-comments-api").call(run_input={})
for row in client.dataset(run["defaultDatasetId"]).iterate_items():
print(row)

JavaScript

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('vonsensey/reddit-scraper-posts-comments-api').call({});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Make, Zapier, LangChain, CrewAI

The Apify app in Make and Zapier has a Run an Actor module: pick vonsensey/reddit-scraper-posts-comments-api. In LangChain and CrewAI the Apify tool wrappers take the same Actor id. A daily schedule needs nothing but the Console: Schedules → Create → this Actor → cron, and the dataset fills on its own.

Run it without configuring anything — Scrape the top posts of a subreddit, a ready-made example you can start as-is or copy.

Use cases

  • Research what people actually say. Posts and comments in one flat table, with score and upvote ratio on every post row.
  • Track a topic inside a community. Keywords matched in a community's feed, re-run on a schedule with Only new to catch new mentions.
  • Mine a community. Pull a subreddit's posts, about 1,000 per feed and more with Deep coverage, for product research, sentiment or training data.
  • Read the conversation. The first page of every thread (about 25 comments) in the same table, with a free row saying how many comments Reddit counts when there are more.

Run it on a schedule

A one-off pull answers a question; a schedule answers it every day without you. Open Schedules in the Apify Console, point a cron at this Actor, and the dataset keeps filling on its own — no server, no cron box, no babysitting. Everything here is built to be re-run: you are billed per row delivered, and the FAQ below says exactly what a scheduled run that finds nothing new costs.

FAQ

Do I need a Reddit API key or login?

No. No key, no OAuth, no account.

Can I paste any Reddit URL?

Community URLs, in any sort and time window, and in-community search URLs, yes. Single post, comment, redd.it, user, About and site-wide search URLs each come back as a free blocked row: this Actor reads them through Reddit's JSON API, which answers only visitors who pass Reddit's verification page, and it never gets past that page.

Are scores and upvote ratios included?

Yes, on every post row, with no flag to enable.

Can I get comments as well as posts?

Yes, the first page of each thread (about 25 comments), in the same run and in the same table, with a type column separating them.


Something wrong, or a field you need that is missing? Open an issue on the Issues tab — it is read and it gets fixed. If this saved you time, a rating on the Store page helps the next person find it.