Reddit Scraper: Subreddit Posts, Comments, Search - No Login avatar

Reddit Scraper: Subreddit Posts, Comments, Search - No Login

Pricing

from $0.80 / 1,000 post scrapeds

Go to Apify Store
Reddit Scraper: Subreddit Posts, Comments, Search - No Login

Reddit Scraper: Subreddit Posts, Comments, Search - No Login

Extract Reddit posts, comments and search results. Up to 1,000 posts per subreddit across hot/new/top/rising with full post text, comment threads, and keyword search. Flat JSON, 33 always-present keys, agent-ready. No login or credentials. SFW only. No start fee, $1.00 per 1,000 posts.

Pricing

from $0.80 / 1,000 post scrapeds

Rating

0.0

(0)

Developer

Santhej Kallada

Santhej Kallada

Maintained by Community

Actor stats

0

Bookmarked

14

Total users

6

Monthly active users

11 days ago

Last modified

Share

Reddit Scraper: Subreddit Posts, Comments, Search — No Login

Scrape posts, comment threads and keyword search results from public Reddit communities into one flat dataset — 33 keys on every row, no login, no API keys, and no start fee: a run that returns nothing costs $0.00.

  • $1.00 per 1,000 posts (search results bill on the same event)
  • $0.30 per 1,000 comments
  • $0.002 per comment thread that actually returned a comment
  • $0.00 to start a run

What does Reddit Scraper do?

Reddit Scraper collects public Reddit content without an account, an app or an API key. Give it one or more communities (subreddits), a search keyword, or both, and it returns:

  • Community listings — posts from each community, sorted hot, new, top, rising, controversial or best, with an optional time window. Text posts include their full body text.
  • Keyword search — site-wide, or scoped to each community you name, with the same sort and time window.
  • Comment threads — optionally, the comments of every post it collected, appended to the same dataset with their nesting depth and parent links.

Everything lands in one table: posts and comments share a single 33-key record shape, distinguished by type and sourceType, so a research task is one call instead of three tools stitched together. Every key is present on every row; a value the source does not report is null, never a missing key.

What you can get per source

SurfacePer source
Community listingUp to 1,000 posts. Measured: 997 hot and 974 new from one community; a top/year listing ended at 497 because the community's feed ends there.
Keyword searchUp to ~179 results per keyword.
Comment threadThe first page of each thread — typically up to ~25 comments, top-level plus some replies.

Asking for more than a ceiling does not fail — it returns what the source has. maxCommentsPerPost is an upper bound, not a guarantee.


Use cases

  • Market and product research — track what a community says about a product, a release or a competitor.
  • Lead and pain-point discovery — find threads where people ask for a tool like yours ("best CRM for a small business", "alternative to …") and read the replies.
  • Trend monitoring — schedule a daily new pull on a set of communities and diff it.
  • Sentiment and NLP corpora — post bodies plus comment threads, already flat and de-duplicated.
  • Support and bug triage — find complaint threads about your product across relevant communities.
  • AI agent workflows — a small, cheap, predictable call an LLM agent can make many times per task.

How to use Reddit Scraper

  1. Open the Input tab.
  2. Add one or more communities — startups, r/startups, /r/startups and https://www.reddit.com/r/startups/ all work — and/or a search keyword.
  3. Pick a sort and, for top, controversial and search, a time range.
  4. Set Max posts per source (default 100). Each community listing and each keyword search is a separate source.
  5. Optional: switch on Include comment threads and set Max comments per post and Minimum comments before fetching a thread.
  6. Click Start. When the run finishes, open the Output tab — the Posts and Comments views show the two row types — and export as JSON, CSV, Excel, XML or HTML, or read the dataset from the API.

Example input

{
"subreddits": ["startups", "r/SaaS"],
"searchQuery": "pricing",
"sort": "top",
"timeRange": "month",
"maxPostsPerSource": 50,
"includeComments": true,
"maxCommentsPerPost": 25,
"minCommentsToFetch": 5
}

This is 4 sources — 2 community listings plus the keyword searched once inside each community — so up to 200 posts, plus up to 25 comments for every post with at least 5 comments.

Site-wide keyword search only:

{
"searchQuery": "best crm for small business",
"sort": "best",
"timeRange": "year",
"maxPostsPerSource": 50
}

Input configuration

FieldTypeDefaultDescription
subredditsarray["programming"]Communities to collect. Accepts programming, r/programming, /r/programming or a full reddit.com URL.
searchQuerystring—Keyword. Site-wide on its own; with subreddits, runs once per community in addition to each community's listing.
sortstringhothot, new, top, rising, controversial, best. Listings serve controversial as top. Searches run hot/best/rising as relevance, top as top, new as new and controversial as most comments — use best for on-topic results, because a top search ranks every loosely matching post by score. sortUsed on each row echoes what ran.
timeRangestringallhour, day, week, month, year, all. Applies to top, controversial and search. On listings hour is served as day.
maxPostsPerSourceinteger1001–1,000. Per source, not per run — see below.
includeCommentsbooleanfalseAppends comment rows for every collected post to the same dataset.
maxCommentsPerPostinteger501–500. Upper bound per post; a thread returns its first page of comments (typically up to ~25).
minCommentsToFetchinteger1Skip threads on posts declaring fewer comments than this — your direct lever on the per-thread charge.
proxyConfigurationobjectResidential USAdvanced. Leave as is.

If you supply neither subreddits nor searchQuery (for example an empty list), the run collects nothing, charges nothing, and says so in its status message.

What counts as a source

maxPostsPerSource is a ceiling per source. A source is one community listing, one keyword search inside one community, or one site-wide search:

InputSourcesRows at maxPostsPerSource: 8
1 community1 listing8
keyword only1 site-wide search8
1 community + keyword1 listing + 1 search16
2 communities + keyword2 listings + 2 searches32

Every row records its origin in sourceType (subreddit, search or comments) and sourceQuery, and a post found by both a listing and a search is written — and billed — once.


Output

One dataset, one record shape, every key present on every row.

Example output

A post row from a community listing:

{
"type": "post",
"id": "t3_1wnfg3d",
"postId": "t3_1wnfg3d",
"parentId": null,
"depth": null,
"subreddit": "startups",
"subredditPrefixed": "r/startups",
"subredditId": "t5_2qh26",
"author": "example_founder",
"authorId": "t2_g6mey3dcb",
"title": "Things you wished you did before launching? [I will not promote]",
"body": "Are there any best practices and/or wisdom that other founders would like to share about things they wished they had done before launching their app? …",
"url": "https://www.reddit.com/r/startups/comments/1wnfg3d/things_you_wished_you_did_before_launching_i_will/",
"permalink": "https://www.reddit.com/r/startups/comments/1wnfg3d/things_you_wished_you_did_before_launching_i_will/",
"domain": "self.startups",
"postType": "text",
"score": 15,
"upvoteRatio": 1,
"numComments": 14,
"numCrossposts": null,
"awardCount": 0,
"flair": "I will not promote",
"isNsfw": null,
"isSpoiler": null,
"isOriginalContent": null,
"isPinned": null,
"createdAt": "2026-09-22T17:10:38.320Z",
"createdTimestamp": 1790097038320,
"rank": 1,
"sourceType": "subreddit",
"sourceQuery": "startups",
"sortUsed": "hot",
"scrapedAt": "2026-09-23T03:52:42.694Z"
}

A comment row (a reply, so it has a parentId):

{
"type": "comment",
"id": "t1_pbex50o",
"postId": "t3_1wnfg3d",
"parentId": "t1_pbes7m8",
"depth": 1,
"subreddit": "startups",
"subredditPrefixed": "r/startups",
"subredditId": null,
"author": "example_founder",
"authorId": null,
"title": null,
"body": "Basically having posthog set up?",
"url": null,
"permalink": "https://www.reddit.com/r/startups/comments/1wnfg3d/comment/pbex50o/",
"domain": null,
"postType": null,
"score": 2,
"upvoteRatio": null,
"numComments": null,
"numCrossposts": null,
"awardCount": null,
"flair": null,
"isNsfw": null,
"isSpoiler": null,
"isOriginalContent": null,
"isPinned": null,
"createdAt": "2026-09-22T18:19:33.790Z",
"createdTimestamp": 1790101173790,
"rank": 2,
"sourceType": "comments",
"sourceQuery": "startups",
"sortUsed": "best",
"scrapedAt": "2026-09-23T03:52:42.694Z"
}

Field reference

#FieldTypeDescription
1typestringpost or comment.
2idstringt3_ for posts, t1_ for comments. Primary de-duplication key.
3postIdstringt3_ id of the post; equals id on post rows.
4parentIdstringDirect parent comment id on replies. null on posts and top-level comments.
5depthintegerComment nesting depth, 0 = top level. null on posts.
6subredditstringBare community name.
7subredditPrefixedstringr/-prefixed name.
8subredditIdstringt5_ id.
9authorstringUsername without u/.
10authorIdstringt2_ id.
11titlestringPost title, HTML entities decoded. null on comments.
12bodystringPost self-text, comment text, or the search snippet on search rows. null when the post has no text (link, image and title-only posts).
13urlstringOutbound/media URL. Equals the permalink for text posts.
14permalinkstringAbsolute link to the post or the specific comment.
15domainstringLink host, or self.{community} for text posts.
16postTypestringtext, link, image, video, crosspost, … as the site reports it.
17scoreintegerNet upvotes at extraction time. null when the site hides it.
18upvoteRationumber0–1.
19numCommentsintegerDeclared comment count on the post.
20numCrosspostsintegerCrosspost count.
21awardCountintegerAward count.
22flairstringPost flair.
23isNsfwbooleanAdult marker. Adult rows are never written.
24isSpoilerbooleanSpoiler marker.
25isOriginalContentbooleanOC marker.
26isPinnedbooleanPinned/stickied in the community.
27createdAtstringISO 8601 UTC.
28createdTimestampintegerEpoch milliseconds, derived from createdAt so the two never disagree.
29rankinteger1-based position within its own listing, search or thread.
30sourceTypestringsubreddit, search or comments.
31sourceQuerystringThe community or keyword that produced the row.
32sortUsedstringThe sort actually applied, echoed back.
33scrapedAtstringISO 8601 UTC extraction time, identical for every row in a run.

Which fields are populated

The record shape is identical everywhere; the populated set depends on the surface. Measured on 2026-09-23 runs (+ populated · – always null):

FieldCommunity listingKeyword searchComments
type id postId permalink rank sourceType sourceQuery sortUsed scrapedAt+++
subreddit subredditPrefixed score createdAt createdTimestamp+++
author+++
parentId depth––+ (parentId on replies)
title numComments++–
authorId++–
subredditId+––
body+ (text posts)+ (snippet, most rows)+
url domain postType upvoteRatio awardCount+––
flair+ (when the post has one)––
numCrossposts isNsfw isSpoiler isOriginalContent isPinned–––

A community listing row therefore carries 26 of the 33 fields. Why some are null: Reddit's older server-rendered site now sends every signed-out visitor to a sign-in page (since early September 2026), so all data comes from the current site, which does not publish crosspost counts or the spoiler / OC / pinned markers to signed-out visitors. The fields stay in the record so a parser written once keeps working.

A search result whose author the search page does not show is skipped rather than written with a blank author, so a search can return slightly fewer rows than requested. Skipped rows are never billed.


How much does it cost to scrape Reddit?

Pay per event. No start fee.

EventPriceWhen it fires
Post scraped$0.001 ($1.00 / 1,000)Per post row written to the dataset. Search results bill on this same event.
Comment scraped$0.0003 ($0.30 / 1,000)Per comment row written to the dataset.
Comment thread fetched$0.002Once per post whose thread returned at least one comment. Empty and deleted threads are free.

Apify Silver and Gold subscribers pay 10% and 20% less per event.

Cost formula you can compute before calling:

total = 0.001 * posts + 0.002 * threadsFetched + 0.0003 * comments

Worked examples

JobCost
25 posts (typical agent call)$0.025
100 hot posts from one community$0.10
1,000 posts from one community$1.00
3 communities × 150 posts$0.45
10 posts + their threads (10 threads, 118 comments) — a measured run$0.0654
100 posts + threads (100 threads, 20 comments each)$0.90
A run that returns no rows$0.00

Why no start fee matters. Many scrapers charge a flat fee when a run starts, before any data exists. For an agent making small, frequent calls, that fee is most of the bill:

This ActorTypical start-fee pricing
Start fee$0.00$0.02 – $0.09 before any data is returned
25 posts$0.025the start fee plus the per-post price
A run that returns 0 rows$0.00the start fee, every time

Billing rules

  1. Charges fire only after rows are written. A failed or empty run is free.
  2. Promoted posts and cross-source duplicates are dropped before they are written, so you are never billed for them.
  3. minCommentsToFetch is a direct lever on the per-thread charge — raise it to skip low-value threads.
  4. maxPostsPerSource and maxCommentsPerPost bound the row events, so your maximum bill is knowable in advance. The run log prints the maximum post charge before collecting anything.

SFW communities only — by design

Adult and quarantined communities are refused, not filtered afterwards:

  1. Every requested community is checked against Reddit's own community flags before any harvesting starts. Age-restricted, quarantined, banned, private and non-existent communities are skipped with a clear reason, cost nothing, and are listed in the run summary (communitiesRefused).
  2. A search reaches communities you did not name; their threads are only requested after the same check.
  3. Search results Reddit marks as adult are dropped before they are written and before they are billed.

There is no toggle to switch this off.


Reliability: what a run tells you

SituationWhat you get
Rows collectedSUCCEEDED, the rows, and a summary in the key-value store record OUTPUT
A source really has nothing in it (e.g. a keyword with no matches)SUCCEEDED, zero rows for it, emptySources incremented, $0.00 for it
Nothing to collect — no community or keyword given, or every community refusedSUCCEEDED, zero rows, $0.00, the reason as the run's status message and in OUTPUT.inputError
One source could not be collected, others wereSUCCEEDED with partial: true and the source counted in failedSources
Nothing collected and no source was emptyFAILED, with a message naming the cause and the summary still written to OUTPUT

A short or empty dataset is never presented as complete: read partial and failedSources in OUTPUT before treating a dataset as complete. The summary also includes posts, comments, rows, threadsFetched, communitiesRefused, requiredFieldFill and estimatedChargeUsd.

The Actor manages its own request rate, rotates to a fresh residential exit address when one is refused, and retries transient failures. Every network wait has a ceiling, so a stalled connection ends as a retry or a clear failure instead of a run that sits silent until its timeout.

Run memory

The Actor defaults to 2048 MB, which is also its maximum. The extraction path renders no page, so 2 GB is ample; raising it would only increase compute cost, and lowering it risks a large run running out of memory. If you call the Actor from the API or a scheduler, leave memoryMbytes unset so the default applies.


Data use, privacy and compliance — please read

Output contains personal data (usernames, user ids, authored text, timestamps, community membership).

  1. You are the data controller. You are responsible for establishing a lawful basis (GDPR Art. 6), providing notice where required (Art. 14), and setting a retention period.
  2. No profiling of individuals. This Actor must not be used to build person-level profiles or datasets about identified or identifiable individuals. It has no user-profile surface and will not get one.
  3. Honour deletions. Content deleted or removed on Reddit after extraction must stop being used.
  4. Special-category warning (GDPR Art. 9). Membership of a community can itself reveal health status, sexual orientation, religious belief or political opinion. The "manifestly made public" exception is assessed per item and cannot be applied wholesale by a bulk collector.
  5. Adult and quarantined communities are refused, with no opt-in.
  6. Respect Reddit's terms and the rights of the people whose content you collect. If you are unsure whether your intended use is lawful in your jurisdiction, take advice before running at scale.

FAQ

Do I need a Reddit account, an app, or API keys? No. There is nothing to configure beyond the input fields.

What does a run that returns nothing cost? Nothing. There is no start fee and charges only fire on rows that were written.

I got zero rows and the run says SUCCEEDED. Is that real? Yes. Either a source genuinely had nothing in it (counted in emptySources), or the input named nothing the Actor can collect — the run's status message and OUTPUT.inputError then say exactly why (for example "age- restricted community"). If sources could not be collected, the run says so in failedSources / partial, or reports FAILED.

How many posts can I get from one community? Up to 1,000 per listing. To go further, run again with a different sort — overlap between sorts is typically partial, and id is stable, so de-duplicating is trivial.

Why are my top search results off-topic? Reddit's search matches loosely, and top orders every match by score, so highly upvoted posts that share one common word with your keyword can rank first. Use sort: "best" (relevance) with a timeRange instead.

Why does my search stop at about 179 results? That is how far a single keyword's result list goes. Split a broad topic into several narrower keywords, or scope the search to specific communities.

Why did I get fewer comments than I asked for? Each thread returns its first page of comments — typically up to ~25, top-level plus some replies — and maxCommentsPerPost is an upper bound on that. Deleted and collapsed replies reduce the count further.

Why are numCrossposts, isSpoiler, isOriginalContent and isPinned always null? Reddit does not publish them to signed-out visitors of its current site, and its older site now requires sign-in. See Which fields are populated.

Can I scrape adult or quarantined communities? No. They are refused before any data is fetched, and there is no setting to change that.

Can I get comments without posts? Comments are always attached to the posts they belong to — set includeComments: true and filter the dataset on type == "comment". Both live in the same table, joined by postId.

Can I export to CSV or Excel? Yes — JSON, CSV, XLSX, XML and HTML from the run page or the API. The fixed 33-key shape means the columns are always the same.

Can an AI agent call this? Yes, that is the design target: nine input fields, published ceilings, a computable cost formula, no pagination state to manage. One call in, complete dataset out.


Tags

reddit · reddit scraper · subreddit scraper · reddit comments · reddit search · social media · sentiment analysis · market research · lead generation · ai agent · no login