# Reddit Scraper — Posts, Comments, Subreddits & Users (`datasiphon/reddit-scraper`) Actor

Reddit scraper for posts, full comment threads (up to 25,000 per post), subreddits and users — no API key, no login. Search, sort and date filters, history past Reddit's 1,000-item cap, flat JSON/CSV rows. Pay per result from $1.20 per 1,000.

- **URL**: https://apify.com/datasiphon/reddit-scraper.md
- **Developed by:** [Kashif Ali](https://apify.com/datasiphon) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.20 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper — posts, comments, subreddits & users (no API key)

**Reddit Scraper** extracts **posts**, **full comment threads** (up to 25,000 comments per post), **subreddits** and **users** into flat JSON/CSV rows — **no Reddit API key, no login**. Search, sort and **date filters** let you go beyond Reddit's ~1,000-item listing cap. **From $1.20 per 1,000 results.**

### What does the Reddit Scraper do?

Give it subreddit names, search keywords, or post/subreddit/user URLs and it returns structured Reddit data for **market research**, **brand monitoring**, **sentiment analysis**, **lead discovery**, **dataset building for AI/LLMs** and academic research.

### Features

- **Scrape Reddit posts and comments** by subreddit, keyword search or URL; full **comment trees** with depth, scores and authors.
- **Subreddit and user data** (subscribers, descriptions, user profiles).
- **Date range filters** (`postedAfter`, `postedBefore`) and **sort / time** options; **history past the 1,000-item cap** via an archive source.
- NSFW toggle, `maxItems` cost cap, `maxCommentsPerPost`.
- Niche research helpers: `discoverRelatedSubreddits` and `classify` (pain-signal and advice-request flags).
- Every record tagged with its **source** (live or archive) and typed fields.

### How to scrape Reddit

1. Add **subreddits**, **searches** or **start URLs**.
2. Choose whether to include comments and set `maxItems`.
3. Run and export JSON, CSV or Excel.

### Example output (real run)

```json
{
  "dataType": "post",
  "id": "1wur4n5",
  "url": "https://www.reddit.com/r/webdev/comments/1wur4n5/vite_asset_bundling/",
  "title": "Vite Asset Bundling",
  "body": "Specifically, I'm grabbing JSON for a trading card game gallery from an endpoint, then dyn…",
  "postType": "self",
  "author": "KaedenCraft",
  "authorId": "t2_71etbnhy",
  "authorPremium": false,
  "subreddit": "webdev",
  "subredditId": "t5_2qs0q",
  "subredditSubscribers": 3315160,
  "subredditType": "public",
  "score": 2,
  "upvoteRatio": 1,
  "numComments": 4,
  "createdAt": "2026-10-01T05:27:16+00:00",
  "flair": "Question",
  "domain": "self.webdev",
  "contentUrl": "https://www.reddit.com/r/webdev/comments/1wur4n5/vite_asset_bundling/",
  "thumbnail": "self",
  "isNsfw": false,
  "isSpoiler": false,
  "isLocked": false,
  "isStickied": false,
  "isOriginalContent": false,
  "isCrosspostable": false,
  "contestMode": false,
  "awards": 0,
  "gilded": 0,
  "crossposts": 0,
  "source": "live",
  "scoreSettled": true
}
```

### Pricing

Pay per event: **$1.20 per 1,000 results** ($0.0012 each). Empty sources and errors are free.

### Limitations

Live Reddit through the unblocker can take 30–90 seconds per request; archive-backed runs are fast but scores can lag ~36 h on fresh posts. Deleted/removed content and private subreddits are not returned.

### Use cases

Brand and competitor monitoring · customer pain-point and idea research · sentiment analysis · AI training datasets · community and trend tracking.

### FAQ

**Do I need a Reddit API key?** No.
**Is it fast?** Archive-backed runs are fast; live Reddit through the unblocker can take 30–90 seconds per request (see technical notes below).
**Can I get comments?** Yes — set `includeComments`; up to 25,000 per post.

### Bulk Reddit comments scraper

Extract Reddit posts and comments at scale — a comments scraper with nested replies, advanced filtering across communities and subreddits, bulk URL input, and scrape-and-export to JSON, CSV or Excel without login.

### Is it legal? No login, no cookies

This Actor only reads **publicly available** pages. It never asks for a cookie, password or account and does not access private data. You are responsible for using the data in line with the source site's terms and applicable law (GDPR/CCPA for personal data).

### Related scrapers by the same author

[YouTube Email Scraper](https://apify.com/datasiphon/youtube-email-scraper) · [YouTube Transcript Scraper](https://apify.com/datasiphon/youtube-transcript-scraper) · [Reddit Scraper](https://apify.com/datasiphon/reddit-scraper) · [Airbnb Scraper](https://apify.com/datasiphon/airbnb-scraper) · [Amazon Product Scraper](https://apify.com/datasiphon/amazon-product-scraper) · [LinkedIn Jobs Scraper](https://apify.com/datasiphon/linkedin-jobs-scraper)

### Technical notes (original README, kept)

## Reddit Scraper Elite — posts, complete comment trees & history past Reddit's 1,000-item cap

Scrapes Reddit posts, subreddits, users and **full comment trees** (up to 25,000 comments per post) into flat, typed records. Every record carries a `source` field (`live` or `archive`) and a `scoreSettled` flag, so you always know where a number came from and whether it can be trusted.

### Two sources, one output shape

| | **Live Reddit** | **Archive (Arctic Shift)** |
|---|---|---|
| Freshness | real-time | comments ~minutes, posts ~30–60 min behind |
| Scores / comment counts | current | read 0–1 for ~36 h after posting (`scoreSettled: false`), exact afterwards |
| Depth of listing | Reddit's ~1,000-item cap | **no cap** — page back by date |
| Global keyword search | ✅ | ❌ needs *Search inside subreddit* |
| Comment trees | tree + "more" expansion | one call (≤25,000), paged fallback for huge threads |
| Availability | needs an IP Reddit accepts | works from Apify datacenter IPs |

`source: auto` (default) uses live Reddit when it is reachable and falls back to the archive. `live` fails cleanly (no charge) if Reddit is unreachable. `archive` skips live entirely.

#### Live Reddit — working (2026-09-22, build 0.2.20)

Reddit 403s Apify's datacenter *and* residential proxy pools outright (verified directly with real, working residential IPs — the block is on the proxy network's ASN, not IP quality). The one pool that actually gets through is **`UNBLOCKER`** (proxies through a real unblocking backend) — verified end-to-end: live posts with real scores/timestamps, live comment trees with correct nesting, live subreddit metadata.

**This is now the default** when `source` is `auto` or `live` and no `proxyConfiguration`/Reddit OAuth credentials are given -- no setup needed, live data works the same way it does on every other Reddit actor. An explicit `proxyConfiguration` (even `{"useApifyProxy": false}`) always overrides this.

**The honest tradeoff: it's slow.** 30-90 seconds per request (real headless-browser-backed unblocking, not a quick IP swap) and uses Apify's UNBLOCKER unit quota. A 5-post fetch takes ~60-80s; comments add one request per post. Budget `maxItems`/`maxCommentsPerPost` accordingly, or use `source: archive` for fast, cheap, unlimited-history results when live freshness doesn't matter for the task.

### Input (essentials)

- `startUrls` — post, subreddit or user URLs · `subreddits` — names · `searches` (+ `searchInSubreddit`)
- `includeComments`, `maxCommentsPerPost` (≤25,000) · `maxPostsPerSource` · `maxItems` (hard cap on billed results)
- `postedAfter` / `postedBefore` (ISO dates) — on the archive these reach past the 1,000-item cap
- `sort` / `time` (live only; the archive is newest-first) · `includeNsfw`

### Output records (`dataType`: `post` | `comment` | `subreddit` | `user`)

Posts (41 fields): id, url, title, body, postType, author/authorId/authorFlairText/authorPremium, subreddit/subredditId/subredditType, score, upvoteRatio, numComments, createdAt, editedAt, flair, domain, contentUrl, thumbnail, postHint, **previewImages**, **galleryImages** (full-res URLs, in order — verified against a real 8-image gallery post), isNsfw/isSpoiler/isLocked/isStickied/isOriginalContent/isCrosspostable/contestMode, distinguished, suggestedSort, awards, gilded, crossposts, removedBy, `source`, `scoreSettled`.
Comments (27 fields): id, postId, parentId, **depth**, body, author/authorId/authorFlairText/authorPremium, subreddit/subredditId, score, controversiality, createdAt, editedAt, isSubmitter, isStickied, isCollapsed/collapsedReason, distinguished, awards, gilded, **isDeleted** (removed/deleted bodies are kept and flagged), `source`, `scoreSettled`.

### Niche/idea-discovery mode (added 2026-09-22, from real market-research workflows)

Two videos describing Reddit-based startup-idea research (subreddit-graph tools like "Ena", GummySearch) named specific needs. Here's what's real and what isn't:

- **`discoverRelatedSubreddits`**: samples recent posters in a seed subreddit, finds what else they post in, ranks by how many distinct posters overlap (not just raw post count, so one prolific poster can't skew it). Surfaces subreddits you wouldn't have searched for directly. Combine with `relatedMinSubscribers`/`relatedMaxSubscribers` for the "10k-100k sweet spot, growing-but-not-huge" screening pattern from GummySearch. One archive call per sampled author (`relatedSampleAuthors`, default 20) — slower than a normal run, budget for it.
- **`classify: true`**: adds `hasPainSignal`/`isAdviceRequest`/`mentionsSolution`/`mentionsMoney` to posts/comments — cheap keyword heuristics, not AI. This is deliberately not a GummySearch/gigabrain-style LLM summarizer: the actor's job is complete, well-labeled raw data; feed it to Claude/ChatGPT yourself for pattern-finding and summarization, which is cheaper and more flexible than a fixed prompt baked into the actor.
- **NOT buildable from any data source available to this actor:** subreddit subscriber growth-over-time (no historical snapshots exist in the archive or its metadata — verified by inspecting the raw subreddit object). Global keyword-to-subreddit discovery, i.e. "search all of Reddit for X, see which subreddits talk about it" without picking a subreddit first (the archive's `query` param requires `subreddit` or `author` — confirmed by testing; this is the same live-Reddit blocker described above, since live Reddit's `/search.json` supports it natively).
- **Out of scope, not Reddit data:** Amazon Books/product reviews, Google Trends, Glimpse, Meta Ads Library, Exa, Perplexity/Gemini market-sizing, YouTube comments. Separate tools/actors, not this one.

Each run also writes a `SUMMARY` record: items stored per source, skipped targets with reasons (skips are never charged), warnings, `nsfwFiltered` and `duplicatesSkipped` counts.

### Pricing

**Not yet configured.** The code emits one `result` charge event per stored post/comment/subreddit/user (skips are never charged), but the pay-per-event price must still be set on the actor in Apify Console — until then the platform ignores the events (`WARN ... does not use the pay-per-event pricing`) and `maxItems` is the only cost cap.

### Not included (deliberately)

AI sentiment/intent labels, MCP connectors and 75-field parity — other Reddit actors do these; this one focuses on completeness and provenance.

### Performance note

`auto`/`live` mode skips the live-reachability check entirely for runs that have no live-capable target (e.g. `discoverRelatedSubreddits`-only) -- no point spending 30-90s proving live works when nothing in the run would use it.

### Robustness

- A post with `gallery_data` explicitly `null` (not missing) no longer crashes the run -- fixed and stress-tested against 796 real posts, build 0.2.22.
  No single malformed record, one bad comment tree, an unreachable seed subreddit, or an invalid proxy config can crash the whole run -- each is isolated to the smallest scope it affects (one post, one target, one seed) and counted/explained in `SUMMARY` rather than silently dropped or misreported. `SUMMARY` fields: `skipped` (with reasons), `warnings`, `nsfwFiltered`, `duplicatesSkipped`, `malformedRecords`, `commentFetchErrors`/`commentFetchErrorSamples`.

### Known limits

- An invalid `postedAfter`/`postedBefore` (not a valid ISO date) is dropped with a warning, not a crash.
- On the live source, `postedAfter`/`postedBefore` are applied client-side, so a narrow window on a busy subreddit can return fewer posts than `maxPostsPerSource`. (Live path unverified.)
- Subreddit names are cleaned of `r/` prefixes, slashes and full URLs (`r/foo`, `r/foo/`, `foo`, a reddit.com link all resolve the same way); a target that genuinely doesn't exist appears in `SUMMARY.skipped` with `not_found`, never a silent zero.
- NSFW-excluded (`includeNsfw: false`, the default) and duplicate posts are dropped from the stored count but still consume the source's per-target fetch budget — so a target with NSFW/duplicate posts in it can return fewer than `maxPostsPerSource` even inside the window. `SUMMARY.nsfwFiltered`/`duplicatesSkipped` say how many and why; the run does not auto-fetch replacements.
- The archive is a free community service with no uptime guarantee; huge threads occasionally time out server-side (handled by a windowed fallback, but slower).
- Archive comment counts include removed comments, so they can exceed Reddit's displayed count.
- Respect Reddit's terms and applicable law when using collected data; do not collect data for purposes you are not entitled to.

### Verified results (Apify platform, 256 MB, build 0.2.20, 2026-09-22)

| Test | Result |
|---|---|
| r/programming, June-2026 window, 250 items | 249 posts + subreddit record, 12 s, 0 duplicates |
| Post with 161 reported comments | 169 comments (8 more than Reddit's displayed count — the archive retains removed/deleted comments; 1 flagged `isDeleted`), max depth 13, 10 s |
| r/AskReddit thread, 15,264 reported comments | 15,264 unique comments, max depth 16, 0 orphans, every parent present, 861 s (~14 min) |
| Archive keyword search in a subreddit | 30/30 posts matched the keyword |
| Global search / live-only without a usable live path | skipped or failed cleanly, nothing charged |
| **Live Reddit path, real posts** | **working** — 5 live posts via UNBLOCKER, real scores (197/26/107/0) and same-day timestamps, 82 s |
| **Live Reddit path, comments** | **working** — 1 post + 30 live comments, correct depth nesting, 130 s |
| **Live path, plain default (no config at all)** | **working automatically** — `auto` mode with zero explicit proxy config now gets live data by default |

***

### Pricing & cost estimator

Pay-per-event: **$1.50 per 1,000 records**. You are charged only for rows that contain data; empty or failed targets are free. A 100-item run costs about 10–60 seconds of platform time on the default memory and typically under $0.50 in total. Set `maxItems` / the per-link cap to control spend.

### Example output (real run)

```json
{
  "dataType": "post",
  "id": "1wur4n5",
  "url": "https://www.reddit.com/r/webdev/comments/1wur4n5/vite_asset_bundling/",
  "title": "Vite Asset Bundling",
  "body": "Specifically, I'm grabbing JSON for a trading card game gallery from an endpoint, then dyn…",
  "postType": "self",
  "author": "KaedenCraft",
  "authorId": "t2_71etbnhy",
  "authorPremium": false,
  "subreddit": "webdev",
  "subredditId": "t5_2qs0q",
  "subredditSubscribers": 3315160,
  "subredditType": "public",
  "score": 2,
  "upvoteRatio": 1,
  "numComments": 4,
  "createdAt": "2026-10-01T05:27:16+00:00",
  "flair": "Question",
  "domain": "self.webdev",
  "contentUrl": "https://www.reddit.com/r/webdev/comments/1wur4n5/vite_asset_bundling/",
  "thumbnail": "self",
  "isNsfw": false,
  "isSpoiler": false,
  "isLocked": false,
  "isStickied": false,
  "isOriginalContent": false,
  "isCrosspostable": false,
  "contestMode": false,
  "awards": 0,
  "gilded": 0,
  "crossposts": 0,
  "source": "live",
  "scoreSettled": true
}
```

### Limitations

Deleted/removed content and private or quarantined subreddits are not returned. Very deep comment trees can be capped by `maxCommentsPerPost`. Reddit can rate-limit; the actor retries and reports skipped sources in the run summary.

### No login, no cookies

This actor only reads public Reddit pages and the public JSON endpoints. It never asks for a cookie, password or account, and it does not access private data. Make sure your use of the scraped data complies with the source site's terms and applicable law (GDPR/CCPA for personal data).

# Actor input Schema

## `startUrls` (type: `array`):

Post, subreddit or user URLs (reddit.com / old.reddit.com).

## `subreddits` (type: `array`):

Subreddit names, e.g. programming or r/programming.

## `searches` (type: `array`):

Keyword searches. Live source: all of Reddit or one subreddit. Archive source: requires 'Search inside subreddit'.

## `searchInSubreddit` (type: `string`):

Restrict searches to one subreddit (required for archive-source search).

## `source` (type: `string`):

auto = live Reddit when reachable, else archive (each record is tagged with its source). live = fail if Reddit unreachable. archive = Arctic Shift only: full history and complete comment trees, but scores lag ~36h.

## `includeComments` (type: `boolean`):

Fetch the full comment tree for every post (each comment is one billed result).

## `maxCommentsPerPost` (type: `integer`):

Upper bound per post; trees up to 25,000 comments are supported.

## `maxPostsPerSource` (type: `integer`):

Posts collected per input target.

## `maxItems` (type: `integer`):

Hard cap on billed results for the whole run.

## `sort` (type: `string`):

Ordering on the live source. The archive is always newest-first.

## `time` (type: `string`):

Applies to top/controversial/relevance on the live source.

## `postedAfter` (type: `string`):

ISO date, e.g. 2026-01-01. Works on both sources; on the archive it reaches past Reddit's 1,000-item listing cap.

## `postedBefore` (type: `string`):

ISO date, e.g. 2026-06-30.

## `includeNsfw` (type: `boolean`):

Include 18+ posts and comments.

## `redditClientId` (type: `string`):

Enables the OAuth live path if you have registered a Reddit app.

## `redditClientSecret` (type: `string`):

Paired with the client ID.

## `proxyConfiguration` (type: `object`):

Used for the live source. Reddit blocks Apify's datacenter and residential proxy pools outright -- the UNBLOCKER group (default here) genuinely reaches Reddit's JSON API, verified 2026-09-22. It's slow (30-90s per request, it proxies through a real unblocking backend) and priced per unit, so live mode is much slower and slightly costed compared to the archive source -- budget maxItems/maxCommentsPerPost accordingly.

## `classify` (type: `boolean`):

Adds hasPainSignal, isAdviceRequest, mentionsSolution, mentionsMoney to each post/comment — cheap keyword heuristics (no AI), a screening signal not ground truth.

## `discoverRelatedSubreddits` (type: `array`):

Seed subreddit names. Finds subreddits with overlapping audiences by sampling recent posters and aggregating what else they post in — surfaces subreddits you wouldn't have searched for directly.

## `relatedSampleAuthors` (type: `integer`):

How many recent posters to sample per seed subreddit. One archive call per author — higher is slower and more thorough.

## `relatedMinSubscribers` (type: `integer`):

Only keep discovered subreddits at or above this size.

## `relatedMaxSubscribers` (type: `integer`):

Only keep discovered subreddits at or below this size (e.g. 100000 for the '10k-100k sweet spot' niche-discovery pattern).

## Actor input object example

```json
{
  "source": "auto",
  "includeComments": false,
  "maxCommentsPerPost": 500,
  "maxPostsPerSource": 100,
  "maxItems": 100,
  "sort": "new",
  "time": "all",
  "includeNsfw": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "UNBLOCKER"
    ]
  },
  "classify": false,
  "relatedSampleAuthors": 20
}
```

# Actor output Schema

## `results` (type: `string`):

Every row (posts, comments, subreddits, users) as JSON/CSV/Excel.

## `posts` (type: `string`):

Only dataType=post rows.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "UNBLOCKER"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("datasiphon/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["UNBLOCKER"],
    } }

# Run the Actor and wait for it to finish
run = client.actor("datasiphon/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "UNBLOCKER"
    ]
  }
}' |
apify call datasiphon/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,datasiphon/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/h2fCbZ8xycpmmNqFp/builds/H199cUDt8foGhqoJw/openapi.json
