# Reddit Scraper - Posts, Comments, Search, Users & Communities (`zaver.api/reddit-scraper`) Actor

Scrape Reddit without the Reddit API: posts, comments, search results, user profiles and communities from any URL, subreddit, username or keyword. No API key, no login, proxies included, no start fee. Export to CSV, Excel or JSON.

- **URL**: https://apify.com/zaver.api/reddit-scraper.md
- **Developed by:** [Zaver](https://apify.com/zaver.api) (community)
- **Categories:** Social media, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.99 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper — Posts, Comments, Search, Users & Communities (No API Key)

**Scrape Reddit without the API.** Paste any Reddit URL, community, username or keyword and export posts, comments, user profiles and communities to **CSV, Excel or JSON**. No Reddit account, no OAuth app, no API key, no proxies to buy. **Pay per result with no start fee.**

> ⭐ One Actor for every Reddit page · posts + full comment trees · keyword search · user and community data · proxies included · pay only per row delivered · API and MCP ready

### What does Reddit Scraper do?

Reddit Scraper is an all-in-one **Reddit data extractor**. It reads the same public pages a logged-out visitor sees and turns them into clean, flat rows. You can mix every input type in a single run:

| You paste | You get |
|---|---|
| A community, such as `r/webscraping` | Its posts, sorted by new, hot, top, rising or controversial |
| A post URL | The post and its comments, nested replies included |
| A username, such as `u/spez` | The profile, the user's posts and, optionally, their comments |
| A keyword, such as `best crm for startups` | Matching posts, communities or users from Reddit search |
| A Reddit search URL or share link | Whatever that link points to |

Every row carries a `type` (`post`, `comment`, `user`, `community`) and a `source` telling you which input produced it, so a run with fifty inputs still exports as one tidy table.

### What data can you scrape from Reddit?

#### Post fields

| Field | Description |
|---|---|
| `title`, `text` | Post title and full self-text |
| `url`, `id` | Permalink and Reddit post ID |
| `subreddit`, `subredditSubscribers` | Community name and its member count |
| `author`, `authorId` | Username and account ID |
| `score`, `upvoteRatio` | Net upvotes and upvote percentage |
| `numComments`, `crossposts`, `awards` | Engagement counts |
| `createdAt` | ISO 8601 timestamp (UTC) |
| `flair` | Post flair text |
| `link`, `domain` | Outbound link for link posts |
| `imageUrls`, `videoUrl`, `thumbnail` | Media URLs, galleries included |
| `isNSFW`, `isSpoiler`, `isStickied`, `isLocked`, `isSelf`, `isVideo` | Flags |

#### Comment fields

| Field | Description |
|---|---|
| `text` | Full comment body |
| `author`, `authorId`, `isSubmitter` | Who wrote it, and whether they are the original poster |
| `score`, `controversiality`, `awards` | Engagement |
| `depth`, `parentId`, `parentType` | Position in the reply tree, so you can rebuild the thread |
| `postId`, `postTitle`, `postUrl`, `subreddit` | The post the comment belongs to |
| `createdAt`, `isEdited`, `isStickied`, `distinguished` | Timestamp and moderation flags |

#### User fields

`username`, `displayName`, `bio`, `totalKarma`, `postKarma`, `commentKarma`, `createdAt`, `accountAgeDays`, `followers`, `isVerified`, `hasVerifiedEmail`, `isPremium`, `isModerator`, `isEmployee`, `avatar`.

#### Community fields

`name`, `title`, `description`, `subscribers`, `activeUsers`, `createdAt`, `communityType`, `language`, `icon`, `isNSFW`.

### A Reddit API alternative

Since Reddit changed its API pricing, the official Data API needs a registered app, OAuth credentials and, for commercial use, a paid agreement. Free access is rate limited and listings stop at 1,000 items. PRAW, the Python wrapper most tutorials use, needs those same credentials and inherits the same limits.

This Actor is a **Reddit API alternative** for posts, comments, search results, user profiles and communities. It works **without the Reddit API**: there is no app to register, no client ID or secret, no token refresh and no API pricing to negotiate. You pay per row you export and nothing else.

| | Official Reddit API or PRAW | This Actor |
|---|---|---|
| Setup | Register an app, manage OAuth tokens | Paste links or keywords |
| Commercial use | Paid agreement with Reddit | Pay per result |
| Rate limits | Yours to manage | Handled for you |
| Proxies and retries | Not applicable, but you build the pipeline | Included |
| Output | Raw API objects | Flat rows: CSV, Excel, JSON |

### Why use this Reddit scraper

#### No API key, no login, no rate-limit math

Reddit's official API needs an approved app, OAuth tokens and a paid tier for commercial volume. This Actor needs none of that. Paste links and press Start.

#### One price for everything, and no start fee

Most Reddit scrapers charge per result **plus** a fee every time a run starts, and some charge extra for comments. Here every row is the same price and starting a run is free, which matters when you schedule small runs every hour.

#### Proxies are included

Reddit blocks datacenter traffic. Every request here goes out through residential proxies, rotated automatically. You do not configure them and you do not pay for them separately.

#### You only pay for rows you receive

Billing happens right after each batch reaches your dataset. Stop a run halfway and you pay for half. Private, banned or empty targets cost nothing.

#### Built for pipelines

Flat rows, stable field names, ISO timestamps, and a `source` on every row. It drops straight into a spreadsheet, a warehouse, a vector store or an LLM prompt.

### How to scrape Reddit

1. Open the Actor and paste your inputs into **Reddit URLs, communities or users**, one per line. Add keywords under **Search keywords** if you want search results.
2. Choose **Sort** and **Time range** for community listings and searches.
3. Set **Max posts per target**. Turn on **Include comments for listed posts** if you need comments for every post in a listing.
4. Press **Start**. Rows appear in the dataset as they are collected.
5. Export as CSV, Excel, JSON or XML, or read the dataset through the API.

> 💡 A post URL always returns its comments, up to **Max comments per post**. Set that field to `0` if you only want the post itself.

### Input example

```json
{
  "startUrls": [
    "r/webscraping",
    "u/spez",
    "https://www.reddit.com/r/AskReddit/comments/1wrj0s2/"
  ],
  "searches": [
    "best crm for startups"
  ],
  "sort": "new",
  "time": "month",
  "maxPostsPerTarget": 100,
  "includeComments": true,
  "maxCommentsPerPost": 50
}
```

| Field | Type | Description |
|---|---|---|
| `startUrls` | array | Any mix of Reddit links and short names, one per line: a post URL, a community (`r/python` or its URL, optionally with `/top/?t=week`), a user (`u/spez` or the profile URL), a search URL, or a share link. Anything that is not a link is treated as a search keyword. |
| `searches` | array | Keywords to search across all of Reddit, one per line. |
| `searchType` | string | What keyword searches return. Default `"posts"`. |
| `sort` | string | Order for community listings and searches. `relevance` and `comments` apply to searches only; `rising` and `controversial` to community listings only. Default `"new"`. |
| `time` | string | Time window for `top`, `controversial` and searches. Default `"all"`. |
| `maxPostsPerTarget` | integer | Maximum posts to collect per community, user or search. Reddit itself stops a single listing at about 1,000 posts. Default `100`. |
| `includeComments` | boolean | Also collect comments for every post found in a community, user or search. Post URLs you paste directly always return their comments. Default `false`. |
| `maxCommentsPerPost` | integer | Maximum comments per post, nested replies included. Set 0 to skip comments even for post URLs. Default `100`. |
| `commentSort` | string | Which comments come first when a post has more than your limit. Default `"top"`. |
| `includeNSFW` | boolean | Include posts, comments and communities marked 18+. Off by default. Default `false`. |
| `maxItems` | integer | Hard cap on billable rows for the whole run. 0 means no cap. Default `0`. |

### Output example

A post row:

```json
{
  "type": "post",
  "id": "1wab3cd",
  "url": "https://www.reddit.com/r/webscraping/comments/1wab3cd/how_do_you_handle_rate_limits/",
  "title": "How do you handle rate limits at scale?",
  "text": "We are collecting about two million pages a day and ...",
  "subreddit": "webscraping",
  "author": "data_wrangler",
  "score": 184,
  "upvoteRatio": 0.97,
  "numComments": 63,
  "createdAt": "2026-09-27T15:10:03Z",
  "flair": "Getting started",
  "isSelf": true,
  "isNSFW": false,
  "subredditSubscribers": 91250,
  "source": "r/webscraping"
}
```

A comment row from the same run:

```json
{
  "type": "comment",
  "id": "pcdofpa",
  "url": "https://www.reddit.com/r/webscraping/comments/1wab3cd/how_do_you_handle_rate_limits/pcdofpa/",
  "text": "Rotate sessions before you hit the limit, not after.",
  "author": "proxy_pat",
  "score": 41,
  "depth": 1,
  "parentId": "pcdnx2q",
  "parentType": "comment",
  "isSubmitter": false,
  "postId": "1wab3cd",
  "postTitle": "How do you handle rate limits at scale?",
  "subreddit": "webscraping",
  "createdAt": "2026-09-27T16:02:44Z",
  "source": "r/webscraping"
}
```

### Pricing

Pay per result. **No start fee, no monthly rental, no minimum.**

| Event | Charged for | Price |
|---|---|---|
| `result-scraped` | each post, comment, user or community row delivered | **$1.99 / 1,000** |

| You export | You pay |
|---|---|
| 1,000 results | **$1.99** |
| 10,000 results | **$19.90** |
| 100,000 results | **$199.00** |

You are charged right after each batch lands in your dataset, so a run you stop early costs only what it delivered. Targets that fail, are private, or return nothing are free. Proxies are included in the price - there is nothing else to pay for.

### Use cases

#### Market and product research

Pull every post from the communities where your customers talk. Read what they praise, what they complain about and which products they name, in their own words.

#### Brand and competitor monitoring

Schedule a keyword search for your brand and your competitors every hour. New mentions arrive in the same dataset, ready for a Slack alert or a dashboard.

#### AI training data and RAG

Reddit threads are question-and-answer pairs written by people. Export posts with their comment trees, keep `depth` and `parentId`, and you have conversation data for fine-tuning, evaluation or retrieval.

#### Sentiment analysis

Collect comments on a product launch, a game update or a policy change, then score them with your own model. `score` and `createdAt` let you weight by visibility and chart the trend.

#### Content and SEO research

Find the questions people keep asking in your niche. Sort by top for the year to see what resonated, then write the answer.

#### Academic and social research

Study how communities react to events. Rows carry stable IDs and UTC timestamps, so datasets are reproducible.

### FAQ

**Do I need a Reddit account or API key?**
No. It works without the Reddit API. The Actor reads public pages without logging in.

**Is scraping Reddit legal?**
This is general information, not legal advice. The Actor collects only what any logged-out visitor can see, and never touches private communities, private messages or login-protected pages. Whether your use is lawful depends on where you are, what you collect and what you do with it. Usernames and posts can be personal data under laws such as GDPR and CCPA, and Reddit's User Agreement restricts automated collection and some commercial uses of its content. Collect only what you need, do not republish personal data, honour deletion requests, and ask a lawyer before building a commercial product on the data.

**How many posts can I get from one community?**
Reddit stops any single listing at about 1,000 posts. To go wider, run the same community with several sorts (`new`, `top` for the year, `top` for the month, `controversial`) or add keyword searches restricted to that community. Duplicates are easy to remove on `id`.

**How many comments can I get from one post?**
Up to 10,000 per post in this Actor. It expands "load more comments" and "continue this thread" branches automatically. For very large threads, the [Reddit Comments Scraper](https://apify.com/zaver.api/reddit-comments-scraper) goes up to 50,000 and does not charge for the post row.

**Can I scrape private communities or deleted content?**
No. Only content visible to a logged-out visitor is returned. Removed comments are skipped and not billed.

**Does it return NSFW content?**
Only if you turn on **Include NSFW content**. It is off by default, and skipped rows are not billed.

**Can I search inside one community?**
Yes. Paste a search URL such as `https://www.reddit.com/r/sales/search/?q=crm`, or use the [Reddit Search Scraper](https://apify.com/zaver.api/reddit-posts-search-scraper), which takes a list of communities.

**How fast is it?**
A listing page holds 100 posts and a comment page up to 500 comments, so a typical run delivers a few thousand rows per minute.

**Why did one of my inputs return an error row?**
The community or user may be private, banned, suspended or misspelled. Error rows name the input and the reason, and are never billed.

**Can I run it on a schedule?**
Yes. Use Apify schedules, the API, or an MCP client.

### Integrations

- **Apify API** - start runs and fetch datasets from Python, Node.js, cURL or any HTTP client.
- **MCP server** - call this Actor as a tool from Claude, Cursor, ChatGPT agents and any MCP client.
- **No-code** - Make, Zapier, n8n, Pipedream, Google Sheets, Airtable, Slack, webhooks.
- **Schedules** - run hourly, daily or weekly and append to the same dataset for monitoring.
- **Exports** - JSON, JSONL, CSV, Excel (XLSX), XML, HTML table, RSS.

### Related Reddit scrapers

- [Reddit Comments Scraper](https://apify.com/zaver.api/reddit-comments-scraper) - every comment and nested reply from any post
- [Reddit Search Scraper](https://apify.com/zaver.api/reddit-posts-search-scraper) - posts, communities and users by keyword
- [Reddit Lead Finder](https://apify.com/zaver.api/reddit-lead-finder) - people asking for what you sell, scored by buying intent

***

*Not affiliated with, endorsed by or sponsored by Reddit, Inc. This Actor reads only publicly available pages, does not log in and does not access private communities, private messages or deleted content. You are responsible for using the data in line with applicable laws, Reddit's terms and the privacy rights of the people whose public posts you collect.*

# Actor input Schema

## `startUrls` (type: `array`):

Any mix of Reddit links and short names, one per line: a post URL, a community (<code>r/python</code> or its URL, optionally with <code>/top/?t=week</code>), a user (<code>u/spez</code> or the profile URL), a search URL, or a share link. Anything that is not a link is treated as a search keyword.

## `searches` (type: `array`):

Keywords to search across all of Reddit, one per line.

## `searchType` (type: `string`):

What keyword searches return.

## `sort` (type: `string`):

Order for community listings and searches. <code>relevance</code> and <code>comments</code> apply to searches only; <code>rising</code> and <code>controversial</code> to community listings only.

## `time` (type: `string`):

Time window for <code>top</code>, <code>controversial</code> and searches.

## `maxPostsPerTarget` (type: `integer`):

Maximum posts to collect per community, user or search. Reddit itself stops a single listing at about 1,000 posts.

## `includeComments` (type: `boolean`):

Also collect comments for every post found in a community, user or search. Post URLs you paste directly always return their comments.

## `maxCommentsPerPost` (type: `integer`):

Maximum comments per post, nested replies included. Set 0 to skip comments even for post URLs.

## `commentSort` (type: `string`):

Which comments come first when a post has more than your limit.

## `includeNSFW` (type: `boolean`):

Include posts, comments and communities marked 18+. Off by default.

## `maxItems` (type: `integer`):

Hard cap on billable rows for the whole run. 0 means no cap.

## Actor input object example

```json
{
  "startUrls": [
    "r/webscraping"
  ],
  "searchType": "posts",
  "sort": "new",
  "time": "all",
  "maxPostsPerTarget": 100,
  "includeComments": false,
  "maxCommentsPerPost": 100,
  "commentSort": "top",
  "includeNSFW": false,
  "maxItems": 0
}
```

# Actor output Schema

## `results` (type: `string`):

All dataset items.

## `resultsCsv` (type: `string`):

CSV download.

## `runSummary` (type: `string`):

Counts for the run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "r/webscraping"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("zaver.api/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["r/webscraping"] }

# Run the Actor and wait for it to finish
run = client.actor("zaver.api/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "r/webscraping"
  ]
}' |
apify call zaver.api/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,zaver.api/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/GdxgAVhbMHRmmcXY8/builds/4trlr3HyTd1E7ecTu/openapi.json
