# Reddit Scraper – Posts, Comments & Search (`mscraper/reddit-scraper`) Actor

Extract Reddit posts and comments from subreddits, searches and post URLs. Get text, scores, timestamps, media links and reply relationships. Export JSON, CSV or Excel. Provider access included; pay per saved result.

- **URL**: https://apify.com/mscraper/reddit-scraper.md
- **Developed by:** [mscraper](https://apify.com/mscraper) (community)
- **Categories:** Social media
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.49 / 1,000 reddit results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Reddit Scraper do?

**Collect Reddit posts and comments as structured data** from subreddit pages, keyword searches, and individual discussion URLs. Extract post text, comment bodies, scores, timestamps, media links, author names, and parent/reply relationships from [Reddit](https://www.reddit.com/).

Start with the default input to collect 10 recent posts from r/programming. Provider access is included: you do not need a Reddit account, an API key, cookies, or a proxy subscription. Run through Apify Console or the API, schedule recurring runs, and connect the dataset to your existing workflow.

### Why use Reddit Scraper?

- **Topic research:** collect recent or top posts from a specific subreddit.
- **Brand and product research:** search post titles and text for keywords, optionally within one community.
- **Discussion analysis:** collect comments from selected posts and retain the post and parent IDs.
- **Data pipelines:** export consistent records without parsing Reddit pages yourself.

Posts and comments share one dataset and are distinguished by `dataType`. Missing provider fields are `null`; media collections are arrays. Counts and scores are snapshots, and can change after the run.

### How to use Reddit Scraper

1. Open the **Input** tab and enter a subreddit URL, one or more post URLs, or search terms.
2. Choose the sort order and maximum total results. The default collects posts only.
3. To include discussion replies, turn off **Skip comments** and set **Maximum comments per post**. Increase the total result limit to leave room for both posts and comments.
4. Set your maximum run charge in Apify, then start the Actor.
5. Open the **Dataset** output to inspect and export the results. **Run summary** reports saved counts, provider attempts, warnings, and why the run stopped.

### Input

The Input tab contains all supported options. A small starting input:

```json
{
    "startUrls": [{ "url": "https://www.reddit.com/r/programming/" }],
    "sort": "new",
    "maxItems": 10,
    "maxPostCount": 10,
    "skipComments": true,
    "maxRequests": 10,
    "maxRetries": 1
}
```

For keyword search, leave `startUrls` empty and use:

```json
{
    "startUrls": [],
    "searches": ["typescript"],
    "searchCommunityName": "programming",
    "sort": "top",
    "time": "month",
    "maxItems": 10,
    "skipComments": true,
    "maxRequests": 3
}
```

| Option                              | Meaning                                                                              |
| ----------------------------------- | ------------------------------------------------------------------------------------ |
| `startUrls`                         | HTTPS Reddit subreddit or post URLs. Profile URLs are not supported.                 |
| `searches`, `searchCommunityName`   | Post keyword searches, optionally restricted to a subreddit.                         |
| `sort`                              | Subreddits: `new`, `hot`, `top`. Searches also support `relevance` and `comments`.   |
| `time`                              | `hour`, `day`, `week`, `month`, `year`, or `all`; applied only with `top` sorting.   |
| `maxItems`                          | Global cap on saved posts **plus** comments across all sources.                      |
| `maxPostCount`                      | Global cap on saved posts.                                                           |
| `maxComments`                       | Maximum saved comments per post, including returned replies.                         |
| `skipComments`                      | Default `true`; set `false` to retrieve comments.                                    |
| `commentSort`                       | `top`, `new`, `confidence`, `controversial`, `old`, or `qa`.                         |
| `postDateLimit`, `commentDateLimit` | Oldest accepted date: ISO date or a relative value such as `7 days`.                 |
| `includeNSFW`                       | Include returned posts marked NSFW. Does not bypass access restrictions.             |
| `maxRequests`, `maxRetries`         | Bound provider calls and transient retries. Every retry counts toward `maxRequests`. |

Date and NSFW filters apply to the data returned by the provider. Restrictive filters can consume requests without producing records. Earlier sources may consume the global result budget before later sources are visited. `time` is not an exact local date filter.

### Output

An illustrative post record (synthetic values; shortened for readability):

```json
{
    "dataType": "post",
    "id": "t3_example123",
    "parsedId": "example123",
    "url": "https://www.reddit.com/r/programming/comments/example123/example/",
    "username": "sample_user",
    "title": "A discussion about TypeScript",
    "communityName": "r/programming",
    "parsedCommunityName": "programming",
    "body": "Example post text.",
    "upVotes": 42,
    "numberOfComments": 12,
    "createdAt": "2026-09-25T12:00:00.000Z",
    "scrapedAt": "2026-09-27T00:00:00.000Z",
    "imageUrls": [],
    "videoUrls": []
}
```

Comments use `dataType: "comment"`, IDs starting with `t1_`, and `postId`/`parentId` to identify their discussion and immediate parent. A parent comment may be outside your requested result limit. The dataset contains a full flat record for each item; an absent title on a comment is `null`.

You can download the dataset in **JSON, HTML, CSV, or Excel**. Use JSON when you need media arrays and explicit null values. Exporting the same dataset again does not create new scraped-result events.

### Data table

| Fields                                             | Type                   | Description                                                                          |
| -------------------------------------------------- | ---------------------- | ------------------------------------------------------------------------------------ |
| `dataType`, `id`, `parsedId`                       | string                 | Record type, full Reddit ID and ID without prefix.                                   |
| `url`                                              | string or null         | Reddit permalink.                                                                    |
| `username`, `userId`, `authorFlair`                | string or null         | Available author information. Deleted authors may be absent.                         |
| `title`, `body`, `html`                            | string or null         | Title, text and provider HTML. Treat HTML as untrusted content before displaying it. |
| `communityName`, `parsedCommunityName`, `category` | string or null         | Community with/without `r/`; category retains the subreddit name.                    |
| `link`, `flair`                                    | string or null         | Post destination and flair.                                                          |
| `postId`, `parentId`                               | string or null         | Comment relationships.                                                               |
| `numberOfComments`, `numberOfReplies`              | integer or null        | Counts supplied by the provider, not inferred from partial results.                  |
| `upVotes`, `upVoteRatio`                           | integer/number or null | Provider `ups` and ratio; not an independent count of all votes.                     |
| `isVideo`, `isAd`, `over18`                        | boolean or null        | Provider flags; unknown flags remain null.                                           |
| `thumbnailUrl`, `imageUrls`, `videoUrls`           | URL or URL arrays      | Available media links. Media files are not downloaded.                               |
| `createdAt`, `scrapedAt`                           | ISO timestamp or null  | Creation time when available; extraction time is always present.                     |

### Shared usage protection

All cloud requests, including retries, reserve a shared provider allowance before dispatch. Defaults: 140,000 provider units per 30-day safety window, 1,000 units per paid user per 24-hour window, and 3 units per Free user per 30-day window. All Free users share a pool of 100 units. Windows begin with first usage and do not mirror subscription resets. Concurrent runs share the same counters. Contact the developer for larger allowances. These safety caps can stop a run before its requested limits.

### Pricing / Cost estimation

**$0.49 per 1,000 saved results**, plus **$0.01 per Actor start at the default 256 MB memory**. One result is one unique post or one unique comment saved during the run. Provider access is included; no separate provider subscription is required.

| Saved results            | Estimated charge at default memory |
| ------------------------ | ---------------------------------: |
| 10 posts                 |                           $0.01490 |
| 1 post + 20 comments     |                           $0.02029 |
| 100 posts + 900 comments |                           $0.50000 |

**New empty-check event:** $0.00025 per provider unit ($0.25/1,000 units) for a successful response containing a validated empty list. Provider errors, ambiguous not-found errors and malformed responses are not empty checks. Nonempty pages, filtered records and duplicates do not trigger this event. This fee applies only when `empty-check-unit` is present in the run's actual pricing snapshot. Existing pricing snapshots remain supported with a 40-request cap and no empty-check fee during the pricing transition.

Apify charges the start event per GB of allocated memory, with a minimum of one event. Increasing memory above 1 GB increases the start charge. The Pricing tab and your run's displayed pricing are authoritative.

The Actor checks whether another result is affordable **before fetching another page and before writing each record**. It stops when no further result fits your maximum charge. Duplicate and filtered records are not saved or charged as results. Empty or failed runs can still incur the start charge; results saved before a later provider error remain available and billable. Running the same input again creates a new run and can save/charge those records again.

### Tips and advanced options

- Start with 10–25 results and inspect the output before increasing limits.
- Keep comments off when you only need posts. For comments, set both a per-post cap and a total result cap.
- Use `top` with a time window for bounded keyword research; use date limits for local filtering.
- Runs share the provider's account-level request allowance. Waiting for an available request slot is expected during concurrent runs; increasing memory does not increase provider throughput.
- `request-budget`, `item-limit`, `post-limit`, and `charge-limit` in the summary identify intentional stops. A successful bounded run does not mean that every available Reddit result was retrieved.
- If a later page fails, the Actor keeps earlier saved results and reports the failure. Inspect the summary before assuming a dataset is complete.

### FAQ, limitations, and support

**Does this retrieve every post or comment?** No. Provider search, pagination, Reddit access, deletions, and your configured limits determine coverage. The comments V2 endpoint returned two distinct pages of 50 comments in validation; this is not a completeness guarantee. A tested second page of subreddit posts returned a provider error, so deep historical subreddit collection should not be assumed reliable.

**Can I search comments or scrape user profiles?** This version searches posts and retrieves comments under selected posts. Keyword search over all Reddit comments, user profiles, and community discovery are not supported. Unsupported input fields are rejected before a provider request.

**Is this a drop-in replacement for other Reddit Actors?** It uses familiar post/comment field names but supports its own documented subset. Check input modes, limits, and nullable output fields before switching an existing workflow.

**Can I access private, deleted, or restricted content?** The Actor does not promise access to content unavailable through its provider and does not bypass private-community restrictions.

**Can I resume a stopped run?** Treat a stopped or failed run as partial and preserve its dataset. Start a new run for additional work; automatic exact-cursor recovery is not currently promised. Deduplication applies within a run.

**Where can I get help?** Use the Actor's **Issues** tab with a run ID and a redacted input. Do not include credentials. Custom output mappings or additional modes can be discussed there.

This Actor is independently developed by mscraper and is not affiliated with or endorsed by Reddit. Use collected data in accordance with applicable terms and privacy requirements, and avoid collecting data you are not entitled to process.

### Related scrapers for community and market research

Extend Reddit discussion research with public social content and domain-level traffic estimates.

- [TikTok Scraper](https://apify.com/mscraper/tiktok-scraper) — Collect videos, creator profiles and comments to explore how a topic appears in short-form content.
- [X (Twitter) Scraper](https://apify.com/mscraper/x-scraper) — Retrieve public profiles, timeline posts and follower lists for accounts relevant to your research.
- [Similarweb Quick Scraper](https://apify.com/mscraper/similarweb-quick-scraper) — Check estimated visits, traffic sources and top countries for websites mentioned in Reddit discussions.

Run each Actor separately and combine its exported data in your own workflow. Each Actor has its own input, output and pricing.

# Actor input Schema

## `startUrls` (type: `array`):

Reddit subreddit or post URLs. Leave empty when using search terms. Profile URLs are not supported.

## `searches` (type: `array`):

Search Reddit posts.

## `searchCommunityName` (type: `string`):

Optional community name, with or without r/.

## `ignoreStartUrls` (type: `boolean`):

Use only search terms.

## `searchPosts` (type: `boolean`):

Enable post searches. Other search record types are not implemented.

## `maxItems` (type: `integer`):

Global cap on saved posts plus comments.

## `maxPostCount` (type: `integer`):

Global post cap across all sources. Comments also share maxItems.

## `sort` (type: `string`):

Subreddit URLs support hot, new, top. Search supports all listed modes.

## `time` (type: `string`):

Applied only with top sorting; the provider rejects time for other search modes.

## `postDateLimit` (type: `string`):

Past ISO date or duration, such as 7 days. Local filtering may require extra API pages.

## `includeNSFW` (type: `boolean`):

When false, excludes posts flagged over\_18. Does not force the provider to retrieve age-restricted content.

## `skipComments` (type: `boolean`):

Retrieve only posts; reduces API requests.

## `maxComments` (type: `integer`):

Saved comments per post, including replies returned by the provider.

## `commentSort` (type: `string`):

Sorting for the verified V2 comments endpoint.

## `commentDateLimit` (type: `string`):

Past ISO date or duration, such as 7 days. Local filtering may require extra API pages.

## `maxRequests` (type: `integer`):

Total paid API attempts, including retries. The run stops when exhausted.

## `maxRetries` (type: `integer`):

Retries only network errors, HTTP 429 and 5xx. All retries consume maxRequests.

## `requestsPerMinute` (type: `integer`):

One request at a time. This is a per-run limit; concurrent actors share the provider subscription.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/programming/"
    }
  ],
  "searches": [],
  "searchCommunityName": "",
  "ignoreStartUrls": false,
  "searchPosts": true,
  "maxItems": 10,
  "maxPostCount": 10,
  "sort": "new",
  "time": "all",
  "postDateLimit": "",
  "includeNSFW": false,
  "skipComments": true,
  "maxComments": 10,
  "commentSort": "top",
  "commentDateLimit": "",
  "maxRequests": 10,
  "maxRetries": 1,
  "requestsPerMinute": 40
}
```

# Actor output Schema

## `dataset` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.reddit.com/r/programming/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mscraper/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.reddit.com/r/programming/" }] }

# Run the Actor and wait for it to finish
run = client.actor("mscraper/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/programming/"
    }
  ]
}' |
apify call mscraper/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mscraper/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/qtLncOOYJdbtHEQaW/builds/yThF5v2D6OD3zAbv0/openapi.json
