# Reddit Scraper (`calm_builder/reddit-scraper`) Actor

Scrape Reddit posts, comments, users and subreddits without login or API key. Use subreddits, keyword search, usernames or any Reddit link. Deep mode collects far more than Reddit's 1,000-post limit. Full comment threads, nested or flat. Pay only per result.

- **URL**: https://apify.com/calm_builder/reddit-scraper.md
- **Developed by:** [Coder](https://apify.com/calm_builder) (community)
- **Stats:** 3 total users, 2 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper

Scrape Reddit **posts, comments, user profiles and subreddits** without a Reddit account, API key or cookies. Add subreddits, keyword searches, usernames or any Reddit link, mix them in one run, and download clean structured data as JSON, CSV or Excel.

### What makes it different

- **More than Reddit normally shows.** Reddit shows about 1,000 posts per sort and about 250 per search. Turn on **Deep mode** to collect several times as many unique posts from big subreddits and searches.
- **Complete comment threads.** Nested replies included. Threads with thousands of comments come back nearly in full (98% of an 8,786-comment thread in about half a minute in our tests).
- **Rich data.** About 47 fields per post and 34 per comment: score, upvote ratio, flair, author flair, awards, edited date, locked/stickied/archived, images, gallery, video URL, links in the text, crosspost info and more.
- **Fair billing.** You pay only for rows you receive. Duplicates, deleted placeholders and results removed by your filters are never charged.
- **Reliable.** Built to keep working when Reddit changes things, with health checks on every part.

### What you can scrape

| Source | Example | You get |
|---|---|---|
| Subreddit | `technology`, `r/AskReddit`, `https://www.reddit.com/r/Cooking/top/?t=week` | Posts (hot, new, top, rising, controversial) |
| Search | `iphone 17`, `"exact phrase" subreddit:apple` | Posts, comments or subreddits |
| User | `spez`, `u/spez` | Profile, posts and/or comments |
| Post URL | `https://www.reddit.com/r/technology/comments/1x00kmg/...` | The post (+ its comments) |
| Any link | old.reddit.com, redd.it, share links `/r/name/s/...` | Detected automatically |

### How to use it

1. Add one or more **subreddits**, **search terms**, **usernames** or **Reddit URLs**.
2. Set **Maximum results per source**, and turn on **Deep mode** if you need more than Reddit normally shows.
3. Turn on **Include comments** to get the discussion under every post, as one row per comment or as threads.
4. Optional: date range (pick a date or a period like the last 7 days), minimum score, strict keyword match, NSFW on/off.
5. Click **Start** and download the results in any format.

### Output

You get **five tables**:

| Table | Contains |
|---|---|
| **All results** | Every row of the run with a `dataType` column (`post`, `comment`, `user`, `subreddit`). Never empty. |
| **Posts** | Only posts, from every source |
| **Comments** | Only comments and replies |
| **Users** | Only user profiles |
| **Subreddits** | Only subreddit info |

Use the clean tables for CSV/Excel exports and Google Sheets: each has only its own columns. Every row has an `input` column that tells you which subreddit, search, user or URL it came from.

Example post:

```json
{
  "dataType": "post",
  "id": "1x00kmg",
  "url": "https://www.reddit.com/r/technology/comments/1x00kmg/...",
  "title": "Man discovers his parents' coffee machine used 1TB of data",
  "body": null,
  "postType": "link",
  "link": "https://www.example.com/article",
  "score": 21450,
  "upvoteRatio": 0.95,
  "commentCount": 1712,
  "flair": "Hardware",
  "author": "example_user",
  "subreddit": "technology",
  "subredditSubscribers": 17000000,
  "createdAt": "2026-10-05T14:02:11Z",
  "nsfw": false,
  "input": "r/technology"
}
```

Example comment:

```json
{
  "dataType": "comment",
  "id": "pegh5t5",
  "postId": "1x00kmg",
  "postTitle": "Man discovers his parents' coffee machine used 1TB of data",
  "parentId": "1x00kmg",
  "parentType": "post",
  "depth": 0,
  "body": "The place I'm renting has a samsung smart fridge...",
  "score": 7094,
  "controversiality": 0,
  "author": "example_user",
  "isSubmitter": false,
  "createdAt": "2026-10-06T08:16:05Z"
}
```

#### Comment format

- **One row per comment** (default): every comment and reply is its own row, linked by `parentId`, `parentType`, `depth`, `postId` and `replyCount`. Best for spreadsheets and analysis.
- **Threads**: one row per top-level comment with its replies nested inside a `replies` array (replies to replies too), like on Reddit. Best for reading conversations or feeding AI. Best exported as JSON.

Both formats cost the same: you pay per comment, whether it is its own row or nested in a thread.

User rows include karma (post, comment, total), account creation date, bio, avatar, verified/moderator/premium flags. Subreddit rows include description, members, weekly visitors, weekly contributions, creation date, NSFW and type.

### Pricing

Pay per result: each post, comment, user profile and subreddit is charged once, when it is saved. Use **Maximum results per source** and **Maximum comments per post** to control the cost, and set a maximum cost per run in the run options.

### Tips

- **More than 1,000 posts from one subreddit?** Turn on Deep mode and raise the maximum.
- **Only new posts since your last run?** Sort by New and set Posted after to the last day. The run stops as soon as it reaches older posts, so scheduled monitoring stays cheap.
- **Search returns loosely related posts?** Turn on Strict keyword match.
- **Only subreddit info or user profiles?** Set Maximum results per source to 0.
- **A user's comments without their profile?** Turn off Include user profile.
- **Brand monitoring:** search your brand name, sort by New, schedule the run daily and connect the dataset to Slack, Google Sheets or a webhook.

### FAQ

**Do I need a Reddit account or API key?** No.

**Can it scrape private or banned subreddits?** No. Only content visible to anyone without logging in.

**Why fewer results than my maximum?** The subreddit, user or search has fewer posts than that, or Reddit's normal limit was reached. Turn on Deep mode to go further.

**In deep mode, are results in my chosen sort?** The posts Reddit shows for your sort come first, in that order. The extra posts deep mode adds come after them, without a guaranteed order; sort the dataset by score or date if you need one.

**Is scraping Reddit legal?** This Actor collects only publicly visible data. You are responsible for how you use it, including personal data rules such as GDPR. When in doubt, ask a lawyer.

### Support

Found a problem or need a field that is missing? Open an issue on the Issues tab and it will be looked at quickly.

# Actor input Schema

## `subreddits` (type: `array`):

Subreddit names or links, one per line. Examples: technology, r/AskReddit, https://www.reddit.com/r/Cooking/

## `searches` (type: `array`):

Keywords searched across all of Reddit, one search per line. Reddit search operators work: "exact phrase", subreddit:name, author:name, site:domain.com, flair:name.

## `users` (type: `array`):

Reddit users, one per line. Examples: spez, u/spez, https://www.reddit.com/user/spez/. Each user gives a profile row plus their posts and/or comments (see 'User content').

## `startUrls` (type: `array`):

Any Reddit link, one per line: a post, a subreddit (with sort, e.g. /r/technology/top/?t=week), a user profile, a search results page or a share link (/r/name/s/...). old.reddit.com, new.reddit.com and redd.it links work too.

## `maxPostsPerSource` (type: `integer`):

How many posts to collect from each subreddit, search or user (comments or subreddits for those search types). Set 0 to get only the subreddit info and user profiles, without posts. Reddit normally shows about 1,000 posts per sort and about 250 per search; turn on deep mode to collect more.

## `deepMode` (type: `boolean`):

Reddit normally shows about 1,000 posts per subreddit sort and about 250 per search. Deep mode collects many more unique posts (often several times as many on big subreddits and searches). Your time range and filters still apply. Results in your chosen sort come first, in Reddit's order; the extra posts are added after them without a guaranteed order. Takes longer. Use it with a high 'Maximum results per source'.

## `sort` (type: `string`):

Order of subreddit posts. A sort in a subreddit URL takes priority.

## `searchSort` (type: `string`):

Order of search results.

## `time` (type: `string`):

Used by the Top and Controversial sorts and by search.

## `searchFor` (type: `string`):

What the search terms should find.

## `searchSubreddit` (type: `string`):

Optional. Limit the search terms to one subreddit, e.g. technology.

## `strictKeywordMatch` (type: `boolean`):

Keep only results whose title or text contains every word of the search term. Reddit search can return loosely related posts; this removes them, and you are not charged for removed results.

## `includeComments` (type: `boolean`):

Also collect the comments of every post found, including all nested replies. Choose the format below.

## `maxCommentsPerPost` (type: `integer`):

0 = all comments. Large threads with thousands of comments are supported.

## `commentSort` (type: `string`):

Which comments come first when you collect only part of a thread.

## `commentFormat` (type: `string`):

How comments are saved. 'One row per comment': every comment and reply is its own row, linked by parentId, depth and postId; best for spreadsheets and analysis. 'Threads': one row per top-level comment with its replies nested inside (replies to replies too), like on Reddit; best for reading conversations or AI. Applies to comments of posts; comment search and user comments are always one row per comment. Both formats cost the same: you pay per comment.

## `userContent` (type: `string`):

What to collect for each username.

## `includeUserProfile` (type: `boolean`):

Add one profile row per username (karma, join date, bio, avatar). Turn off to get only their posts or comments.

## `includeCommunityInfo` (type: `boolean`):

Add one row per subreddit with its description, members, weekly visitors, weekly contributions, creation date and more. Set 'Maximum results per source' to 0 to get only this info.

## `postedAfter` (type: `string`):

Only results posted on or after this date. Pick a date, or a relative period such as the last 7 days.

## `postedBefore` (type: `string`):

Only results posted on or before this date. Pick a date, or a relative period such as 30 days ago.

## `minScore` (type: `integer`):

Only posts and comments with at least this score.

## `includeNsfw` (type: `boolean`):

Turn off to skip posts marked NSFW.

## Actor input object example

```json
{
  "subreddits": [
    "technology"
  ],
  "maxPostsPerSource": 100,
  "deepMode": false,
  "sort": "hot",
  "searchSort": "relevance",
  "time": "all",
  "searchFor": "posts",
  "strictKeywordMatch": false,
  "includeComments": false,
  "maxCommentsPerPost": 100,
  "commentSort": "top",
  "commentFormat": "flat",
  "userContent": "posts",
  "includeUserProfile": true,
  "includeCommunityInfo": false,
  "includeNsfw": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `posts` (type: `string`):

No description

## `comments` (type: `string`):

No description

## `users` (type: `string`):

No description

## `subreddits` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "technology"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("calm_builder/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "subreddits": ["technology"] }

# Run the Actor and wait for it to finish
run = client.actor("calm_builder/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "technology"
  ]
}' |
apify call calm_builder/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,calm_builder/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/k28xIFJ523q96myo1/builds/eoqD3kuf1bggHRU3M/openapi.json
