# Reddit Scraper - Posts, Comments, Subreddits & Users (`seemuapps/reddit-scraper`) Actor

Scrape Reddit posts, comments, subreddit stats and user activity from subreddit, post, profile or search URLs and keywords - export to CSV, JSON or Excel.

- **URL**: https://apify.com/seemuapps/reddit-scraper.md
- **Developed by:** [Seemu Scraping](https://apify.com/seemuapps) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper - Posts, Comments, Subreddits & Users

Scrape Reddit posts, comment threads, subreddit stats and user activity from any mix of subreddit, post, user profile and search URLs, or plain keywords. No login, no Reddit API key, no browser. Every result is a flat row, so it exports cleanly to CSV, Excel, JSON or Google Sheets.

### What you get

- **Posts**: post ID, URL, subreddit, author, title, body text, score, upvote ratio, comment count, created date, flair, outbound link and domain, plus self-post, video, NSFW, spoiler, locked and pinned flags
- **Comments**: comment ID, post ID and title, parent ID, reply depth, author, body, score, created date, permalink, and whether the commenter is the original poster (OP)
- **Subreddits**: name, description, subscriber count, weekly active users, weekly contributions, creation date, icon and banner image
- **User activity**: a user's posts and comments, with the same fields as above
- **Search results**: posts matching any keyword, across all of Reddit or inside one subreddit
- Filters for sort order (hot, new, top, rising, relevance, most comments), time window (past hour to all time), keywords, and comment depth
- Pagination across runs, so you can keep collecting a subreddit or search in batches

### Use cases

- **Market research and brand monitoring**: track what people say about your product, competitors or industry
- **Lead generation**: find people asking for recommendations in niche subreddits
- **Sentiment analysis and NLP datasets**: collect post and comment text at scale for classification or fine-tuning
- **Content and trend research**: find the top posts and discussions in any community over a given time window
- **Community analytics**: compare subreddit size, activity and engagement

### How to use

1. Add one or more **Reddit URLs**. You can mix:
   - Subreddits: `https://www.reddit.com/r/webscraping/` or `r/webscraping`
   - Posts: `https://www.reddit.com/r/AskReddit/comments/abc123/...`, comment permalinks or `redd.it` short links
   - User profiles: `https://www.reddit.com/user/spez` or `u/spez`
   - Search pages: `https://www.reddit.com/search/?q=web+scraping` or `https://www.reddit.com/r/webscraping/search/?q=proxy`
2. Or type **Search keywords**, one per line. You can limit them to one subreddit with **Search within subreddit**.
3. Choose **Sort by** and **Time filter**.
4. Set **Max posts per source** (default 50; 0 = no limit).
5. Turn on **Include comments** to scrape the comment thread of every post. Use **Max comments per post** and **Max reply depth** to control how much you get.
6. Run the Actor. Results appear in the **Output** tab, with views for Posts, Comments and Subreddits.

#### Continuing a run (pagination)

When a run scrapes a single subreddit, search, or user (with posts-only or comments-only selected), it saves a cursor so the next run can pick up where it stopped.

After the run finishes, open the **Key-value store** tab, copy the `NEXT_PAGE_ID` value and paste it into **Page ID** on your next run. If `NEXT_PAGE_ID` is `null`, you've fetched everything.

### Output format

Every row has a `recordType` of `post`, `comment` or `subreddit`. All rows share the same columns, and columns that don't apply to a row are empty, so one CSV holds everything. Comments link to their post through `postId` and to their parent comment through `parentId`.

A post row:

```json
{
  "recordType": "post",
  "sourceType": "subreddit",
  "source": "https://www.reddit.com/r/webscraping/",
  "id": "1q6pxwn",
  "url": "https://www.reddit.com/r/webscraping/comments/1q6pxwn/just_started_web_scraping_is_this_a_good_start/",
  "subreddit": "webscraping",
  "author": "franik33",
  "title": "Just Started Web Scraping — Is This a Good Start?",
  "body": "Hi everyone, I started getting into web scraping about 3–4 days ago...",
  "score": 33,
  "upvoteRatio": 0.92,
  "numComments": 53,
  "createdAt": "2026-01-07T20:01:30.000Z",
  "flair": null,
  "linkUrl": null,
  "isSelf": true,
  "over18": false
}
```

A comment row:

```json
{
  "recordType": "comment",
  "id": "nyfjrfq",
  "postId": "1q6pxwn",
  "postTitle": "Just Started Web Scraping — Is This a Good Start?",
  "parentId": "1q6pxwn",
  "depth": 0,
  "author": "hasdata_com",
  "body": "Nice work for 3 days in...",
  "score": 19,
  "isOp": false,
  "createdAt": "2026-01-08T17:16:35.000Z",
  "url": "https://www.reddit.com/r/webscraping/comments/1q6pxwn/comment/nyfjrfq/"
}
```

A subreddit row:

```json
{
  "recordType": "subreddit",
  "subreddit": "webscraping",
  "description": "The first rule of web scraping is...",
  "subscribers": null,
  "weeklyActiveUsers": 19779,
  "weeklyContributions": 398,
  "createdAt": "2014-04-06T10:50:02.243Z",
  "iconUrl": "https://styles.redditmedia.com/..."
}
```

If a source returns nothing, for example a subreddit that doesn't exist or is banned or private, a user with no public activity, or a deleted post, you get one free `notice` row with a `message` explaining why. The rest of the run continues.

### Tips

- **Sort and time filter**: the time filter applies to searches, user activity and subreddit listings sorted by Top or New. Like on Reddit, Hot and Rising listings ignore it.
- **User profiles**: posts and comments come from Reddit's search index, so very old or removed activity may be missing.
- **Comments**: threads are returned in Reddit's order, with top comments and their replies first. Replies hidden behind "load more replies" links are not expanded.
- **Keyword filter**: use **Only keep posts containing** to drop loosely related search results.
- **Duplicates**: a post that appears in several sources during a run is returned only once.

### FAQ

**Do I need a Reddit account or API key?** No. The Actor only reads public Reddit content.

**Can I scrape private or banned subreddits?** No. Only public content is available. Those sources return a free notice row.

**How do I export to Excel or Google Sheets?** Open the run's **Output** tab and choose CSV, Excel or JSON, or use the Google Sheets integration.

# Actor input Schema

## `startUrls` (type: `array`):

Any mix of subreddit URLs (https://www.reddit.com/r/webscraping or r/webscraping), post URLs (including comment permalinks and redd.it links), user profile URLs (https://www.reddit.com/user/spez or u/spez) and Reddit search URLs (https://www.reddit.com/search/?q=...).

## `searches` (type: `array`):

Keywords or phrases to search Reddit posts for, one per line. Each keyword is its own source and results are de-duplicated across the run.

## `searchSubreddit` (type: `string`):

Restrict keyword searches to one subreddit, e.g. 'technology' or 'r/technology'. Leave empty to search all of Reddit.

## `sort` (type: `string`):

Sort order for subreddit listings, searches and user activity. Subreddits use hot/new/top/rising; searches use relevance/new/top/most comments (hot maps to relevance, rising to new). User profiles use new or top.

## `timeFilter` (type: `string`):

Only return posts from this window. Applies to searches, user activity, and subreddit listings sorted by Top or New (Hot and Rising listings ignore it, as on Reddit).

## `maxItems` (type: `integer`):

Maximum posts to return for each subreddit, search or user (user profiles return up to this many posts and this many comments). 0 = no limit, keep going until the list ends or the run times out.

## `includeComments` (type: `boolean`):

Also scrape the comment thread of every post. Each comment becomes its own row (recordType 'comment') linked to its post by postId.

## `maxCommentsPerPost` (type: `integer`):

Maximum comments to return per post when comments are included, in thread order (top comments with their replies first). 0 = all comments that Reddit loads without expanding collapsed branches.

## `maxCommentDepth` (type: `integer`):

How many levels of replies to keep. 1 = top-level comments only, 2 = plus direct replies, and so on. 0 = no depth limit.

## `includeSubredditInfo` (type: `boolean`):

For every subreddit URL, add one row (recordType 'subreddit') with its description, subscriber count, weekly active users, weekly contributions, icon and banner.

## `userContent` (type: `string`):

What to scrape for user profile URLs: the user's posts, their comments, or both.

## `filterKeywords` (type: `array`):

Optional. Keep only posts whose title or text contains at least one of these words (case-insensitive). User comments are matched on their text. Leave empty to keep everything.

## `pageId` (type: `string`):

Paste NEXT\_PAGE\_ID from the previous run's Key-value store to continue where it stopped. Only used when the run has exactly one subreddit, search, or user (posts-only or comments-only) source.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.reddit.com/r/webscraping/"
  ],
  "sort": "hot",
  "timeFilter": "all",
  "maxItems": 50,
  "includeComments": false,
  "maxCommentsPerPost": 20,
  "maxCommentDepth": 0,
  "includeSubredditInfo": true,
  "userContent": "both"
}
```

# Actor output Schema

## `results` (type: `string`):

One row per post, comment or subreddit, identified by recordType. Posts: id, url, subreddit, author, title, body, score, upvoteRatio, numComments, createdAt, flair, linkUrl and flags. Comments: id, postId, parentId, depth, author, body, score, isOp. Subreddits: subscribers, weeklyActiveUsers, description, icon. Notice rows explain sources that returned nothing and are free.

## `nextPageId` (type: `string`):

NEXT\_PAGE\_ID record in the default key-value store. Paste into Page ID on the next run to resume a single-source run; null when the list is exhausted.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.reddit.com/r/webscraping/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("seemuapps/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": ["https://www.reddit.com/r/webscraping/"] }

# Run the Actor and wait for it to finish
run = client.actor("seemuapps/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.reddit.com/r/webscraping/"
  ]
}' |
apify call seemuapps/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,seemuapps/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/mdL7QbzgUdzXyQayj/builds/y8IPDp6dUOT3aRBT6/openapi.json
