# Reddit Scraper ⚡ Advanced Data, Best Value (`claygenius/reddit-scraper`) Actor

Scrape Reddit posts, comments, subreddits and user profiles from any Reddit link or keyword search. No login, no browser, flat spreadsheet-ready rows, low cost per result.

- **URL**: https://apify.com/claygenius/reddit-scraper.md
- **Developed by:** [Muhammad Shamshad Aslam](https://apify.com/claygenius) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper

Paste any Reddit link or type a keyword, get clean rows back: **posts, comments, communities and user profiles**, each with score, author, dates, flair, media links and the post every comment belongs to. No login, no browser, no cookies to paste. One form, one dataset, low cost per result.

Works with the links people actually copy: a post, a subreddit, a user, a Reddit search page, a `redd.it` short link or a mobile share link. Mix them freely with keyword searches; everything lands in the same dataset.

### Why this one

- 🔗 **Paste anything** — post, subreddit (`r/pasta`, `/r/pasta/top/?t=week`), user (`u/spez`, `/user/spez/comments`), search page, community profile (`/r/pasta/about`), `redd.it` and `/s/` share links, old.reddit and new.reddit hosts
- 🔍 **Keyword search built in** — find posts, communities or users; restrict to one or more subreddits; full Reddit search syntax (`"exact phrase"`, `title:`, `author:`, `-excluded`)
- 💬 **Comments done properly** — score, depth, parent comment, top-level flag, OP flag, and the post title and URL on every comment row. Pick Top / Best / New / Controversial / Old / Q\&A order. Choose separate rows (spreadsheets, Clay) or comments nested inside the post (JSON, AI pipelines)
- 🖼️ **Media resolved** — `postType` (text / link / image / video / gallery / poll / crosspost), direct `mediaUrl`, every `galleryImages` entry, thumbnail, external link and its domain
- 📅 **Filters that save requests** — date range, minimum score, NSFW switch. On chronological listings the scraper stops paging as soon as it passes your start date
- 🧾 **Nothing silently dropped** — private, banned, deleted or malformed inputs are listed in the key-value store record `FAILED_URLS` and summarised in the log
- ✅ **Duplicate-free** — the same post reached through two URLs or a URL and a search is saved once
- ✅ **No login, no proxy fiddling** — reads Reddit as a logged-out visitor; the default proxy setting just works

### Input

| Field | Type | Default | Description |
|---|---|---|---|
| `startUrls` | array | 1 post + 1 subreddit | Reddit links, one per line. See the list above for what each kind of link returns |
| `searchQueries` | array | empty | Keywords, one per line. Each is its own search |
| `searchType` | select | `posts` | What keywords should find: `posts`, `communities` or `users` |
| `searchInSubreddits` | array | empty | Restrict post searches to these subreddits (`pasta` or `r/pasta`). Each keyword is searched in each subreddit |
| `sort` | select | `relevance` | Order for searches and for subreddit URLs without a sort in the path: `relevance`, `hot`, `new`, `top`, `rising`, `comments` |
| `timeRange` | select | `all` | `hour`, `day`, `week`, `month`, `year`, `all`; used by Top, Relevance and Most comments |
| `postedAfter` | date | empty | Keep posts created on or after this date (YYYY-MM-DD, UTC) |
| `postedBefore` | date | empty | Keep posts created on or before this date |
| `minScore` | integer | `0` | Keep posts with at least this score. 0 = everything |
| `includeNsfw` | boolean | `true` | Off = skip posts and communities marked 18+ |
| `includeComments` | boolean | `true` | Scrape comments. One extra request per post found through a subreddit, search or user page |
| `maxCommentsPerPost` | integer | `20` | Comments per post in the chosen order, replies included. 0 = all, including ones behind "load more" |
| `commentsSort` | select | `top` | `top`, `best`, `new`, `controversial`, `old`, `qa` |
| `commentsOutput` | select | `rows` | `rows` = one row per comment. `nested` = comments array inside the post row |
| `maxResultsPerSource` | integer | `50` | Posts per subreddit / search / user page, and separately comments per user page. 0 = everything Reddit exposes (about 1,000 per listing) |
| `maxItems` | integer | `0` | Stop after this many rows in total. 0 = no limit. A simple cost cap |
| `proxyConfig` | proxy | Residential | Keep the default. Reddit blocks datacenter IPs |

#### Example input

```json
{
  "startUrls": [
    "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
    "https://www.reddit.com/r/pasta/top/?t=month",
    "https://www.reddit.com/user/spez/"
  ],
  "searchQueries": ["carbonara"],
  "searchInSubreddits": ["pasta", "cooking"],
  "sort": "new",
  "postedAfter": "2025-01-01",
  "minScore": 10,
  "maxCommentsPerPost": 20,
  "maxResultsPerSource": 100
}
```

### Output

Every row has a `type` (`post`, `comment`, `community` or `user`), an `id`, a `url`, and a `source` naming the input URL or search that produced it. Fields Reddit does not provide for an item are `null`.

#### Post

```json
{
  "type": "post",
  "id": "vwi6jx",
  "url": "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
  "title": "Pasta Peperoni and Ricotta cheese (how to make Peperoni sauce very very creamy)…",
  "text": null,
  "author": "Cooking_Vito_e_Daisy",
  "authorFlair": null,
  "subreddit": "pasta",
  "subredditUrl": "https://www.reddit.com/r/pasta/",
  "subredditSubscribers": 1279942,
  "score": 304,
  "upvoteRatio": 0.99,
  "numComments": 21,
  "numCrossposts": 1,
  "awards": 0,
  "createdAt": "2022-07-11T13:15:17.000Z",
  "createdTimestamp": 1657545317,
  "editedAt": null,
  "postType": "video",
  "flair": "Homemade Dish",
  "isNsfw": false,
  "isSpoiler": false,
  "isPinned": false,
  "isLocked": false,
  "isSelfPost": false,
  "distinguished": null,
  "linkUrl": "https://v.redd.it/htn04py9vxa91",
  "linkDomain": "v.redd.it",
  "mediaUrl": "https://v.redd.it/htn04py9vxa91/DASH_1080.mp4?source=fallback",
  "thumbnailUrl": "https://b.thumbs.redditmedia.com/nJH9DcMFOlQN4PF_wzeYLLwWANta2S4z7UPDsu4WidQ.jpg",
  "galleryImages": [],
  "crosspostOf": null,
  "source": "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/"
}
```

With `commentsOutput` = `nested`, the post row also carries `comments` (an array of comment objects with the fields below, minus the post fields) and `commentsScraped`.

#### Comment

```json
{
  "type": "comment",
  "id": "ifpvu1o",
  "url": "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/ifpvu1o/",
  "text": "For homemade dishes such as lasagna, spaghetti, mac and cheese etc. please type out a basic recipe…",
  "author": "AutoModerator",
  "authorFlair": null,
  "score": 1,
  "awards": 0,
  "createdAt": "2022-07-11T13:16:07.000Z",
  "createdTimestamp": 1657545367,
  "editedAt": null,
  "depth": 0,
  "parentCommentId": null,
  "isTopLevel": true,
  "postId": "vwi6jx",
  "postTitle": "Pasta Peperoni and Ricotta cheese (how to make Peperoni sauce very very creamy)…",
  "postUrl": "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
  "subreddit": "pasta",
  "isOp": false,
  "isPinned": true,
  "distinguished": "moderator",
  "controversiality": 0,
  "source": "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/"
}
```

#### Community

```json
{
  "type": "community",
  "id": "2qoor",
  "name": "pasta",
  "url": "https://www.reddit.com/r/pasta/",
  "title": "Pasta",
  "description": "For lovers of pasta. Homemade pasta, pasta making, pasta dishes…",
  "sidebar": "As mentioned on the [Don Geronimo Show](…)…",
  "subscribers": 1279942,
  "activeUsers": null,
  "createdAt": "2008-11-17T00:16:04.000Z",
  "createdTimestamp": 1226880964,
  "isNsfw": false,
  "communityType": "public",
  "language": "en",
  "iconUrl": "https://styles.redditmedia.com/t5_2qoor/styles/communityIcon_mo430f2zes281.png?width=256&s=…",
  "bannerUrl": "https://styles.redditmedia.com/t5_2qoor/styles/mobileBannerImage_nvwb892e3ele1.png?width=4000&s=…",
  "source": "https://www.reddit.com/r/pasta/about/"
}
```

#### User

```json
{
  "type": "user",
  "id": "1w72",
  "name": "spez",
  "url": "https://www.reddit.com/user/spez/",
  "displayName": "spez",
  "bio": "Reddit CEO",
  "postKarma": 184487,
  "commentKarma": 756493,
  "totalKarma": 940980,
  "createdAt": "2005-06-06T04:00:00.000Z",
  "createdTimestamp": 1118030400,
  "isPremium": true,
  "isMod": true,
  "isEmployee": true,
  "isVerified": true,
  "hasVerifiedEmail": true,
  "isNsfw": false,
  "isSuspended": false,
  "avatarUrl": "https://styles.redditmedia.com/t5_3k30p/styles/profileIcon_uj015iwx9s7g1.png?…",
  "bannerUrl": "https://b.thumbs.redditmedia.com/KWeEpVxXOGLoloMbM0IxGt9EiKPXizpwFgcSeWqtpZM.png",
  "source": "https://www.reddit.com/user/spez/"
}
```

**Field notes**

- `text` is the post body or comment body as Reddit markdown. Link, image, video and gallery posts have no body, so `text` is `null` and the content sits in `linkUrl`, `mediaUrl` or `galleryImages`.
- `postType` is one of `text`, `link`, `image`, `video`, `gallery`, `poll`, `crosspost`. For videos `mediaUrl` is the direct MP4; for images the full-size file; for galleries the first image, with all of them in `galleryImages`.
- `score` is upvotes minus downvotes as Reddit reports it; `upvoteRatio` is the share of upvotes.
- Comments whose body is `[removed]` or `[deleted]` are skipped. `depth` is 0 for a top-level comment; `parentCommentId` is `null` for those and the parent's `id` for replies.
- A **subreddit URL** returns its posts. Add `/about` to get the community profile row instead. A **user URL** returns the profile row plus posts and comments; `/user/name/submitted` gives posts only and `/user/name/comments` comments only.
- Filters (`postedAfter`, `postedBefore`, `minScore`, `includeNsfw`) apply to posts found through subreddits, searches and user pages, and to a user's comment history. A post URL you pasted yourself is always scraped.
- `activeUsers` is only filled when Reddit publishes the number for that community.

### What counts as a result

Only rows written to the dataset are charged. Inputs that are private, banned, deleted or not a Reddit link cost nothing; they are listed in the key-value store record `FAILED_URLS` under `notFound`, `private`, `failed` and `invalid`, and summarised at the end of the log.

### Use cases

- **Market and audience research** — pull every post about a product, brand or problem from the subreddits your customers use
- **Lead generation** — find people asking for recommendations, then hand the author list to a contact-finding step
- **Content and SEO** — mine top posts and comment threads for the questions people actually ask
- **Brand and reputation monitoring** — run daily on a keyword search sorted by New with `postedAfter` set to yesterday
- **AI training and analysis** — nested comment layout gives one JSON document per thread
- **Community analytics** — track subscribers, active users and posting volume across subreddits
- **Clay, n8n, Make, Zapier** — pass URLs or keywords in, get flat rows back

### Tips

- Start with the prefilled input and the default limits, then raise `maxResultsPerSource` and `maxCommentsPerPost`.
- Want only posts? Turn `includeComments` off: one request per 100 posts instead of one per post.
- Monitoring a subreddit? Use the `/new/` URL (or `sort` = `new`) with `postedAfter`; the scraper stops paging at your date instead of walking the whole listing.
- Reddit exposes roughly 1,000 posts per listing and 250 per search. To go deeper, split by time range (`timeRange` = `month`, `year`) or by subreddit.
- Export as CSV, Excel or JSON from the dataset tab, or pull via the Apify API.

# Actor input Schema

## `startUrls` (type: `array`):

Any Reddit link, one per line:<br>• <b>Post</b> → the post and its comments<br>• <b>Subreddit</b> (reddit.com/r/pasta, r/pasta, reddit.com/r/pasta/top/?t=week) → its posts, in the order the URL says<br>• <b>User</b> (reddit.com/user/spez) → profile, posts and comments; add /submitted or /comments for just one<br>• <b>Search page</b> copied from Reddit (reddit.com/search/?q=…) → the search results<br>• <b>Community profile</b> (reddit.com/r/pasta/about) → subscribers, description, icon<br>Short links (redd.it/…), share links (/r/…/s/…), old.reddit.com and new.reddit.com all work. Duplicates are fetched once.

## `searchQueries` (type: `array`):

Search Reddit for these terms, one per line. Each term is its own search. Reddit search syntax works ("exact phrase", title:carbonara, author:name, -excluded). Leave empty to only use URLs.

## `searchType` (type: `string`):

What the keywords should find. Posts can also have their comments scraped (section 3).

## `searchInSubreddits` (type: `array`):

Restrict the keyword search to these subreddits, e.g. pasta or r/pasta, one per line. Each keyword is searched in each subreddit separately. Empty = all of Reddit. Ignored for community and user searches.

## `sort` (type: `string`):

Order for keyword searches and for subreddit URLs that don't name a sort themselves (reddit.com/r/pasta/ uses this setting; reddit.com/r/pasta/new/ always uses New). Subreddits fall back to Hot for the search-only options.

## `timeRange` (type: `string`):

Used with the Top, Relevance and Most comments sorts, the same as the time menu on Reddit.

## `postedAfter` (type: `string`):

Only keep posts created on or after this date (YYYY-MM-DD, UTC). On New-sorted listings the scraper also stops paging once it reaches older posts, so this saves requests too.

## `postedBefore` (type: `string`):

Only keep posts created on or before this date (YYYY-MM-DD, UTC).

## `minScore` (type: `integer`):

Only keep posts with at least this score (upvotes minus downvotes). 0 = keep everything.

## `includeNsfw` (type: `boolean`):

Off = posts and communities marked 18+ are skipped.

## `includeComments` (type: `boolean`):

For a post URL: that post's comments. For posts found through subreddits, searches and user pages: one extra request per post. Turn off for the fastest post-only run.

## `maxCommentsPerPost` (type: `integer`):

Comments saved per post, in the order of the sort below (replies count too). 0 = every comment, including ones behind "load more" — big threads can have thousands.

## `commentsSort` (type: `string`):

Same options as the comment sort menu on Reddit. Top gives you the highest-scored comments first.

## `commentsOutput` (type: `string`):

Separate rows are best for spreadsheets, Clay and filtering by comment. Nested keeps one row per post with its comments inside, best for JSON and AI pipelines.

## `maxResultsPerSource` (type: `integer`):

Posts per subreddit, search or user page (and separately, comments per user page). 0 = everything Reddit exposes, which is about 1,000 per listing.

## `maxItems` (type: `integer`):

Stop once this many rows (posts + comments + communities + users) are saved across all sources. 0 = no limit. Useful as a cost cap.

## `proxyConfig` (type: `object`):

Residential proxies are recommended: Reddit rate-limits and blocks datacenter IPs quickly.

## Actor input object example

```json
{
  "startUrls": [
    "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
    "https://www.reddit.com/r/pasta/"
  ],
  "searchQueries": [],
  "searchType": "posts",
  "searchInSubreddits": [],
  "sort": "relevance",
  "timeRange": "all",
  "minScore": 0,
  "includeNsfw": true,
  "includeComments": true,
  "maxCommentsPerPost": 10,
  "commentsSort": "top",
  "commentsOutput": "rows",
  "maxResultsPerSource": 5,
  "maxItems": 0,
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Scraped posts, comments, communities and users with score, author, dates, media links and the post each comment belongs to

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
        "https://www.reddit.com/r/pasta/"
    ],
    "searchQueries": [],
    "searchType": "posts",
    "searchInSubreddits": [],
    "sort": "relevance",
    "timeRange": "all",
    "minScore": 0,
    "includeNsfw": true,
    "includeComments": true,
    "maxCommentsPerPost": 10,
    "commentsSort": "top",
    "commentsOutput": "rows",
    "maxResultsPerSource": 5,
    "maxItems": 0,
    "proxyConfig": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("claygenius/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [
        "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
        "https://www.reddit.com/r/pasta/",
    ],
    "searchQueries": [],
    "searchType": "posts",
    "searchInSubreddits": [],
    "sort": "relevance",
    "timeRange": "all",
    "minScore": 0,
    "includeNsfw": True,
    "includeComments": True,
    "maxCommentsPerPost": 10,
    "commentsSort": "top",
    "commentsOutput": "rows",
    "maxResultsPerSource": 5,
    "maxItems": 0,
    "proxyConfig": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("claygenius/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://www.reddit.com/r/pasta/comments/vwi6jx/pasta_peperoni_and_ricotta_cheese_how_to_make/",
    "https://www.reddit.com/r/pasta/"
  ],
  "searchQueries": [],
  "searchType": "posts",
  "searchInSubreddits": [],
  "sort": "relevance",
  "timeRange": "all",
  "minScore": 0,
  "includeNsfw": true,
  "includeComments": true,
  "maxCommentsPerPost": 10,
  "commentsSort": "top",
  "commentsOutput": "rows",
  "maxResultsPerSource": 5,
  "maxItems": 0,
  "proxyConfig": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call claygenius/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,claygenius/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ng7Em4HEDOzGitIw5/builds/d604ah3RKLacHzpvN/openapi.json
