# Reddit Scraper: Posts & All Comments, No API Key (`kalilfagundes.w/reddit-scraper`) Actor

Scrape Reddit posts, complete comment threads, search results and user profiles without an API key. Expands every "load more comments", 4x more comments than first-page scrapers. Export to CSV, Excel or JSON. For data scientists, engineers and researchers.

- **URL**: https://apify.com/kalilfagundes.w/reddit-scraper.md
- **Developed by:** [Kalil Fagundes](https://apify.com/kalilfagundes.w) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper: Posts, Full Comment Threads & Search (No API Key)

Reddit Scraper extracts Reddit posts, **complete comment threads**, search results and user profiles without a Reddit account, API key or OAuth app. It expands every **"load more comments"** and **"continue this thread"** link, so you get the whole conversation, not the first ~500 comments. Export Reddit data to CSV, Excel, JSON or Google Sheets, or call it from Python, JavaScript, n8n, Make, Zapier and AI agents over MCP.

**$1.55 per 1,000 posts, $1.00 per 1,000 comments.** Built for software engineers, data scientists, ML and NLP engineers, academic researchers and market analysts.

### Reddit Scraper in 30 seconds

- **What it scrapes:** subreddit posts (hot, new, top, rising, controversial), Reddit search results, user profiles, and full comment threads from any post URL.
- **What makes it different:** it opens every hidden comment. On a 2,873-comment r/AskReddit thread it returned **1,970 comments, 4x the 499 a first-page scraper gets**.
- **No Reddit API needed:** no account, no API key, no OAuth app, no rate-limit approval.
- **Output:** one clean JSON item per post or comment, with the reply tree (`parentId`, `depth`), media URLs, and analysis flags such as deleted, removed, moderator and controversial.
- **Formats:** JSON, CSV, Excel, XML, HTML, JSONL, plus API, webhooks and integrations.
- **Speed:** about 2,000 comments a minute from a single thread.
- **Price:** $1.55 per 1,000 posts and $1.00 per 1,000 comments. No start fee, no compute or proxy charges.

### Why most Reddit scrapers miss most of the comments

When Reddit serves a thread it does not send every comment. It sends the first ~500 and hides the rest behind two kinds of links:

- **"load more comments"**: whole branches of replies, folded away
- **"continue this thread"**: deep reply chains, cut off after a few levels

A scraper that only reads the page Reddit sends gets that first slice and stops. Nothing in the output warns you: the dataset looks complete, but on any popular thread the biggest and deepest part of the discussion is missing. Sentiment scores, topic models, training sets and research findings built on that data rest on a biased sample, skewed toward the top-voted, top-level comments.

### How this scraper gets every Reddit comment

Reddit Scraper follows **every "load more comments" and every "continue this thread" link**, in batches of up to 100 comments per request, until the thread is exhausted or your limit is reached. Each comment comes with `parentId` and `depth`, so you can rebuild the exact reply tree.

**Measured on a real r/AskReddit thread in October 2026 (Reddit counter: 2,873 comments):**

| Method | Comments returned |
|---|---|
| First page only (what most scrapers return) | 499 |
| **Reddit Scraper, all hidden comments expanded** | **1,970** |

That is **4x more data from the same thread**. The remaining difference to Reddit's counter is comments that were removed or deleted, which Reddit no longer serves to anyone. Every returned comment linked back to a parent in the dataset, with no orphans and correct depths.

### Reddit Scraper vs. the Reddit API and other options

| | Reddit Scraper | Official Reddit API | First-page scrapers | Copy and paste |
|---|---|---|---|---|
| Account, API key or app approval | Not needed | Required | Usually not needed | Not needed |
| Comments behind "load more comments" | **All** | Extra calls you write yourself | Missing | Click by hand |
| Reply tree (`parentId`, `depth`) | Yes | Yes | Often missing | No |
| Export to CSV, Excel, JSON | Built in | Code it yourself | Varies | Manual |
| Cost | $1.00–$1.55 per 1,000 items | Paid for commercial use since 2023 | Varies | Your time |

### What can you do with Reddit data?

- **Brand monitoring and social listening:** track mentions of your brand, product or competitors across all of Reddit, with exact-phrase search and a `searchQuery` tag on every match. Schedule a daily run to follow sentiment over time.
- **Market research and customer pain points:** find what people complain about, ask for and compare in their own words. The long-tail replies where users describe workarounds rarely sit in the first 500 comments.
- **Sentiment and opinion analysis:** analyse the full discussion, not a sample biased toward the most upvoted takes. `controversiality` flags polarizing comments; `isDeleted` and `isRemoved` let you drop placeholders.
- **AI training data, fine-tuning and RAG:** complete conversation trees for dialogue datasets, reply-pair extraction, classifier training and retrieval corpora for LLMs. `postTitle` on every comment keeps the context.
- **Academic research:** reproducible collection of whole threads for discourse, polarization, misinformation and online community studies.
- **Content and trend research:** see which posts, formats and topics take off in a subreddit, with scores, upvote ratios, comment counts and community size.

#### Why it fits your work

- **Data scientists and analysts:** sentiment and opinion analysis on the whole thread, with stable, typed fields ready for pandas, R or BigQuery.
- **ML and NLP engineers:** full reply trees for fine-tuning and RAG, with `parentId` and `depth` for reply-pair extraction.
- **Software engineers:** clean JSON with stable field names, an API you can call from any language, and a schema that does not change between runs.
- **Academic researchers:** whole threads collected the same way every time, with the time of collection in `scrapedAt`.
- **Market and product researchers:** the real problems, workarounds and competitor complaints buried deep in threads.

### What does Reddit Scraper do?

| Mode | What you give it | What you get |
|---|---|---|
| 📋 **Subreddit Posts** | One or more subreddits | Posts sorted by hot, new, top, rising or controversial, optionally with their full comment threads |
| 🔍 **Search Reddit** | One keyword or a list of them | Matching posts from all of Reddit or one subreddit, deduplicated across queries, optionally with full comment threads |
| 👤 **User Profile** | One or more usernames | A user's posts, comments, or both |
| 💬 **Post Comments** | One or more post URLs | The post and its complete comment thread, every hidden comment included |

### What data can I extract from Reddit?

| Posts | Comments |
|---|---|
| Title, body text, author | Comment text, author |
| Score, upvote ratio, comment count | Score, controversial flag |
| Subreddit, subreddit member count | Subreddit, **title of the post it belongs to** |
| Created and edited dates | Created and edited dates |
| Permalink, external link, domain | Permalink |
| **Image, gallery and video URLs** (full size, every gallery image) | `postId`, `parentId`, `depth` to rebuild the reply tree |
| Flair, author flair, post type, thumbnail | Whether the author is the post's OP |
| NSFW, spoiler, pinned, locked, archived, OC, crosspost parent | Moderator or admin comment |
| Deleted or removed by moderators | Deleted or removed by moderators |
| **The search query that found it** (Search mode) | Time it was scraped |

Every item also carries `scrapedAt`, the time it was collected, so repeated runs of the same search can be compared over time.

**Fields built for analysis:**

- `isDeleted` / `isRemoved` flag the `[deleted]` and `[removed]` placeholders, so you can drop them before sentiment analysis or model training.
- `distinguished` marks official moderator and admin posts, such as the AutoModerator notice pinned on most threads, so they do not skew your results.
- `controversiality` is Reddit's own flag for comments with many up and down votes: a ready-made signal for polarizing opinions.
- `postTitle` on every comment keeps the context when you export comments to a spreadsheet or feed them to an LLM.
- `searchQuery` tells you which of your queries found each post when you search several phrases in one run.

### How to scrape Reddit, step by step

1. Click **Try for free** (or **Start**) and pick a **Scraping Mode**.
2. Fill in the field for that mode: subreddits, search terms, usernames or post URLs.
3. Choose whether to **Extract Comments** and set **Max Results**.
4. Click **Start** and wait for the run to finish.
5. Download the results from the **Output** tab as JSON, CSV, Excel, XML or HTML.

#### How to scrape a subreddit

Choose **Subreddit Posts**, enter one or more subreddit names (`python` or `r/python` both work), pick a sort and, for `top`, a time range such as `year`. Leave **Extract Comments** on to get each post's comments too.

#### How to get all comments from a Reddit post

Choose **Post Comments**, paste the post URL (www, old or plain reddit.com links all work), set **Max Comments per Post** to `0` and **Max Results** above the size of the thread. Every hidden comment is included.

#### How to search Reddit by keyword

Choose **Search Reddit** and enter a keyword, or several in **Search Queries List**. Wrap phrases in double quotes, such as `"looking for an alternative to"`, for exact matching. Use **Limit Search to Subreddit** to search one community only, and sort by `new`, or `top` with a time range, for recent posts.

#### How to scrape a Reddit user's posts and comments

Choose **User Profile**, enter usernames (with or without `u/`) and pick posts, comments or both.

#### How to export Reddit data to CSV, Excel or Google Sheets

After the run, open the **Output** tab and pick a format: CSV and Excel open directly in spreadsheets. For Google Sheets, use the Apify Google Sheets integration or the n8n workflow included with the source. Filter on the `type` column to keep only posts or only comments.

#### How to download Reddit images and videos

Every post has `mediaUrls`: full-size image links, every image of a gallery in order, and Reddit-hosted video files. Links to YouTube and other sites are in `externalUrl`.

#### How to monitor Reddit automatically

Save your input as a task and add an Apify **Schedule** (for example daily), sorting by `new`. Connect a webhook, Slack, email, n8n or Make to get each run's results, and deduplicate on `id` across runs.

#### How to use Reddit data in Python, pandas or an LLM

Call the Actor with the Apify Python client (example below), load the items into a pandas DataFrame, or connect the Actor to Claude, ChatGPT-compatible clients, Cursor or VS Code through Apify's MCP server so an AI agent can search Reddit on its own.

### How much does it cost to scrape Reddit?

| Event | Price |
|---|---|
| Post saved to the dataset | **$1.55 per 1,000** ($0.00155 each) |
| Comment saved to the dataset | **$1.00 per 1,000** ($0.001 each) |

Nothing else is billed: no start fee, no charge for compute time or residential proxies. Comments are priced lower on purpose: full threads are what this scraper is for, and they are where the volume is.

| You scrape | You pay |
|---|---|
| 100 posts, no comments | $0.16 |
| 10 posts with 1,990 comments | $2.01 |
| One 2,000-comment thread in full | $2.00 |
| 1,000 posts with 9,000 comments | $10.55 |

Set a **maximum cost per run** in the run options and the scraper stops at that amount: it never scrapes results you would not be charged for.

To spend less:

- Uncheck **Extract Comments** when you only need posts.
- Keep **Max Results** close to what you need. Posts and comments both count toward it.
- Lower **Max Comments per Post** if you only need the top of each thread.

### Input

Only `mode` plus the field for that mode is required. Everything else has sensible defaults.

| Field | Default | What it does |
|---|---|---|
| `includeComments` | `true` | Also fetch each post's comments in Subreddit and Search mode |
| `mode` | `subreddit_posts` | `subreddit_posts`, `search`, `user_profile` or `post_comments` |
| `subreddits` | | Subreddit names, with or without `r/` |
| `sort` | `hot` | `hot`, `new`, `top`, `rising`, `controversial` |
| `timeFilter` | `week` | `hour`, `day`, `week`, `month`, `year`, `all`. Only used when the sort is `top` |
| `searchQuery` | | One search term. Wrap a phrase in double quotes for an exact match |
| `searchQueriesList` | | Several search terms in one run. Overrides `searchQuery` |
| `searchSubreddit` | | Limit the search to one subreddit |
| `searchSort` | `relevance` | `relevance`, `hot`, `top`, `new`, `comments` |
| `usernames` | | Usernames, with or without `u/` |
| `userContentType` | `overview` | `overview` (posts and comments), `submitted`, `comments` |
| `postUrls` | | Full post URLs (www, old or plain reddit.com) |
| `maxCommentsPerPost` | `100` | Comments per post. `0` means no limit |
| `maxResults` | `100` | Total items for the whole run, posts and comments together (1 to 10,000) |
| `includeNsfw` | `false` | Include adult posts and their comments |
| `proxyConfiguration` | Residential | Apify Proxy settings. Residential IPs are strongly recommended |

#### Input examples

**Top posts of the year from a subreddit, with comments**

```json
{
    "mode": "subreddit_posts",
    "subreddits": ["enem", "vestibular"],
    "sort": "top",
    "timeFilter": "year",
    "includeComments": true,
    "maxCommentsPerPost": 0,
    "maxResults": 2000
}
```

**Brand monitoring with several exact phrases**

```json
{
    "mode": "search",
    "searchQueriesList": ["\"notion alternative\"", "\"switched from notion\""],
    "searchSort": "new",
    "includeComments": false,
    "maxResults": 300
}
```

**Every comment of a few threads**

```json
{
    "mode": "post_comments",
    "postUrls": ["https://www.reddit.com/r/AskReddit/comments/1wrj0s2/"],
    "maxCommentsPerPost": 0,
    "maxResults": 5000
}
```

**A user's recent comments**

```json
{
    "mode": "user_profile",
    "usernames": ["spez"],
    "userContentType": "comments",
    "maxResults": 200
}
```

### Output

Each item is either a `post` or a `comment`, so you can filter on `type`. Comments follow their post in the dataset.

#### Post example

`searchQuery` only appears on posts found in Search mode.

```json
{
    "type": "post",
    "id": "1w6o67w",
    "subreddit": "enem",
    "title": "Um pessoal desse sub nos últimos dias",
    "author": "example_user",
    "selftext": "",
    "url": "https://www.reddit.com/r/enem/comments/1w6o67w/um_pessoal_desse_sub_nos_ultimos_dias/",
    "externalUrl": "https://i.redd.it/5znjedyv6enh1.jpeg",
    "mediaUrls": [
        "https://i.redd.it/5znjedyv6enh1.jpeg"
    ],
    "score": 2071,
    "upvoteRatio": 0.99,
    "numComments": 37,
    "subredditSubscribers": 412000,
    "created": "2026-09-03T23:46:38+00:00",
    "editedAt": "",
    "isNSFW": false,
    "isSpoiler": false,
    "isPinned": false,
    "isLocked": false,
    "isArchived": false,
    "isDeleted": false,
    "isRemoved": false,
    "distinguished": "",
    "flair": "Humor",
    "domain": "i.redd.it",
    "isVideo": false,
    "thumbnail": "",
    "postHint": "image",
    "isOriginalContent": false,
    "authorFlair": "",
    "crosspostParent": "",
    "mediaOnly": false,
    "isGallery": false,
    "scrapedAt": "2026-10-02T19:42:58+00:00",
    "searchQuery": "\"enem 2026\""
}
```

#### Comment example

```json
{
    "type": "comment",
    "id": "p7ok2xq",
    "postId": "1w6o67w",
    "postTitle": "Um pessoal desse sub nos últimos dias",
    "parentId": "p7oj9bd",
    "subreddit": "enem",
    "author": "another_user",
    "body": "Literalmente eu na semana da prova.",
    "score": 57,
    "controversiality": 0,
    "created": "2026-09-04T00:12:51+00:00",
    "editedAt": "",
    "depth": 1,
    "isSubmitter": false,
    "isDeleted": false,
    "isRemoved": false,
    "distinguished": "",
    "url": "https://www.reddit.com/r/enem/comments/1w6o67w/comment/p7ok2xq/",
    "scrapedAt": "2026-10-02T19:42:58+00:00"
}
```

`parentId` is the comment being replied to, or the post id for a top-level comment (`depth` 0). Comments Reddit shows on first load come in thread order; comments loaded from "load more comments" links come after them, so use `parentId` to rebuild the exact tree.

### Limits and tips

- **About 1,000 posts per listing.** Reddit stops paginating any listing (a subreddit sort, a search, a user profile) at roughly 1,000 items. That is a Reddit limit, not one of this Actor. To go further, run the same subreddit with several sorts (`new`, `top` with different `timeFilter` values, `controversial`) or several search terms, then deduplicate on `id`.
- **Comments are not affected by that limit.** A single thread can return thousands of comments.
- **Comment counts differ from Reddit's.** `numComments` includes removed and deleted comments, which Reddit no longer serves, so a thread usually returns fewer comments than that number.
- **One long thread cannot use up a run.** In Subreddit and Search mode, each post may use at most a fifth of `maxResults` for its comments, so a run always reaches several posts. The log says when this lowers `maxCommentsPerPost`. Post Comments mode has no such cap.
- **`maxResults` is shared across queries** in `searchQueriesList` and is used in order, so later queries get nothing once it runs out. Allow at least 25 per query.
- **Quote your search phrases.** Unquoted, Reddit matches loose words and returns mostly unrelated posts.
- **`relevance` favours old, popular posts.** For recent results use `new`, or `top` with a `timeFilter`.
- **Only public content.** Private, quarantined-by-invite and banned subreddits, and deleted accounts, return nothing.

### FAQ

#### Is it legal to scrape Reddit?

Reddit Scraper only reads public pages that any logged-out visitor can see. Scraping public data is generally allowed, but you are responsible for how you use it: follow Reddit's terms, data protection laws such as GDPR and LGPD, and do not collect personal data without a legitimate reason. If unsure, ask a lawyer.

#### Can I scrape Reddit without an API key?

Yes. Reddit Scraper needs no Reddit account, API key or OAuth app. It reads Reddit's public pages through a real headless browser.

#### What is the best alternative to the Reddit API?

For collecting public posts and comments, a scraper like this one avoids app approval, rate-limit negotiations and Reddit's paid commercial API access, which started in 2023. You pay only per result, and you get full comment threads that would otherwise take many extra API calls.

#### Is there an alternative to Pushshift?

Pushshift is no longer publicly available. Reddit Scraper covers the common Pushshift use cases for current data: subreddit posts, keyword search, user histories and full comment threads. It cannot fetch posts beyond Reddit's ~1,000-item listing limit, so it is not a full historical archive.

#### How do I get all comments from a Reddit thread?

Use **Post Comments** mode with `maxCommentsPerPost` set to `0` and a `maxResults` larger than the thread. The Actor follows every "load more comments" and "continue this thread" link, so you get the whole thread, not the ~500 comments Reddit shows on first load.

#### Why do other Reddit scrapers return fewer comments for the same thread?

Reddit only sends the first ~500 comments of a thread and hides the rest behind "load more comments" and "continue this thread" links. Scrapers that do not open those links stop there without telling you. This one opens all of them.

#### Why is the comment count lower than the number Reddit shows?

Reddit's counter includes removed and deleted comments, which it no longer serves to anyone. Every comment Reddit still serves is returned.

#### How many Reddit posts can I scrape?

Up to 10,000 results per run, posts and comments together. Each Reddit listing (one subreddit sort, one search, one profile) stops at about 1,000 posts, a limit set by Reddit. Combine sorts, time ranges and search terms to collect more, then deduplicate on `id`.

#### Can I scrape Reddit posts from a specific date range?

Not by exact dates: Reddit does not offer a date filter. Use `top` with a time range (`hour`, `day`, `week`, `month`, `year`, `all`), or `new` and filter on the `created` field.

#### How fast is it?

About 2,000 comments a minute from a single thread, and up to 25 posts per request when comments are off. Requests are paced to avoid blocks.

#### Can I scrape private or NSFW subreddits?

Private subreddits: no, only public content. NSFW posts are skipped by default; turn on **Include NSFW Content** to include them and their comments.

#### Does it download images and videos?

It returns the links: `mediaUrls` has full-size images, every image of a gallery and Reddit-hosted videos. Download them with any tool or script.

#### Can AI agents like Claude or ChatGPT use this scraper?

Yes. Add it to any MCP client (Claude, Cursor, VS Code and others) through Apify's MCP server, as shown below, and the agent can search Reddit and read full threads by itself.

#### Do I need proxies?

Yes. Reddit blocks most datacenter IPs. Keep the default **Residential** Apify Proxy group; its cost is already included in the price.

#### What happens with a subreddit, user or URL that doesn't exist?

The Actor logs a warning, skips it and carries on with the rest of the run. A run that finds nothing at all ends with a status message explaining the likely causes.

#### Can I schedule runs or get results through the API?

Yes. Use Apify **Schedules** for recurring runs, and the API, webhooks or integrations (n8n, Make, Zapier, Google Sheets) to collect the dataset.

### Use the Reddit Scraper API from code

**JavaScript**

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: '<APIFY_TOKEN>' });
const run = await client.actor('kalilfagundes.w/reddit-scraper').call({
    mode: 'subreddit_posts',
    subreddits: ['enem'],
    sort: 'top',
    timeFilter: 'month',
    maxResults: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.length);
```

**Python**

```python
from apify_client import ApifyClient

client = ApifyClient("<APIFY_TOKEN>")
run = client.actor("kalilfagundes.w/reddit-scraper").call(run_input={
    "mode": "search",
    "searchQuery": "\"redação nota 1000\"",
    "searchSort": "new",
    "maxResults": 200,
})
items = client.dataset(run["defaultDatasetId"]).list_items().items
print(len(items))
```

**MCP (Claude, Cursor, VS Code and other AI clients)**

```json
{
    "mcpServers": {
        "reddit-scraper": {
            "url": "https://mcp.apify.com?tools=kalilfagundes.w/reddit-scraper",
            "headers": { "Authorization": "Bearer <APIFY_TOKEN>" }
        }
    }
}
```

An n8n workflow that searches Reddit for complaint phrases and logs them to Google Sheets is included in the source under `examples/n8n`.

### Changelog

- **1.5**: richer output. Posts get `mediaUrls` (full-size images, every gallery image, Reddit videos), `subredditSubscribers`, `isLocked`, `isArchived` and, in Search mode, `searchQuery`. Comments get `postTitle` and `controversiality`. Both get `isDeleted`, `isRemoved`, `distinguished`, `editedAt` (a date, replacing the old `edited` number) and `scrapedAt`. Removed `awards` and `isPromoted`, which Reddit no longer fills (awards were discontinued in 2023).
- **1.4**: follows "load more comments" and "continue this thread" links, so threads are complete instead of stopping at ~500 comments. New `parentId` field on comments. A run migrated to another server mid-run no longer repeats items. Posts and comments are priced separately, and runs stop at the user's maximum cost.
- **1.3**: "Extract Comments" moved to the top of the form and enabled by default.

### Feedback

Found a bug or want a feature? Open an issue on the **Issues** tab.

# Actor input Schema

## `includeComments` (type: `boolean`):

Also fetch the comments of each post in Subreddit Posts and Search mode. Uncheck to collect only the posts: one request per 25 posts instead of one extra request per post, so runs are much faster and use far less residential proxy traffic. Comments count toward maxResults, so raise maxResults when this is on; no single post may use more than a fifth of it. Post Comments mode always extracts comments, and User Profile mode uses User Content Type instead.

## `mode` (type: `string`):

What type of Reddit data to scrape.

## `subreddits` (type: `array`):

List of subreddit names to scrape (without r/ prefix). Example: python, webdev, machinelearning

## `sort` (type: `string`):

How to sort subreddit posts.

## `timeFilter` (type: `string`):

Time range filter. Applies ONLY when the sort is Top: that means sort='top' in Subreddit Posts mode, or searchSort='top' in Search mode. With any other sort value this field is accepted but has no effect.

## `searchQuery` (type: `string`):

Keyword or phrase to search across Reddit. IMPORTANT: Reddit matches loose words by default, not phrases. To match an exact phrase, wrap it in double quotes: "looking for an alternative to". (When sending input as JSON, the quotes must be escaped.) An unquoted multi-word query returns largely unrelated posts. Use searchQueriesList for multiple queries in one run.

## `searchQueriesList` (type: `array`):

Run multiple search queries in a single job. Results are merged and deduplicated by post ID, and this field overrides Search Query when provided. IMPORTANT: wrap each phrase in double quotes for exact-phrase matching, for example "looking for an alternative to" as one entry and "is there a tool that" as another. (When sending input as JSON, the inner quotes must be escaped.) Unquoted phrases return loosely matched, mostly irrelevant posts. Note that maxResults is a shared budget across all queries and is consumed in order, so later queries return nothing once it is used up. Ideal for brand monitoring, competitor research, and topic mapping.

## `searchSubreddit` (type: `string`):

Optional: restrict search to a specific subreddit (without r/ prefix). Leave empty to search all of Reddit.

## `searchSort` (type: `string`):

How to sort search results. Relevance, the default, favours highly upvoted posts, which are often several years old. For recent results use Top together with timeFilter, or New for the newest posts regardless of engagement.

## `usernames` (type: `array`):

List of Reddit usernames to scrape (without u/ prefix).

## `userContentType` (type: `string`):

What type of content to scrape from user profiles.

## `postUrls` (type: `array`):

Full Reddit post URLs to extract comments from.

## `maxCommentsPerPost` (type: `integer`):

Maximum number of comments to extract per post. Set to 0 for no per-post limit. In Subreddit Posts and Search mode, comments count toward maxResults and a single post may use at most a fifth of it, so one long thread cannot fill the run before the second post is reached. The run log says so when that lowers this number. Post Comments mode is not capped this way: there the thread is the target.

## `maxResults` (type: `integer`):

Maximum number of results for the whole run. This is a shared total, not a per-query limit: when using searchQueriesList the budget is consumed one query at a time, so allow at least 25 per query you expect results from. Up to 10,000.

## `includeNsfw` (type: `boolean`):

When enabled, adult/NSFW content is included in results. Default is off, which filters out NSFW posts and the comments on them, so a filtered post never leaves its thread behind.

## `proxyConfiguration` (type: `object`):

Proxy settings. Residential proxies are strongly recommended, Reddit blocks most datacenter IPs.

## Actor input object example

```json
{
  "includeComments": true,
  "mode": "subreddit_posts",
  "subreddits": [
    "python"
  ],
  "sort": "hot",
  "timeFilter": "week",
  "searchQuery": "machine learning",
  "searchQueriesList": [],
  "searchSort": "relevance",
  "usernames": [
    "spez"
  ],
  "userContentType": "overview",
  "postUrls": [
    "https://www.reddit.com/r/Python/comments/1r19hu1/after_25_years_using_orms_i_switched_to_raw/"
  ],
  "maxCommentsPerPost": 100,
  "maxResults": 100,
  "includeNsfw": false,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset containing all scraped posts and comments

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "python"
    ],
    "searchQuery": "machine learning",
    "searchQueriesList": [],
    "usernames": [
        "spez"
    ],
    "postUrls": [
        "https://www.reddit.com/r/Python/comments/1r19hu1/after_25_years_using_orms_i_switched_to_raw/"
    ],
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("kalilfagundes.w/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": ["python"],
    "searchQuery": "machine learning",
    "searchQueriesList": [],
    "usernames": ["spez"],
    "postUrls": ["https://www.reddit.com/r/Python/comments/1r19hu1/after_25_years_using_orms_i_switched_to_raw/"],
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("kalilfagundes.w/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "python"
  ],
  "searchQuery": "machine learning",
  "searchQueriesList": [],
  "usernames": [
    "spez"
  ],
  "postUrls": [
    "https://www.reddit.com/r/Python/comments/1r19hu1/after_25_years_using_orms_i_switched_to_raw/"
  ],
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call kalilfagundes.w/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,kalilfagundes.w/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/RgAUgINRBfzDFeuv6/builds/DDOwZWflc5hrSKhHA/openapi.json
