# Reddit Scraper: Posts, Comments, Search & Users (`ecclesiasteslabs/reddit-scraper`) Actor

Scrape Reddit posts and comments from subreddits, keyword searches, user profiles and post links. Clean, flat JSON/CSV with full text. No login, no API key. Lightweight and reliable, at $1 per 1,000 results.

- **URL**: https://apify.com/ecclesiasteslabs/reddit-scraper.md
- **Developed by:** [ecclesiasteslabs](https://apify.com/ecclesiasteslabs) (community)
- **Categories:** Social media, Lead generation, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 2 bookmarks
- **User rating**: 5.00 out of 5 stars

## Pricing

from $0.70 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper: Posts, Comments, Search & Users

Get Reddit posts and comments as clean, flat data. Give it subreddits, search keywords, usernames or post links, and get back one row per post or comment, ready for spreadsheets, dashboards, or AI pipelines.

- **No login, no Reddit API key**
- **Four sources in one run:** subreddits, keyword search (all of Reddit or inside your subreddits), user profiles, single post links
- **Inputs combine the way you'd expect:** subreddit `medicine` + search `AI` = AI posts in r/medicine; user + search = that user's posts about the topic
- **Full text** of every post and comment, converted from HTML to plain text
- **One flat shape** for posts and comments, with a `type` column to filter on
- **Safe defaults:** comments are off and the cap is 100 results, so a run never grows into a surprise bill
- **Lightweight:** no browser, so runs finish in seconds instead of timing out

#### What can you use it for?

- **Brand and competitor monitoring:** search for your brand or a competitor's name and get every new mention.
- **Lead generation:** find people asking for a tool or service like yours (`"looking for" crm`, `"any alternative to" notion`).
- **Market research:** read what a niche community talks about, in bulk.
- **AI and RAG datasets:** clean text with permalinks and dates, ready to chunk and embed.
- **Sentiment analysis:** feed post and comment text into your own model.

#### How to use it

1. Fill in at least one of **Subreddits**, **Search queries**, **Users** or **Post URLs**.
2. Pick a **Sort** (hot, new, top...) and a **Time filter** (works with every sort).
3. Set **Maximum results**. It's shared fairly between your sources, and you only pay for results actually saved.
4. Turn on **Include comments** if you want the discussion too.
5. Click **Start**, then download as JSON, CSV, Excel or HTML, or read it through the API.

#### Input example

```json
{
  "searchQueries": ["\"notion alternative\""],
  "sort": "new",
  "timeFilter": "week",
  "maxItems": 200,
  "maxPostsPerSource": 50,
  "includeComments": true,
  "maxCommentsPerPost": 10
}
```

More examples:

- Latest 500 posts in two subreddits: `{"subreddits": ["startups", "SaaS"], "sort": "new", "maxItems": 500, "maxPostsPerSource": 250}`
- One thread with up to 200 comments: `{"postUrls": ["https://www.reddit.com/r/Python/comments/1wwv97o/"], "includeComments": true, "maxCommentsPerPost": 200, "maxItems": 201}`
- Top posts of the month mentioning a keyword in one subreddit: `{"subreddits": ["SaaS"], "searchQueries": ["pricing"], "sort": "top", "timeFilter": "month"}`

#### Output example

```json
{
  "type": "post",
  "id": "1wwv97o",
  "url": "https://www.reddit.com/r/Python/comments/1wwv97o/my_python_magic_is_gone/",
  "subreddit": "Python",
  "author": "RedYad2",
  "title": "My Python magic is gone",
  "text": "I've been coding for a long time, and I've been developing with Python for years...",
  "linkUrl": null,
  "thumbnailUrl": null,
  "createdAt": "2026-10-03T19:08:43Z",
  "postId": "1wwv97o",
  "postTitle": "My Python magic is gone",
  "source": "r/python",
  "scrapedAt": "2026-10-04T13:46:12Z"
}
```

Comments have the same columns, with `type: "comment"`, the comment's own `url`, and `postId`/`postTitle` pointing to the post they belong to.

#### Output fields

| Field | Description |
|---|---|
| `type` | `post` or `comment` |
| `id` | Reddit id |
| `url` | Permanent link |
| `subreddit` | Subreddit name |
| `author` | Username or `[deleted]` |
| `title` | Post title (posts only) |
| `text` | Plain text of the post or comment |
| `linkUrl` | External link of link posts |
| `thumbnailUrl` | Preview image, when available |
| `createdAt` | Posted time, ISO 8601 UTC |
| `postId`, `postTitle` | The post (for comments: the parent post) |
| `source` | Which of your inputs produced the row |
| `scrapedAt` | When it was collected |

#### What this scraper does *not* return

To stay fast and reliable, this actor reads Reddit's public feeds. Those feeds don't include **vote scores, upvote ratio, comment counts, flair, or comment reply depth**. If you need to rank posts by votes, this isn't the right tool. If you need the text of what people are saying, it is, and at a fraction of the price.

#### How much does it cost?

**$1.00 per 1,000 results** (posts or comments). No start fee, no monthly rent. Duplicates are never charged. Set a maximum cost per run in Apify and the scraper stops when it hits it.

| Results | Price |
|---|---|
| 100 | $0.10 |
| 1,000 | $1.00 |
| 10,000 | $10.00 |

Paid Apify plans get a discount: **Bronze $0.90, Silver $0.80, Gold and above $0.70 per 1,000 results.**

Apify's free plan includes $5 of monthly credit, enough for about 5,000 results.

#### Run it from your code

Get your API token in Apify Console → Settings → API & Integrations.

**Python** (`pip install apify-client`)

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("ecclesiasteslabs/reddit-scraper").call(run_input={"subreddits": ["medicine"], "searchQueries": ["artificial intelligence"], "timeFilter": "month", "maxItems": 100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

**JavaScript / Node.js** (`npm install apify-client`)

```javascript
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('ecclesiasteslabs/reddit-scraper').call({"subreddits": ["medicine"], "searchQueries": ["artificial intelligence"], "timeFilter": "month", "maxItems": 100});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items);
```

**HTTP (one call, returns the results)**

```bash
curl -X POST "https://api.apify.com/v2/acts/ecclesiasteslabs~reddit-scraper/run-sync-get-dataset-items?token=YOUR_APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"subreddits": ["medicine"], "searchQueries": ["artificial intelligence"], "timeFilter": "month", "maxItems": 100}'
```

#### Integrations

- **AI agents (MCP):** connect Claude, ChatGPT, Cursor or any MCP client to [Apify's MCP server](https://mcp.apify.com) and the agent can call this actor directly. Every input has a plain-language description and safe defaults, so `{"searchQueries": ["your topic"]}` is enough.
- **No-code:** use the Apify modules in **Make**, **Zapier** or **n8n** to send results to Google Sheets, Airtable, Slack or a CRM.
- **Schedules:** run it daily or hourly in Apify Console → Schedules, and get only new results with a short time filter.

#### FAQ

**Is scraping Reddit legal?** This actor only collects publicly visible posts and comments, with no login. You are responsible for how you use the data, including Reddit's terms and privacy laws like GDPR. Don't use personal data for spam.

**Why are there fewer search results than I expected?** Reddit's own search returns a limited number of results per query. Split broad topics into several narrower queries, or add a time filter.

**How many posts can I get from a subreddit?** Reddit serves about 1,000 posts per listing (per sort). Use several sorts or time filters to collect more.

**Something broke?** Open an issue on the Issues tab. We run a health check every day and fix breakages fast.

# Changelog

This Actor's version history is a separate document: https://apify.com/ecclesiasteslabs/reddit-scraper/changelog.md

# Actor input Schema

## `subreddits` (type: `array`):

Subreddits to use. Accepts `python`, `r/python`, or a full link. **Alone:** returns the subreddit's posts. **With search queries:** the queries are searched only inside these subreddits (e.g. subreddit `medicine` + query `AI` = AI posts in r/medicine).

## `searchQueries` (type: `array`):

Keywords to search for, e.g. `chatgpt alternatives` or `"my brand"` (quotes = exact phrase). Each query is a separate search. Where they apply: inside the subreddits above if you list any; otherwise to the users below (only their posts containing a query are kept); otherwise all of Reddit.

## `users` (type: `array`):

Reddit usernames whose submitted posts to scrape. Accepts `spez`, `u/spez`, or a profile link. If you also enter search queries, only this user's posts containing a query are kept.

## `postUrls` (type: `array`):

Links to individual Reddit posts, e.g. `https://www.reddit.com/r/python/comments/abc123/title/` or `https://redd.it/abc123`. Returns the post, plus its comments if 'Include comments' is on. Search queries and the time filter don't apply to post links.

## `sort` (type: `string`):

Order of posts. Subreddits: hot, new, top, rising, controversial. Search: relevance, hot, top, new, comments. Users: hot, new, top, controversial. If a source doesn't support the chosen sort, the closest one is used and the log says so.

## `timeFilter` (type: `string`):

Only return posts from this time window. Works with every sort: for search, top and controversial Reddit applies it; for hot, new and rising the scraper filters by post date. 'All time' = no filter.

## `maxItems` (type: `integer`):

Total cap on results (posts + comments) for the whole run. It's shared fairly between all your sources (subreddits, searches, users, post links), and anything a source doesn't use goes to the next one. You only pay for results actually saved.

## `maxPostsPerSource` (type: `integer`):

Maximum posts from each single source (one subreddit, one search, one user). Reddit's feeds usually stop at about 1,000 posts per listing, and search often returns fewer.

## `includeComments` (type: `boolean`):

If on, also saves the comments of every post found (highest-voted threads first). Each comment counts as one result. Comments come in thread order; Reddit's feed does not include vote scores or reply depth.

## `maxCommentsPerPost` (type: `integer`):

Only used when 'Include comments' is on. Maximum comments saved for each post (top-voted first).

## `proxyConfiguration` (type: `object`):

Residential proxy is the default and recommended: Reddit rate-limits shared datacenter IPs quickly. Leave as is unless you know you need something else.

## Actor input object example

```json
{
  "subreddits": [
    "technology"
  ],
  "sort": "hot",
  "timeFilter": "all",
  "maxItems": 50,
  "maxPostsPerSource": 25,
  "includeComments": false,
  "maxCommentsPerPost": 20,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "technology"
    ],
    "maxItems": 50,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ]
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("ecclesiasteslabs/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": ["technology"],
    "maxItems": 50,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
    },
}

# Run the Actor and wait for it to finish
run = client.actor("ecclesiasteslabs/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "technology"
  ],
  "maxItems": 50,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ]
  }
}' |
apify call ecclesiasteslabs/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ecclesiasteslabs/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/4kCIbAy1lbUVwtsMo/builds/qdhnjisYxqDaPCIXZ/openapi.json
