# Reddit Subreddit Scraper (`jungle_synthesizer/reddit-subreddit-scraper`) Actor

Extract posts from any public subreddit — title, author, score, flair, awards, post type, and optional comment threads.

- **URL**: https://apify.com/jungle\_synthesizer/reddit-subreddit-scraper.md
- **Developed by:** [BowTiedRaccoon](https://apify.com/jungle_synthesizer) (community)
- **Categories:** Social media, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.60 / 1,000 record scrapeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Subreddit Scraper

Extract posts from any public subreddit — title, author, score, flair, awards, post type,
and optional top-level comment threads.

### Features

- Scrape one or more subreddits in a single run
- Sort by **hot**, **new**, **top** (with a time range), or **rising**
- Automatically widens across sort modes to keep filling toward your requested count
- Optional per-post comment extraction (author, body, score, timestamp)
- Clean, typed output — numbers stay numbers, booleans stay booleans, arrays stay arrays

### Usage

#### Basic

```json
{
  "subreddits": ["technology"],
  "sort": "hot",
  "maxItems": 25
}
```

#### Multiple subreddits, top posts of the week

```json
{
  "subreddits": ["technology", "science", "worldnews"],
  "sort": "top",
  "topTimeRange": "week",
  "maxItems": 50
}
```

#### With comments

```json
{
  "subreddits": ["technology"],
  "sort": "hot",
  "maxItems": 10,
  "includeComments": true,
  "maxCommentsPerPost": 10
}
```

Comment extraction loads each post's own page, so it's slower than the default listing-only
mode — expect roughly one extra page load per post.

### Input Parameters

| Parameter            | Type    | Default | Description                                                                |
|----------------------|---------|---------|----------------------------------------------------------------------------|
| `subreddits`         | Array   | —       | One or more subreddit names, without the `r/` prefix. Required.            |
| `sort`               | String  | `hot`   | Listing sort: `hot`, `new`, `top`, or `rising`.                            |
| `topTimeRange`       | String  | `day`   | Time window for `top` sort: `hour`, `day`, `week`, `month`, `year`, `all`. |
| `maxItems`           | Integer | 25      | Maximum posts to return, across all subreddits combined. Required.         |
| `includeComments`    | Boolean | `false` | Fetch each post's top-level comments.                                      |
| `maxCommentsPerPost` | Integer | 10      | Top-level comments to keep per post when `includeComments` is on.          |

### Output

Each record is one post:

```json
{
  "id": "t3_1wqnnai",
  "post_title": "Amazon's delivery drones are annoying an entire Texas neighborhood",
  "post_author": "AdSpecialist6598",
  "post_url": "https://www.reddit.com/r/technology/comments/1wqnnai/amazons_delivery_drones/",
  "post_content": "https://www.techspot.com/news/113979-amazon-delivery-drones.html",
  "post_type": "link",
  "subreddit": "technology",
  "score": 2498,
  "num_comments": 582,
  "created_utc": "2026-09-26T10:58:30.886000+0000",
  "flair": "Business",
  "is_nsfw": false,
  "is_pinned": false,
  "awards": [],
  "top_comments": [],
  "url": "https://www.reddit.com/r/technology/comments/1wqnnai/amazons_delivery_drones/",
  "scrapedAt": "2026-09-26T14:43:53.723Z"
}
```

`post_content` holds the outbound link for link/image/video posts, or the self-text for text
posts (when `includeComments` is on — self-text isn't shown on the listing view). `post_type`
is one of `text`, `link`, `image`, `video`, or `poll`.

With `includeComments` on, `top_comments` is populated:

```json
"top_comments": [
  {
    "comment_author": "TwilitRose",
    "comment_body": "Amazon Prime now offers same-day tinnitus.",
    "comment_score": 677,
    "comment_created_utc": "2026-09-26T11:06:34.453000+0000"
  }
]
```

### Notes

- A subreddit's `hot`/`new`/`rising` view surfaces a currently-active slice of its front page
  rather than a fixed page size — if your requested sort alone doesn't reach `maxItems`, the
  actor automatically pulls from the other sort modes for the same subreddit to keep filling.
  Ask for more subreddits, or a broader `topTimeRange`, for reliably larger runs.
- Reddit has not published a separate upvote/downvote breakdown since ~2016 — `score` is the
  net score shown on the page.
- Posts removed by moderators or deleted by their author may show `null` for `post_author`.

# Actor input Schema

## `sp_intended_usage` (type: `string`):

What will this data feed? E.g. lead lists, KYB checks, price tracking.

## `sp_improvement_suggestions` (type: `string`):

Provide any feedback or suggestions for improvements.

## `sp_contact` (type: `string`):

We'll personally help with your use case. No spam.

## `subreddits` (type: `array`):

One or more subreddit names to scrape (without the r/ prefix, e.g. "technology").

## `sort` (type: `string`):

Listing sort order.

## `topTimeRange` (type: `string`):

Time window for the "top" sort (ignored for other sorts).

## `maxItems` (type: `integer`):

Maximum number of posts to scrape (across all subreddits combined).

## `includeComments` (type: `boolean`):

Fetch each post's top comments. Slower — one extra page load per post.

## `maxCommentsPerPost` (type: `integer`):

Top-level comments to keep per post when Include Comments is on.

## Actor input object example

```json
{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "subreddits": [
    "technology"
  ],
  "sort": "hot",
  "topTimeRange": "day",
  "maxItems": 25,
  "includeComments": false,
  "maxCommentsPerPost": 10
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "subreddits": [
        "technology"
    ],
    "maxItems": 25
};

// Run the Actor and wait for it to finish
const run = await client.actor("jungle_synthesizer/reddit-subreddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sp_intended_usage": "Describe your intended use...",
    "sp_improvement_suggestions": "Share your suggestions here...",
    "sp_contact": "Share your email here...",
    "subreddits": ["technology"],
    "maxItems": 25,
}

# Run the Actor and wait for it to finish
run = client.actor("jungle_synthesizer/reddit-subreddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sp_intended_usage": "Describe your intended use...",
  "sp_improvement_suggestions": "Share your suggestions here...",
  "sp_contact": "Share your email here...",
  "subreddits": [
    "technology"
  ],
  "maxItems": 25
}' |
apify call jungle_synthesizer/reddit-subreddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,jungle_synthesizer/reddit-subreddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/nipdV6eMCh1sOxyOS/builds/cccBBZixJkgYGqcTU/openapi.json
