# Reddit Scraper — Posts, Comments, Search & Profiles (`inovaflow/reddit-scraper`) Actor

Scrape Reddit posts and comments from subreddits, keyword searches, post URLs and user profiles. One clean row per post (title, text, author, score, upvote ratio, comments, date, flair, media, NSFW) and per comment (parent, depth). Date filter. No login or API key.

- **URL**: https://apify.com/inovaflow/reddit-scraper.md
- **Developed by:** [inovaflow](https://apify.com/inovaflow) (community)
- **Categories:** Social media, Lead generation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Reddit posts and full comment threads as clean rows — from subreddits, keyword searches, post URLs and user profiles. No login, no API key, $1 per 1,000 posts.**

If you follow a market on Reddit, you know the routine: a dozen subreddits to check, a handful of searches for your product or your competitor, and the one thread everybody links to with 2,000 comments you will never scroll through. Reddit's official API now wants an app, a key and an approval, and a scraper that fails on a busy day is no help either.

We needed this for our own market research, so we built it to be boring and dependable: every run reads Reddit's own JSON (the same data the site shows), rotates to a fresh connection the moment Reddit slows one down, and turns a private, banned or missing community into a line in the status message instead of a failed run. Now we're sharing it.

### Who it's for

- **Founders and product teams** — *"What are people in r/SaaS saying about onboarding this month?"*
- **Marketers and brand teams** — *"Every new post that mentions our product or a competitor, daily."*
- **Researchers and analysts** — *"All 2,000 comments of this thread, with who replied to whom."*
- **AI and data pipelines** — *"Fresh posts from 20 communities as flat JSON, on a schedule."*
- **Lead-gen and community managers** — *"Who keeps asking for a tool like ours?"*

### What you can scrape

| Source | What to enter | You get |
|---|---|---|
| **Subreddits** | `startups`, `r/SaaS` or a subreddit URL | Posts sorted by new, hot, top, rising or controversial (top/controversial with a time window) |
| **Keyword search** | `notion alternative`, `"cold email" agency` — all of Reddit, or inside one subreddit | Matching posts, newest / most relevant / top / most commented first |
| **Post URLs** | Full links, `redd.it` short links, `/s/` share links or post ids | The post plus its comment thread — "load more" branches expanded up to your cap |
| **User profiles** | `spez`, `u/spez` or a profile URL | What the user posted, commented, or both |

Mix them freely. Any Reddit URL works in any field — a subreddit link pasted into "Post URLs" is still read as a subreddit.

**Each post row:** subreddit, title, full text, author, score, upvote ratio, comment count, date (ISO), flair, post kind (text, link, image, gallery, video…), link, direct media URLs (gallery images in order, the video file), NSFW / spoiler / locked / pinned flags and the Reddit link.

**Each comment row** (when you ask for comments): comment id, post id and title, parent id, depth (0 = reply to the post), author, text, score, date, "by OP" flag and the link. Rows come in reading order, so the thread can be rebuilt exactly.

### Options that matter

- **Maximum posts per source** — a cap for each subreddit, search and profile. Reddit itself lists up to ~1,000 posts per feed and ~250 per search; for more, split a search into narrower terms or time windows.
- **Only posts after** — a date (`2026-09-01`) or a window (`24 hours`, `7 days`). With the `new` sort, reading stops at the first older post, so you pay only for what is newer. Great for daily monitoring.
- **Also get comments for every post** — off by default; post URLs always come with their comments.
- **Maximum comments per post** and **comment order** (best, top, new, controversial, old, Q\&A) — the order decides which comments you get when a thread is bigger than your cap.
- **Include NSFW posts** — on by default, every post carries an `nsfw` flag.

### How to set it up

1. Add at least one subreddit, search, post URL or username.
2. Set the maximum posts per source (and "only posts after" if you monitor).
3. Turn on comments if you need them, and set the per-post cap.
4. Run it — or schedule it daily with "only posts after" = `24 hours`.

And that's it. Results land in the dataset with two ready views, **Posts** and **Comments**, and a run summary (`OUTPUT`) lists every source with its count and Reddit's reason for anything it refused.

### Pricing

Pay per row delivered: **$1.00 per 1,000 posts** and **$0.50 per 1,000 comments**, plus a small start fee per run. Proxy traffic is included — you never pay for bandwidth, retries or blocked requests. Duplicates (the same post found by two sources), rows removed by your filters and empty runs cost nothing.

> **Tip:** for monitoring, use the `new` sort with "Only posts after" — the run stops reading as soon as it reaches posts you already have, so a daily check of ten subreddits usually costs a few cents.

### Notes

This Actor reads only public Reddit content, the same way a logged-out visitor sees it — no account, no login, nothing private. Private and quarantined communities are skipped with a note. Use the data in line with Reddit's terms and privacy law; don't use it to harass or profile individuals.

Found a bug or missing a field? Open an issue on the **Issues** tab. Need something custom (sentiment, alerts to Slack, a CRM feed)? Ask there too.

# Actor input Schema

## `subreddits` (type: `array`):

Community names, one per line — `startups`, `r/SaaS` or a URL like `https://www.reddit.com/r/startups/top/?t=week` (a sort in the URL wins over the sort below).

## `sort` (type: `string`):

Which feed to read for subreddits and user profiles. `new` = newest first (best for monitoring and for the "posted after" filter).

## `timeFilter` (type: `string`):

Limits `top` and `controversial` feeds and keyword searches to a period. Ignored by `new`, `hot` and `rising`.

## `searchQueries` (type: `array`):

Search terms, one per line — `notion alternative`, `"cold email" agency`. Reddit's search operators work (`subreddit:SaaS`, `author:name`, `title:crm`, quotes, `OR`).

## `searchSubreddit` (type: `string`):

Leave empty to search all of Reddit. Set a community name (e.g. `SaaS`) to search only there.

## `searchSort` (type: `string`):

`new` = newest matches first, `relevance` = Reddit's best matches, `top` = most upvoted, `comments` = most discussed.

## `postUrls` (type: `array`):

Links to posts, one per line — full URLs, `redd.it` short links, `/s/` share links or bare post ids. Each gives the post row plus its comment thread, up to the comment limit below.

## `usernames` (type: `array`):

Usernames or profile URLs, one per line — `spez`, `u/spez`, `https://www.reddit.com/user/spez/`. Returns what the user posted (or commented — see below).

## `userContent` (type: `string`):

Posts the user submitted, their comments, or both.

## `maxPostsPerSource` (type: `integer`):

Cap for each subreddit, each search and each profile. Reddit itself lists at most ~1,000 posts per feed and ~250 per search. Also caps what you pay.

## `postedAfter` (type: `string`):

Skip older posts: a date (`2026-09-01`) or a window (`24 hours`, `7 days`, `2 weeks`). With the `new` sort, reading stops at the first older post — you pay only for what is newer.

## `includeComments` (type: `boolean`):

Off = post rows only (comments come only for the post URLs above). On = each post from subreddits, searches and profiles also gets its comment thread.

## `maxCommentsPerPost` (type: `integer`):

Cap on comment rows per thread. "Load more" branches are expanded until the cap is reached.

## `commentSort` (type: `string`):

Which comments come first — this decides which ones you get when a thread is bigger than the cap.

## `includeNsfw` (type: `boolean`):

Off = posts marked NSFW (18+) are skipped. On = they are delivered with `nsfw: true`.

## `maxConcurrency` (type: `integer`):

How many threads are read at once. The default suits almost every run.

## `proxyConfiguration` (type: `object`):

Reddit blocks datacenter IPs, so the Actor always uses Apify's residential proxy (you can pick its country here, or supply your own proxy URLs). Proxy traffic is included in the price.

## Actor input object example

```json
{
  "subreddits": [
    "startups",
    "SaaS",
    "Entrepreneur"
  ],
  "sort": "new",
  "timeFilter": "all",
  "searchQueries": [
    "notion alternative",
    "best crm for startups"
  ],
  "searchSubreddit": "SaaS",
  "searchSort": "new",
  "postUrls": [
    "https://www.reddit.com/r/startups/comments/1wmyn43/how_do_new_startups_actually_find_their_first/"
  ],
  "usernames": [
    "spez"
  ],
  "userContent": "posts",
  "maxPostsPerSource": 25,
  "postedAfter": "7 days",
  "includeComments": false,
  "maxCommentsPerPost": 500,
  "commentSort": "confidence",
  "includeNsfw": true,
  "maxConcurrency": 5,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

One row per post: subreddit, title, text, author, score, upvote ratio, comment count, date, flair, link, media and NSFW flag.

## `comments` (type: `string`):

One row per comment: post, parent, depth, author, text, score and date.

## `summary` (type: `string`):

Posts and comments delivered and charged, the per-source report (what Reddit refused and why), filters applied and request statistics.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "startups",
        "SaaS"
    ],
    "maxPostsPerSource": 25,
    "proxyConfiguration": {
        "useApifyProxy": true,
        "apifyProxyGroups": [
            "RESIDENTIAL"
        ],
        "apifyProxyCountry": "US"
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("inovaflow/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": [
        "startups",
        "SaaS",
    ],
    "maxPostsPerSource": 25,
    "proxyConfiguration": {
        "useApifyProxy": True,
        "apifyProxyGroups": ["RESIDENTIAL"],
        "apifyProxyCountry": "US",
    },
}

# Run the Actor and wait for it to finish
run = client.actor("inovaflow/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "startups",
    "SaaS"
  ],
  "maxPostsPerSource": 25,
  "proxyConfiguration": {
    "useApifyProxy": true,
    "apifyProxyGroups": [
      "RESIDENTIAL"
    ],
    "apifyProxyCountry": "US"
  }
}' |
apify call inovaflow/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,inovaflow/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/uKT6iah7ptyPPqnHI/builds/UbDMQh3C1zl6SgvSN/openapi.json
