# Subreddit Scraper - Reddit Posts Past 1,000, Dates, No Login (`benthepythondev/subreddit-posts-scraper`) Actor

Returns the posts of any list of subreddits: title, text, score, comment count, author, flair and media links. Reads past the roughly 950 posts one Reddit list shows by combining its orders, cuts out a date range, and has an only-new mode for monitoring. No login, no API key.

- **URL**: https://apify.com/benthepythondev/subreddit-posts-scraper.md
- **Developed by:** [Ben](https://apify.com/benthepythondev) (community)
- **Categories:** Social media, AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.55 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## 📋 Subreddit Scraper

Returns the posts of any list of subreddits without a Reddit login or API key. One row per post: title, text, score, upvote ratio, comment count, author, flair, media links and date. It reads past the roughly 950 posts one Reddit list shows, cuts out a date range, and has an only-new mode that turns a schedule into a feed of fresh posts.

**Price:** until October 18, 2026 $3.00 per 1,000 posts on the Apify Free plan ($2.55 from Gold up); from October 19 $1.50 ($1.20 from Gold up). A subreddit that is private or does not exist costs nothing. Export to JSON, CSV or Excel, run on a schedule, call via API, or connect to Make, Zapier or n8n.

### 🔎 What is the Subreddit Scraper?

It reads subreddit feeds through Reddit's API, so no browser runs and 512 MB of memory is enough. A default run of ten posts takes about six seconds; a thousand posts take about a minute.

Reddit ends every feed early. The "new" feed of r/python stopped after 950 posts, which reached back seven and a half months. The same subreddit's other feeds (top of all time, of the year, of the month, hot, controversial, rising) hold partly other posts, so when you ask for more than one feed gives, the Actor reads them as well and saves every post once: 2,452 different posts for r/python in a test from our own server.

#### What data does it extract?

- **Post:** id, title, text (plain and Markdown), link and permalink, flair, type flags (self, video, adult, spoiler, pinned, locked)
- **Engagement:** score, upvote ratio, number of comments, awards
- **Author and subreddit**
- **Media:** preview images with sizes, thumbnail, linked domain
- **Dates:** creation time (UTC)
- **For AI use:** word count and estimated token count of the text

#### What users of Reddit scrapers ask for, and what this Actor does

The issue pages of the four most used Reddit Actors on Apify were read on October 4, 2026:

| Asked for there | Here |
|---|---|
| "Max post limit is 900", "what are the actual limits per subreddit?" | Reddit's other feeds are read as well: 2,452 different posts of one subreddit instead of 950 |
| A start and end date | `postedAfter` and `postedBefore`. The newest-first feed is read first and left when the date is reached: August of two subreddits came back as 136 posts in 19 seconds |
| Several subreddits in one run | `subreddits` is a list; each gets its own limit |
| Runs blocked by Reddit (403), slow runs | Reads Reddit's API instead of its web pages; a private or missing subreddit is named in the status message and the run still succeeds |
| Bills that grow beyond the limit, double charges | A limit is exact, a post is saved once, and a run stops at its maximum charge with what it has saved |

### ⬇️ Input

| Field | Type | Default | What it does |
|---|---|---|---|
| `subreddits` | array | | Names or links: `python`, `r/python`, `https://www.reddit.com/r/python/` |
| `sort` | string | `hot` | `hot`, `new`, `top`, `rising`, `controversial`: the first feed that is read |
| `timeFilter` | string | `all` | For top and controversial: `hour`, `day`, `week`, `month`, `year`, `all` |
| `maxPostsPerSubreddit` | integer | 25 | Posts per subreddit, up to 10,000 |
| `postedAfter` | string | | A day (YYYY-MM-DD) or a period back from now (`12 hours`, `7 days`) |
| `postedBefore` | string | | Upper end of a date range (YYYY-MM-DD) |
| `onlyNew` | boolean | false | Skip everything an earlier run with the same monitor name delivered |
| `monitorId` | string | | Name of that memory |

#### Example input

This week's top posts of three subreddits:

```json
{
  "subreddits": ["python", "webscraping", "dataengineering"],
  "sort": "top",
  "timeFilter": "week",
  "maxPostsPerSubreddit": 50
}
```

Everything posted in August:

```json
{
  "subreddits": ["webscraping"],
  "postedAfter": "2026-08-01",
  "postedBefore": "2026-08-31",
  "maxPostsPerSubreddit": 2000
}
```

As many posts of a subreddit as Reddit gives:

```json
{
  "subreddits": ["python"],
  "sort": "new",
  "maxPostsPerSubreddit": 10000
}
```

A feed of new posts for a schedule:

```json
{
  "subreddits": ["AskReddit", "webscraping"],
  "maxPostsPerSubreddit": 100,
  "onlyNew": true,
  "monitorId": "my-subreddits"
}
```

Input names of other Reddit Actors are read as well: `subredditUrls`, `startUrls` (subreddit links), `subreddit`, `maxItems`, `maxPostCount`, `maxPostsCount`, `postDateLimit` and `sinceDate`.

### ⬆️ Output

One row per post (the text is shortened here):

```json
{
  "id": "1wx5qzu",
  "title": "I built a site that logs and classifies scrapers that visit it.",
  "url": "https://www.reddit.com/r/webscraping/comments/1wx5qzu/i_built_a_site_that_logs_and_classifies_scrapers/",
  "permalink": "https://www.reddit.com/r/webscraping/comments/1wx5qzu/i_built_a_site_that_logs_and_classifies_scrapers/",
  "selftext": "I run The Crawler Zoo, a site that identifies every bot that visits and puts it on display as a live exhibit. Since this sub builds the things it …",
  "selftext_markdown": "I run The Crawler Zoo, a site that identifies every bot that visits and puts it on display as a live exhibit. Since this sub builds the things it …",
  "author": "Time_Instruction_955",
  "subreddit": "webscraping",
  "subreddit_id": "t5_318ly",
  "score": 12,
  "upvote_ratio": 1,
  "num_comments": 1,
  "is_self": true,
  "is_video": false,
  "post_hint": null,
  "domain": "self.webscraping",
  "thumbnail": null,
  "images": [],
  "created_utc": "2026-10-04T03:38:43",
  "total_awards_received": 0,
  "link_flair_text": "Bot detection 🤖",
  "over_18": false,
  "spoiler": false,
  "stickied": false,
  "locked": false,
  "word_count": 294,
  "token_count": 447,
  "scraped_at": "2026-10-04T15:08:18.526979+00:00"
}
```

The run also writes a `SUMMARY` record with the number of posts per subreddit and the subreddits that were not available.

### 💰 What a run costs

| Apify plan | Per 1,000 posts until October 18, 2026 | From October 19, 2026 |
|---|---|---|
| Free | $3.00 | $1.50 |
| Bronze | $2.85 | $1.40 |
| Silver | $2.70 | $1.30 |
| Gold, Platinum, Diamond | $2.55 | $1.20 |

Apify's standard start event ($0.00005) is the only other charge; proxies are included. Not charged: a subreddit that is private or does not exist, a post outside your date limits, and a post already delivered in only-new mode.

### ⏱️ Measured on Apify (October 4, 2026)

| Run | Posts | Time |
|---|---|---|
| Sample input (one subreddit) | 10 | 6 s |
| The first 958 posts of one subreddit | 958 | 58 s |
| August of two subreddits | 136 | 19 s |
| Only-new watch on two subreddits, second run right after the first | 1 new | 8 s |

All at 512 MB.

### 🔔 A feed of new posts

Turn on `onlyNew`, give the watch a `monitorId` and put the run on a schedule. The first run delivers the newest posts up to your limit. Later runs read each subreddit newest first and leave it at the first page without anything new, so a quiet subreddit costs nothing. Posts from before the oldest post of the first run stay outside the watch. The memory is a key-value store named `reddit-subreddit-monitor` in your own account; delete a record there to start over.

### 🤖 For AI agents

Smallest useful call:

```json
{"subreddits": ["python"], "sort": "top", "timeFilter": "week", "maxPostsPerSubreddit": 25}
```

Each row is one post with `title`, `selftext`, `score`, `num_comments`, `author`, `subreddit`, `created_utc` and `url`. `subreddits` takes several entries. Add `"postedAfter": "7 days"` for recent posts. A subreddit that is not available returns no rows; the status message and the `SUMMARY` record name it. No credentials are needed.

### 💡 Use cases

- 📡 **Community monitoring:** every new post of the subreddits your customers use.
- 🧠 **Datasets for models:** thousands of posts with text, scores and token counts.
- 📊 **Trend research:** what a community voted to the top this week, this month, this year.
- 🗂️ **Archives of a period:** all posts between two dates.

### ⚠️ Limits, stated plainly

- **Reddit's feeds are the source.** A large subreddit gives about 2,000 to 2,500 different posts over all of its feeds; a full history of every post is not available this way. For that, see the Reddit Archive Scraper.
- **Beyond the first feed the order is mixed.** The first feed keeps the order you chose; posts from the further feeds follow it.
- **A date range reaches as far as the "new" feed.** In a busy subreddit that feed covers days, in a quiet one months. Older ranges are filled from the top and controversial feeds only as far as those posts appear there.
- **Posts only.** For the comments, pass a row's `url` to the Reddit Comments Scraper.
- **Scores are the values at the time of the run.** Deleted and removed posts are not returned.

### ❓ FAQ

**Do I need a Reddit account or API key?** No.

**How many posts can I get from one subreddit?** About 950 from one feed, and about 2,000 to 2,500 when the Actor combines the feeds. Set `maxPostsPerSubreddit` to the number you want.

**How do I get only recent posts?** Set `postedAfter` to a day or a period such as `7 days`.

**Can I read private subreddits?** No. They are reported as not available and cost nothing.

**Can I run it every hour?** Yes, with `onlyNew`. A subreddit without new posts is left after one request.

**How do I call it from code?** With the Apify API or the Python and JavaScript clients; every run returns a dataset you can fetch as JSON or CSV. It also works as a tool through Apify's MCP server.

**Is it legal?** The Actor reads public Reddit posts. Posts carry usernames and can contain personal data: GDPR, CCPA and similar rules apply to how you store and use them, and Reddit's terms apply to you as well.

### 🔗 You might also like

- [Reddit Search Scraper](https://apify.com/benthepythondev/reddit-search-scraper): posts by keyword across Reddit
- [Reddit Comments Scraper](https://apify.com/benthepythondev/reddit-comments-scraper): every comment and reply of a post
- [Reddit Scraper](https://apify.com/benthepythondev/reddit-scraper): posts with nested comment trees, users and search in one Actor
- [Reddit Archive Scraper](https://apify.com/benthepythondev/reddit-archive-scraper): historical posts and comments by exact date range

**Keywords:** subreddit scraper, reddit posts scraper, scrape subreddit, reddit subreddit posts, reddit feed export, reddit posts by date, reddit top posts, subreddit monitoring, reddit data for ai, reddit posts csv, no api key reddit scraper, reddit community data

# Actor input Schema

## `subreddits` (type: `array`):

Names or links of subreddits, one per line: python, r/python or https://www.reddit.com/r/python/. A subreddit that is private or does not exist is reported and costs nothing.

## `sort` (type: `string`):

The order of the first list that is read. One list shows about 950 posts; for more, the Actor reads the subreddit's other lists as well and saves every post once. With a date limit it starts with New.

## `timeFilter` (type: `string`):

Reddit's time range for the Top and Controversial orders.

## `maxPostsPerSubreddit` (type: `integer`):

Posts to save for each subreddit. A large subreddit gives about 2,000 to 2,500 different posts over all of Reddit's lists.

## `postedAfter` (type: `string`):

A day (YYYY-MM-DD) or a period back from now (12 hours, 7 days, 2 months). Posts from before are not saved and not charged.

## `postedBefore` (type: `string`):

Upper end of a date range (YYYY-MM-DD), that day included.

## `onlyNew` (type: `boolean`):

For scheduled runs: reads newest first and skips everything an earlier run with the same monitor name delivered, so you pay only for what is new. The memory is a key-value store in your own account.

## `monitorId` (type: `string`):

Name of the memory used by the option above. Give each watch its own name; without a name the list of inputs is the name.

## Actor input object example

```json
{
  "subreddits": [
    "webscraping"
  ],
  "sort": "hot",
  "timeFilter": "all",
  "maxPostsPerSubreddit": 10,
  "onlyNew": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "webscraping"
    ],
    "maxPostsPerSubreddit": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("benthepythondev/subreddit-posts-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": ["webscraping"],
    "maxPostsPerSubreddit": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("benthepythondev/subreddit-posts-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "webscraping"
  ],
  "maxPostsPerSubreddit": 10
}' |
apify call benthepythondev/subreddit-posts-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,benthepythondev/subreddit-posts-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/G72y5thlMADvPLLph/builds/fLFtnMxpnmbGjzJ3R/openapi.json
