# Reddit Scraper (`aurenic/reddit-scraper`) Actor

Extract Reddit posts, comments, search results, and user history via the Arctic Shift archive with RSS fallback. No OAuth, no API key, no login.

- **URL**: https://apify.com/aurenic/reddit-scraper.md
- **Developed by:** [Aurenic](https://apify.com/aurenic) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.20 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper

Extract Reddit posts, comments, search results, and user history via the Arctic Shift archive with RSS fallback. No OAuth, no API key, no login.

### What does Reddit Scraper do?

Scrape Reddit in four modes:

- **Subreddit posts** — pass one or more subreddits, get posts with title, author, score, upvote ratio, comment count, flair, body, and permalink. Sort by hot, new, top, rising, or controversial, with a time filter.
- **Keyword search** — search across all of Reddit or within one subreddit.
- **Post comments** — pass a post URL or ID, get the comment tree.
- **User history** — pass a username, get their post and comment history.

Every record includes the fields that matter for real analysis — score, upvote ratio, comment count, awards — not just the bare title and URL that RSS-only scrapers return.

**Why this works in 2026.** Reddit deprecated unauthenticated `.json` access in May 2026 — every request to `reddit.com/...json` now returns HTTP 403, from datacenter *and* residential IPs. Self-service API app creation was also shut down under Reddit's Responsible Builder Policy. The actor sidesteps all of it by using two openly accessible backends:

1. **Arctic Shift** — a free, keyless, community-run Reddit archive with an HTTP API. Full post and comment data including scores and vote ratios. Covers May 2025 to present with a ~48-hour indexing lag.
2. **Reddit RSS feeds** — still open in 2026, used as a fallback for the freshest posts that haven't been indexed by Arctic Shift yet.

Neither requires OAuth, an API key, or account approval.

### Output fields

#### Posts

| Field | Description |
|---|---|
| id | Reddit post ID |
| subreddit | Subreddit name |
| title | Post title |
| author | Username |
| body | Post body / selftext (truncated to 20,000 chars) |
| score | Net upvotes |
| upvote\_ratio | Upvote ratio (0–1) |
| num\_comments | Comment count |
| created\_utc | ISO 8601 timestamp |
| permalink | Reddit comment page URL |
| url | External URL for link posts, else permalink |
| domain | Link domain |
| isSelf | Text post flag |
| over18 | NSFW flag |
| linkFlair | Post flair |
| thumbnail | Thumbnail URL |
| awards | Total awards received |
| source | `arctic-shift` or `rss-fallback` |

#### Comments

| Field | Description |
|---|---|
| id | Comment ID |
| postId | Parent post ID |
| parentId | Parent comment ID (or post ID for top-level) |
| subreddit / author | Context |
| body | Comment text (truncated to 10,000 chars) |
| score | Net upvotes |
| created\_utc | ISO 8601 timestamp |
| permalink | Direct comment URL |
| isSubmitter | Whether commenter is the OP |
| depth | Comment tree depth |

### Who is it for?

- **Brand monitoring teams** tracking product mentions across thousands of subreddits
- **Lead generation agencies** finding high-intent discussions ("what tool should I use for…")
- **Market researchers** measuring sentiment on any topic, product, or event
- **AI and RAG builders** ingesting Reddit corpora for training and grounding
- **Data journalists** pulling community discussions by keyword and time window
- **Growth teams** monitoring competitor subreddits for sentiment shifts

### Pricing

**$1.10 per 1,000 results.** No subscription.

| Results | Cost |
|---|---|
| 100 | $0.11 |
| 1,000 | $1.10 |
| 10,000 | $11.00 |

### How to use it

1. Pick a **Mode**.
2. Enter the mode's inputs: subreddits, search query, post URLs/IDs, or usernames.
3. Set **Sort** and **Time Filter** for posts and search.
4. Set **Max Items** (default 200).
5. Click **Start**.

### Output example

```json
{
  "recordType": "post",
  "id": "1dxyz42",
  "subreddit": "python",
  "title": "What's the best way to handle async retries in 2026?",
  "author": "some_user",
  "body": "I've been using tenacity but wondering if there's something better...",
  "score": 342,
  "upvote_ratio": 0.94,
  "num_comments": 87,
  "created_utc": "2026-09-15T14:32:10.000Z",
  "permalink": "https://www.reddit.com/r/python/comments/1dxyz42/whats_the_best_way/",
  "url": "https://www.reddit.com/r/python/comments/1dxyz42/whats_the_best_way/",
  "domain": "self.python",
  "isSelf": true,
  "over18": false,
  "linkFlair": "Discussion",
  "thumbnail": "",
  "awards": 3,
  "source": "arctic-shift",
  "scrapedAt": "2026-09-21T12:00:00.000Z"
}
```

### Technical details

- **No OAuth, no API key, no login.** Reddit's `.json` endpoints are dead as of May 2026 — the actor uses openly accessible backends instead.
- **Primary source: Arctic Shift** (`arctic-shift.photon-reddit.com`) — a free, keyless, community archive with full post and comment data including scores.
- **Fallback source: Reddit RSS** — still returns 200 in 2026 for subreddit feeds and search, used when Arctic Shift has not yet indexed very recent posts (~48-hour lag).
- **Self-throttled** at ~700ms between Arctic Shift requests to respect the free service.
- **No proxy required.** Arctic Shift accepts datacenter IPs.
- **Two extraction paths** — if Arctic Shift returns no data, RSS takes over transparently and records are tagged `source: rss-fallback`.

### Known limits

- **Arctic Shift has ~48-hour indexing lag.** Very recent posts may not yet be indexed. Enable **Use RSS Fallback** to get them (scores will be null on RSS records).
- **Arctic Shift is a free volunteer-run service.** No uptime SLA. The actor retries on 5xx with exponential backoff and falls back to RSS.
- **RSS records lack scores and upvote ratios.** RSS carries no vote data. Records tagged `source: rss-fallback` will have `score: null`, `upvote_ratio: null`.
- **Post body is truncated at 20,000 characters.** Comment body at 10,000.
- **Arctic Shift does not index private or quarantined subreddits.** Only public data is available.

### FAQ

**Do I need a Reddit account?** No. Neither backend requires authentication.

**Do I need a proxy?** No. Arctic Shift accepts datacenter IPs.

**Why is `score` null on some records?** RSS fallback records carry no vote data — the RSS feed doesn't include scores. Arctic Shift records always have scores.

**Why are recent posts missing?** Arctic Shift has ~48-hour indexing lag. Enable RSS fallback for fresh posts, or wait and re-run.

**Is this legal / allowed?** The actor uses two publicly accessible data sources: a community-run archive API that requires no authentication, and Reddit's own open RSS feeds. No login, no terms-of-service circumvention.

**How do I export data?** After a run, go to Storage → Export as JSON, CSV, Excel.

### Support

Open an issue on the Actor's page for bugs or feature requests.

# Actor input Schema

## `mode` (type: `string`):

What to scrape from Reddit.

## `subreddits` (type: `array`):

Subreddit names without the r/ prefix (e.g. python, MachineLearning).

## `searchQuery` (type: `string`):

Keyword to search. Used in search mode.

## `searchSubreddit` (type: `string`):

Optional. Restrict search to one subreddit.

## `postUrls` (type: `array`):

Reddit post URLs or IDs to fetch comments for. Used in comments mode.

## `usernames` (type: `array`):

Reddit usernames (without u/ prefix). Used in user mode.

## `sort` (type: `string`):

Sort order for posts.

## `timeFilter` (type: `string`):

Restrict posts to a time window (applies to top / controversial / search).

## `includeComments` (type: `boolean`):

Not implemented in v0.1 — reserved for future use.

## `maxCommentsPerPost` (type: `integer`):

Hard cap on comments per post in comments mode.

## `maxItems` (type: `integer`):

Hard cap on total records per run.

## `useRssFallback` (type: `boolean`):

If Arctic Shift returns no data (recent posts within 48h indexing lag), fall back to Reddit's open RSS feeds. Scores/upvotes will be null on RSS records.

## Actor input object example

```json
{
  "mode": "posts",
  "subreddits": [
    "python"
  ],
  "searchQuery": "",
  "searchSubreddit": "",
  "postUrls": [],
  "usernames": [],
  "sort": "top",
  "timeFilter": "week",
  "includeComments": false,
  "maxCommentsPerPost": 100,
  "maxItems": 200,
  "useRssFallback": true
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "python"
    ],
    "postUrls": [],
    "usernames": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("aurenic/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": ["python"],
    "postUrls": [],
    "usernames": [],
}

# Run the Actor and wait for it to finish
run = client.actor("aurenic/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "python"
  ],
  "postUrls": [],
  "usernames": []
}' |
apify call aurenic/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,aurenic/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/5M6P5zOZteenadUuN/builds/safWsGhWi3beBK6Te/openapi.json
