# Reddit Scraper — Russian-Language Subreddits (`bovi/reddit-scraper-ru`) Actor

Scrape posts, comments, and user activity from Russian-language and CIS-focused subreddits — r/pikabu, r/AskARussian, r/PlanetRussia, r/russia, and more — via the Arctic Shift archive. No OAuth, no proxy, no 1000-post cap. 25+ fields per record including score and flair.

- **URL**: https://apify.com/bovi/reddit-scraper-ru.md
- **Developed by:** [Vitalii Bondarev](https://apify.com/bovi) (community)
- **Categories:** Social media, Marketing, AI
- **Stats:** 4 total users, 3 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 reddit scraper — russian-language subreddits

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Built for brand monitors, growth researchers, and AI agent pipelines that need Russian-language and CIS-focused Reddit data at scale without OAuth limits.

**Pricing: $1.50 per 1,000 posts · $0.50 per 1,000 comments (when `includeComments=true`)**

**Reddit Scraper — Russian-Language Subreddits** lets you scrape Reddit posts, comments, and user activity from any public subreddit, with a focus on Russian-language and CIS-focused Reddit research (such as pikabu, AskARussian, and PlanetRussia) — no API key, no OAuth, no proxy required. Returns 25+ fields per record including score, upvote ratio, flair, author, and timestamps. Backed by the **Arctic Shift Reddit archive** for unlimited historical depth — no 1000-post-per-subreddit cap that live Reddit imposes. MCP-callable for AI agents. Pay only per result scraped.

### Why This Reddit Scraper Beats the Alternatives

| | This scraper | trudax/reddit-scraper-lite | practicaltools/apify-reddit-api |
|---|:---:|:---:|:---:|
| Price | **$1.50/1000** | $3.40/1000 | $2.00/1000 |
| No proxy cost to buyer | ✓ | ✗ | ✗ |
| Historical data (no 1000-post cap) | ✓ | ✗ | ✗ |
| No OAuth API dependency | ✓ | ✓ | ✗ |
| parse_confidence in every record | ✓ | ✗ | ✗ |
| 25+ fields | ✓ | ✓ | partial |
| Comments included | ✓ | partial | ✗ |

**Key advantage:** Competitors hitting live Reddit directly require residential proxy to avoid 403s — that cost passes to you. This actor uses Arctic Shift (free Reddit archive API) as its backend, so **you pay only for results, not proxy overhead**.

### Reddit Data Fields

| Field | Posts | Comments |
|---|:---:|:---:|
| id | ✓ | ✓ |
| type (`post`/`comment`) | ✓ | ✓ |
| subreddit | ✓ | ✓ |
| title | ✓ | — |
| body | ✓ | ✓ |
| author | ✓ | ✓ |
| score | ✓ | ✓ |
| upvote_ratio | ✓ | — |
| num_comments | ✓ | — |
| created_utc (ISO-8601) | ✓ | ✓ |
| permalink | ✓ | ✓ |
| url | ✓ | ✓ |
| is_self | ✓ | — |
| over_18 (NSFW) | ✓ | — |
| flair_text | ✓ | ✓ |
| domain | ✓ | — |
| subreddit_subscribers | ✓ | ✓ |
| parent_id | — | ✓ |
| depth | — | ✓ |
| is_submitter (OP?) | — | ✓ |
| parse_confidence | ✓ | ✓ |
| warnings | ✓ | ✓ |
| scraped_at | ✓ | ✓ |

#### What `parse_confidence` Means

Every Reddit record includes a score from 0.0 to 1.0:

- **1.0** — all fields parsed cleanly
- **0.9–0.95** — minor field missing (e.g. deleted author)
- **< 0.5** — critical issue (missing ID, no data returned)

`warnings` lists machine-readable codes explaining any deductions — broken scrapes are visible in your dataset, not silently hidden.

### Reddit Scraper Use Cases

- **Brand monitoring** — track keyword mentions across niche Russian-language subreddits
- **Lead generation** — find users asking questions your product solves in CIS communities
- **Sentiment analysis** — bulk-export posts and comments for NLP pipelines
- **Competitor research** — monitor product-related subreddits
- **Content strategy** — find top-performing posts by score or comment count
- **AI agent memory** — feed recent subreddit discussion into agent context

### How to Use Reddit Scraper

#### Scrape Reddit Subreddit Posts

```json
{
  "subreddits": ["pikabu", "AskARussian"],
  "sort": "new",
  "maxItems": 200,
  "includeComments": false
}
```

#### Scrape Reddit Posts + Comments Together

```json
{
  "subreddits": ["PlanetRussia"],
  "sort": "new",
  "maxItems": 100,
  "includeComments": true,
  "maxCommentsPerPost": 25
}
```

#### Scrape Reddit User Activity

```json
{
  "users": ["ivan_ivanov", "some_username"],
  "maxItems": 50
}
```

#### Scrape via Reddit URL

```json
{
  "urls": ["https://www.reddit.com/r/pikabu/"],
  "maxItems": 200
}
```

### Input Parameters

| Parameter | Type | Default | Description |
|---|---|---|---|
| `subreddits` | string\[] | — | Subreddit names (e.g. `pikabu`, `r/AskARussian`) |
| `urls` | string\[] | — | Reddit subreddit or profile URLs |
| `users` | string\[] | — | Usernames to scrape (e.g. `ivan_ivanov`) |
| `sort` | `new`/`old` | `new` | Sort order |
| `maxItems` | integer | 100 | Max posts per subreddit or user |
| `includeComments` | boolean | false | Also scrape comments |
| `maxCommentsPerPost` | integer | 50 | Max comments per post |

### Sample Output

```json
{
  "type": "post",
  "id": "1d2e3f4",
  "subreddit": "pikabu",
  "title": "Как работает этот новый закон?",
  "body": "Есть ли тут юристы, которые могут объяснить...",
  "author": "user123",
  "score": 847,
  "upvote_ratio": 0.97,
  "num_comments": 62,
  "created_utc": "2026-05-20T14:32:11+00:00",
  "permalink": "/r/pikabu/comments/1d2e3f4/как_работает_этот_новый_закон/",
  "url": "https://www.reddit.com/r/pikabu/comments/1d2e3f4/",
  "flair_text": "Вопрос",
  "subreddit_subscribers": 1200000,
  "parse_confidence": 1.0,
  "warnings": [],
  "scraped_at": "2026-06-05T09:00:00+00:00"
}
```

### Pricing — Pay Per Reddit Post or Comment

**$1.50 per 1,000 posts** · **$0.50 per 1,000 comments** (when `includeComments=true`) — PPE, no per-run fee. No proxy cost — Reddit data is fetched via Arctic Shift at no additional infrastructure charge. First $5 Apify credit covers ~3,300 post records.

### Data Source & Freshness

This actor fetches from **Arctic Shift** (`arctic-shift.photon-reddit.com`), a community-maintained Reddit archive based on historical data dumps. Data is updated continuously with an approximate 36-hour lag on engagement metrics (score, num_comments) for very recent posts. Historical data goes back years with no per-subreddit post cap.

Arctic Shift is a free service with no uptime SLA. The `parse_confidence` and `warnings` fields in every record surface any API anomalies so you can filter them downstream.

### Use with AI Agents (MCP)

This Reddit scraper is callable as a **tool by AI agents** (Claude Desktop, Cursor, VS Code, n8n, LangGraph, CrewAI, or any MCP-compatible client) via Apify's hosted Model Context Protocol server.

```json
{
  "mcpServers": {
    "apify": {
      "command": "npx",
      "args": [
        "mcp-remote",
        "https://mcp.apify.com/?tools=bovi/reddit-scraper-ru",
        "--header",
        "Authorization: Bearer <YOUR_APIFY_TOKEN>"
      ]
    }
  }
}
```

Keep `maxItems` low (e.g. 25) when calling from agents to limit token volume.

### Frequently Asked Questions

**Does this Reddit scraper need an API key?**
No. It uses Arctic Shift (a community Reddit archive), not the official Reddit API. No OAuth, no app registration.

**Why is there a 36-hour lag?**
Arctic Shift syncs from Reddit data dumps continuously. Very recent posts (< 36h) may have slightly outdated score and num_comments — all other fields are accurate.

**Can I get more than 1000 posts from a subreddit?**
Yes. Unlike live Reddit, Arctic Shift has no 1000-post cap. Use `maxItems` to control volume; the actor paginates via timestamps.

**Is residential proxy needed?**
No — this actor does not hit live Reddit endpoints. No proxy cost to you.

***

### Brand Monitoring & Incremental Scraping

Use `sinceDate` and Apify schedules to run this actor daily and get only new posts for ongoing brand-monitoring workflows. Set `includeComments=true` and a low `maxCommentsPerPost` for lightweight recurring runs that track sentiment changes over time.

***

*Not affiliated with Reddit. Data retrieved from Arctic Shift, a community-maintained public Reddit archive.*

### Integrations

Built for social-listening and research teams tracking communities, trends, and sentiment at scale — the JSON/dataset output drops into the tools you already run, no glue code:

- **n8n / Make / Zapier** — trigger a run or pipe every new dataset item into 500+ apps (Google Sheets, Airtable, Slack, HubSpot, your database) with no code: [n8n](https://docs.apify.com/platform/integrations/n8n), [Make](https://docs.apify.com/platform/integrations/make), [Zapier](https://docs.apify.com/platform/integrations/zapier).
- **Webhooks** — fire your own endpoint the moment a run finishes, to push results straight into your pipeline ([docs](https://docs.apify.com/platform/integrations/webhooks)).
- **MCP server** — expose this actor as a tool to Claude, Cursor, or any [MCP client](https://mcp.apify.com) so an AI agent can pull this data mid-conversation ([guide](https://blog.apify.com/how-to-use-mcp/)).
- **API & SDKs** — fetch the dataset as JSON, CSV, or Excel through the Apify REST API or the Python / JS SDKs.

See all [Apify integrations](https://apify.com/integrations).

### More scrapers from our toolkit

Building a data pipeline? These actors pair well with this one — each runs on your own Apify account with the same pay-per-result pricing, no subscription:

- [Social Media Finder](https://apify.com/bovi/social-media-finder)
- [Telegram Channel Scraper](https://apify.com/bovi/telegram-channel-scraper)
- [Truth Social Scraper](https://apify.com/bovi/truth-social-scraper)
- [Twitch Scraper](https://apify.com/bovi/twitch-scraper)
- [Douyin Scraper](https://apify.com/bovi/douyin-scraper)
- [Maigret Username Osint](https://apify.com/bovi/maigret-username-osint)

Chain any of them together from the **Integrations** tab (the *Run succeeded* trigger) to build a multi-step workflow — one actor's output feeds the next.

### Use it from your existing tools

#### Use with Claude Desktop / Cursor / Cline (MCP)

Load this actor as a tool in your AI assistant. Call it directly from your AI assistant via the Apify MCP server — no Store browsing needed. Paste this into your MCP client config (e.g. `claude_desktop_config.json`) and restart the client:

```json
{
  "mcpServers": {
    "apify-reddit-scraper-ru": {
      "command": "npx",
      "args": [
        "-y",
        "@apify/actors-mcp-server",
        "--tools",
        "bovi/reddit-scraper-ru"
      ],
      "env": {
        "APIFY_TOKEN": "YOUR_APIFY_TOKEN"
      }
    }
  }
}
```

Replace `YOUR_APIFY_TOKEN` with your own Apify API token (free at apify.com → Settings → Integrations). Curated to a handful of tools so the agent selects reliably.

#### Works with Clay

Run this actor as an HTTP enrichment step inside a Clay table:

- **Method:** `POST`
- **URL:** `https://api.apify.com/v2/acts/bovi~reddit-scraper-ru/run-sync-get-dataset-items?token={{apify_token}}`
- **Body (JSON):** map your Clay columns to the actor input (see the Input section above), e.g. `{"subreddits": "{{clay_column}}"}`

The run finishes synchronously and returns the dataset rows straight into your Clay table. It runs on Apify's cloud under your own token and usage. Synchronous runs must complete within 300 seconds.

### Usage statistics

This Actor creates a small, content-free summary at the end of each run. It is used only to monitor reliability and improve this Actor. A copy is saved as `USAGE_STATS` in your own Apify key-value store, so you can see the exact record created for your run.

Set `disableUsageStats` to `true` in the input to opt out. Nothing is sent then; your `USAGE_STATS` record only says that statistics were disabled.

Only these fields are recorded:

- schema version, Actor name and build number;
- UTC start and finish hour (not a precise timestamp);
- run duration, number of results and time to the first result, each as a coarse range;
- whether the result was empty, the end status, and an error type from a fixed list;
- memory setting and counts of charged events;
- names of the input fields you set, never their values;
- the selected option for input fields that offer a fixed list of choices (for example a sort order).

We do not collect input text, search terms, URLs, domains, usernames, email addresses, names, proxy credentials, tokens, scraped records, output items, raw error messages, stack traces, or hashes of any of those values. Records are kept for no longer than 13 months, used only as aggregated operational statistics, and never sold or shared.

#### Additional fields (Phase 2)

This Actor also records your Apify user ID, whether Apify marks the account as paying, the size range of list inputs, the selected country when the input offers a fixed list of countries, and one category from a fixed Actor taxonomy. We use these fields only for aggregate reliability, repeat-use and cross-Actor analysis; reports suppress any cell with fewer than five distinct users.

The same `disableUsageStats: true` input flag turns these fields off too. The user ID is removed after 13 months; we do not export, sell, share, or attempt to re-identify this data.

#### Run-outcome signals (v2)

To learn whether a run did what it was asked to do, the record also holds a few more coarse ranges and yes/no flags. None of them contains content:

- the result limit you asked for (a range, when the input has one) and what share of it was delivered;
- results delivered per input item you listed (a range);
- output quality as ranges: how fully the result fields were filled, the share of rows that look like errors, the share of duplicate rows, and how many different fields appeared. These are counted in memory while results are saved; no result content is kept;
- how the run was started (console, API, schedule, webhook, another Actor);
- how it ended: stopped by you, timed out, reached the requested limit, stopped by the charge limit, and how many times the platform moved the run;
- if this Actor reports it: how many items to process worked or failed (ranges) and one failure reason from a fixed list;
- a short code made from the names of the input fields you set, never their values.

#### Repeat-run fingerprint (v2)

When your Apify user ID is recorded (see above), the record also holds an 8-character one-way code made from your input (proxy settings left out) and this Actor's name. It only lets us see that the same account ran the same input again soon after an unsatisfying run; we never see the input itself. It is stored only in the database, never published, and reports use it in aggregate with the same five-user minimum. It is the one exception to the statement above that no hashes are collected, and `disableUsageStats: true` turns it off.

# Actor input Schema

## `subreddits` (type: `array`):

Subreddit names to scrape posts from. Each item is a subreddit name (with or without 'r/' prefix). Example: \['python', 'learnpython', 'machinelearning']. Use this for topic/community-based sampling.

## `urls` (type: `array`):

Full Reddit URLs to scrape. Supports subreddit listing pages and user profile pages. Example: \['https://www.reddit.com/r/python/', 'https://www.reddit.com/user/spez/']. Use when you have a direct link rather than a subreddit name.

## `users` (type: `array`):

Reddit usernames to scrape posts and comments from. Each item is a username (with or without 'u/' prefix). Example: \['spez', 'gallowboob']. Use for user activity research.

## `maxItems` (type: `integer`):

Maximum number of posts (or comments) to return per subreddit or user. Keep low (25–100) for agent mid-conversation sampling; use higher values (500+) for bulk research exports.

## `sort` (type: `string`):

Sort order for results. 'new' returns most recent posts first (default, best for current sentiment). 'old' returns oldest posts first.

## `searchQuery` (type: `string`):

Filter posts by title keyword. Works in combination with subreddits — e.g. set subreddits=\['python'] and searchQuery='async' to find Python posts about async. Leave empty to get all posts in the subreddit.

## `includeComments` (type: `boolean`):

Set to true to also collect comments, in addition to posts. Comments appear as separate records with type='comment'. Increases result count and run cost. Default: false (posts only).

## `maxCommentsPerPost` (type: `integer`):

When includeComments is true: maximum number of comments to collect per post. Default: 50. Lower this to reduce token volume when using results with an AI agent.

## `proxyConfiguration` (type: `object`):

Optional Apify proxy settings. This actor uses Arctic Shift (a free Reddit archive API) and does not require proxy for most runs. Enable only if your network blocks the archive endpoint.

## `disableUsageStats` (type: `boolean`):

See the README Usage statistics section for details.

## Actor input object example

```json
{
  "subreddits": [
    "pikabu",
    "AskARussian",
    "PlanetRussia",
    "russia",
    "russian"
  ],
  "maxItems": 100,
  "sort": "new",
  "includeComments": false,
  "maxCommentsPerPost": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  },
  "disableUsageStats": false
}
```

# Actor output Schema

## `results` (type: `string`):

Dataset containing Reddit Scraper records (type, subreddit, author, title, body, score, num_comments, upvote_ratio, created_utc, permalink, parse_confidence, scraped_at).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "pikabu",
        "AskARussian",
        "PlanetRussia",
        "russia",
        "russian"
    ],
    "maxItems": 100,
    "sort": "new",
    "includeComments": false,
    "maxCommentsPerPost": 50,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("bovi/reddit-scraper-ru").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": [
        "pikabu",
        "AskARussian",
        "PlanetRussia",
        "russia",
        "russian",
    ],
    "maxItems": 100,
    "sort": "new",
    "includeComments": False,
    "maxCommentsPerPost": 50,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("bovi/reddit-scraper-ru").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "pikabu",
    "AskARussian",
    "PlanetRussia",
    "russia",
    "russian"
  ],
  "maxItems": 100,
  "sort": "new",
  "includeComments": false,
  "maxCommentsPerPost": 50,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call bovi/reddit-scraper-ru --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bovi/reddit-scraper-ru"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hHjSDhm2DNHkbNox1/builds/Es78SgJBsFYSg2Kpw/openapi.json
