# Reddit Scraper (`mlg14/reddit-scraper`) Actor

Collect public Reddit posts, comments, community details, and user activity from archived JSON records. Search within communities and export structured data.

- **URL**: https://apify.com/mlg14/reddit-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Social media, Automation
- **Stats:** 3 total users, 2 monthly users, 66.7% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Scraper

Scrape public Reddit posts, comments, community details, and user activity into structured records. Export Reddit data to JSON, CSV, or a spreadsheet through the dataset, with community search and direct post URLs as a practical Reddit API alternative for archived public content.

This actor reads archived public Reddit records rather than authenticated account data. The archive may lag live pages, and a post that has since been edited or removed can still appear in its earlier captured state. Each row includes a `dataType` so posts, comments, communities, and users remain easy to separate after export.

### What data can you extract from Reddit?

One dataset can contain four kinds of row. A community URL adds community metadata followed by posts. A direct post URL adds the post and its available comments. A user URL adds a profile summary and public posts; a user comments URL adds the profile summary and comments. A search term can find titles and comment text within the community you specify.

| Field | Description | Example or availability |
| --- | --- | --- |
| `dataType` | Row kind | `post`, `comment`, `community`, `user` |
| `id` | Full item identifier | `t3_1wq2vmx` |
| `parsedId` | Post or comment identifier without prefix | `1wq2vmx` |
| `url` | Public discussion, community, or profile URL | `https://www.reddit.com/r/Python/` |
| `username` | Public account name for a post, comment, or user | Present when captured |
| `userId` | Public account identifier | Present when captured |
| `title` | Post headline or community title | `Pre-parsing change detection` |
| `communityName` | Community with the `r/` prefix | `r/Python` |
| `parsedCommunityName` | Community without the prefix | `Python` |
| `body` | Post text or comment text | Empty after known removal markers |
| `html` | HTML version of a text body | Often empty in the archive |
| `numberOfComments` | Count shown on a post | `0` |
| `upVotes` | Reported up vote count | `1` |
| `upVoteRatio` | Reported up vote ratio | `1` |
| `score` | Reported score | `1` |
| `authorFlair` | Author flair text | May be empty |
| `flair` | Post flair text | `Discussion` |
| `isVideo` | Video post marker | `false` |
| `isAd` | Advertising marker | `false` |
| `over18` | Mature content marker | `false` when known |
| `imageUrls` | Captured preview, gallery, or image links | `[]` if none |
| `videoUrls` | Captured video fallback links | `[]` if none |
| `outboundUrl` | Linked destination when distinct from the discussion URL | May be empty |
| `thumbnail` | Thumbnail URL | May be empty |
| `locked` | Reply lock marker | `false` |
| `stickied` | Pinned item marker | `false` |
| `spoiler` | Spoiler marker | `false` |
| `archived` | Archived post marker | `false` |
| `createdAt` | Original creation time in UTC | `2026-09-25T17:46:25Z` |
| `retrievedAt` | Time the archive captured the record | `2026-09-25T17:46:51Z` |
| `scrapedAt` | Time this actor saved the row | UTC timestamp |
| `parentId` | Parent post or comment ID on comment rows | `t3_...` or `t1_...` |
| `postId` | Parent post ID on comment rows | `t3_...` |
| `category` | Community name on comment rows | `Python` |
| `numberOfReplies` | Reply count if the source supplies one | Usually empty |
| `name` | Full community identifier | `t5_...` |
| `headerImage` | Community header image URL | May be empty |
| `iconUrl` | Community icon URL | May be empty |
| `description` | Public community description | May be empty |
| `numberOfMembers` | Subscriber count at capture time | Changes over time |
| `userIcon` | User icon URL if available | Usually empty |
| `postKarma` | Archived user post karma | Available on some user summaries |
| `commentKarma` | Archived user comment karma | Available on some user summaries |

Fields are shared across row types, so a post does not contain user profile karma and a community does not contain comment text. Empty values are represented as `null`, except media lists, which are empty arrays. Counts and scores reflect the archived snapshot, not a promise about the current live page. The `retrievedAt` field helps you judge how old the snapshot is before using those numbers.

### How to scrape Reddit

1. Enter one or more public Reddit URLs in **Start URLs**, or add search terms and a **Search community**. Community, post, and user URLs are supported. Search terms are ignored when URLs are present.
2. Set the total item limit and the per-target post and comment limits. A direct post URL can produce one post row and multiple comment rows, so leave room for both.
3. Choose whether to include mature content, collect post comments, or include community details. For incremental collection, add a post or comment date limit.
4. Run the actor and check the dataset row count. Download the resulting dataset as JSON or CSV, or connect it to a spreadsheet workflow.

A community URL such as `https://www.reddit.com/r/Python/` is the simplest starting point. It can produce a community row and a sequence of posts. A post URL under `/r/<community>/comments/<post-id>/` is useful when you need the comment tree for one discussion. A user URL under `/user/<name>/` returns a user row and public posts. To collect that user's comments instead, use `/user/<name>/comments/`.

### Input

| Name | Type | Default | Description |
| --- | --- | --- | --- |
| `startUrls` | URL list | Example community URL | Public community, post, or user URLs. These take precedence over `searches`. |
| `searches` | String list | Empty | Terms searched in post titles and optionally comment bodies. Requires `searchCommunityName` for text search. |
| `searchCommunityName` | String | Empty | Community for keyword search, without `r/`. The form suggests an example. |
| `searchPosts` | Boolean | `true` | Search titles for each term. |
| `searchComments` | Boolean | `false` | Search comment bodies for each term. |
| `searchCommunities` | Boolean | `false` | Look up an exact community name matching a term. |
| `searchUsers` | Boolean | `false` | Look up an exact username matching a term. |
| `skipComments` | Boolean | `false` | Skip comments on direct post URLs. |
| `skipUserPosts` | Boolean | `false` | Skip public posts on user profile URLs. |
| `skipCommunity` | Boolean | `false` | Skip community metadata on community URLs. |
| `includeNSFW` | Boolean | `false` | Include records marked as mature. |
| `maxItems` | Integer | `100` | Total dataset cap across targets; `0` removes the overall cap. |
| `maxPostCount` | Integer | `100` | Post cap per community, user, or search term. |
| `maxComments` | Integer | `100` | Comment cap per post, user comment URL, or search term. |
| `maxCommunitiesCount` | Integer | `10` | Cap for exact community lookups. |
| `maxUserCount` | Integer | `10` | Cap for exact user lookups. |
| `postDateLimit` | ISO date | Empty | Skip posts older than this date. |
| `commentDateLimit` | ISO date | Empty | Skip comments older than this date. |
| `proxyConfiguration` | Object | Platform proxy | Optional network settings for the public source. |

For a recent community sample:

```json
{
  "startUrls": [{"url": "https://www.reddit.com/r/Python/"}],
  "maxItems": 35,
  "maxPostCount": 35,
  "skipCommunity": true,
  "includeNSFW": false
}
```

For a focused keyword query, remove `startUrls` and use `searches` with `searchCommunityName`. The source requires a community for post title and comment body text search. Exact community and user lookups can be requested with their corresponding switches, but they should not be treated as broad discovery across all names. When a date limit is set, the actor reads newest records first and stops once it crosses the limit.

### Output example

This shortened row came from a successful 35-item community run. The omitted fields in the actual row include media arrays, flags, and available counts. It represents one archive capture, so its scores and dates should be interpreted at that time.

```json
{
  "dataType": "post",
  "id": "t3_1wq2vmx",
  "parsedId": "1wq2vmx",
  "url": "https://www.reddit.com/r/Python/comments/1wq2vmx/preparsing_change_detection/",
  "title": "Pre-parsing change detection",
  "communityName": "r/Python",
  "parsedCommunityName": "Python",
  "numberOfComments": 0,
  "upVotes": 1,
  "score": 1,
  "flair": "Discussion",
  "over18": false,
  "createdAt": "2026-09-25T17:46:25Z",
  "retrievedAt": "2026-09-25T17:46:51Z"
}
```

The run returned 35 rows. All 35 had `dataType`, `id`, `url`, `title`, `username`, `communityName`, and `createdAt`. That verifies the community post path for this example; it does not imply every optional field will be populated for every community or time period.

### Use cases

- **Community monitoring:** Save recent discussions from a defined community and compare new post IDs across scheduled runs. The ID lets you avoid counting repeated posts.
- **Topic research:** Search titles and comments inside a community, then review discussion text and timestamps together. A specific community keeps the query focused and makes coverage easier to explain.
- **Conversation analysis:** Start from a direct post URL to pair the original post with archived comments. Use `postId` and `parentId` to connect comment rows back to the discussion and their immediate parent.
- **Content inventory:** Export post titles, text, links, media URLs, scores, and flair for a community. Separate post content from outbound links using `url` and `outboundUrl`.
- **Public activity review:** Use a public user URL to examine archived posts or a user comments URL for comments. Karma totals may be present on the profile summary, but profile fields are thinner than post and comment fields.
- **Historical comparison:** Keep `createdAt`, `retrievedAt`, and `scrapedAt` in downstream records. The three timestamps distinguish publication, archive capture, and your own export.

### How much does it cost to scrape Reddit?

There is no published per-1,000-result price for this actor yet. One successful validation run produced 35 rows with platform usage of approximately **$0.00005**. That is one observed run, not a guaranteed unit price: runtime, source response time, retries, memory, proxy settings, and the number of target URLs can change usage.

For planning only, dividing that single observed usage by 35 yields about **$0.00143 per 1,000 rows**. At exactly that same rate, 100 rows would be about $0.00014; 1,000 rows would be about $0.00143; and 5,000 rows would be about $0.00714. These are arithmetic examples, not a Store charge or a promise about larger runs. A larger export may require more pages, and an unavailable source may spend time retrying before it returns no rows.

Set `maxItems` to avoid collecting more rows than you need. Set `maxPostCount` and `maxComments` as well when you have several targets: the first controls post rows per target and the second controls comment rows per target. If you run the actor regularly, keep the exported IDs and compare them with prior runs so you only process new records in your own workflow.

### Tips for best results

Use a precise community URL when you know where a topic is discussed. This gives the actor a direct collection path and avoids relying on text matching. Use a post URL when comments are the main goal; a community listing alone returns posts, not all comments in every post. A user comments URL is different from a user profile URL and intentionally selects comment activity.

Start with `maxItems` between 30 and 100 so you can inspect coverage and field fill before raising limits. The source supports timestamp-based pagination for posts and comments; the actor moves to older records until it reaches your limit, the date cutoff, or the source stops returning rows. For longer history, run separate narrow community or user targets rather than assuming a single request will cover everything. The source can rate limit or return a temporary error, so rerun later if a source error is reported.

Use `postDateLimit` or `commentDateLimit` for recurring collection. An ISO date such as `2026-09-01` is accepted. The actor interprets a date without a timezone as UTC. Keep a small overlap between scheduled windows in your downstream process, then deduplicate by `dataType` and `id`. This protects against records arriving in the archive after their creation time.

Inspect `retrievedAt` before treating a count as current. Scores, subscriber counts, and comment counts are snapshots. For media exports, check both `imageUrls` and `videoUrls`; a link post may have an `outboundUrl` without either media list being populated. For text exports, `body: null` can mean that the archived record contained a removal marker, while some source records simply have no text body.

### Limits

Live Reddit pages, their JSON listings, RSS feeds, and several internal routes returned access blocks or a human-verification page during discovery. The actor therefore uses a public archive of Reddit records. The archive is independent of live page availability, but it may lag, miss records, or retain an older state after edits or removals. It is unsuitable when you need a guaranteed live score, exact current member count, or a complete record of a fast-changing discussion.

Keyword text search is restricted to one community. Unrestricted global keyword search, live relevance sorting, and a complete popular-feed view are not available from the selected source. Community and user name lookups are exact matches. Post and comment pagination is time based, so very dense periods with many records at exactly the same second may require separate time windows to avoid gaps. A direct post URL can be looked up by its post ID, but comments are limited by `maxComments` and the archive's coverage.

Some leader-style fields are unavailable in many archive records. HTML text, a user icon, a user description, reply counts, and current profile flags are often empty. The actor leaves these values empty rather than inventing them. Public records that are deleted or removed after archive capture can still be present; evaluate them before redistribution. Private, login-only, and account-specific content is outside this actor's scope.

### Use with connected agents

The input and output schemas make the actor usable through a connected workflow that can supply JSON input and read dataset rows. A useful request is: “Collect the 50 newest archived posts from `r/Python`, exclude mature posts, and summarize recurring questions with post URLs.” Another is: “For this public discussion URL, collect up to 100 comments and group them by parent ID.” Include the community, limits, and freshness requirement in the request so the result can be interpreted correctly.

### FAQ

#### Is it appropriate to collect this data?

The actor handles public records. Respect the site's terms, applicable privacy law, and the expectations of participants. Do not use public usernames or comments for harassment, sensitive profiling, or other personal-data misuse. Recheck content before redistributing a historical archive copy, especially when a live post may have been edited or removed.

#### Do I need a login or account credentials?

No account credentials are accepted. The selected source exposes archived public records. It cannot see private communities, private messages, account settings, or content that requires a logged-in session.

#### Do I need a proxy?

The actor has a proxy setting and uses the platform's proxy path by default. The successful validation run used that setup. A proxy does not make live Reddit pages available here; the actor collects from the archived JSON source. Source rate limits or outages remain possible.

#### How fast is a run?

The 35-row validation run finished in about half a minute, including startup. A run with more pages, comments, targets, or retries can take longer. Use row caps and a date window to keep a scheduled job predictable.

#### Can I schedule and monitor collection?

Yes. Schedule runs at an interval that fits your use case and monitor their status and dataset row counts. Store previous IDs in your destination to identify new records. A successful run with fewer rows than expected should prompt a check of the community activity, date cutoff, and source freshness.

#### Can I export to a spreadsheet?

Yes. Download the dataset as CSV and open it in a spreadsheet, or pass JSON rows to an integration. Filter on `dataType` before building tables because the four row kinds have different fields.

#### Why is a field empty or a count outdated?

The archive only contains what it captured. A text body may have been removed, a media preview may not have been recorded, or a profile field may never have been in the source. Counts can change after `retrievedAt`. Treat nulls as missing source data, not as zero.

#### Can I search every community at once?

No. Text search needs `searchCommunityName`. This constraint keeps the query bounded and matches the selected source's search behavior. Supply a list of community URLs if your goal is to monitor several known communities.

### Integrations

Use the dataset API to retrieve JSON rows after a run, download CSV for a spreadsheet, or send completed run data to a webhook. Scheduling and workflow connections can trigger repeat collection and compare IDs across runs. Keep `dataType` in every downstream record so fields from different row kinds are interpreted correctly.

### Support

Open an issue on the Issues tab; we reply within 24h and add fields on request.

# Actor input Schema

## `startUrls` (type: `array`):

Public Reddit community, post, or user URLs. Search terms are ignored when URLs are supplied.

## `searches` (type: `array`):

Title and comment text queries. A community is required for keyword search.

## `searchCommunityName` (type: `string`):

Community to search, without r/. Required when using search terms.

## `searchPosts` (type: `boolean`):

Return matching post titles for each search term.

## `searchComments` (type: `boolean`):

Return matching comment bodies for each search term.

## `searchCommunities` (type: `boolean`):

Look up a community whose name exactly matches each search term.

## `searchUsers` (type: `boolean`):

Look up a user whose name exactly matches each search term.

## `skipComments` (type: `boolean`):

Do not collect comments when a direct post URL is supplied.

## `skipUserPosts` (type: `boolean`):

Do not collect posts from a user URL.

## `skipCommunity` (type: `boolean`):

Do not add a community metadata row when collecting a community URL.

## `includeNSFW` (type: `boolean`):

Include records marked as mature content.

## `maxItems` (type: `integer`):

Hard cap on all dataset rows. Set 0 for no overall cap.

## `maxPostCount` (type: `integer`):

Maximum post rows from each community or user URL or search term.

## `maxComments` (type: `integer`):

Maximum comment rows from each post URL, user comment URL, or search term.

## `maxCommunitiesCount` (type: `integer`):

Maximum exact community matches for a search term.

## `maxUserCount` (type: `integer`):

Maximum exact user matches for a search term.

## `postDateLimit` (type: `string`):

Only return posts created on or after this ISO date or timestamp.

## `commentDateLimit` (type: `string`):

Only return comments created on or after this ISO date or timestamp.

## `proxyConfiguration` (type: `object`):

Optional proxy settings for the public data source.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/Python/"
    }
  ],
  "searchCommunityName": "Python",
  "searchPosts": true,
  "searchComments": false,
  "searchCommunities": false,
  "searchUsers": false,
  "skipComments": false,
  "skipUserPosts": false,
  "skipCommunity": false,
  "includeNSFW": false,
  "maxItems": 100,
  "maxPostCount": 100,
  "maxComments": 100,
  "maxCommunitiesCount": 10,
  "maxUserCount": 10,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `items` (type: `string`):

Public post, comment, community, and user rows in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.reddit.com/r/Python/"
        }
    ],
    "searchCommunityName": "Python"
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/reddit-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.reddit.com/r/Python/" }],
    "searchCommunityName": "Python",
}

# Run the Actor and wait for it to finish
run = client.actor("mlg14/reddit-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/Python/"
    }
  ],
  "searchCommunityName": "Python"
}' |
apify call mlg14/reddit-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/reddit-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3lUmUmE4fBgcgejIL/builds/0QCD20iiSOYJxLojp/openapi.json
