# Reddit Search Scraper (`mlg14/reddit-search-scraper`) Actor

Search public Reddit posts by keyword, subreddit, or URL from public archives; export post metadata with optional archived comments.

- **URL**: https://apify.com/mlg14/reddit-search-scraper.md
- **Developed by:** [MLG Data](https://apify.com/mlg14) (community)
- **Categories:** Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Search Scraper

Search public Reddit posts by phrase, community, or direct URL and export structured Reddit data to CSV, JSON, or Excel. The actor is designed as a Reddit search data source for post content, public engagement counts, and optional comments.

**Current access status:** The latest remote test returned no records. Reddit returned HTTP 403 for the search JSON route on both datacenter and residential connections, including after a browser challenge attempt. The actor fails visibly when every target is blocked. It is not ready for production collection until a public route is accessible again and a golden run passes.

### What data can you extract from Reddit?

A post and each collected comment are separate dataset rows. `kind` distinguishes them. All rows use flat keys, so the same dataset can be exported without expanding nested objects. A null value means the source did not provide that field or the field does not apply to that record type. The descriptions below define the intended output contract; field availability has not been validated by a successful live run.

| Field | Description | Example shape |
|---|---|---|
| `kind` | `post` or `comment` | `"post"` |
| `id` | Stable identifier for the post or comment | `"abc123"` |
| `query` | Search phrase that found the post | `"coffee grinder"` |
| `title` | Post title; null for comments | `"Choosing a grinder"` |
| `body` | Post text or comment body | `"Looking for suggestions..."` |
| `author` | Public username as returned by Reddit | `"example_user"` |
| `subreddit` | Community name | `"Coffee"` |
| `score` | Public vote score | `42` |
| `upvoteRatio` | Post upvote share, if supplied | `0.91` |
| `numComments` | Public post comment count | `18` |
| `createdAt` | Creation timestamp in UTC | `"2026-09-01T12:00:00Z"` |
| `url` | Canonical Reddit page URL | `"https://www.reddit.com/r/Coffee/comments/abc123/example/"` |
| `permalink` | Relative page path | `"/r/Coffee/comments/abc123/example/"` |
| `outboundUrl` | Destination linked from a post | `"https://www.reddit.com/..."` |
| `flair` | Public post flair text | `"Question"` |
| `domain` | Link destination domain | `"self.Coffee"` |
| `thumbnail` | Public thumbnail URL when present | `"https://..."` |
| `postHint` | Content hint returned for a post | `"image"` |
| `isSelf` | Whether the post is text only | `true` |
| `isVideo` | Whether the post is a video | `false` |
| `isGallery` | Whether the post is a gallery | `false` |
| `isNsfw` | Whether the post is marked adult | `false` |
| `isSpoiler` | Whether the post is marked as a spoiler | `false` |
| `isLocked` | Whether replies are locked | `false` |
| `isStickied` | Whether the item is pinned | `false` |
| `edited` | Edit timestamp or false, as returned | `false` |
| `postId` | Parent post identifier on comment rows | `"abc123"` |
| `postUrl` | Parent post URL on comment rows | `"https://www.reddit.com/..."` |
| `parentId` | Parent post or comment fullname | `"t3_abc123"` |
| `depth` | Comment nesting level | `0` |

The `url` field points to a Reddit page. For a link post, `outboundUrl` points to its destination, which may be an external site. A text post may have a Reddit URL in both places. `score` and `numComments` can change after collection, so retain the run date when comparing exports. The `query` field records the search phrase for keyword inputs; it is empty for broad community listings and direct post links.

### How to scrape Reddit

1. Enter a phrase in `queries`, a community in `subredditName`, or one or more Reddit links in `urls`.
2. Choose a sort and timeframe when searching. Set `maxPosts` for a per-target cap and `maxItems` for a cap across the entire dataset.
3. Enable `scrapeComments` if discussion text is needed, and set `maxComments` to bound each thread.
4. Run the actor. If the site accepts the requests, inspect the dataset and export CSV, JSON, or Excel. If the site blocks every request, the run fails with a clear error and produces no rows.

Start with a small search and verify the first records before scheduling larger runs. Search ranking can change between runs, and the same post can be found through several queries. The actor deduplicates posts and comments by identifier within one run. Keep `kind` together with `id` when combining datasets from different runs, because a post ID and a comment ID are separate record types.

### Input

| Name | Type | Default | Description |
|---|---|---|---|
| `queries` | string array | none | Global post search phrases. |
| `subredditName` | string | none | Community name without `r/`. |
| `subredditKeywords` | string array | none | Phrases to search inside the selected community. |
| `urls` | string array | none | Direct post, search, community, or user-submitted URLs. Takes priority over other targets. |
| `sort` | string | `relevance` | Global ranking: relevance, hot, top, new, or comments. |
| `timeframe` | string | `all` | Global search period: all, year, month, week, day, or hour. |
| `subredditSort` | string | `relevance` | Ranking for community keyword searches. |
| `subredditTimeframe` | string | `all` | Time period for community keyword searches. |
| `maxPosts` | integer | `100` | Maximum accepted posts per input target. |
| `maxItems` | integer | `0` | Total post and comment row cap; zero means no cap. |
| `scrapeComments` | boolean | `false` | Collect public comments from accepted posts. |
| `maxComments` | integer | `100` | Comment cap per post, subject to the initial thread response. |
| `dateFrom` | date or timestamp | none | Earliest UTC post time to keep. |
| `dateTo` | date or timestamp | none | Latest UTC post time to keep. |
| `commentDateFrom` | date or timestamp | none | Earliest UTC comment time to keep. |
| `commentDateTo` | date or timestamp | none | Latest UTC comment time to keep. |
| `includeNsfw` | boolean | `false` | Include posts marked adult. |
| `strictTokenFilter` | boolean | `false` | Require every query word in title, body, or outbound URL. |
| `proxyConfiguration` | object | enabled | Connection settings; datacenter is attempted before residential. |

At least one target is needed: a query, a community, or a direct URL. When `urls` is present, it takes priority. A community without keywords uses its newest listing. The post date fields filter output after Reddit returns results; they do not make Reddit search enumerate every historical post in that range. The `timeframe` options are Reddit search controls. Enter dates as `YYYY-MM-DD` or an ISO timestamp. A plain start date begins at midnight UTC, while a plain end date includes that day's final second.

A bounded example input:

```json
{
  "queries": ["coffee grinder"],
  "sort": "new",
  "timeframe": "month",
  "maxPosts": 40,
  "maxItems": 40,
  "scrapeComments": false
}
```

This is the golden input. Its latest run returned zero rows because Reddit blocked the structured search route, so this input should not be treated as a successful sample. The requested minimum remains 30 posts to catch a regression that silently returns only a few rows.

### Output example

No genuine output item is available from the golden run. The run returned HTTP 403 and an empty dataset. A sample row is intentionally omitted until a remote run produces one. The field table above documents the expected schema, while the current access status describes what has actually been verified.

When the route becomes accessible, every accepted post is written to the default dataset. With comment collection enabled, comments follow their parent post as additional rows. The `kind` and `postId` fields allow a consumer to separate posts from replies and join each reply to its thread. JSON keeps numbers, booleans, and nulls in their native form. Tabular exports expose the same keys as columns.

### Use cases

- Community research: collect matching public posts from one or several communities, then group them by `subreddit` and `query`.
- Product feedback review: search for a product category, inspect post text, and compare recurring questions or complaints.
- Discussion analysis: include comments to inspect public replies alongside the post that started a thread.
- Trend monitoring: schedule the same query with `sort=new` and retain stable IDs to distinguish newly seen posts from previously exported posts.
- Content discovery: use `flair`, `postHint`, `isVideo`, and `isGallery` to separate different public post formats.
- Engagement review: compare public `score`, `upvoteRatio`, and `numComments` across an exported collection, with the understanding that these values change over time.

These workflows depend on public access. The current 403 response prevents collection from this account, so none of these uses is presently verified end to end. Check an initial dataset before building a downstream process around it. For research requiring complete historical coverage, Reddit search is an imperfect source even when accessible: ranking, moderation, deletion, and result windows can omit posts.

### How much does it cost to scrape Reddit?

The configured event price is **$1.00 per 1,000 saved rows**. Posts and comments both count as rows. At that rate, 40 saved rows cost **$0.04**, 100 saved rows cost **$0.10**, and 1,000 saved rows cost **$1.00** in result events. These are arithmetic examples, not measured successful runs. Additional platform usage or connection costs may apply according to the account configuration. The failed golden run saved zero rows, although the remote run still consumed platform resources.

Use `maxItems` to cap the number of emitted rows, especially when comments are enabled. `maxPosts` controls posts per search target, while `maxComments` controls the initial set of comments per accepted post. A 40-post run with comments enabled can produce more than 40 rows unless `maxItems` also limits the total. If budget control matters, set both limits explicitly and review the actual result count after each run.

### Tips for best results

Use a specific query before trying a broad one. A short query may mix unrelated meanings, and a strict token filter can remove relevant posts that use different wording. Compare a small unfiltered run with a filtered run before depending on strict matching. For a known community, enter its name and one or more community keywords. Without community keywords, the actor requests the newest public listing rather than search results.

Use `sort=new` for recent monitoring, and combine it with a timeframe when appropriate. The separate `dateFrom` and `dateTo` settings remove records outside your exact date window after retrieval. They cannot recover posts missing from Reddit's returned pages. Increase `maxPosts` only after checking whether pagination exposes more unique posts. Search results may repeat, and the actor skips duplicate identifiers inside the same run.

Comment collection adds a thread request for each accepted post. Keep `maxComments` modest for a first run. The actor reads the initial public comment tree and follows replies included in that response; it does not expand additional comment placeholders. A post's public `numComments` may therefore exceed the number of comment rows saved. Removed comments, unavailable threads, and blocked thread requests can reduce the count further. If a thread request fails after the post was saved, the post remains in the dataset and the failure is logged.

### Limits

**Current blocking:** Search pages returned challenges and structured search routes returned HTTP 403 in remote probes. The golden run confirmed the block on both datacenter and residential connections. This is the primary open issue. The actor does not treat a challenge page as a valid result. A run with no retrievable targets ends in failure rather than reporting a successful empty dataset.

**Result coverage:** Reddit search does not provide an unlimited historical scan. A query, sort, and timeframe combination can expose only a practical window of results. The actor follows JSON pagination cursors until the requested cap, an empty page, or a repeated cursor. It does not fan out across alternative sorts or automatically divide large date ranges. Some posts can be absent because of ranking, deletion, community restrictions, or moderation.

**Access boundaries:** Private communities, login-only content, removed material, and unavailable posts are outside the public-data contract. The actor does not use a user account. Public authorship fields can be deleted or anonymized by the site, and fields such as flair, thumbnails, media hints, and upvote ratio are optional. Engagement counts are snapshots, not permanent facts.

**Comments:** The initial thread response may omit deeper replies behind expansion placeholders. `maxComments` is a ceiling, not a guarantee. A direct comment permalink is treated as its parent thread. Separate comment date filters apply only when comment collection is enabled.

**Filters:** An exact post date range is checked after retrieval. If ranking does not surface enough posts from the selected dates, a high `maxPosts` value may still return fewer accepted rows than expected. `includeNsfw=false` excludes flagged posts; it does not inspect content for additional sensitive material.

### Automated workflows

The dataset can be consumed through a dataset endpoint or downloaded after a run. A scheduled run can reuse the same input for periodic monitoring. For repeat collections, upsert by `kind` and `id`, and store each run's timestamp separately if engagement changes matter. A webhook can notify a downstream process when a run ends. Treat a failed run as an access failure and avoid interpreting zero rows as zero matching Reddit posts.

Two example workflow requests are: “Find recent public posts about coffee grinders and return their titles, communities, URLs, and comment counts,” and “Collect posts from one community this week, include up to 20 public replies per post, and export the dataset.” These describe intended usage only; the current blocked state must be resolved first.

### FAQ

**Is this collecting public data?** The actor is designed for publicly accessible posts and comments. It does not sign in or request private communities. Use exported content in line with applicable law, site terms, and privacy obligations. Avoid using public usernames for unwanted contact or profiling.

**Do I need to configure a proxy?** The default connection setup first tries datacenter access and escalates after blocks. The latest test still received 403 responses after escalation. Changing the proxy setting is not a demonstrated fix.

**How fast is a run?** No successful collection time has been measured. Runtime depends on site access, number of pages, and whether each post requires a comment request. The failed 40-post test spent about a minute attempting the blocked source.

**Can I schedule repeated runs?** Yes, the input is suitable for a recurring schedule, but scheduling a currently blocked source will repeat the failure. First confirm a successful small run.

**Can I export to a spreadsheet?** Dataset rows can be downloaded as CSV or Excel as well as JSON. The latest golden dataset is empty, so export currently contains no useful records.

**Why is a field empty?** Some values apply only to posts or only to comments. Other fields are optional in Reddit's public response. A null is distinct from a zero score or a false flag.

**Why are there fewer comments than the displayed count?** The first thread response can omit replies behind expansion placeholders, and the actor respects the configured comment cap. Removed or inaccessible replies also reduce exported rows.

### Integrations

Use dataset exports, an API endpoint, scheduling, and webhooks to move successful records into a reporting or storage workflow. Use the stable record identifiers for upserts, and monitor failed runs so a site block does not appear as an ordinary quiet period. Integrations should only be enabled after a live run produces the expected rows.

### Support

Open an issue with the run ID, input with sensitive values removed, the observed error, and the expected result. The current reproducible issue is HTTP 403 on anonymous Reddit JSON search requests.

# Actor input Schema

## `queries` (type: `array`):

Search phrases across public Reddit posts. Ignored when urls are supplied.

## `subredditName` (type: `string`):

Optional community name without r/. With no subredditKeywords, collect its newest posts.

## `subredditKeywords` (type: `array`):

Search phrases to run within subredditName. Leave empty for the newest feed.

## `urls` (type: `array`):

Optional post, search, subreddit, or user-submitted URLs. These take priority over queries and subredditName.

## `sort` (type: `string`):

Global archive order. New uses creation time; top and comments use archived counts. Relevance and hot currently use newest order.

## `timeframe` (type: `string`):

Reddit's time window for global searches.

## `subredditSort` (type: `string`):

Order the returned community archive page: newest for relevance/new, archived score for top/hot, or archived comment count for comments.

## `subredditTimeframe` (type: `string`):

Client-side time window applied to returned community archive posts.

## `maxPosts` (type: `integer`):

Maximum posts per target. Community archive currently exposes at most 100 per target.

## `maxItems` (type: `integer`):

Hard cap across posts and comments. Zero means no total cap.

## `scrapeComments` (type: `boolean`):

Collect public comments from each saved post.

## `maxComments` (type: `integer`):

Maximum comments from the initial public thread response for each post.

## `dateFrom` (type: `string`):

Keep posts on or after this UTC date (YYYY-MM-DD) or ISO timestamp.

## `dateTo` (type: `string`):

Keep posts on or before this UTC date (YYYY-MM-DD) or ISO timestamp.

## `commentDateFrom` (type: `string`):

Keep comments on or after this UTC date or ISO timestamp.

## `commentDateTo` (type: `string`):

Keep comments on or before this UTC date or ISO timestamp.

## `includeNsfw` (type: `boolean`):

Include posts marked as adult content.

## `strictTokenFilter` (type: `boolean`):

Keep a post only when every query word appears in its title, body, or outbound URL.

## `proxyConfiguration` (type: `object`):

Connection settings for public archive requests; datacenter is tried before residential escalation.

## Actor input object example

```json
{
  "queries": [
    "coffee grinder"
  ],
  "sort": "relevance",
  "timeframe": "all",
  "subredditSort": "relevance",
  "subredditTimeframe": "all",
  "maxPosts": 100,
  "maxItems": 0,
  "scrapeComments": false,
  "maxComments": 100,
  "includeNsfw": false,
  "strictTokenFilter": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `results` (type: `string`):

All scraped items in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "coffee grinder"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("mlg14/reddit-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "queries": ["coffee grinder"] }

# Run the Actor and wait for it to finish
run = client.actor("mlg14/reddit-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "coffee grinder"
  ]
}' |
apify call mlg14/reddit-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,mlg14/reddit-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8WlQ45gatMiDQDsiQ/builds/PFdKjMxjUY3tspf7p/openapi.json
