# Reddit Posts & Search Scraper (`coregent/reddit-posts-search-scraper`) Actor

Search public Reddit posts by keyword, scrape subreddit feeds, or paste post and search URLs. One deduplicated row per post with title, text, author, community, score, comments, flair, media, timestamps and source lineage. One global limit. No Reddit login or API key.

- **URL**: https://apify.com/coregent/reddit-posts-search-scraper.md
- **Developed by:** [Delowar Munna](https://apify.com/coregent) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 post results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Reddit Posts & Search Scraper

![Reddit Posts & Search Scraper](https://raw.githubusercontent.com/coregentdevspace/reddit-posts-search-scraper-assets/main/thumbnail-reddit-posts-search-scraper.png "Reddit Posts & Search Scraper — subreddit feeds, Reddit search and post URLs in, one deduplicated post row out")

Search public Reddit posts by keyword, scrape subreddit feeds, or paste known Reddit post/search
URLs. Every match becomes one clean, deduplicated post row with text, author, community, engagement,
flair, media, timestamps and source lineage. No Reddit account, cookies or API key required.

- **One global limit.** `maxResults` bounds the whole run, across every subreddit, query and URL.
- **One row per post, whichever source found it.** A post matched by two queries and its subreddit
  is delivered once, with every matching query and community on the row.
- **Built for scheduled runs.** Skip lists for IDs and URLs, a `publishedAfter` boundary that stops
  paging a `new` feed the moment it reaches old posts, and a chronological sort by default.
- **Honest about depth.** Reddit's public search returns at most about 250 posts per query or
  community for one sort and time window. The run says so per source instead of paging in circles.

Comments are a separate Actor: [Reddit Comments Scraper](https://apify.com/coregent/reddit-comments-scraper)
takes the `url` or `postId` from these rows. User histories belong to the Reddit User Scraper.

### Quick start

**Subreddits and keyword searches together:**

```json
{
    "subreddits": ["SaaS", "Entrepreneur"],
    "searchQueries": ["best CRM", "HubSpot alternative"],
    "maxResults": 200,
    "sortBy": "new",
    "timeFilter": "month",
    "includeNsfw": false,
    "enrichPostDetails": true
}
```

Four sources, 200 posts at most, newest first, deduplicated across all four. The run summary in the
key-value store says how many candidates each source produced and where each one stopped.

**Reddit URLs, with filters:**

```json
{
    "startUrls": [
        "https://www.reddit.com/r/python/search/?q=asyncio&restrict_sr=1",
        "https://www.reddit.com/r/AskReddit/comments/1wjydiy/",
        "https://redd.it/1wjgckp"
    ],
    "maxResults": 100,
    "minScore": 10,
    "postTypes": ["text"],
    "excludeKeywords": ["hiring"],
    "includeNsfw": false,
    "enrichPostDetails": true
}
```

A search restricted to r/Python plus two direct posts. The search URL carries its own query;
the two post URLs are looked up directly. Posts under 10 points, non-text posts and posts mentioning
"hiring" are dropped before they are saved, and are not charged.

**A daily scheduled run:**

```json
{
    "subreddits": ["personalfinance", "legaladvice"],
    "sortBy": "new",
    "publishedAfter": "24h",
    "maxResults": 500,
    "skipPostIds": ["1wquuj6", "1wp6zmh"],
    "includeNsfw": false,
    "enrichPostDetails": true
}
```

Each community stops paging once it reaches posts older than a day, and the IDs from the previous
run are skipped. See [Incremental scheduled runs](#incremental-scheduled-runs).

### Search Reddit

Put keywords or Reddit search syntax in `searchQueries`, one query per line:

```json
{ "searchQueries": ["open source LLM", "\"vector database\" subreddit:MachineLearning"] }
```

Each query is a site-wide Reddit search. `sortBy` chooses the ordering Reddit applies (`relevance`,
`hot`, `top`, `new`, `comments`), `timeFilter` the window. A search URL in `startUrls` works the same
way and keeps the sort and window it carries.

### Scrape subreddits

Put community names in `subreddits` in any spelling — `python`, `r/Python`,
`https://www.reddit.com/r/python/` — or a subreddit URL with a sort segment such as
`https://www.reddit.com/r/python/top/?t=week`, which keeps that sort and window for that source
only. `rising` and `controversial` are feed orderings; where a surface lacks them the run substitutes
`hot` / `top` and says so in the log.

`r/all` and `r/popular` are aggregate feeds rather than communities and are not supported.

### Direct posts and search URLs

`startUrls` accepts every common form of a post reference — full post URLs on any Reddit host,
`redd.it` short links, `/r/<sub>/s/` share links, `t3_` fullnames and bare IDs — plus subreddit URLs
and search URLs (`https://www.reddit.com/search/?q=…`, `https://www.reddit.com/r/python/search/?q=…&restrict_sr=1`).
Anything else (a user profile, a wiki page, a multireddit) is reported as invalid with a reason, never guessed at.

### Input reference

| Field | Default | What it does |
|---|---|---|
| `subreddits` | `[]` | Communities to scrape (names, `r/name`, or URLs) |
| `searchQueries` | `[]` | Site-wide keyword searches |
| `startUrls` | `[]` | Post, subreddit or search URLs |
| `maxResults` | `500` | **Global** ceiling on unique posts delivered by the run |
| `sortBy` | `new` | `new`, `hot`, `top`, `relevance`, `comments`, `rising`, `controversial` |
| `timeFilter` | `all` | `hour` … `all`; applies to `top`, `relevance`, `comments` |
| `publishedAfter` | — | ISO date or relative window (`24h`, `7d`, `2w`); stops a `new` feed at the boundary |
| `publishedBefore` | — | Upper date bound |
| `includeNsfw` | `false` | Include over-18 posts |
| `minScore`, `minComments` | — | Engagement floors |
| `flairs`, `authors` | `[]` | Allow-lists (case-insensitive) |
| `includeKeywords`, `excludeKeywords` | `[]` | Substring filters on title and body |
| `postTypes` | `[]` | `text`, `link`, `image`, `video`, `gallery`, `poll`, `crosspost` |
| `skipPostIds`, `skipUrls` | `[]` | Incremental skip lists — never saved, never charged |
| `enrichPostDetails` | `true` | Full post details on every row — body, flair, post type, media and flags — at no extra charge |
| `maxResultsPerSource` | `250` | Fairness cap per source |
| `maxPagesPerSource` | `20` | Crawl ceiling per source (7 posts per page on Reddit's public search) |
| `maxConcurrency` | `3` | Sources paged in parallel (1–5) |
| `requestDelayMs` | — | Extra pacing on top of the built-in rate control |
| `debugMode` | `false` | Verbose logging |
| `proxyConfiguration` | Apify Datacenter | See the proxy policy below |

At least one of `subreddits`, `searchQueries` or `startUrls` is required. Every filter keeps a post
whose value Reddit does not publish — an unknown score is not a low score.

### Global max and cost control

You are charged once per unique post saved to the dataset, and nothing else: not duplicates across
sources, not posts your filters removed, not skip-list matches, not posts discovered past the limit,
not invalid inputs, not sources that failed, not empty searches, not the run summaries. `maxResults`
is the only number you need to bound a run. If your Apify account has a per-run spending limit, the
run stops fetching the moment it is met rather than working on unpaid.

### Incremental scheduled runs

1. Sort by `new` (the default).
2. Set `publishedAfter` to the interval — `24h` for a daily schedule, `7d` for weekly. Each `new`
   feed stops paging as soon as it reaches posts older than that, so a quiet subreddit costs one page.
3. For exact de-duplication across runs, pass the previous run's `postId` values as `skipPostIds`
   (or its `url` values as `skipUrls`). Skipped posts are neither saved nor charged.

Overlapping sources within one run are deduplicated automatically; the `matchedQueries`,
`matchedSubreddits` and `sourceInputs` fields on each row say which inputs found it.

### Output fields

One flat row per post, `recordType: "post"`.

| Field | Meaning |
|---|---|
| `postId`, `postFullId` | Reddit's ID and `t3_` fullname — the dedupe key |
| `url`, `permalink` | Canonical post URL and its path |
| `title`, `text` | Title and body text (plain text; `""` for a post with no body) |
| `authorUsername`, `authorId` | Author (`null` when deleted) |
| `subreddit`, `subredditPrefixed`, `subredditId` | Community |
| `flair` | Post flair text |
| `score`, `upvoteRatio`, `commentCount`, `awardCount` | Engagement |
| `createdAt`, `editedAt` | ISO timestamps |
| `isNsfw`, `isSpoiler`, `isStickied`, `isLocked`, `isArchived`, `isQuarantinedSubreddit`, `isDeletedOrRemoved` | Flags |
| `postType` | `text`, `link`, `image`, `video`, `gallery`, `poll`, `crosspost` |
| `outboundUrl`, `domain`, `thumbnailUrl`, `mediaUrls`, `crosspostParentId` | Link and media |
| `detailLevel` | `full` (all fields) or `listing` (see below) |
| `matchedQueries`, `matchedSubreddits`, `sourceInputs` | Which inputs found the post |
| `capturedSequence`, `scrapedAt` | Capture order and time |

`null` always means "not published by the source", never a guess.

**Listing rows.** Reddit's public listing carries the identity, engagement and timestamp fields
only. With `enrichPostDetails` on (the default) the Actor reads every source through a search
service that returns full records, looks up direct post URLs one by one, and every row is
`detailLevel: "full"`. With it off, or when the detail service is unavailable for a run, rows are
`detailLevel: "listing"` and `text`, `flair`, `postType`, `upvoteRatio`, `editedAt`, the media
fields and the stickied/locked flags are `null`. The run log says which applied.

### Dataset views and sample records

![Reddit Posts & Search Scraper — Posts view, table (one row per post with community, author, score, comments, flair and post type)](https://raw.githubusercontent.com/coregentdevspace/reddit-posts-search-scraper-assets/main/reddit-posts-search-scraper-output-posts-table-view.png "Reddit Posts & Search Scraper — Posts view, table")

The dataset has six views. They are projections of the same dataset, so no row appears twice. Below
is one real record for each view, from a 150-post run over r/personalfinance, r/legaladvice, two
keyword searches, an r/Python search URL and two direct posts. Usernames are anonymised.

#### Posts — every field

The full schema. This is a text post found by the r/Python search URL:

```json
{
    "recordType": "post",
    "postId": "1w43div",
    "postFullId": "t3_1w43div",
    "url": "https://www.reddit.com/r/Python/comments/1w43div/when_do_you_prefer_asynciosemaphore_over_an/",
    "permalink": "/r/Python/comments/1w43div/when_do_you_prefer_asynciosemaphore_over_an/",
    "title": "When do you prefer asyncio.Semaphore over an asyncio.Queue for limiting concurrency?",
    "text": "I've been thinking about concurrency control in asyncio.\n\nA common pattern for limiting concurrent work is:\n\n    sem = asyncio.Semaphore(10)\n\n    async with sem:\n        await do_work()\n\nBut in many cases, couldn't the same problem be modeled by putting work into an asyncio.Queue and running a fixed number of worker tasks?\n\nI'm curious how experienced Python developers decide between the two approaches.\n\nAre there real-world situations where a semaphore is clearly the better abstraction than a worker queue? Are there meaningful differences in cancellation behavior, backpressure, fairness, task lifetime, or code complexity?\n\nI'd especially be interested in examples from production async Python code.",
    "authorUsername": "example_author",
    "authorId": "t2_example",
    "subreddit": "Python",
    "subredditPrefixed": "r/Python",
    "subredditId": "t5_2qh0y",
    "flair": "Discussion",
    "score": 64,
    "upvoteRatio": 0.8809523809523809,
    "commentCount": 13,
    "createdAt": "2026-09-01T06:08:20.940Z",
    "editedAt": null,
    "isNsfw": false,
    "isSpoiler": false,
    "isStickied": false,
    "isLocked": false,
    "isArchived": false,
    "isQuarantinedSubreddit": false,
    "isDeletedOrRemoved": false,
    "postType": "text",
    "outboundUrl": null,
    "domain": "self.Python",
    "thumbnailUrl": null,
    "mediaUrls": [],
    "awardCount": 0,
    "crosspostParentId": null,
    "detailLevel": "full",
    "matchedQueries": [
        "asyncio"
    ],
    "matchedSubreddits": [],
    "sourceInputs": [
        "https://www.reddit.com/r/python/search/?q=asyncio&restrict_sr=1"
    ],
    "capturedSequence": 32,
    "scrapedAt": "2026-09-27T00:36:57.450Z"
}
```

#### Search results

Rows found by a query, with the queries that matched:

```json
{
    "matchedQueries": [
        "HubSpot alternative"
    ],
    "postId": "1wp6zmh",
    "title": "a competitor shut down and I pulled an all-nighter. I have no idea what I'm doing",
    "subredditPrefixed": "r/SaaS",
    "authorUsername": "example_author",
    "score": 122,
    "commentCount": 67,
    "createdAt": "2026-09-24T16:53:07.516Z",
    "flair": null,
    "postType": "text",
    "url": "https://www.reddit.com/r/SaaS/comments/1wp6zmh/a_competitor_shut_down_and_i_pulled_an_allnighter/"
}
```

#### Subreddit posts

Rows read from a community feed:

```json
{
    "matchedSubreddits": [
        "personalfinance"
    ],
    "postId": "1wquuj6",
    "title": "Deceased relative had 30,000 employee stock options I found in old SEC filings. How do I trace what happened to them?",
    "authorUsername": "example_author",
    "score": 835,
    "commentCount": 100,
    "createdAt": "2026-09-26T16:30:25.431Z",
    "flair": "Employment",
    "postType": "text",
    "isStickied": false,
    "url": "https://www.reddit.com/r/personalfinance/comments/1wquuj6/deceased_relative_had_30000_employee_stock/"
}
```

#### High engagement

Score, upvote ratio, comments and awards, compact. This is a direct post URL:

```json
{
    "score": 10468,
    "upvoteRatio": 0.9604797483287456,
    "commentCount": 6980,
    "awardCount": 7,
    "title": "What warning signs of a worsening economy are you seeing in your line of work that the rest of us wouldn’t notice?",
    "subredditPrefixed": "r/AskReddit",
    "authorUsername": "example_author",
    "createdAt": "2026-09-18T18:33:12.419Z",
    "url": "https://www.reddit.com/r/AskReddit/comments/1wjydiy/what_warning_signs_of_a_worsening_economy_are_you/"
}
```

#### Links & media

Outbound URL, domain, thumbnail and media URLs for link and media posts:

```json
{
    "postType": "image",
    "title": "awesome-jev-projects - a curated directory of ~600 open-source tools built on the idea of using a fast, cheap typed-decision model for the small choices in agent loops instead of burning a full reasoning LLM on every branch",
    "outboundUrl": "https://i.redd.it/svq1kaiagbrh1.png",
    "domain": "i.redd.it",
    "thumbnailUrl": "https://preview.redd.it/svq1kaiagbrh1.png?width=140&height=124&auto=webp&s=7015e3a4f80369130d591a51f65e52a687c3ddff",
    "mediaUrls": [
        "https://preview.redd.it/svq1kaiagbrh1.png?auto=webp&s=0286533ed115c57287e89965f46c0ee74b8e694b",
        "https://preview.redd.it/svq1kaiagbrh1.png?width=108&auto=webp&s=1656173e50ad559a030fe2db72d0eaa5697c2213",
        "https://preview.redd.it/svq1kaiagbrh1.png?width=216&auto=webp&s=48c5e7142c51d3e514bfdec27727689def7088d1",
        "https://preview.redd.it/svq1kaiagbrh1.png?width=320&auto=webp&s=73b7be4cd11fdb42d37291998b58b5364465fa50",
        "https://preview.redd.it/svq1kaiagbrh1.png?width=640&auto=webp&s=022797265457091698183d9ad1cd84eef6dda302"
    ],
    "subredditPrefixed": "r/BestGitHubRepos",
    "score": 51,
    "createdAt": "2026-09-23T18:50:48.706Z",
    "url": "https://www.reddit.com/r/BestGitHubRepos/comments/1woeo4l/awesomejevprojects_a_curated_directory_of_600/"
}
```

#### Provenance

Which inputs produced each row, in capture order. This row came from a `redd.it` short link:

```json
{
    "capturedSequence": 16,
    "postId": "1wjgckp",
    "sourceInputs": [
        "https://redd.it/1wjgckp"
    ],
    "matchedQueries": [],
    "matchedSubreddits": [],
    "detailLevel": "full",
    "scrapedAt": "2026-09-27T00:36:54.494Z"
}
```

### Pricing

Pay per result: one `post-result` event per unique post saved, with full post details included.
There is no start fee and no charge for an empty run. The rate per post depends on your Apify
plan. This README deliberately quotes no figures — the Pricing tab shows the live rate for each plan.

### Limits and Reddit search depth

- Reddit's public search returns at most about **250 posts per query or community** for one sort
  and time window (measured: 36 pages of 7, then no next page). A source that hits it is reported
  as `truncationReason: "search-depth-limit"` in `SOURCE_SUMMARY`. To go deeper, split the source:
  `sortBy: "top"` with `timeFilter: "week"`, `"month"`, `"year"` are three different 250s.
- `maxResultsPerSource` defaults to that ceiling.
- The search surfaces have no `rising` or `controversial` ordering. With full post details on (the
  default) those sorts are substituted with `hot` / `top` and the run log says so; with details off
  they are honoured where a subreddit's own feed supports them.
- Private, quarantined-behind-a-wall and banned communities answer "not found".

### No-login / proxy posture

The Actor asks Reddit's public, server-rendered search surface as a plainly named client — no
account, no cookies, no API key, no browser. It paces itself to Reddit's published logged-out
budget and rests an exit that draws a challenge page.

#### 🚦 Proxy policy

Use **Apify Datacenter** proxy (the default) or **no proxy**. Both work for Reddit's public search
surface at this Actor's conservative concurrency.

**Apify Residential proxy is not supported.** The run fails at startup if `apifyProxyGroups`
includes `RESIDENTIAL`: in pay-per-event Actors, residential bandwidth is billed to the developer
rather than to the run, and Reddit serves the same results to datacenter addresses. If you
genuinely need residential routing, supply your own provider under **Custom proxy URLs** — that
traffic goes through your account and is honoured in full:

```
http://user:pass@proxy.iproyal.com:12321
http://user:pass@proxy.brightdata.com:22225
http://user:pass@proxy.oxylabs.io:7777
```

### Troubleshooting

| Symptom | Where to look | Likely cause |
|---|---|---|
| Fewer posts than `maxResults` | `SOURCE_SUMMARY.truncationReason` | `search-depth-limit` (Reddit's ceiling), `publishedAfter`, `maxResultsPerSource`, or the sources simply ran out |
| A source is `failed: not-found` | run log | the community does not exist, is private, or the search URL had no query |
| A source is `failed: blocked` / `challenge` | `RUN_SUMMARY.blocks`, `challenges` | Reddit refused the exit; `RETRY_INPUT` in the key-value store is a ready-to-run input for those sources |
| `detailLevel: "listing"` rows | run log | enrichment off, or the detail service was unavailable for the run |
| `text` is `""` | — | the post has no body (a title-only or link post); `null` would mean unknown |
| Rows have `matchedQueries: []` | `sourceInputs` | the post came from a subreddit or a direct URL, not a search |

### API and integrations

Call the Actor with the JSON input above through the Apify API, a webhook or a schedule. Small
runs are fine synchronously (`run-sync-get-dataset-items`); use asynchronous runs for hundreds of
posts. Rows are pushed as they are found, so a webhook on `ACTOR.RUN.SUCCEEDED` sees the whole
dataset and a partial run keeps what it captured. Three key-value records accompany every run:
`RUN_SUMMARY` (counts, stop reason, economics), `SOURCE_SUMMARY` (one record per source) and,
when relevant, `RETRY_INPUT`.

### Responsible use

This Actor collects public Reddit content only. You are responsible for lawful use, for privacy and
data-protection obligations, for Reddit's terms, and for how you handle personal data downstream.
Deleted or removed content is not reconstructed from anywhere.

### Changelog

- **1.0** — Initial release: subreddits, search queries and Reddit URLs in; one deduplicated post
  row out; global limit, incremental skip lists, `publishedAfter` boundary, full post details,
  six dataset views, pay per result.

# Actor input Schema

## `subreddits` (type: `array`):

Communities to scrape, one per line: python, r/Python, or https://www.reddit.com/r/python/. A subreddit URL with a sort segment (https://www.reddit.com/r/python/top/?t=week) keeps that sort and time window. r/all and r/popular are aggregate feeds, not communities, and are not supported.

## `searchQueries` (type: `array`):

Keywords or Reddit search syntax, one per line — best CRM, "HubSpot alternative", subreddit:python asyncio. Each query is a site-wide Reddit search.

## `startUrls` (type: `array`):

Direct post URLs (any host, redd.it short links, share links, t3\_ IDs or bare IDs), subreddit URLs, or Reddit search URLs (https://www.reddit.com/search/?q=vector+database, https://www.reddit.com/r/python/search/?q=asyncio\&restrict\_sr=1). User profiles belong to the Reddit User Scraper; comment threads to the Reddit Comments Scraper.

## `maxResults` (type: `integer`):

Global ceiling on unique posts delivered by this run, across every subreddit, query and URL. The run stops as soon as this many posts have been saved. Posts beyond it are never fetched and never charged.

## `sortBy` (type: `string`):

Order in which each source is read — and therefore which posts a limit keeps. 'new' is chronological and the right choice for scheduled runs. 'relevance' and 'comments' are search orderings; 'rising' and 'controversial' exist on subreddit feeds only and are substituted with 'hot' / 'top' where a surface lacks them (the run log says so).

## `timeFilter` (type: `string`):

Reddit's own time window for 'top', 'relevance' and 'comments' sorts.

## `publishedAfter` (type: `string`):

Keep only posts created on or after this moment. ISO date (2026-09-01, 2026-09-01T12:00:00Z) or a relative window: 24h, 7d, 2w, 3m, 1y. On a 'new'-sorted source the run stops paging as soon as it reaches older posts, which is what makes scheduled incremental runs cheap.

## `includeNsfw` (type: `boolean`):

Include posts flagged over-18. Off by default.

## `minScore` (type: `integer`):

Keep only posts with at least this score (upvotes minus downvotes).

## `minComments` (type: `integer`):

Keep only posts with at least this many comments.

## `publishedBefore` (type: `string`):

Keep only posts created on or before this moment (ISO date, or a relative window such as 7d).

## `flairs` (type: `array`):

Keep only posts whose flair matches one of these values (case-insensitive). Posts with no flair are dropped; posts whose flair is unknown are kept.

## `authors` (type: `array`):

Keep only posts by these usernames (case-insensitive, with or without u/).

## `includeKeywords` (type: `array`):

Keep only posts whose title or body contains at least one of these terms (case-insensitive substring match).

## `excludeKeywords` (type: `array`):

Drop posts whose title or body contains any of these terms.

## `postTypes` (type: `array`):

Keep only these kinds of post. Leave empty for all.

## `skipPostIds` (type: `array`):

Post IDs (with or without the t3\_ prefix) already collected by an earlier run. They are neither saved nor charged.

## `skipUrls` (type: `array`):

Post URLs already collected by an earlier run, in any form. Neither saved nor charged.

## `enrichPostDetails` (type: `boolean`):

Reddit's public listing carries title, author, community, score, comment count and timestamp. Leave this on to add the body text, flair, post type, media, outbound link, upvote ratio and edited/locked/stickied flags to every row, at no extra charge. With it on, the 'rising' and 'controversial' sorts are substituted with 'hot' / 'top'. Switch off for listing fields only.

## `maxResultsPerSource` (type: `integer`):

Fairness cap so one broad subreddit or query cannot fill the whole run. Reddit's public search returns at most ~250 posts per query or subreddit for a given sort and time window; split a source by time window to go deeper.

## `maxPagesPerSource` (type: `integer`):

Crawl ceiling per source. A page is 7 posts on Reddit's public search surface.

## `maxConcurrency` (type: `integer`):

How many sources to page at once (1–5). Each runs on its own proxy session with its own Reddit rate budget.

## `requestDelayMs` (type: `integer`):

Extra pacing on top of the built-in rate control. Raise it if Reddit starts answering with 429s.

## `debugMode` (type: `boolean`):

Verbose logging of pages, cursors and skips.

## `proxyConfiguration` (type: `object`):

Apify Datacenter proxy (the default) or no proxy both work. Apify RESIDENTIAL proxy is refused and fails the run at startup: residential bandwidth is billed to the Actor, while Reddit serves the same search results to datacenter addresses. To route through your own residential provider, use Custom proxy URLs — that traffic goes through your account and is honoured.

## Actor input object example

```json
{
  "subreddits": [
    "SaaS",
    "Entrepreneur"
  ],
  "searchQueries": [
    "best CRM",
    "HubSpot alternative"
  ],
  "startUrls": [],
  "maxResults": 200,
  "sortBy": "new",
  "timeFilter": "all",
  "includeNsfw": false,
  "flairs": [],
  "authors": [],
  "includeKeywords": [],
  "excludeKeywords": [],
  "postTypes": [],
  "skipPostIds": [],
  "skipUrls": [],
  "enrichPostDetails": true,
  "maxResultsPerSource": 250,
  "maxPagesPerSource": 20,
  "maxConcurrency": 3,
  "debugMode": false,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `posts` (type: `string`):

One flat row per unique post, whichever source found it.

## `searchResults` (type: `string`):

Posts found by a search query, with the queries that matched.

## `subredditPosts` (type: `string`):

Posts read from a subreddit feed.

## `highEngagement` (type: `string`):

Compact projection for finding the strongest posts.

## `linksMedia` (type: `string`):

Compact fields for link, image, video and gallery posts.

## `provenance` (type: `string`):

Which inputs produced each row, in capture order.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "subreddits": [
        "SaaS",
        "Entrepreneur"
    ],
    "searchQueries": [
        "best CRM",
        "HubSpot alternative"
    ],
    "maxResults": 200,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("coregent/reddit-posts-search-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "subreddits": [
        "SaaS",
        "Entrepreneur",
    ],
    "searchQueries": [
        "best CRM",
        "HubSpot alternative",
    ],
    "maxResults": 200,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("coregent/reddit-posts-search-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "subreddits": [
    "SaaS",
    "Entrepreneur"
  ],
  "searchQueries": [
    "best CRM",
    "HubSpot alternative"
  ],
  "maxResults": 200,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call coregent/reddit-posts-search-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,coregent/reddit-posts-search-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/gmlTgcdaTqsWG4rQ4/builds/gx5an8TQIgOK7qUWQ/openapi.json
