# Reddit Scraper — Posts, Comments, Communities & Users (`nice_dev/reddit-posts-scraper`) Actor

Scrape Reddit posts, full comment threads, communities and user profiles by community, keyword, user or URL. 179 fields per row including score, upvote ratio and flair. Export to JSON, CSV or Excel.

- **URL**: https://apify.com/nice\_dev/reddit-posts-scraper.md
- **Developed by:** [Nice Dev](https://apify.com/nice_dev) (community)
- **Categories:** Social media, Marketing, MCP servers
- **Stats:** 2 total users, 1 monthly users, 50.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.07 / 1,000 posts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### 💬 What is Reddit Posts Scraper?

**Reddit Posts Scraper** extracts **posts, comments, communities and user profiles from [Reddit](https://www.reddit.com)**: the whole **title and body text**, the **score, upvote ratio and comment count**, the **author**, the **flair**, every **image and video link**, and the full **comment tree** — one row per comment, with the post, the parent and the depth that rebuild the thread. It works as a **Reddit API alternative**: run it on a schedule, call it from your code, or plug it into Make, Zapier or n8n.

Type a **search term** (`rust async`), name a **community** (`programming`), paste any Reddit URL, or give a **username** — the four add up in a single run. Click **Start** and download the rows in JSON, CSV or Excel. No login, nothing to install, and it is **fast: about 1,000 posts in 40 seconds**, for **$0.10 per 1,000 posts** (less on paid plans).

### 📋 What data can you extract from Reddit?

Every row carries the same **227 fields**, whatever its kind, so a CSV export stays square: a field a kind of row does not have is `null`, never missing. The `dataType` column says what the row is — `post`, `comment`, `community` or `user`.

| Category | What you get |
| --- | --- |
| 📝 **Post** | title, whole text, link, community — `Good Tools Are Invisible` in `programming` |
| 👍 **Votes** | score, upvote ratio, number of comments, crossposts — `128`, `0.94`, `37` |
| 💬 **Comments** | one row per comment, with the post, the comment it answers and how deep it sits: the whole thread in a spreadsheet |
| 👤 **Author** | username, flair, badges; on a profile row, karma, avatar, bio, cake day and trophies |
| 🏘️ **Communities** | description, member count, icons and banners, what posting allows; as options, rules in full, pinned posts and wiki pages |
| 🖼️ **Images and videos** | every image at its largest size, the hosted video and its length, the outbound link |
| 🏷️ **Flair and flags** | flair with its colours, adult content, spoiler, pinned, locked, removed and why |
| 📈 **Computed for you** | age in hours, score and comments per hour, total engagement |
| 🕒 **Dates** | published, edited, account created — `2026-09-18T07:45:54.000Z`, what the date filters read |

Every field, with what it holds, is listed in the **Output** section below.

### ✅ Why use Reddit Posts Scraper?

- 🚀 **Fast, and you pay per row**: about 1,000 posts in 40 seconds, and you pay for the rows you get (plus a small fee per row a filter you set checks).
- 🧵 **The whole comment thread, one row per comment**: `postId`, `parentId`, `parentKind`, `depth` and `replyCount` rebuild the tree in any spreadsheet — no nested column nobody can open.
- 🧩 **Four sources in one run**: search terms, communities, usernames and pasted URLs add up. No mode to choose, no second run to launch.
- 📊 **Figures Reddit does not give**: age in hours, score per hour, comments per hour, total engagement and the comment-to-score ratio, computed on every row.
- 🔔 **Monitoring built in**: tick **New items only**, schedule the Actor, and each run returns — and charges — only what it has never delivered.
- 🗂️ **Communities and profiles too**: member counts, rules, karma and account age, in the same dataset and the same columns.
- 🔌 API, scheduling, webhooks and integrations (Make, Zapier, n8n, Google Sheets, Slack, Airtable…), plus JSON/CSV/Excel export, through the Apify platform.

### 🚀 How to scrape Reddit

1. Create a free Apify account.
2. Open **Reddit Posts Scraper** and type a **Search term** (for example `rust async`) or a community in **Communities** (for example `programming`).
3. Or paste your own Reddit URLs into **Reddit URLs**: a community, a sorted feed, a search results page, a user profile or a single post. Every filter the URL carries is kept and pagination is automatic.
4. Tick **Include comments** if you want the threads, set **Maximum results** (100 by default, 0 = no limit), then click **Start**.
5. Download the dataset in JSON, CSV, Excel or through the API.

### 💰 How much does it cost to scrape Reddit?

This Actor uses **pay per event** pricing. You pay per row, and only for the rows you get:

| What you get | Free plan | Bronze | Silver | Gold |
| --- | --- | --- | --- | --- |
| Post — per 1,000 rows | $0.100 | $0.098 | $0.089 | $0.068 |
| Comment — per 1,000 rows | $0.17 | $0.15 | $0.12 | $0.09 |
| Community — per 1,000 rows | $1.39 | $1.38 | $1.29 | $1.19 |
| User profile — per 1,000 rows | $1.39 | $0.99 | $0.98 | $0.97 |
| Community rules read (`includeRules`, `rules-read`) — per 1,000 communities | $1.89 | $1.88 | $1.87 | $1.86 |
| Pinned posts read (`includeStickiedPosts`, `stickied-posts-read`) — per 1,000 communities | $30.00 | $29.99 | $29.98 | $29.97 |
| Wiki list or page read (`includeWiki`, `wiki-page-read`) — per 1,000 reads | $0.90 | $0.89 | $0.88 | $0.87 |

Examples on the free plan: 1,000 posts cost **$0.10**; 500 posts with 2,000 of their comments cost **$0.39**; 100 communities cost **$0.139**.

Plus **$0.001 per run start**. The filters you set (dates, keywords, authors, flair, figures…) are applied by the Actor on every row it reads, kept or not: each row checked is charged a small filter fee (`filter-check`, see the Pricing tab). The two filters every run applies — adult content left out, and the kinds of row you did not ask for — are free. If Reddit turns away the default connection, the run moves to another one on its own; each request made that way is charged a small fee (`residential-fallback`, see the Pricing tab). The community options are charged on their own, only when you tick them and only for what they read (table above): `rules-read` per community whose rules are read, `stickied-posts-read` per community whose pinned posts are read, `wiki-page-read` per wiki list or page read — with 10 wiki pages, at most 11 reads per community. Platform usage (compute, proxy) is included in the price.

### ⚙️ Input

Search the whole of Reddit for a term:

```json
{
    "query": "rust async",
    "maxItems": 200
}
```

Two communities, the comment threads, only what is recent and only what has not been delivered before:

```json
{
    "subreddits": ["programming", "rust"],
    "searchComments": true,
    "includeComments": true,
    "maxCommentsPerPost": 50,
    "maxItemsPerQuery": 100,
    "postedAfter": "7 days",
    "onlyNew": true,
    "stateKey": "rust-daily"
}
```

Or your own URLs, a user and a single post:

```json
{
    "startUrls": [
        { "url": "https://www.reddit.com/r/programming/top/?t=week" },
        { "url": "https://www.reddit.com/r/rust/search?q=tokio" }
    ],
    "usernames": ["spez"],
    "postUrls": ["https://www.reddit.com/r/programming/comments/1wjjqc7/"],
    "maxItems": 500
}
```

Give at least one source: a search term, a community, a URL, a user or a post. With none, the Actor runs an example instead of stopping — this week's top of r/programming and a search for "web scraping", at most 20 rows — and says so in the log.

| Field | Notes |
| --- | --- |
| `query`, `searchQueries`, `strictSearch` | Free-text search over the whole of Reddit, or inside the communities below. `searchQueries` adds more terms, one search each; every term is searched in every community (max 500 searches per run). `strictSearch` wraps the term in quotes so Reddit matches the whole phrase. |
| `subreddits`, `mergeSubreddits` | Community names without the r/ prefix, for example `programming`. With a search term, each community is searched separately; alone, each one is read as a feed. Without a search term, `mergeSubreddits` reads them all as a single feed: far fewer requests, but Reddit's per-feed ceiling of about 1,000 posts then applies to the group as a whole. |
| `startUrls`, `ignoreStartUrls` | Any Reddit URL: a community, a sorted feed, a search results page (`https://www.reddit.com/r/rust/search?q=tokio`), a user profile or a single post. Filters carried by the URL are kept. `ignoreStartUrls` keeps them saved in the form but skips them for this run. |
| `usernames` | Accounts to read, without the u/ prefix, for example `spez`. Their posts, their comments when comment rows are wanted, and their profile when profile rows are wanted. |
| `postUrls` | Single posts to read with their thread, for example `https://www.reddit.com/r/programming/comments/1wjjqc7/`. |
| `searchPosts`, `searchComments`, `searchCommunities`, `searchUsers`, `includeTrophies` | Which kinds of row the run returns. They mix freely in one dataset and the dataType column tells them apart. Posts only by default. With `includeTrophies`, each user profile row also lists its trophies (one more request per profile). |
| `includeRules`, `includeStickiedPosts`, `includeWiki`, `maxWikiPages`, `maxWikiPageChars` | A full profile of each community row, off by default. `includeRules`: its rules, the titles in the rules column and, in the ruleDetails column, each rule's text, order, what it applies to and report reason (one more request per community). `includeStickiedPosts`: the posts pinned at its top (one more request). `includeWiki`: the name of every wiki page, and the text (Markdown) of up to `maxWikiPages` of them, index first, each cut at `maxWikiPageChars` characters (one request for the list, plus one per page read; settings pages are never read). |
| `sort` | Order the results come in: `new`, `hot`, `top`, `rising`, `controversial` for a community feed; `relevance`, `comments`, `new`, `top` for a search. An order a listing does not accept falls back to the closest one it does. |
| `timeFilter` | Window Reddit itself applies to the top and controversial orders, and to every search: `hour`, `day`, `week`, `month`, `year` or `all`. |
| `maxItems`, `maxItemsPerQuery` | Stop after this many rows for the whole run, and for each source separately. 0 means no limit. Set the per-source cap when you run several searches, so one busy community cannot eat the whole budget. |
| `maxPosts`, `maxComments`, `maxCommunities`, `maxUsers` | Caps per kind of row, on top of the global one. 0 means no separate cap. |
| `includeComments`, `maxCommentsPerPost`, `maxCommentDepth`, `commentSort`, `expandMoreComments` | Read the thread of each post, one row per comment. The two caps bound how many comments and how deep; the order is `top`, `new`, `confidence`, `controversial`, `old` or `qa`. With `expandMoreComments`, the comments Reddit folds behind "load more comments" are unfolded too, in their place in the thread (one more request per 100 of them). Off by default: it costs one more request per post and comments are charged on their own. |
| `postedAfter`, `postedBefore`, `commentedAfter`, `commentedBefore` | Publication date range: a date such as `2026-09-01`, or a period before now such as `7 days` or `24 hours`. The last two apply to comment rows only; without them the first two apply to comments as well. |
| `includeKeywords`, `excludeKeywords` | Keep only the rows whose title or text holds one of these words, and drop the ones that hold one of those. Case, accents and invisible characters are ignored, so `benchmark` matches a title written any way. |
| `flairs`, `onlyWithFlair`, `onlyWithMedia` | Keep only the posts with one of these flairs, for example `Discussion`; or any flair at all; or only the posts that carry an image or a video. |
| `authors`, `excludeAuthors` | Keep only the rows written by these accounts, for example `spez`, or drop the ones written by them, for example `AutoModerator`. |
| `includeNsfw`, `excludeStickied`, `skipRemoved` | Adult content is left out by default. You can also drop the posts pinned at the top of a community, and the rows whose text Reddit has removed or whose author deleted it. |
| `minScore`, `minComments`, `minAwards` | Thresholds on the figures; 0, the default, turns each one off — comments scored below 0 are then kept too: set `1` to keep only what scores 1 or more. A row whose score Reddit hides cannot be proven to reach the threshold, so it is dropped (it costs only its filter fee). |
| `onlyNew`, `stateKey`, `resetState` | Monitoring: only the rows never delivered under this memory key, for example `rust-daily`. `resetState` forgets the memory once. |
| `customLabels` | Free key/value pairs copied onto every row, so rows from several runs can be told apart in one spreadsheet. |
| `proxyConfiguration`, `maxConcurrency`, `maxRequestsPerMinute`, `minRequestIntervalMs`, `maxRequestRetries`, `debugLog` | Advanced. The proxy is included in the price: leave the default (the residential proxy is not available). The four numbers pace the run, and the last switch makes the log verbose. |

### 📦 Output

One row per post, comment, community or user profile. Here is a real post row, shortened to the columns most people use — the run returns all 227:

```json
{
    "dataType": "post",
    "id": "1wjjqc7",
    "fullId": "t3_1wjjqc7",
    "url": "https://www.reddit.com/r/programming/comments/1wjjqc7/",
    "title": "~2min bites C# courses that keep you sharp on code reviews",
    "body": "Lately, I've felt the desire to share more of what I've learned. And I must say, teaching is a skill that's much harder to develop than I expected… [shortened]",
    "postType": "link",
    "author": "Coding-Mojo",
    "authorId": "t2_2ihnd0fiwc",
    "subreddit": "programming",
    "subredditPrefixed": "r/programming",
    "subredditSubscribers": 6921153,
    "score": 0,
    "upvoteRatio": 0.24,
    "numComments": 9,
    "totalAwards": 0,
    "ageHours": 52.24,
    "scorePerHour": 0,
    "engagementTotal": 9,
    "wordCount": 93,
    "flair": null,
    "isNsfw": false,
    "isStickied": false,
    "isSelf": false,
    "domain": "youtu.be",
    "linkUrl": "https://youtu.be/J27gidcV8OA",
    "images": [],
    "hasMedia": false,
    "mediaType": "embed",
    "publishedAt": "2026-09-18T07:45:54.000Z",
    "isEdited": false,
    "searchTerm": null,
    "sourceUrl": "https://www.reddit.com/r/programming/new?page=1",
    "scrapedAt": "2026-09-20T12:00:00.000Z"
}
```

You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

#### All 227 fields

| Fields | What they give you |
| --- | --- |
| `dataType`, `id`, `fullId`, `url`, `permalink`, `postUrl`, `postId`, `parentId`, `parentKind`, `depth`, `replyCount` | **Identity and links.** The row kind (post, comment, community or user), its Reddit ids and the pages a human can open. A comment row also carries the post it belongs to, the thing it answers and how deep it sits. |
| `title`, `body`, `bodyHtml`, `postTitle`, `postAuthor`, `postScore`, `titleLength`, `bodyLength`, `wordCount` | **Text.** The title and the whole body of a post or a comment, as text and as HTML, with their lengths and word count. |
| `author`, `authorId`, `authorFlair`, `authorFlairBackgroundColor`, `authorFlairCssClass`, `authorFlairTextColor`, `authorFlairType`, `authorFlairRichtext`, `authorFlairTemplateId`, `authorPremium`, `authorPatreonFlair`, `authorIsBlocked`, `isSubmitter`, `distinguished` | **Author.** Who wrote it, their account id, their flair in that community and the badges Reddit shows next to the name. |
| `subreddit`, `subredditPrefixed`, `subredditId`, `subredditSubscribers`, `subredditType` | **Community of the row.** Which community the row comes from, its id, its member count at the time of the run and whether it is public or restricted. |
| `score`, `upVotes`, `downVotes`, `upvoteRatio`, `scoreHidden`, `numComments`, `numCrossposts`, `numDuplicates`, `controversiality`, `viewCount` | **Votes and replies.** Score, up and down votes, the ratio of upvotes, the number of comments, crossposts and duplicates. |
| `totalAwards`, `gilded`, `awards`, `allAwardings`, `topAwardedType` | **Awards.** How many awards the row got, which ones, and the raw award objects for anything else. |
| `aiAnalysis`, `ageHours`, `scorePerHour`, `commentsPerHour`, `engagementTotal`, `commentToScoreRatio`, `outboundUrlHost` | **Computed for you.** Figures Reddit does not give: the age in hours, score and comments per hour, total engagement, the comment-to-score ratio and the host a post links to. |
| `flair`, `flairBackgroundColor`, `flairCssClass`, `flairTextColor`, `flairType`, `flairRichtext`, `flairTemplateId`, `postType`, `contentCategories`, `discussionType`, `suggestedSort` | **Flair and kind of post.** The flair with its colours, and what the post really is: text, link, image, gallery, video or poll. |
| `isNsfw`, `isSpoiler`, `isStickied`, `isLocked`, `isArchived`, `isHidden`, `isQuarantined`, `isOriginalContent`, `isSelf`, `isVideo`, `isGallery`, `isMeta`, `isCrosspostable`, `isRobotIndexable`, `isCreatedFromAdsUi`, `contestMode`, `allowLiveComments`, `mediaOnly`, `sendReplies`, `noFollow`, `pinned`, `collapsed` | **Flags.** Every true or false Reddit publishes on a row: adult content, spoiler, pinned, locked, archived, original content, and the rest. |
| `collapsedReason`, `unrepliableReason`, `removedByCategory`, `removalReason`, `banReason`, `bannedBy`, `approvedBy`, `modNote`, `modReasonTitle`, `whitelistStatus`, `treatmentTags` | **Moderation.** Why a row was collapsed, removed or banned, and by whom, when Reddit says so publicly. |
| `domain`, `linkUrl`, `urlOverriddenByDest`, `thumbnail`, `images`, `galleryCount`, `videoUrl`, `videoDuration`, `hasMedia`, `mediaType`, `media`, `secureMedia`, `mediaMetadata`, `galleryData`, `preview`, `postHint`, `crosspostParent` | **Link and media.** The outbound link and its domain, the thumbnail, every image at its largest size, the hosted video with its length, and the raw media objects for embeds. |
| `publishedAt`, `editedAt`, `isEdited`, `createdAt` | **Dates.** When it was published, when it was last edited, and — on a community or a profile row — when the account was created. Always ISO 8601 in UTC. |
| `name`, `displayName`, `description`, `descriptionHtml`, `publicDescription`, `publicDescriptionHtml`, `membersCount`, `communityIcon`, `iconImg`, `bannerImage`, `bannerBackgroundImage`, `mobileBannerImage`, `headerImg`, `headerTitle`, `primaryColor`, `keyColor`, `lang`, `advertiserCategory`, `submissionType`, `submitText`, `suggestedCommentSort`, `emojisEnabled`, `commentScoreHideMins`, `rules`, `ruleDetails`, `wikiPageNames`, `wikiPages`, `stickiedPosts`, `similarCommunities`, `bannerBackgroundColor`, `bannerSize`, `headerSize`, `iconSize`, `emojisCustomSize`, `linkFlairPosition`, `userFlairPosition`, `submitLinkLabel`, `submitTextLabel`, `submitTextHtml` | **Community rows.** What a community row holds: its description, member count, icons and banners, colours, language and settings. Three options add more: its rules in full (title, text, what they apply to, report reason), the posts pinned at its top, and its wiki (every page name, and the text of the pages you choose). |
| `allowImages`, `allowVideos`, `allowVideogifs`, `allowGalleries`, `allowPolls`, `allowDiscovery`, `spoilersEnabled`, `linkFlairEnabled`, `originalContentTagEnabled`, `wikiEnabled`, `restrictPosting`, `restrictCommenting`, `acceptFollowers`, `allOriginalContent`, `allowPredictions`, `allowPredictionContributors`, `allowPredictionsTournament`, `allowTalks`, `allowedMediaInComments`, `commentContributionSettings`, `collapseDeletedComments`, `communityReviewed`, `disableContributorRequests`, `freeFormReports`, `hasMenuWidget`, `hideAds`, `isCrosspostableSubreddit`, `publicTraffic`, `shouldArchivePosts`, `shouldShowMediaInCommentsSetting`, `showMedia`, `showMediaPreview`, `userFlairEnabledInSr`, `canAssignLinkFlair`, `canAssignUserFlair` | **What a community allows.** The posting rules of a community: images, videos, galleries, polls, discovery, spoilers, flairs, wiki, followers. |
| `username`, `totalKarma`, `linkKarma`, `commentKarma`, `awardeeKarma`, `awarderKarma`, `isGold`, `isMod`, `isEmployee`, `isVerified`, `hasVerifiedEmail`, `hideFromRobots`, `snoovatarImg`, `profileTitle`, `profileDescription`, `profileUrl`, `bio`, `isCakeDay`, `hasSubscribed`, `previousNames`, `isSuspended`, `trophies` | **User profile rows.** What a profile row holds: the four karma counters, the badges, the avatar, the profile text and the cake day. |
| `rank`, `searchTerm`, `sourceUrl`, `customLabels`, `scrapedAt` | **Every row.** The search term that found it, the source it came from, your own labels and the moment it was read. |

### 💡 Tips

#### How to get more results

Reddit stops any single feed at about 1,000 posts, whatever the cap you set. To go past that, split the run: several communities instead of one, several search terms, or the same community read with `sort` set to `new`, `top` and `hot` in three runs, or `timeFilter` set to `week` and then `month`. Set `maxItems` to 0 to take everything a source has.

#### How to reduce costs

You pay per row, so the levers are the caps (`maxItems`, `maxItemsPerQuery`, and the four per-kind caps), the filters — a row a filter drops costs only its small filter fee — and `onlyNew` for recurring runs, which never charges the same row twice. Comments are charged on their own: leave **Include comments** off when you only want the posts.

#### The comment tree, in a spreadsheet

Each comment is a row of its own. `postId` says which post it belongs to, `parentId` and `parentKind` say what it answers (the post itself, or another comment), `depth` says how deep it sits and `replyCount` how many direct replies it got. Sort by `postId` then by the order the rows came in and you have the thread as it reads on the site. Use `maxCommentDepth` to keep only the top-level answers, and `maxCommentsPerPost` to bound the cost of a very busy thread.

#### Monitoring: only the new rows

Tick **New items only** (`onlyNew`) and schedule the Actor. The first run returns everything; each later run skips what it has already delivered: those rows are not saved, not charged, and the post's thread is not even opened. The memory lives in a named key-value store of your account (`reddit-posts-scraper-seen`, up to 150,000 rows per key) and is only updated with rows that really reached the dataset, so a failed run never hides anything. Give each schedule its own `stateKey` — two schedules sharing a key would hide each other's rows — and tick `resetState` once to start over. In the order `new` (the default), a search stops once it meets 200 rows in a row that you already have (fewer when **Maximum results** is lower); posts pinned at the top of a community are not counted in that run. Any other order is read as you chose it, without that early stop.

#### Filter by publication date

`postedAfter` and `postedBefore` take a date (`2026-09-01`, the whole day included, UTC) or a period before now (`7 days`, `2 weeks`, `1 month`, `24 hours`). They read `publishedAt` on a post and a comment, `createdAt` on a community and a profile; a row without a date is dropped as soon as a bound is set. `commentedAfter` and `commentedBefore` narrow the comment rows alone, which is how you get a recent discussion on an old post. Filtered-out rows are not saved and do not count in `maxItems` (each row checked costs the filter fee, see Pricing), and the run summary says how many there were. With `postedAfter` in the order `new` (the default), a search stops at the first page that is entirely too old; any other order is read on, its rows too old dropped.

### 🔌 Integrations and API

Call the Actor through the Apify API, the JavaScript or Python clients, or connect it with integrations and webhooks (Make, Zapier, n8n, Google Sheets, Slack, Airtable…). The dataset can be fetched as JSON or CSV from any tool.

### 🤖 Use with AI agents (MCP)

AI agents (Claude, ChatGPT, Cursor…) can find and run this Actor through the [Apify MCP server](https://mcp.apify.com), billed to their Apify account like any run. It returns one item per Reddit post, comment, community or user profile (`dataType` says which). Actor id: `nice_dev/reddit-posts-scraper`; MCP server with this Actor only: `https://mcp.apify.com/?tools=fetch-actor-details,nice_dev/reddit-posts-scraper`.

Smallest input, for a cheap first call:

```json
{
    "subreddits": ["programming"],
    "maxItems": 10
}
```

Key output fields: `dataType`, `url`, `title`, `body`, `author`, `subreddit`, `score`, `numComments`, `publishedAt` (and `postId` on a comment).

Cost per 1,000 rows: $0.10 for posts, $0.17 for comments, $1.39 for communities or user profiles, plus $0.001 per run start (Gold plan: $0.068, $0.09, $1.19, $0.97); the filters you set, the fallback connection and the community options (rules, pinned posts, wiki) cost extra, see the pricing section above. Cap each call with `maxItems` and, through the API, with the run option `maxTotalChargeUsd`.

### ❓ FAQ

#### Is it legal to scrape Reddit?

The Actor only reads what Reddit shows publicly to any anonymous visitor. It logs in to nothing, uses no account, and solves no captcha. Results contain content written by people under a pseudonym, and the profile rows contain personal data protected by GDPR: do not store it without a legitimate reason, and keep in mind that a pseudonym can still identify someone. You are responsible for using the data in compliance with Reddit's Terms of Use, its User Agreement and applicable law — in particular if you train a model on it or republish it. This Actor is not affiliated with Reddit.

#### Does it need a login or a proxy?

No login and no account. The proxy is included in the price: leave the default setting (the residential proxy is not available). A request Reddit turns away moves the run to a new IP with a new session and is tried again at once, without using up its retries (10 times at most per request), as long as the run has an IP left: 8 in all with the default setting, 10 with your own proxies. Past them, or without a proxy, it waits before each retry: 5 seconds, doubled at each try, 255 seconds at most.

#### Is the data safe to open in Excel or to show on a web page?

Titles and texts are what the authors wrote, copied as they are. A text can begin with `-`, `+`, `=` or `@`: Excel and Google Sheets may read such a cell of a CSV file as a formula or as a number. The Actor leaves the text as it is, so that the JSON and the API give the real value — when you open a CSV, import these columns as text. Every URL column holds an http(s) address or `null`, never anything a browser would execute. `bodyHtml`, `descriptionHtml` and the raw media columns hold the site's own HTML, which the Actor does not sanitize: escape every field like any text written by a stranger before you put it on a web page.

#### Known limitations

- Reddit stops any single feed or search at about 1,000 posts. The **Tips** section says how to go past it.
- Comment rows only come from the threads when **Include comments** is on; without it, comments only come from searches that ask for them.
- The thread of a very busy post is cut by `maxCommentsPerPost` and `maxCommentDepth`. The "load more comments" placeholders are followed only with `expandMoreComments`, up to 10 requests (1,000 folded comments) per post, as long as the post's own time allows; "continue this thread" links are not.
- `onlyNew` remembers row ids, not their content: a post whose score changed is not returned again.
- Two runs sharing the same `stateKey` at the same time may both return the same new row.
- A deleted account is written as `[deleted]` and a removed text as `[removed]`: that is what Reddit itself serves.

**A run that reaches its timeout** stops itself about 45 seconds before it: no new page is asked, what it read is saved and, with `onlyNew`, remembered, and the run ends *Succeeded* with "Stopped before the run's timeout". Resurrect it to go on from there, or give the next run a longer timeout (Run options).

**A run the platform stops without warning** (out of memory)

- Resurrect it: it goes on from where it stood at most a minute before the stop. What it had read since is read again, and the rows already saved are skipped: none is delivered or charged twice, and `maxItems` still counts them.
- With `onlyNew`, the memory is saved once a minute: resurrect the stopped run and the rows it had saved meanwhile join the memory; leave it stopped for good, and the next run may return up to a minute of them once more.

#### Something doesn't work?

The last line of the log counts the rows saved, the rows filtered out, the posts no longer on Reddit (deleted while the run was reading them) and the requests that failed after every retry. Those requests and the deleted posts are listed, with the reason, in the `FAILED_REQUESTS` record of the run's key-value store. A run that saved nothing and had failed requests fails, and its last message gives the cause — a private, members-only or banned community says so, instead of "run it again". A community name that does not exist is answered by Reddit with an empty page: the run then ends "No listings found", so check the spelling.

If Reddit changes its pages, you are told instead of paying for blank rows. A results page that counts rows but gives none the Actor can read is an error (listed in `FAILED_REQUESTS`), never a quiet "No listings found". If the first 20 rows read all lack their author, their community, their score or their publication date, the run saves nothing more, stops and fails, and its last message names the missing field: at most those first rows are charged. A row that `postedAfter` / `postedBefore` drops because it has no date at all counts among those 20.

### 🛟 Support

Open an issue in the **Issues** tab with a link to your run: the run log and the `FAILED_REQUESTS` record of the key-value store show exactly which URLs failed and why.

# Actor input Schema

## `startUrls` (type: `array`):

Paste any Reddit URL: a community (`https://www.reddit.com/r/programming/`), a sorted feed (`https://www.reddit.com/r/news/top/`), a search results page, a user profile (`https://www.reddit.com/user/spez/`) or a single post. Each URL is scraped with its own sort and time filter when the URL carries them. Max 1 000 URLs.

## `ignoreStartUrls` (type: `boolean`):

Keep the URLs saved in the form but skip them for this run, so you can try the other sources without deleting them.

## `subreddits` (type: `array`):

Community names without the r/ prefix, for example `programming` or `MachineLearning`. Scraped with the sort and time filter chosen below; crossed with the search terms above (max 500 searches per run). Several communities can be merged into a single faster feed with the Merge communities option.

## `query` (type: `string`):

Search Reddit for this term. Combined with Communities it searches inside those communities only; on its own it searches the whole site. For example `web scraping` or `rust async`.

## `searchQueries` (type: `array`):

Several search terms in one run. They are added to the single Search term above. Every term is crossed with every community, up to 500 searches per run.

## `strictSearch` (type: `boolean`):

Require the whole search term to appear as written, instead of letting Reddit match on separate words.

## `usernames` (type: `array`):

Reddit usernames without the u/ prefix. Their posts and, when comments are enabled, their comments are collected. For example `spez` — `u/spez` and a full profile URL are accepted too.

## `authors` (type: `array`):

Keep only items written by these usernames, without the u/ prefix. Leave empty to keep every author. For example `spez`.

## `postUrls` (type: `array`):

Direct links to single posts. Useful to re-scrape a known list of threads with their full comment tree. For example `https://www.reddit.com/r/programming/comments/1wjjqc7/`.

## `searchPosts` (type: `boolean`):

Include posts in the results.

## `searchComments` (type: `boolean`):

Include comments as their own rows, with parentId and depth so you can rebuild the thread. Comments of a post are only fetched when Include comments is on.

## `searchCommunities` (type: `boolean`):

Include one row per community with its subscriber count, description and settings (its rules too with the option below).

## `searchUsers` (type: `boolean`):

Include one row per user profile with karma, account age and verification flags.

## `includeTrophies` (type: `boolean`):

On each user profile row, add its trophies (name, description, date, icon). One more request per profile.

## `includeRules` (type: `boolean`):

On each community row, add its rules: the title of each (rules) and, in ruleDetails, its full text, its place in the order, what it applies to (posts, comments or all) and the reason shown when reporting it. One more request per community, charged on its own (rules-read, see Pricing).

## `includeStickiedPosts` (type: `boolean`):

On each community row, add the posts its moderators pinned at the top (title, id, link, author, date, score, comments). One more request per community, charged on its own (stickied-posts-read, see Pricing).

## `includeWiki` (type: `boolean`):

On each community row, add its wiki: the name of every page, and the text of up to "Maximum wiki pages" of them (index first). One request for the list, plus one per page read, each charged on its own (wiki-page-read, see Pricing).

## `maxWikiPages` (type: `integer`):

With "Include wiki": wiki pages whose text is read for each community, one request each. 0 = the page names only. Settings pages (sidebar, stylesheet, bots) are never read.

## `maxWikiPageChars` (type: `integer`):

With "Include wiki": longest text kept for a wiki page; a longer page is cut and marked "truncated".

## `sort` (type: `string`):

Order of the feed. Relevance, comments and top apply to searches; hot, new, top, rising and controversial apply to communities. A sort carried by a pasted URL wins over this setting.

## `timeFilter` (type: `string`):

Only used by the Top and Controversial sorts and by searches. For any other sort, use the Posted after / Posted before filters below.

## `maxItems` (type: `integer`):

Stop the run after this many rows in total, counting posts, comments, communities and profiles. 0 means no limit.

## `maxItemsPerQuery` (type: `integer`):

Cap applied to each source separately — a community, a search term in each community, a user's posts, a user's comments, a pasted URL — so one busy source cannot eat the whole run. 0 means no per-source cap. Reddit itself stops any single feed at about 1000 posts; to go past that, split the run by sort, by time window or by search term.

## `maxPosts` (type: `integer`):

Cap on post rows alone, on top of Maximum results. 0 means no separate cap.

## `maxComments` (type: `integer`):

Cap on comment rows for the whole run, on top of the per-post cap. 0 means no separate cap.

## `maxCommunities` (type: `integer`):

Cap on community rows. 0 means no separate cap.

## `maxUsers` (type: `integer`):

Cap on user profile rows. 0 means no separate cap.

## `mergeSubreddits` (type: `boolean`):

Ask Reddit for all the communities at once instead of one after another, when no search term is given (with one, each community is searched on its own). Much faster, with far fewer requests, but the 1000-post ceiling then applies to the merged feed as a whole, and Maximum results per source no longer separates them.

## `includeComments` (type: `boolean`):

Fetch the comment tree of every post. This costs one extra request per post and is billed as comment results.

## `maxCommentsPerPost` (type: `integer`):

Highest number of comments collected from a single post, replies included: 50 by default. 0 means no limit.

## `maxCommentDepth` (type: `integer`):

How deep into the reply chains to go. 1 keeps only top-level comments; 10, the default, follows almost every thread to the end.

## `expandMoreComments` (type: `boolean`):

Also read the comments Reddit folds behind « load more comments » (one more request per 100 of them, up to 10 per post), within the comments per post above.

## `commentSort` (type: `string`):

Order in which a post's comments are read. This decides which comments you get when Maximum comments per post cuts the tree short.

## `includeNsfw` (type: `boolean`):

Keep posts and communities marked NSFW. Off by default.

## `postedAfter` (type: `string`):

Keep only items published after this moment, read from the post or comment creation date. Absolute (2026-09-01 or a full ISO timestamp) or relative (7 days, 24 hours, 2 weeks). A date such as `2026-09-01`, or a period before now such as `7 days` or `24 hours`.

## `postedBefore` (type: `string`):

Keep only items published before this moment, read from the post or comment creation date. Same formats as Posted after. A date such as `2026-09-30`, or a period before now such as `2 days`.

## `commentedAfter` (type: `string`):

Applies to comment rows only, read from the comment date. Same formats as Posted after. Leave empty to date-filter posts and comments alike. A date such as `2026-09-01`, or a period before now such as `24 hours`.

## `commentedBefore` (type: `string`):

Applies to comment rows only, read from the comment date. Same formats as Posted before. A date such as `2026-09-30`, or a period before now such as `1 hour`.

## `includeKeywords` (type: `array`):

Keep an item only if its title or text contains at least one of these words. Case is ignored. For example `benchmark`.

## `excludeKeywords` (type: `array`):

Drop an item when its title or text contains any of these words. Case is ignored. For example `giveaway` or `upvote this`.

## `flairs` (type: `array`):

Keep only posts carrying one of these flairs. Leave empty to keep every post. For example `Discussion` or `Show & Tell`.

## `onlyWithFlair` (type: `boolean`):

Drop posts with no flair at all. Independent from the Flairs list above.

## `onlyWithMedia` (type: `boolean`):

Keep only posts carrying an image, a gallery or a video.

## `excludeAuthors` (type: `array`):

Usernames to skip, without the u/ prefix. Useful to drop bots and moderators. For example `AutoModerator`.

## `excludeStickied` (type: `boolean`):

Drop the announcement posts moderators pin at the top of a community.

## `minScore` (type: `integer`):

Keep only items whose score reaches this number. 0, the default, turns the filter off: comments scored below 0 are kept too. Set 1 to keep only what scores 1 or more; scores can be negative.

## `minComments` (type: `integer`):

Keep only posts with at least this many comments.

## `minAwards` (type: `integer`):

Keep only items that received at least this many awards.

## `skipRemoved` (type: `boolean`):

Drop posts and comments whose body is \[deleted] or \[removed]. Their metadata is still readable, which is why they are kept by default.

## `onlyNew` (type: `boolean`):

Return only items this Actor has never returned before under the same state key. Ideal for a scheduled run that should not repeat yesterday's posts. In the order New (the default), a search stops early once it meets a long run of items already returned.

## `stateKey` (type: `string`):

Name of the memory used by New items only. Use a different key for each monitoring job so they do not share their history. For example `rust-daily`.

## `resetState` (type: `boolean`):

Wipe the memory of this state key before starting, so everything counts as new again.

## `customLabels` (type: `object`):

Free key/value pairs copied onto every row, for example {"campaign": "q4-launch"}. Handy to tell several scheduled runs apart in one dataset.

## `proxyConfiguration` (type: `object`):

Leave as is. The Actor picks the right network itself and falls back automatically when a request is refused. The residential proxy is not available in this Actor: a run that asks for it stops at once.

## `maxConcurrency` (type: `integer`):

How many requests run at the same time. Raising it makes a run finish sooner but increases the chance of being throttled.

## `maxRequestsPerMinute` (type: `integer`):

Overall speed limit. The default is a comfortable pace; lower it if you run several jobs against Reddit at once.

## `minRequestIntervalMs` (type: `integer`):

Smallest gap between two requests, in milliseconds. Unlike the per-minute rate, this spreads the requests evenly instead of letting them go out in a burst. 2500 = Reddit's sustained quota of 24 requests a minute; 0 = no gap.

## `maxRequestRetries` (type: `integer`):

How many times a failed request is tried again before the Actor gives up on it. A request Reddit turns away is first tried again on a new IP without using up these retries, up to 10 times, as long as the run has an IP left (8 in all with the default proxy, 10 with your own).

## `debugLog` (type: `boolean`):

Write every request and every filtered item to the log. Turn on only when something looks wrong.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/programming/top/?t=week"
    }
  ],
  "ignoreStartUrls": false,
  "subreddits": [
    "programming"
  ],
  "query": "web scraping",
  "searchQueries": [],
  "strictSearch": false,
  "usernames": [],
  "authors": [],
  "postUrls": [],
  "searchPosts": true,
  "searchComments": false,
  "searchCommunities": false,
  "searchUsers": false,
  "includeTrophies": false,
  "includeRules": false,
  "includeStickiedPosts": false,
  "includeWiki": false,
  "maxWikiPages": 10,
  "maxWikiPageChars": 20000,
  "sort": "new",
  "timeFilter": "all",
  "maxItems": 100,
  "maxItemsPerQuery": 0,
  "maxPosts": 0,
  "maxComments": 0,
  "maxCommunities": 0,
  "maxUsers": 0,
  "mergeSubreddits": false,
  "includeComments": false,
  "maxCommentsPerPost": 50,
  "maxCommentDepth": 10,
  "expandMoreComments": false,
  "commentSort": "top",
  "includeNsfw": false,
  "includeKeywords": [],
  "excludeKeywords": [],
  "flairs": [],
  "onlyWithFlair": false,
  "onlyWithMedia": false,
  "excludeAuthors": [],
  "excludeStickied": false,
  "minScore": 0,
  "minComments": 0,
  "minAwards": 0,
  "skipRemoved": false,
  "onlyNew": false,
  "stateKey": "default",
  "resetState": false,
  "customLabels": {},
  "proxyConfiguration": {
    "useApifyProxy": true
  },
  "maxConcurrency": 2,
  "maxRequestsPerMinute": 24,
  "minRequestIntervalMs": 2500,
  "maxRequestRetries": 5,
  "debugLog": false
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.reddit.com/r/programming/top/?t=week"
        }
    ],
    "subreddits": [
        "programming"
    ],
    "query": "web scraping",
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("nice_dev/reddit-posts-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://www.reddit.com/r/programming/top/?t=week" }],
    "subreddits": ["programming"],
    "query": "web scraping",
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("nice_dev/reddit-posts-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.reddit.com/r/programming/top/?t=week"
    }
  ],
  "subreddits": [
    "programming"
  ],
  "query": "web scraping",
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call nice_dev/reddit-posts-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,nice_dev/reddit-posts-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/LcdIwSoKkvizEl4sU/builds/DSQnJZ05cFe84978t/openapi.json
