Reddit Posts & Search Scraper avatar

Reddit Posts & Search Scraper

Pricing

from $1.00 / 1,000 post results

Go to Apify Store
Reddit Posts & Search Scraper

Reddit Posts & Search Scraper

Search public Reddit posts by keyword, scrape subreddit feeds, or paste post and search URLs. One deduplicated row per post with title, text, author, community, score, comments, flair, media, timestamps and source lineage. One global limit. No Reddit login or API key.

Pricing

from $1.00 / 1,000 post results

Rating

0.0

(0)

Developer

Delowar Munna

Delowar Munna

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Reddit Posts & Search Scraper

Search public Reddit posts by keyword, scrape subreddit feeds, or paste known Reddit post/search URLs. Every match becomes one clean, deduplicated post row with text, author, community, engagement, flair, media, timestamps and source lineage. No Reddit account, cookies or API key required.

  • One global limit. maxResults bounds the whole run, across every subreddit, query and URL.
  • One row per post, whichever source found it. A post matched by two queries and its subreddit is delivered once, with every matching query and community on the row.
  • Built for scheduled runs. Skip lists for IDs and URLs, a publishedAfter boundary that stops paging a new feed the moment it reaches old posts, and a chronological sort by default.
  • Honest about depth. Reddit's public search returns at most about 250 posts per query or community for one sort and time window. The run says so per source instead of paging in circles.

Comments are a separate Actor: Reddit Comments Scraper takes the url or postId from these rows. User histories belong to the Reddit User Scraper.

Quick start

Subreddits and keyword searches together:

{
"subreddits": ["SaaS", "Entrepreneur"],
"searchQueries": ["best CRM", "HubSpot alternative"],
"maxResults": 200,
"sortBy": "new",
"timeFilter": "month",
"includeNsfw": false,
"enrichPostDetails": true
}

Four sources, 200 posts at most, newest first, deduplicated across all four. The run summary in the key-value store says how many candidates each source produced and where each one stopped.

Reddit URLs, with filters:

{
"startUrls": [
"https://www.reddit.com/r/python/search/?q=asyncio&restrict_sr=1",
"https://www.reddit.com/r/AskReddit/comments/1wjydiy/",
"https://redd.it/1wjgckp"
],
"maxResults": 100,
"minScore": 10,
"postTypes": ["text"],
"excludeKeywords": ["hiring"],
"includeNsfw": false,
"enrichPostDetails": true
}

A search restricted to r/Python plus two direct posts. The search URL carries its own query; the two post URLs are looked up directly. Posts under 10 points, non-text posts and posts mentioning "hiring" are dropped before they are saved, and are not charged.

A daily scheduled run:

{
"subreddits": ["personalfinance", "legaladvice"],
"sortBy": "new",
"publishedAfter": "24h",
"maxResults": 500,
"skipPostIds": ["1wquuj6", "1wp6zmh"],
"includeNsfw": false,
"enrichPostDetails": true
}

Each community stops paging once it reaches posts older than a day, and the IDs from the previous run are skipped. See Incremental scheduled runs.

Search Reddit

Put keywords or Reddit search syntax in searchQueries, one query per line:

{ "searchQueries": ["open source LLM", "\"vector database\" subreddit:MachineLearning"] }

Each query is a site-wide Reddit search. sortBy chooses the ordering Reddit applies (relevance, hot, top, new, comments), timeFilter the window. A search URL in startUrls works the same way and keeps the sort and window it carries.

Scrape subreddits

Put community names in subreddits in any spelling — python, r/Python, https://www.reddit.com/r/python/ — or a subreddit URL with a sort segment such as https://www.reddit.com/r/python/top/?t=week, which keeps that sort and window for that source only. rising and controversial are feed orderings; where a surface lacks them the run substitutes hot / top and says so in the log.

r/all and r/popular are aggregate feeds rather than communities and are not supported.

Direct posts and search URLs

startUrls accepts every common form of a post reference — full post URLs on any Reddit host, redd.it short links, /r/<sub>/s/ share links, t3_ fullnames and bare IDs — plus subreddit URLs and search URLs (https://www.reddit.com/search/?q=…, https://www.reddit.com/r/python/search/?q=…&restrict_sr=1). Anything else (a user profile, a wiki page, a multireddit) is reported as invalid with a reason, never guessed at.

Input reference

FieldDefaultWhat it does
subreddits[]Communities to scrape (names, r/name, or URLs)
searchQueries[]Site-wide keyword searches
startUrls[]Post, subreddit or search URLs
maxResults500Global ceiling on unique posts delivered by the run
sortBynewnew, hot, top, relevance, comments, rising, controversial
timeFilterallhour … all; applies to top, relevance, comments
publishedAfter—ISO date or relative window (24h, 7d, 2w); stops a new feed at the boundary
publishedBefore—Upper date bound
includeNsfwfalseInclude over-18 posts
minScore, minComments—Engagement floors
flairs, authors[]Allow-lists (case-insensitive)
includeKeywords, excludeKeywords[]Substring filters on title and body
postTypes[]text, link, image, video, gallery, poll, crosspost
skipPostIds, skipUrls[]Incremental skip lists — never saved, never charged
enrichPostDetailstrueFull post details on every row — body, flair, post type, media and flags — at no extra charge
maxResultsPerSource250Fairness cap per source
maxPagesPerSource20Crawl ceiling per source (7 posts per page on Reddit's public search)
maxConcurrency3Sources paged in parallel (1–5)
requestDelayMs—Extra pacing on top of the built-in rate control
debugModefalseVerbose logging
proxyConfigurationApify DatacenterSee the proxy policy below

At least one of subreddits, searchQueries or startUrls is required. Every filter keeps a post whose value Reddit does not publish — an unknown score is not a low score.

Global max and cost control

You are charged once per unique post saved to the dataset, and nothing else: not duplicates across sources, not posts your filters removed, not skip-list matches, not posts discovered past the limit, not invalid inputs, not sources that failed, not empty searches, not the run summaries. maxResults is the only number you need to bound a run. If your Apify account has a per-run spending limit, the run stops fetching the moment it is met rather than working on unpaid.

Incremental scheduled runs

  1. Sort by new (the default).
  2. Set publishedAfter to the interval — 24h for a daily schedule, 7d for weekly. Each new feed stops paging as soon as it reaches posts older than that, so a quiet subreddit costs one page.
  3. For exact de-duplication across runs, pass the previous run's postId values as skipPostIds (or its url values as skipUrls). Skipped posts are neither saved nor charged.

Overlapping sources within one run are deduplicated automatically; the matchedQueries, matchedSubreddits and sourceInputs fields on each row say which inputs found it.

Output fields

One flat row per post, recordType: "post".

FieldMeaning
postId, postFullIdReddit's ID and t3_ fullname — the dedupe key
url, permalinkCanonical post URL and its path
title, textTitle and body text (plain text; "" for a post with no body)
authorUsername, authorIdAuthor (null when deleted)
subreddit, subredditPrefixed, subredditIdCommunity
flairPost flair text
score, upvoteRatio, commentCount, awardCountEngagement
createdAt, editedAtISO timestamps
isNsfw, isSpoiler, isStickied, isLocked, isArchived, isQuarantinedSubreddit, isDeletedOrRemovedFlags
postTypetext, link, image, video, gallery, poll, crosspost
outboundUrl, domain, thumbnailUrl, mediaUrls, crosspostParentIdLink and media
detailLevelfull (all fields) or listing (see below)
matchedQueries, matchedSubreddits, sourceInputsWhich inputs found the post
capturedSequence, scrapedAtCapture order and time

null always means "not published by the source", never a guess.

Listing rows. Reddit's public listing carries the identity, engagement and timestamp fields only. With enrichPostDetails on (the default) the Actor reads every source through a search service that returns full records, looks up direct post URLs one by one, and every row is detailLevel: "full". With it off, or when the detail service is unavailable for a run, rows are detailLevel: "listing" and text, flair, postType, upvoteRatio, editedAt, the media fields and the stickied/locked flags are null. The run log says which applied.

Dataset views and sample records

Reddit Posts & Search Scraper — Posts view, table (one row per post with community, author, score, comments, flair and post type)

The dataset has six views. They are projections of the same dataset, so no row appears twice. Below is one real record for each view, from a 150-post run over r/personalfinance, r/legaladvice, two keyword searches, an r/Python search URL and two direct posts. Usernames are anonymised.

Posts — every field

The full schema. This is a text post found by the r/Python search URL:

{
"recordType": "post",
"postId": "1w43div",
"postFullId": "t3_1w43div",
"url": "https://www.reddit.com/r/Python/comments/1w43div/when_do_you_prefer_asynciosemaphore_over_an/",
"permalink": "/r/Python/comments/1w43div/when_do_you_prefer_asynciosemaphore_over_an/",
"title": "When do you prefer asyncio.Semaphore over an asyncio.Queue for limiting concurrency?",
"text": "I've been thinking about concurrency control in asyncio.\n\nA common pattern for limiting concurrent work is:\n\n sem = asyncio.Semaphore(10)\n\n async with sem:\n await do_work()\n\nBut in many cases, couldn't the same problem be modeled by putting work into an asyncio.Queue and running a fixed number of worker tasks?\n\nI'm curious how experienced Python developers decide between the two approaches.\n\nAre there real-world situations where a semaphore is clearly the better abstraction than a worker queue? Are there meaningful differences in cancellation behavior, backpressure, fairness, task lifetime, or code complexity?\n\nI'd especially be interested in examples from production async Python code.",
"authorUsername": "example_author",
"authorId": "t2_example",
"subreddit": "Python",
"subredditPrefixed": "r/Python",
"subredditId": "t5_2qh0y",
"flair": "Discussion",
"score": 64,
"upvoteRatio": 0.8809523809523809,
"commentCount": 13,
"createdAt": "2026-09-01T06:08:20.940Z",
"editedAt": null,
"isNsfw": false,
"isSpoiler": false,
"isStickied": false,
"isLocked": false,
"isArchived": false,
"isQuarantinedSubreddit": false,
"isDeletedOrRemoved": false,
"postType": "text",
"outboundUrl": null,
"domain": "self.Python",
"thumbnailUrl": null,
"mediaUrls": [],
"awardCount": 0,
"crosspostParentId": null,
"detailLevel": "full",
"matchedQueries": [
"asyncio"
],
"matchedSubreddits": [],
"sourceInputs": [
"https://www.reddit.com/r/python/search/?q=asyncio&restrict_sr=1"
],
"capturedSequence": 32,
"scrapedAt": "2026-09-27T00:36:57.450Z"
}

Search results

Rows found by a query, with the queries that matched:

{
"matchedQueries": [
"HubSpot alternative"
],
"postId": "1wp6zmh",
"title": "a competitor shut down and I pulled an all-nighter. I have no idea what I'm doing",
"subredditPrefixed": "r/SaaS",
"authorUsername": "example_author",
"score": 122,
"commentCount": 67,
"createdAt": "2026-09-24T16:53:07.516Z",
"flair": null,
"postType": "text",
"url": "https://www.reddit.com/r/SaaS/comments/1wp6zmh/a_competitor_shut_down_and_i_pulled_an_allnighter/"
}

Subreddit posts

Rows read from a community feed:

{
"matchedSubreddits": [
"personalfinance"
],
"postId": "1wquuj6",
"title": "Deceased relative had 30,000 employee stock options I found in old SEC filings. How do I trace what happened to them?",
"authorUsername": "example_author",
"score": 835,
"commentCount": 100,
"createdAt": "2026-09-26T16:30:25.431Z",
"flair": "Employment",
"postType": "text",
"isStickied": false,
"url": "https://www.reddit.com/r/personalfinance/comments/1wquuj6/deceased_relative_had_30000_employee_stock/"
}

High engagement

Score, upvote ratio, comments and awards, compact. This is a direct post URL:

{
"score": 10468,
"upvoteRatio": 0.9604797483287456,
"commentCount": 6980,
"awardCount": 7,
"title": "What warning signs of a worsening economy are you seeing in your line of work that the rest of us wouldn’t notice?",
"subredditPrefixed": "r/AskReddit",
"authorUsername": "example_author",
"createdAt": "2026-09-18T18:33:12.419Z",
"url": "https://www.reddit.com/r/AskReddit/comments/1wjydiy/what_warning_signs_of_a_worsening_economy_are_you/"
}

Outbound URL, domain, thumbnail and media URLs for link and media posts:

{
"postType": "image",
"title": "awesome-jev-projects - a curated directory of ~600 open-source tools built on the idea of using a fast, cheap typed-decision model for the small choices in agent loops instead of burning a full reasoning LLM on every branch",
"outboundUrl": "https://i.redd.it/svq1kaiagbrh1.png",
"domain": "i.redd.it",
"thumbnailUrl": "https://preview.redd.it/svq1kaiagbrh1.png?width=140&height=124&auto=webp&s=7015e3a4f80369130d591a51f65e52a687c3ddff",
"mediaUrls": [
"https://preview.redd.it/svq1kaiagbrh1.png?auto=webp&s=0286533ed115c57287e89965f46c0ee74b8e694b",
"https://preview.redd.it/svq1kaiagbrh1.png?width=108&auto=webp&s=1656173e50ad559a030fe2db72d0eaa5697c2213",
"https://preview.redd.it/svq1kaiagbrh1.png?width=216&auto=webp&s=48c5e7142c51d3e514bfdec27727689def7088d1",
"https://preview.redd.it/svq1kaiagbrh1.png?width=320&auto=webp&s=73b7be4cd11fdb42d37291998b58b5364465fa50",
"https://preview.redd.it/svq1kaiagbrh1.png?width=640&auto=webp&s=022797265457091698183d9ad1cd84eef6dda302"
],
"subredditPrefixed": "r/BestGitHubRepos",
"score": 51,
"createdAt": "2026-09-23T18:50:48.706Z",
"url": "https://www.reddit.com/r/BestGitHubRepos/comments/1woeo4l/awesomejevprojects_a_curated_directory_of_600/"
}

Provenance

Which inputs produced each row, in capture order. This row came from a redd.it short link:

{
"capturedSequence": 16,
"postId": "1wjgckp",
"sourceInputs": [
"https://redd.it/1wjgckp"
],
"matchedQueries": [],
"matchedSubreddits": [],
"detailLevel": "full",
"scrapedAt": "2026-09-27T00:36:54.494Z"
}

Pricing

Pay per result: one post-result event per unique post saved, with full post details included. There is no start fee and no charge for an empty run. The rate per post depends on your Apify plan. This README deliberately quotes no figures — the Pricing tab shows the live rate for each plan.

Limits and Reddit search depth

  • Reddit's public search returns at most about 250 posts per query or community for one sort and time window (measured: 36 pages of 7, then no next page). A source that hits it is reported as truncationReason: "search-depth-limit" in SOURCE_SUMMARY. To go deeper, split the source: sortBy: "top" with timeFilter: "week", "month", "year" are three different 250s.
  • maxResultsPerSource defaults to that ceiling.
  • The search surfaces have no rising or controversial ordering. With full post details on (the default) those sorts are substituted with hot / top and the run log says so; with details off they are honoured where a subreddit's own feed supports them.
  • Private, quarantined-behind-a-wall and banned communities answer "not found".

No-login / proxy posture

The Actor asks Reddit's public, server-rendered search surface as a plainly named client — no account, no cookies, no API key, no browser. It paces itself to Reddit's published logged-out budget and rests an exit that draws a challenge page.

🚦 Proxy policy

Use Apify Datacenter proxy (the default) or no proxy. Both work for Reddit's public search surface at this Actor's conservative concurrency.

Apify Residential proxy is not supported. The run fails at startup if apifyProxyGroups includes RESIDENTIAL: in pay-per-event Actors, residential bandwidth is billed to the developer rather than to the run, and Reddit serves the same results to datacenter addresses. If you genuinely need residential routing, supply your own provider under Custom proxy URLs — that traffic goes through your account and is honoured in full:

http://user:pass@proxy.iproyal.com:12321
http://user:pass@proxy.brightdata.com:22225
http://user:pass@proxy.oxylabs.io:7777

Troubleshooting

SymptomWhere to lookLikely cause
Fewer posts than maxResultsSOURCE_SUMMARY.truncationReasonsearch-depth-limit (Reddit's ceiling), publishedAfter, maxResultsPerSource, or the sources simply ran out
A source is failed: not-foundrun logthe community does not exist, is private, or the search URL had no query
A source is failed: blocked / challengeRUN_SUMMARY.blocks, challengesReddit refused the exit; RETRY_INPUT in the key-value store is a ready-to-run input for those sources
detailLevel: "listing" rowsrun logenrichment off, or the detail service was unavailable for the run
text is ""—the post has no body (a title-only or link post); null would mean unknown
Rows have matchedQueries: []sourceInputsthe post came from a subreddit or a direct URL, not a search

API and integrations

Call the Actor with the JSON input above through the Apify API, a webhook or a schedule. Small runs are fine synchronously (run-sync-get-dataset-items); use asynchronous runs for hundreds of posts. Rows are pushed as they are found, so a webhook on ACTOR.RUN.SUCCEEDED sees the whole dataset and a partial run keeps what it captured. Three key-value records accompany every run: RUN_SUMMARY (counts, stop reason, economics), SOURCE_SUMMARY (one record per source) and, when relevant, RETRY_INPUT.

Responsible use

This Actor collects public Reddit content only. You are responsible for lawful use, for privacy and data-protection obligations, for Reddit's terms, and for how you handle personal data downstream. Deleted or removed content is not reconstructed from anywhere.

Changelog

  • 1.0 — Initial release: subreddits, search queries and Reddit URLs in; one deduplicated post row out; global limit, incremental skip lists, publishedAfter boundary, full post details, six dataset views, pay per result.