Reddit Search Scraper
Pricing
$1.00 / 1,000 results
Reddit Search Scraper
Search public Reddit posts by keyword, subreddit, or URL from public archives; export post metadata with optional archived comments.
Pricing
$1.00 / 1,000 results
Rating
0.0
(0)
Developer
MLG Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 days ago
Last modified
Categories
Share
Search public Reddit posts by phrase, community, or direct URL and export structured Reddit data to CSV, JSON, or Excel. The actor is designed as a Reddit search data source for post content, public engagement counts, and optional comments.
Current access status: The latest remote test returned no records. Reddit returned HTTP 403 for the search JSON route on both datacenter and residential connections, including after a browser challenge attempt. The actor fails visibly when every target is blocked. It is not ready for production collection until a public route is accessible again and a golden run passes.
What data can you extract from Reddit?
A post and each collected comment are separate dataset rows. kind distinguishes them. All rows use flat keys, so the same dataset can be exported without expanding nested objects. A null value means the source did not provide that field or the field does not apply to that record type. The descriptions below define the intended output contract; field availability has not been validated by a successful live run.
| Field | Description | Example shape |
|---|---|---|
kind | post or comment | "post" |
id | Stable identifier for the post or comment | "abc123" |
query | Search phrase that found the post | "coffee grinder" |
title | Post title; null for comments | "Choosing a grinder" |
body | Post text or comment body | "Looking for suggestions..." |
author | Public username as returned by Reddit | "example_user" |
subreddit | Community name | "Coffee" |
score | Public vote score | 42 |
upvoteRatio | Post upvote share, if supplied | 0.91 |
numComments | Public post comment count | 18 |
createdAt | Creation timestamp in UTC | "2026-09-01T12:00:00Z" |
url | Canonical Reddit page URL | "https://www.reddit.com/r/Coffee/comments/abc123/example/" |
permalink | Relative page path | "/r/Coffee/comments/abc123/example/" |
outboundUrl | Destination linked from a post | "https://www.reddit.com/..." |
flair | Public post flair text | "Question" |
domain | Link destination domain | "self.Coffee" |
thumbnail | Public thumbnail URL when present | "https://..." |
postHint | Content hint returned for a post | "image" |
isSelf | Whether the post is text only | true |
isVideo | Whether the post is a video | false |
isGallery | Whether the post is a gallery | false |
isNsfw | Whether the post is marked adult | false |
isSpoiler | Whether the post is marked as a spoiler | false |
isLocked | Whether replies are locked | false |
isStickied | Whether the item is pinned | false |
edited | Edit timestamp or false, as returned | false |
postId | Parent post identifier on comment rows | "abc123" |
postUrl | Parent post URL on comment rows | "https://www.reddit.com/..." |
parentId | Parent post or comment fullname | "t3_abc123" |
depth | Comment nesting level | 0 |
The url field points to a Reddit page. For a link post, outboundUrl points to its destination, which may be an external site. A text post may have a Reddit URL in both places. score and numComments can change after collection, so retain the run date when comparing exports. The query field records the search phrase for keyword inputs; it is empty for broad community listings and direct post links.
How to scrape Reddit
- Enter a phrase in
queries, a community insubredditName, or one or more Reddit links inurls. - Choose a sort and timeframe when searching. Set
maxPostsfor a per-target cap andmaxItemsfor a cap across the entire dataset. - Enable
scrapeCommentsif discussion text is needed, and setmaxCommentsto bound each thread. - Run the actor. If the site accepts the requests, inspect the dataset and export CSV, JSON, or Excel. If the site blocks every request, the run fails with a clear error and produces no rows.
Start with a small search and verify the first records before scheduling larger runs. Search ranking can change between runs, and the same post can be found through several queries. The actor deduplicates posts and comments by identifier within one run. Keep kind together with id when combining datasets from different runs, because a post ID and a comment ID are separate record types.
Input
| Name | Type | Default | Description |
|---|---|---|---|
queries | string array | none | Global post search phrases. |
subredditName | string | none | Community name without r/. |
subredditKeywords | string array | none | Phrases to search inside the selected community. |
urls | string array | none | Direct post, search, community, or user-submitted URLs. Takes priority over other targets. |
sort | string | relevance | Global ranking: relevance, hot, top, new, or comments. |
timeframe | string | all | Global search period: all, year, month, week, day, or hour. |
subredditSort | string | relevance | Ranking for community keyword searches. |
subredditTimeframe | string | all | Time period for community keyword searches. |
maxPosts | integer | 100 | Maximum accepted posts per input target. |
maxItems | integer | 0 | Total post and comment row cap; zero means no cap. |
scrapeComments | boolean | false | Collect public comments from accepted posts. |
maxComments | integer | 100 | Comment cap per post, subject to the initial thread response. |
dateFrom | date or timestamp | none | Earliest UTC post time to keep. |
dateTo | date or timestamp | none | Latest UTC post time to keep. |
commentDateFrom | date or timestamp | none | Earliest UTC comment time to keep. |
commentDateTo | date or timestamp | none | Latest UTC comment time to keep. |
includeNsfw | boolean | false | Include posts marked adult. |
strictTokenFilter | boolean | false | Require every query word in title, body, or outbound URL. |
proxyConfiguration | object | enabled | Connection settings; datacenter is attempted before residential. |
At least one target is needed: a query, a community, or a direct URL. When urls is present, it takes priority. A community without keywords uses its newest listing. The post date fields filter output after Reddit returns results; they do not make Reddit search enumerate every historical post in that range. The timeframe options are Reddit search controls. Enter dates as YYYY-MM-DD or an ISO timestamp. A plain start date begins at midnight UTC, while a plain end date includes that day's final second.
A bounded example input:
{"queries": ["coffee grinder"],"sort": "new","timeframe": "month","maxPosts": 40,"maxItems": 40,"scrapeComments": false}
This is the golden input. Its latest run returned zero rows because Reddit blocked the structured search route, so this input should not be treated as a successful sample. The requested minimum remains 30 posts to catch a regression that silently returns only a few rows.
Output example
No genuine output item is available from the golden run. The run returned HTTP 403 and an empty dataset. A sample row is intentionally omitted until a remote run produces one. The field table above documents the expected schema, while the current access status describes what has actually been verified.
When the route becomes accessible, every accepted post is written to the default dataset. With comment collection enabled, comments follow their parent post as additional rows. The kind and postId fields allow a consumer to separate posts from replies and join each reply to its thread. JSON keeps numbers, booleans, and nulls in their native form. Tabular exports expose the same keys as columns.
Use cases
- Community research: collect matching public posts from one or several communities, then group them by
subredditandquery. - Product feedback review: search for a product category, inspect post text, and compare recurring questions or complaints.
- Discussion analysis: include comments to inspect public replies alongside the post that started a thread.
- Trend monitoring: schedule the same query with
sort=newand retain stable IDs to distinguish newly seen posts from previously exported posts. - Content discovery: use
flair,postHint,isVideo, andisGalleryto separate different public post formats. - Engagement review: compare public
score,upvoteRatio, andnumCommentsacross an exported collection, with the understanding that these values change over time.
These workflows depend on public access. The current 403 response prevents collection from this account, so none of these uses is presently verified end to end. Check an initial dataset before building a downstream process around it. For research requiring complete historical coverage, Reddit search is an imperfect source even when accessible: ranking, moderation, deletion, and result windows can omit posts.
How much does it cost to scrape Reddit?
The configured event price is $1.00 per 1,000 saved rows. Posts and comments both count as rows. At that rate, 40 saved rows cost $0.04, 100 saved rows cost $0.10, and 1,000 saved rows cost $1.00 in result events. These are arithmetic examples, not measured successful runs. Additional platform usage or connection costs may apply according to the account configuration. The failed golden run saved zero rows, although the remote run still consumed platform resources.
Use maxItems to cap the number of emitted rows, especially when comments are enabled. maxPosts controls posts per search target, while maxComments controls the initial set of comments per accepted post. A 40-post run with comments enabled can produce more than 40 rows unless maxItems also limits the total. If budget control matters, set both limits explicitly and review the actual result count after each run.
Tips for best results
Use a specific query before trying a broad one. A short query may mix unrelated meanings, and a strict token filter can remove relevant posts that use different wording. Compare a small unfiltered run with a filtered run before depending on strict matching. For a known community, enter its name and one or more community keywords. Without community keywords, the actor requests the newest public listing rather than search results.
Use sort=new for recent monitoring, and combine it with a timeframe when appropriate. The separate dateFrom and dateTo settings remove records outside your exact date window after retrieval. They cannot recover posts missing from Reddit's returned pages. Increase maxPosts only after checking whether pagination exposes more unique posts. Search results may repeat, and the actor skips duplicate identifiers inside the same run.
Comment collection adds a thread request for each accepted post. Keep maxComments modest for a first run. The actor reads the initial public comment tree and follows replies included in that response; it does not expand additional comment placeholders. A post's public numComments may therefore exceed the number of comment rows saved. Removed comments, unavailable threads, and blocked thread requests can reduce the count further. If a thread request fails after the post was saved, the post remains in the dataset and the failure is logged.
Limits
Current blocking: Search pages returned challenges and structured search routes returned HTTP 403 in remote probes. The golden run confirmed the block on both datacenter and residential connections. This is the primary open issue. The actor does not treat a challenge page as a valid result. A run with no retrievable targets ends in failure rather than reporting a successful empty dataset.
Result coverage: Reddit search does not provide an unlimited historical scan. A query, sort, and timeframe combination can expose only a practical window of results. The actor follows JSON pagination cursors until the requested cap, an empty page, or a repeated cursor. It does not fan out across alternative sorts or automatically divide large date ranges. Some posts can be absent because of ranking, deletion, community restrictions, or moderation.
Access boundaries: Private communities, login-only content, removed material, and unavailable posts are outside the public-data contract. The actor does not use a user account. Public authorship fields can be deleted or anonymized by the site, and fields such as flair, thumbnails, media hints, and upvote ratio are optional. Engagement counts are snapshots, not permanent facts.
Comments: The initial thread response may omit deeper replies behind expansion placeholders. maxComments is a ceiling, not a guarantee. A direct comment permalink is treated as its parent thread. Separate comment date filters apply only when comment collection is enabled.
Filters: An exact post date range is checked after retrieval. If ranking does not surface enough posts from the selected dates, a high maxPosts value may still return fewer accepted rows than expected. includeNsfw=false excludes flagged posts; it does not inspect content for additional sensitive material.
Automated workflows
The dataset can be consumed through a dataset endpoint or downloaded after a run. A scheduled run can reuse the same input for periodic monitoring. For repeat collections, upsert by kind and id, and store each run's timestamp separately if engagement changes matter. A webhook can notify a downstream process when a run ends. Treat a failed run as an access failure and avoid interpreting zero rows as zero matching Reddit posts.
Two example workflow requests are: “Find recent public posts about coffee grinders and return their titles, communities, URLs, and comment counts,” and “Collect posts from one community this week, include up to 20 public replies per post, and export the dataset.” These describe intended usage only; the current blocked state must be resolved first.
FAQ
Is this collecting public data? The actor is designed for publicly accessible posts and comments. It does not sign in or request private communities. Use exported content in line with applicable law, site terms, and privacy obligations. Avoid using public usernames for unwanted contact or profiling.
Do I need to configure a proxy? The default connection setup first tries datacenter access and escalates after blocks. The latest test still received 403 responses after escalation. Changing the proxy setting is not a demonstrated fix.
How fast is a run? No successful collection time has been measured. Runtime depends on site access, number of pages, and whether each post requires a comment request. The failed 40-post test spent about a minute attempting the blocked source.
Can I schedule repeated runs? Yes, the input is suitable for a recurring schedule, but scheduling a currently blocked source will repeat the failure. First confirm a successful small run.
Can I export to a spreadsheet? Dataset rows can be downloaded as CSV or Excel as well as JSON. The latest golden dataset is empty, so export currently contains no useful records.
Why is a field empty? Some values apply only to posts or only to comments. Other fields are optional in Reddit's public response. A null is distinct from a zero score or a false flag.
Why are there fewer comments than the displayed count? The first thread response can omit replies behind expansion placeholders, and the actor respects the configured comment cap. Removed or inaccessible replies also reduce exported rows.
Integrations
Use dataset exports, an API endpoint, scheduling, and webhooks to move successful records into a reporting or storage workflow. Use the stable record identifiers for upserts, and monitor failed runs so a site block does not appear as an ordinary quiet period. Integrations should only be enabled after a live run produces the expected rows.
Support
Open an issue with the run ID, input with sensitive values removed, the observed error, and the expected result. The current reproducible issue is HTTP 403 on anonymous Reddit JSON search requests.
