Reddit Scraper – Posts, Comments & Search avatar

Reddit Scraper – Posts, Comments & Search

Pricing

from $0.49 / 1,000 reddit results

Go to Apify Store
Reddit Scraper – Posts, Comments & Search

Reddit Scraper – Posts, Comments & Search

Extract Reddit posts and comments from subreddits, searches and post URLs. Get text, scores, timestamps, media links and reply relationships. Export JSON, CSV or Excel. Provider access included; pay per saved result.

Pricing

from $0.49 / 1,000 reddit results

Rating

0.0

(0)

Developer

mscraper

mscraper

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

What does Reddit Scraper do?

Collect Reddit posts and comments as structured data from subreddit pages, keyword searches, and individual discussion URLs. Extract post text, comment bodies, scores, timestamps, media links, author names, and parent/reply relationships from Reddit.

Start with the default input to collect 10 recent posts from r/programming. Provider access is included: you do not need a Reddit account, an API key, cookies, or a proxy subscription. Run through Apify Console or the API, schedule recurring runs, and connect the dataset to your existing workflow.

Why use Reddit Scraper?

  • Topic research: collect recent or top posts from a specific subreddit.
  • Brand and product research: search post titles and text for keywords, optionally within one community.
  • Discussion analysis: collect comments from selected posts and retain the post and parent IDs.
  • Data pipelines: export consistent records without parsing Reddit pages yourself.

Posts and comments share one dataset and are distinguished by dataType. Missing provider fields are null; media collections are arrays. Counts and scores are snapshots, and can change after the run.

How to use Reddit Scraper

  1. Open the Input tab and enter a subreddit URL, one or more post URLs, or search terms.
  2. Choose the sort order and maximum total results. The default collects posts only.
  3. To include discussion replies, turn off Skip comments and set Maximum comments per post. Increase the total result limit to leave room for both posts and comments.
  4. Set your maximum run charge in Apify, then start the Actor.
  5. Open the Dataset output to inspect and export the results. Run summary reports saved counts, provider attempts, warnings, and why the run stopped.

Input

The Input tab contains all supported options. A small starting input:

{
"startUrls": [{ "url": "https://www.reddit.com/r/programming/" }],
"sort": "new",
"maxItems": 10,
"maxPostCount": 10,
"skipComments": true,
"maxRequests": 10,
"maxRetries": 1
}

For keyword search, leave startUrls empty and use:

{
"startUrls": [],
"searches": ["typescript"],
"searchCommunityName": "programming",
"sort": "top",
"time": "month",
"maxItems": 10,
"skipComments": true,
"maxRequests": 3
}
OptionMeaning
startUrlsHTTPS Reddit subreddit or post URLs. Profile URLs are not supported.
searches, searchCommunityNamePost keyword searches, optionally restricted to a subreddit.
sortSubreddits: new, hot, top. Searches also support relevance and comments.
timehour, day, week, month, year, or all; applied only with top sorting.
maxItemsGlobal cap on saved posts plus comments across all sources.
maxPostCountGlobal cap on saved posts.
maxCommentsMaximum saved comments per post, including returned replies.
skipCommentsDefault true; set false to retrieve comments.
commentSorttop, new, confidence, controversial, old, or qa.
postDateLimit, commentDateLimitOldest accepted date: ISO date or a relative value such as 7 days.
includeNSFWInclude returned posts marked NSFW. Does not bypass access restrictions.
maxRequests, maxRetriesBound provider calls and transient retries. Every retry counts toward maxRequests.

Date and NSFW filters apply to the data returned by the provider. Restrictive filters can consume requests without producing records. Earlier sources may consume the global result budget before later sources are visited. time is not an exact local date filter.

Output

An illustrative post record (synthetic values; shortened for readability):

{
"dataType": "post",
"id": "t3_example123",
"parsedId": "example123",
"url": "https://www.reddit.com/r/programming/comments/example123/example/",
"username": "sample_user",
"title": "A discussion about TypeScript",
"communityName": "r/programming",
"parsedCommunityName": "programming",
"body": "Example post text.",
"upVotes": 42,
"numberOfComments": 12,
"createdAt": "2026-09-25T12:00:00.000Z",
"scrapedAt": "2026-09-27T00:00:00.000Z",
"imageUrls": [],
"videoUrls": []
}

Comments use dataType: "comment", IDs starting with t1_, and postId/parentId to identify their discussion and immediate parent. A parent comment may be outside your requested result limit. The dataset contains a full flat record for each item; an absent title on a comment is null.

You can download the dataset in JSON, HTML, CSV, or Excel. Use JSON when you need media arrays and explicit null values. Exporting the same dataset again does not create new scraped-result events.

Data table

FieldsTypeDescription
dataType, id, parsedIdstringRecord type, full Reddit ID and ID without prefix.
urlstring or nullReddit permalink.
username, userId, authorFlairstring or nullAvailable author information. Deleted authors may be absent.
title, body, htmlstring or nullTitle, text and provider HTML. Treat HTML as untrusted content before displaying it.
communityName, parsedCommunityName, categorystring or nullCommunity with/without r/; category retains the subreddit name.
link, flairstring or nullPost destination and flair.
postId, parentIdstring or nullComment relationships.
numberOfComments, numberOfRepliesinteger or nullCounts supplied by the provider, not inferred from partial results.
upVotes, upVoteRatiointeger/number or nullProvider ups and ratio; not an independent count of all votes.
isVideo, isAd, over18boolean or nullProvider flags; unknown flags remain null.
thumbnailUrl, imageUrls, videoUrlsURL or URL arraysAvailable media links. Media files are not downloaded.
createdAt, scrapedAtISO timestamp or nullCreation time when available; extraction time is always present.

Shared usage protection

All cloud requests, including retries, reserve a shared provider allowance before dispatch. Defaults: 140,000 provider units per 30-day safety window, 1,000 units per paid user per 24-hour window, and 3 units per Free user per 30-day window. All Free users share a pool of 100 units. Windows begin with first usage and do not mirror subscription resets. Concurrent runs share the same counters. Contact the developer for larger allowances. These safety caps can stop a run before its requested limits.

Pricing / Cost estimation

$0.49 per 1,000 saved results, plus $0.01 per Actor start at the default 256 MB memory. One result is one unique post or one unique comment saved during the run. Provider access is included; no separate provider subscription is required.

Saved resultsEstimated charge at default memory
10 posts$0.01490
1 post + 20 comments$0.02029
100 posts + 900 comments$0.50000

New empty-check event: $0.00025 per provider unit ($0.25/1,000 units) for a successful response containing a validated empty list. Provider errors, ambiguous not-found errors and malformed responses are not empty checks. Nonempty pages, filtered records and duplicates do not trigger this event. This fee applies only when empty-check-unit is present in the run's actual pricing snapshot. Existing pricing snapshots remain supported with a 40-request cap and no empty-check fee during the pricing transition.

Apify charges the start event per GB of allocated memory, with a minimum of one event. Increasing memory above 1 GB increases the start charge. The Pricing tab and your run's displayed pricing are authoritative.

The Actor checks whether another result is affordable before fetching another page and before writing each record. It stops when no further result fits your maximum charge. Duplicate and filtered records are not saved or charged as results. Empty or failed runs can still incur the start charge; results saved before a later provider error remain available and billable. Running the same input again creates a new run and can save/charge those records again.

Tips and advanced options

  • Start with 10–25 results and inspect the output before increasing limits.
  • Keep comments off when you only need posts. For comments, set both a per-post cap and a total result cap.
  • Use top with a time window for bounded keyword research; use date limits for local filtering.
  • Runs share the provider's account-level request allowance. Waiting for an available request slot is expected during concurrent runs; increasing memory does not increase provider throughput.
  • request-budget, item-limit, post-limit, and charge-limit in the summary identify intentional stops. A successful bounded run does not mean that every available Reddit result was retrieved.
  • If a later page fails, the Actor keeps earlier saved results and reports the failure. Inspect the summary before assuming a dataset is complete.

FAQ, limitations, and support

Does this retrieve every post or comment? No. Provider search, pagination, Reddit access, deletions, and your configured limits determine coverage. The comments V2 endpoint returned two distinct pages of 50 comments in validation; this is not a completeness guarantee. A tested second page of subreddit posts returned a provider error, so deep historical subreddit collection should not be assumed reliable.

Can I search comments or scrape user profiles? This version searches posts and retrieves comments under selected posts. Keyword search over all Reddit comments, user profiles, and community discovery are not supported. Unsupported input fields are rejected before a provider request.

Is this a drop-in replacement for other Reddit Actors? It uses familiar post/comment field names but supports its own documented subset. Check input modes, limits, and nullable output fields before switching an existing workflow.

Can I access private, deleted, or restricted content? The Actor does not promise access to content unavailable through its provider and does not bypass private-community restrictions.

Can I resume a stopped run? Treat a stopped or failed run as partial and preserve its dataset. Start a new run for additional work; automatic exact-cursor recovery is not currently promised. Deduplication applies within a run.

Where can I get help? Use the Actor's Issues tab with a run ID and a redacted input. Do not include credentials. Custom output mappings or additional modes can be discussed there.

This Actor is independently developed by mscraper and is not affiliated with or endorsed by Reddit. Use collected data in accordance with applicable terms and privacy requirements, and avoid collecting data you are not entitled to process.

Extend Reddit discussion research with public social content and domain-level traffic estimates.

  • TikTok Scraper — Collect videos, creator profiles and comments to explore how a topic appears in short-form content.
  • X (Twitter) Scraper — Retrieve public profiles, timeline posts and follower lists for accounts relevant to your research.
  • Similarweb Quick Scraper — Check estimated visits, traffic sources and top countries for websites mentioned in Reddit discussions.

Run each Actor separately and combine its exported data in your own workflow. Each Actor has its own input, output and pricing.