Reddit Posts, Comments & Subreddit Analytics Scraper avatar

Reddit Posts, Comments & Subreddit Analytics Scraper

Pricing

from $3.00 / 1,000 result founds

Go to Apify Store
Reddit Posts, Comments & Subreddit Analytics Scraper

Reddit Posts, Comments & Subreddit Analytics Scraper

Scrape public Reddit posts, comments, search results, and subreddit stats through provider-backed access. Structured JSON for AI, research, and monitoring. $0.003/result plus usage.

Pricing

from $3.00 / 1,000 result founds

Rating

0.0

(0)

Developer

Khadin Akbar

Khadin Akbar

Maintained by Community

Actor stats

1

Bookmarked

871

Total users

59

Monthly active users

5 days ago

Last modified

Share

Use this Apify Actor to collect public Reddit posts, comments, search results, and subreddit analytics in structured JSON. It accepts subreddit URLs or names, a search query across Reddit, or direct post URLs, and each returned record represents one Reddit item: a post, a comment, or subreddit analytics. Useful fields include author, subreddit, score, upvote ratio, comment count, post type, media URLs, subscriber counts, active users, and timestamps, so the output can support AI workflows, research, monitoring, and reporting. This Actor is available through Apify and can also be called through Apify MCP.

Best fit and connected workflows

This Actor fits workflows that need public Reddit data in a clean dataset rather than page HTML. It is well suited for:

  • subreddit-level research with subredditUrls
  • broad topic monitoring with searchQuery
  • thread extraction from a specific Reddit post with postUrls
  • post-only scraping or post-plus-comment collection
  • subreddit analytics snapshots for audience size and activity context

If you are building an AI workflow, the dataset structure makes it easy to filter by type and route posts, comments, and subreddit analytics into different downstream steps.

Practical scenario

Maya, a product manager, wants to understand how users discuss a new feature in r/Python. She starts with the subreddit name, sets includeComments to true, and limits comment depth so the output stays easy to review. The returned records include post titles, scores, comment bodies, authors, and the subreddit analytics record. Maya uses the post scores and comment threads to decide which discussions to bring into a weekly review, then shares the result set with her team for follow-up reading.

Input fields

FieldTypePurpose
subredditUrlsarray of stringsScrape posts from one or more subreddit URLs or names.
searchQuerystringSearch across all of Reddit by keyword or phrase.
withinSubredditstringRestrict a search to one subreddit.
postUrlsarray of stringsExtract specific posts with their full comment threads.
maxPostsintegerMaximum number of posts to return.
maxCommentsPerPostintegerMaximum number of comments to extract for each post.
commentDepthLimitintegerMaximum nesting depth for comment extraction.
postDateLimitstringIgnore posts created before a given ISO 8601 date or datetime.
includeCommentsbooleanInclude full comment threads for each post.
includeSubredditAnalyticsbooleanInclude subreddit metadata such as subscribers and active users.
sortBystringSort results by hot, new, top, rising, comments, or relevance.
timeFilterstringApply a time window for top or search results.
includeNsfwbooleanInclude posts marked as NSFW.

Focused JSON input example

{
"subredditUrls": [
"https://reddit.com/r/MachineLearning"
],
"maxPosts": 10,
"includeComments": true,
"maxCommentsPerPost": 5,
"commentDepthLimit": 2,
"includeSubredditAnalytics": true,
"sortBy": "hot",
"timeFilter": "week",
"includeNsfw": false
}

Output fields

The dataset can include three record types: post, comment, and subreddit_analytics. Filter on the type field to separate them.

FieldTypeMeaning
typestringRecord type discriminator.
idstring or nullReddit item ID.
titlestring or nullPost title.
bodystring or nullPost body or comment text.
authorstringAuthor username.
subredditstring or nullSubreddit name for posts and comments.
subreddit_subscribersinteger or nullSubscriber count at scrape time.
scoreinteger or nullNet score.
upvote_rationumber or nullUpvote ratio for posts.
num_commentsinteger or nullComment count for posts.
urlstring or nullDirect Reddit URL for the record.
external_urlstring or nullExternal link target for link posts.
media_urlsarray or nullImage or video URLs attached to the post.
flairstring or nullPost flair label.
awards_countintegerAwards received.
post_typestring or nulltext, link, image, or video.
is_nsfwbooleanNSFW flag.
is_spoilerbooleanSpoiler flag.
is_stickiedbooleanSticky/pinned flag.
post_idstring or nullParent post ID for comments.
parent_idstring or nullParent comment or post ID for comments.
depthinteger or nullComment nesting depth.
is_submitterboolean or nullWhether the commenter is the original post author.
namestring or nullSubreddit display name for analytics records.
descriptionstring or nullSubreddit description for analytics records.
subscribersinteger or nullTotal subscribers for analytics records.
active_usersinteger or nullCurrently active users for analytics records.
nsfwboolean or nullNSFW flag for the subreddit.
subreddit_typestring or nullSubreddit visibility type.
icon_urlstring or nullSubreddit icon URL.
banner_urlstring or nullSubreddit banner URL.
created_atstringOriginal Reddit creation timestamp.
source_urlstringReddit JSON API URL used to extract the record.
scraped_atstringTimestamp when the Actor extracted the record.

Illustrative output record

{
"type": "post",
"id": "t3_abc123",
"title": "Python libraries for machine learning discussion in 2026",
"body": "I've been using scikit-learn and comparing other libraries...",
"author": "u/john_doe",
"subreddit": "r/MachineLearning",
"subreddit_subscribers": 2840000,
"score": 1482,
"upvote_ratio": 0.94,
"num_comments": 73,
"url": "https://reddit.com/r/MachineLearning/comments/abc123/python_libraries_discussion/",
"external_url": "https://arxiv.org/abs/2401.12345",
"media_urls": [
"https://i.redd.it/example.jpg"
],
"flair": "Research Paper",
"awards_count": 5,
"post_type": "text",
"is_nsfw": false,
"is_spoiler": false,
"is_stickied": false,
"created_at": "2026-03-20T14:22:00.000Z",
"source_url": "https://www.reddit.com/r/MachineLearning.json?sort=hot&limit=100",
"scraped_at": "2026-03-28T14:30:00.000Z"
}

How it works

This Actor uses owner-managed provider-backed access to read public Reddit data. It works across three modes: subreddit scraping, Reddit search, and specific post URL extraction. The output is structured as JSON and written to the default dataset, with OUTPUT and RUN_SUMMARY records also stored in the key-value store for diagnostics and execution summaries.

The dataset schema is designed for downstream filtering by type, so posts, comments, and subreddit analytics can be processed separately or together.

Pricing

Reddit Posts, Comments & Subreddit Analytics Scraper uses Pay per event plus Apify platform usage. The primary billable event is Result found, which is charged for each extracted Reddit post, comment, or subreddit analytics record. There is also a one-time actor start event.

A simple event-count example: if a run returns one hundred posts and fifty comments, the Actor records one hundred fifty Result found events, plus the start event, and Apify platform usage is added on top. For current pricing details, check the live Pricing tab in the Apify console.

Use with AI agents (MCP)

This Actor is available through Apify MCP and is designed for tool use in AI-assisted workflows. It can be called by an agent that translates a natural-language request into the Actor input fields, then reads the dataset output back into the conversation.

Exact Actor identity: khadinakbar/reddit-posts-comments-scraper

Example agent prompt:

Find posts from r/MachineLearning about foundation models from the last month, include comments up to depth 2, and return subreddit analytics too.

Output interpretation:

  • type = post identifies post records
  • type = comment identifies comment records
  • type = subreddit_analytics identifies subreddit metadata records
  • use post_id and parent_id to reconstruct comment threads
  • use scraped_at for freshness tracking
  • use source_url for provenance and repeatable re-fetching

Scope and pagination guidance:

  • use maxPosts to set the number of posts to collect
  • use maxCommentsPerPost to control comment volume
  • use commentDepthLimit to shape thread depth
  • use postDateLimit to focus on recent content
  • use includeComments and includeSubredditAnalytics to select which record types are included
  • for broad searches, sortBy and timeFilter help narrow the result set

Apify API example

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({
token: process.env.APIFY_TOKEN,
});
const run = await client.actor('khadinakbar/reddit-posts-comments-scraper').call({
searchQuery: 'artificial intelligence regulation',
maxPosts: 20,
sortBy: 'relevance',
timeFilter: 'month',
includeComments: true,
maxCommentsPerPost: 10,
includeSubredditAnalytics: true,
});
const items = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.items.filter((item) => item.type === 'post'));

Best results and outcome guidance

Choose subredditUrls when you already know the community you want to inspect. Use searchQuery when you are tracking a topic across Reddit. Use postUrls when you need a specific thread with its comments. For comment-heavy threads, combine a moderate maxCommentsPerPost with commentDepthLimit to keep the output focused. For monitoring workflows, postDateLimit and sortBy: "new" or sortBy: "top" help keep result sets aligned with the time window you care about.

Continue the workflow

Design note

I found that every dataset record is required to include type and scraped_at, which makes it straightforward to sort records by category and freshness in downstream workflows.

FAQ

How should I choose between subreddit URLs and search queries?
Use subredditUrls for a known community and searchQuery for a topic that spans Reddit.

Can I pull comments without posts?
This Actor is structured around post extraction with optional comment threads. Set includeComments to false for post-only runs, or use postUrls when you want a specific discussion thread.

How can I filter the dataset after the run?
Filter by type to separate post, comment, and subreddit_analytics records.

What is the difference between post_id and parent_id in comments?
post_id points to the parent post. parent_id points to the direct parent item, which can be a post or another comment.

When should I use withinSubreddit?
Use it together with searchQuery when you want a search limited to one subreddit.

Responsible use

This Actor is for public Reddit data. Use it in ways that respect Reddit content access, your rights and approvals, and the intended downstream use of the data. Review the returned records before redistribution, and use the is_nsfw, is_spoiler, and is_stickied fields as part of your own filtering and handling logic where relevant.

Reddit Posts, Comments & Subreddit Analytics Scraper is designed as a focused workflow.