Reddit Posts, Comments & Subreddit Analytics Scraper
Pricing
from $3.00 / 1,000 result founds
Reddit Posts, Comments & Subreddit Analytics Scraper
Scrape public Reddit posts, comments, search results, and subreddit stats through provider-backed access. Structured JSON for AI, research, and monitoring. $0.003/result plus usage.
Pricing
from $3.00 / 1,000 result founds
Rating
0.0
(0)
Developer
Khadin Akbar
Maintained by CommunityActor stats
1
Bookmarked
871
Total users
59
Monthly active users
5 days ago
Last modified
Categories
Share
Use this Apify Actor to collect public Reddit posts, comments, search results, and subreddit analytics in structured JSON. It accepts subreddit URLs or names, a search query across Reddit, or direct post URLs, and each returned record represents one Reddit item: a post, a comment, or subreddit analytics. Useful fields include author, subreddit, score, upvote ratio, comment count, post type, media URLs, subscriber counts, active users, and timestamps, so the output can support AI workflows, research, monitoring, and reporting. This Actor is available through Apify and can also be called through Apify MCP.
Best fit and connected workflows
This Actor fits workflows that need public Reddit data in a clean dataset rather than page HTML. It is well suited for:
- subreddit-level research with
subredditUrls - broad topic monitoring with
searchQuery - thread extraction from a specific Reddit post with
postUrls - post-only scraping or post-plus-comment collection
- subreddit analytics snapshots for audience size and activity context
If you are building an AI workflow, the dataset structure makes it easy to filter by type and route posts, comments, and subreddit analytics into different downstream steps.
Practical scenario
Maya, a product manager, wants to understand how users discuss a new feature in r/Python. She starts with the subreddit name, sets includeComments to true, and limits comment depth so the output stays easy to review. The returned records include post titles, scores, comment bodies, authors, and the subreddit analytics record. Maya uses the post scores and comment threads to decide which discussions to bring into a weekly review, then shares the result set with her team for follow-up reading.
Input fields
| Field | Type | Purpose |
|---|---|---|
subredditUrls | array of strings | Scrape posts from one or more subreddit URLs or names. |
searchQuery | string | Search across all of Reddit by keyword or phrase. |
withinSubreddit | string | Restrict a search to one subreddit. |
postUrls | array of strings | Extract specific posts with their full comment threads. |
maxPosts | integer | Maximum number of posts to return. |
maxCommentsPerPost | integer | Maximum number of comments to extract for each post. |
commentDepthLimit | integer | Maximum nesting depth for comment extraction. |
postDateLimit | string | Ignore posts created before a given ISO 8601 date or datetime. |
includeComments | boolean | Include full comment threads for each post. |
includeSubredditAnalytics | boolean | Include subreddit metadata such as subscribers and active users. |
sortBy | string | Sort results by hot, new, top, rising, comments, or relevance. |
timeFilter | string | Apply a time window for top or search results. |
includeNsfw | boolean | Include posts marked as NSFW. |
Focused JSON input example
{"subredditUrls": ["https://reddit.com/r/MachineLearning"],"maxPosts": 10,"includeComments": true,"maxCommentsPerPost": 5,"commentDepthLimit": 2,"includeSubredditAnalytics": true,"sortBy": "hot","timeFilter": "week","includeNsfw": false}
Output fields
The dataset can include three record types: post, comment, and subreddit_analytics. Filter on the type field to separate them.
| Field | Type | Meaning |
|---|---|---|
type | string | Record type discriminator. |
id | string or null | Reddit item ID. |
title | string or null | Post title. |
body | string or null | Post body or comment text. |
author | string | Author username. |
subreddit | string or null | Subreddit name for posts and comments. |
subreddit_subscribers | integer or null | Subscriber count at scrape time. |
score | integer or null | Net score. |
upvote_ratio | number or null | Upvote ratio for posts. |
num_comments | integer or null | Comment count for posts. |
url | string or null | Direct Reddit URL for the record. |
external_url | string or null | External link target for link posts. |
media_urls | array or null | Image or video URLs attached to the post. |
flair | string or null | Post flair label. |
awards_count | integer | Awards received. |
post_type | string or null | text, link, image, or video. |
is_nsfw | boolean | NSFW flag. |
is_spoiler | boolean | Spoiler flag. |
is_stickied | boolean | Sticky/pinned flag. |
post_id | string or null | Parent post ID for comments. |
parent_id | string or null | Parent comment or post ID for comments. |
depth | integer or null | Comment nesting depth. |
is_submitter | boolean or null | Whether the commenter is the original post author. |
name | string or null | Subreddit display name for analytics records. |
description | string or null | Subreddit description for analytics records. |
subscribers | integer or null | Total subscribers for analytics records. |
active_users | integer or null | Currently active users for analytics records. |
nsfw | boolean or null | NSFW flag for the subreddit. |
subreddit_type | string or null | Subreddit visibility type. |
icon_url | string or null | Subreddit icon URL. |
banner_url | string or null | Subreddit banner URL. |
created_at | string | Original Reddit creation timestamp. |
source_url | string | Reddit JSON API URL used to extract the record. |
scraped_at | string | Timestamp when the Actor extracted the record. |
Illustrative output record
{"type": "post","id": "t3_abc123","title": "Python libraries for machine learning discussion in 2026","body": "I've been using scikit-learn and comparing other libraries...","author": "u/john_doe","subreddit": "r/MachineLearning","subreddit_subscribers": 2840000,"score": 1482,"upvote_ratio": 0.94,"num_comments": 73,"url": "https://reddit.com/r/MachineLearning/comments/abc123/python_libraries_discussion/","external_url": "https://arxiv.org/abs/2401.12345","media_urls": ["https://i.redd.it/example.jpg"],"flair": "Research Paper","awards_count": 5,"post_type": "text","is_nsfw": false,"is_spoiler": false,"is_stickied": false,"created_at": "2026-03-20T14:22:00.000Z","source_url": "https://www.reddit.com/r/MachineLearning.json?sort=hot&limit=100","scraped_at": "2026-03-28T14:30:00.000Z"}
How it works
This Actor uses owner-managed provider-backed access to read public Reddit data. It works across three modes: subreddit scraping, Reddit search, and specific post URL extraction. The output is structured as JSON and written to the default dataset, with OUTPUT and RUN_SUMMARY records also stored in the key-value store for diagnostics and execution summaries.
The dataset schema is designed for downstream filtering by type, so posts, comments, and subreddit analytics can be processed separately or together.
Pricing
Reddit Posts, Comments & Subreddit Analytics Scraper uses Pay per event plus Apify platform usage. The primary billable event is Result found, which is charged for each extracted Reddit post, comment, or subreddit analytics record. There is also a one-time actor start event.
A simple event-count example: if a run returns one hundred posts and fifty comments, the Actor records one hundred fifty Result found events, plus the start event, and Apify platform usage is added on top. For current pricing details, check the live Pricing tab in the Apify console.
Use with AI agents (MCP)
This Actor is available through Apify MCP and is designed for tool use in AI-assisted workflows. It can be called by an agent that translates a natural-language request into the Actor input fields, then reads the dataset output back into the conversation.
Exact Actor identity: khadinakbar/reddit-posts-comments-scraper
Example agent prompt:
Find posts from r/MachineLearning about foundation models from the last month, include comments up to depth 2, and return subreddit analytics too.
Output interpretation:
type = postidentifies post recordstype = commentidentifies comment recordstype = subreddit_analyticsidentifies subreddit metadata records- use
post_idandparent_idto reconstruct comment threads - use
scraped_atfor freshness tracking - use
source_urlfor provenance and repeatable re-fetching
Scope and pagination guidance:
- use
maxPoststo set the number of posts to collect - use
maxCommentsPerPostto control comment volume - use
commentDepthLimitto shape thread depth - use
postDateLimitto focus on recent content - use
includeCommentsandincludeSubredditAnalyticsto select which record types are included - for broad searches,
sortByandtimeFilterhelp narrow the result set
Apify API example
import { ApifyClient } from 'apify-client';const client = new ApifyClient({token: process.env.APIFY_TOKEN,});const run = await client.actor('khadinakbar/reddit-posts-comments-scraper').call({searchQuery: 'artificial intelligence regulation',maxPosts: 20,sortBy: 'relevance',timeFilter: 'month',includeComments: true,maxCommentsPerPost: 10,includeSubredditAnalytics: true,});const items = await client.dataset(run.defaultDatasetId).listItems();console.log(items.items.filter((item) => item.type === 'post'));
Best results and outcome guidance
Choose subredditUrls when you already know the community you want to inspect. Use searchQuery when you are tracking a topic across Reddit. Use postUrls when you need a specific thread with its comments. For comment-heavy threads, combine a moderate maxCommentsPerPost with commentDepthLimit to keep the output focused. For monitoring workflows, postDateLimit and sortBy: "new" or sortBy: "top" help keep result sets aligned with the time window you care about.
Continue the workflow
- Then use Reddit Scraper – Posts, Comments & Subreddits to extend Reddit Posts, Comments & Subreddit Analytics Scraper research with a complementary content contract.
- Then use Reddit Posts Search Scraper to extend Reddit Posts, Comments & Subreddit Analytics Scraper research with a complementary discovery contract.
Design note
I found that every dataset record is required to include type and scraped_at, which makes it straightforward to sort records by category and freshness in downstream workflows.
FAQ
How should I choose between subreddit URLs and search queries?
Use subredditUrls for a known community and searchQuery for a topic that spans Reddit.
Can I pull comments without posts?
This Actor is structured around post extraction with optional comment threads. Set includeComments to false for post-only runs, or use postUrls when you want a specific discussion thread.
How can I filter the dataset after the run?
Filter by type to separate post, comment, and subreddit_analytics records.
What is the difference between post_id and parent_id in comments?
post_id points to the parent post. parent_id points to the direct parent item, which can be a post or another comment.
When should I use withinSubreddit?
Use it together with searchQuery when you want a search limited to one subreddit.
Responsible use
This Actor is for public Reddit data. Use it in ways that respect Reddit content access, your rights and approvals, and the intended downstream use of the data. Review the returned records before redistribution, and use the is_nsfw, is_spoiler, and is_stickied fields as part of your own filtering and handling logic where relevant.
Reddit Posts, Comments & Subreddit Analytics Scraper is designed as a focused workflow.