Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s
Pricing
from $1.60 / 1,000 results
Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s
Lightning-fast Reddit scraping at 10 items in just 1 sec. Extract posts, comments, user profiles, subreddit feeds, and trending communities into structured JSON for research, monitoring, sentiment analysis, competitive intelligence, and lead generation
Pricing
from $1.60 / 1,000 results
Rating
0.0
(0)
Developer
Mikolabs
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 days ago
Last modified
Categories
Share
Reddit Scraper — Complete Suite for Posts, Comments, Users & Subreddits
Extract Reddit posts, full comment threads, user profiles, subreddit feeds, and trending communities at scale. Collect structured JSON data from Reddit for market research, brand monitoring, sentiment analysis, competitive intelligence, lead generation, and content strategy — no coding required.
Overview
Reddit Scraper collects publicly available Reddit posts and, when enabled, comment records with the metadata teams usually need for analysis, enrichment, and automation. Each record can include core content, authorship, engagement metrics, subreddit context, timestamps, canonical URLs, flair data, and media-related fields when available. Reddit is one of the web's richest sources of real-world opinions, product feedback, community discussion, and trend signals, which makes it valuable for research, monitoring, and downstream decision-making.
This actor turns that information into consistent dataset records so teams can replace repetitive manual collection with repeatable runs. It supports six scraping modes — keyword search, subreddit feeds, post threads with nested comments, user profiles, direct URL extraction, and subreddit discovery — all configurable through a simple point-and-click interface. It is designed for production use with stable identifiers, overlap-friendly outputs for deduplication, and reliable scheduled collection workflows.
Whether you need 20 posts for a quick check or thousands of records for a research pipeline, this actor handles it with automatic pagination, deduplication, and reliable rate management.
Why Use This Actor
- Market research and analytics teams: Track conversation volume, engagement, subreddit activity, and topic trends across keywords, communities, and time windows to inform strategic decisions.
- Product and content teams: Discover user pain points, feature requests, language patterns, and high-performing discussion themes for product roadmap planning and editorial calendars.
- Brand monitoring and competitive intelligence teams: Watch competitor mentions, product launch reactions, brand sentiment shifts, category-level discussion spikes, and recurring conversation patterns without manually checking Reddit every day.
- Developers and data engineering teams: Feed Reddit data into ETL pipelines, data warehouses, dashboards, and APIs using structured JSON records that are easy to upsert and model.
- Lead generation and enrichment teams: Identify relevant communities, active buyer discussions, context around buyer interests, brand mentions, and niche topics.
- Academic and social researchers: Collect discussion datasets for discourse analysis, public opinion research, community dynamics studies, and longitudinal content tracking.
- SEO and digital marketing teams: Mine Reddit for content ideas, keyword research, real user language, and high-engagement topics to improve search rankings, ad copy, and editorial strategy.
- Operations and automation teams: Run recurring jobs on a schedule and use stable record keys to deduplicate overlapping results from queries, subreddit searches, and direct URLs.
Pricing & Plans (No Hidden Fees)
Transparent and predictable pricing with no extra proxy costs, no setup fees, and no hidden maintenance charges.
Tiered Pricing Structure
| Tier / Discount Level | Price per 1,000 Results | Effective Savings | Minimum Scrape |
|---|---|---|---|
| No Discount (Standard / Pay-As-You-Go) | $4.00 / 1,000 items | Standard Rate | 1 item |
| 🥉 Bronze Discount | $2.00 / 1,000 items | 50% OFF | 20 items |
| 🥈 Silver Discount | $1.80 / 1,000 items | 55% OFF | 20 items |
| 🥇 Gold Discount | $1.60 / 1,000 items | 60% OFF | 20 items |
Plan Comparison
| Feature | Free Tier | Subscriber / Paid Tier |
|---|---|---|
| Free Daily Allowance | 20 items / run (4 runs / day free) | Unlimited |
| Pricing | $4.00 / 1,000 results (or free allowance) | Down to $1.60 / 1,000 results |
| Additional Fees | $0.00 (No extra fees) | $0.00 (No extra fees) |
| Proxy / Bandwidth Costs | Included ($0.00) | Included ($0.00) |
| All 6 Scraping Modes | ✅ | ✅ |
| AI Sentiment Analysis | ✅ | ✅ |
| AI Content Taxonomy | ✅ | ✅ |
| Advanced Filters | ✅ | ✅ |
| Run Summary Dashboard | ✅ | ✅ |
| Scheduling & Automation | ✅ | ✅ |
Free users can test every feature (search, subreddits, comments, user profiles, AI sentiment, content taxonomy) with up to 20 items per run and 4 runs per day completely free of charge. Upgrade for volume discounts down to $1.60 / 1,000 items with zero hidden fees.
How to Use Reddit Scraper — Step by Step
Step 1: Choose a Scraping Mode
Select one of six modes from the Scraping Mode dropdown at the top of the input form:
| Mode | What It Does | When to Use |
|---|---|---|
| 🔍 Search | Search Reddit by keywords and phrases | Finding posts about a topic, brand, or product across all of Reddit |
| 🌐 Subreddits | Scrape subreddit feeds (Hot, Top, New, Rising) | Monitoring specific communities and collecting feed-based datasets |
| 💬 Posts | Extract specific post threads with nested comments | Deep-diving into individual discussions and comment trees |
| 👤 Users | Collect user profiles, karma, and submission history | Researching influencers, key accounts, or user activity patterns |
| 🔗 Direct URLs | Scrape any Reddit URL you provide | Extracting specific pages, threads, or search results you already found |
| 🧭 Discover | Find popular and trending subreddits | Discovering new communities, niches, and emerging topics |
Step 2: Configure Your Mode
Fill in the settings section that corresponds to your chosen mode:
- Search mode → Enter your keywords in "Search Keywords / Phrases"
- Subreddits mode → Enter subreddit names in "Subreddits to Scrape"
- Users mode → Enter Reddit usernames in "Reddit Usernames"
- Direct URLs mode → Paste Reddit URLs in "Direct Reddit URLs"
- Posts mode → Paste specific post URLs in "Direct Reddit URLs"
- Discover mode → No additional input needed, the actor finds trending communities automatically
Step 3: Set Filters (Optional)
Use the Filters & Content Quality Bounds section to narrow your results by post type (text, image, video, gallery, link), minimum/maximum upvote score, date ranges, flair keywords, title or body keywords, domain restrictions, and NSFW inclusion/exclusion.
Step 4: Enable AI Analytics (Optional)
Toggle AI Sentiment Analysis and/or AI Content Taxonomy Classification to automatically enrich every record with sentiment scores and topic labels. No external API keys required.
Step 5: Set Limits & Run
Configure your item limits in the Limits & Performance section, then click Start to begin the extraction. Results are saved to your Apify dataset as clean JSON, ready for download in JSON, CSV, Excel, XML, or RSS formats.
Input Parameters
Provide any combination of modes, queries, URLs, and filters to control what the actor collects and how focused the results should be.
🎯 Step 1: Mode Selection
| Parameter | Type | Default | Description |
|---|---|---|---|
mode | string | search | Scraping mode. Options: search, subreddits, posts, users, directUrls, discover |
🔍 Step 2 — Search Settings (Mode = search)
Use when you want to discover relevant posts across all of Reddit matching your keywords.
| Parameter | Type | Default | Description |
|---|---|---|---|
queries | string[] | – | Keywords or search phrases to look for across Reddit. Accepts multiple terms. Use this when you want the actor to discover relevant posts for you. |
searchSubreddit | string | – | Optional: restrict search to a single subreddit (without r/). Leave empty to search all of Reddit. |
strictSearch | boolean | false | Make Reddit search follow your keywords more closely, relying more on exact keywords and less on loose semantic matching. This usually returns fewer posts, but keeps results closer to your query. |
strictTokenFilter | boolean | false | Make the actor scan each saved post's title, body, and URL and keep only posts that match all of your query keywords. This reduces output size, but keeps the most accurate results. |
🌐 Step 2 — Subreddit & Feed Settings (Mode = subreddits)
Use when you want to collect posts from one or more subreddit feeds.
| Parameter | Type | Default | Description |
|---|---|---|---|
subreddits | string[] | ["AskReddit", "technology"] | List of subreddit names to scrape (without r/). Accepts multiple subreddits. |
sort | string | hot | Ranking order for subreddit feeds or search results. Options: hot (Trending), new (Latest), top (Most Upvoted), rising, controversial, relevance |
timeframe | string | all | Time window for top and controversial sorts. Options: hour, day, week, month, year, all. Use postsCreatedAfter and postsCreatedBefore for exact record-level date filtering. |
fullSubredditMode | boolean | false | Deep crawl mode — paginates as far back as the Reddit API permits for maximum historical coverage. |
🔗 Step 2 — Direct URLs & Post Links (Mode = posts or directUrls)
Use when you have specific Reddit links to extract.
| Parameter | Type | Default | Description |
|---|---|---|---|
urls | string[] | – | Direct Reddit URLs to scrape, such as post URLs, comment permalinks, subreddit pages, user pages, or Reddit search pages. When provided, URL input takes priority over search queries. |
👤 Step 2 — User Profile Settings (Mode = users)
Use when you want to collect user profiles, karma scores, and submission history.
| Parameter | Type | Default | Description |
|---|---|---|---|
usernames | string[] | ["spez"] | Reddit usernames to extract (without u/). |
userContent | string | all | What activity to collect for user profiles. Options: all (Profile + Posts + Comments), submitted (Posts Only), comments (Comments Only) |
💬 Comments & Discussion Settings
Configure comment extraction for any post-based mode.
| Parameter | Type | Default | Description |
|---|---|---|---|
includeComments | boolean | false | When enabled, the actor also saves comments from each collected post. Useful for sentiment analysis, deeper discussion review, and thread-level context. |
maxCommentsPerPost | integer | 25 | Maximum number of comments to collect per post when comment collection is enabled. Lower values help keep runs faster and datasets smaller. |
commentsSort | string | confidence | Ranking order for comments. Options: confidence (Best), top (Most Upvoted), new, controversial, old, qa (Q&A) |
flattenComments | boolean | false | When true, saves each comment as a separate independent row in the dataset instead of nesting under the parent post. |
⚡ Filters & Content Quality Bounds
Fine-tune results with granular post-level filters. All filters are optional and can be combined.
| Parameter | Type | Default | Description |
|---|---|---|---|
postType | string | all | Filter by media format. Options: all, text (Self Text Only), image, video, gallery, link (External Links Only) |
minScore | integer | – | Only keep posts with at least this upvote score. |
maxScore | integer | – | Only keep posts at or below this upvote score. |
minComments | integer | – | Only keep posts with at least this many comments. |
minUpvoteRatio | number | – | Filter by upvote ratio (e.g. 0.85 for ≥85% positive). |
postsCreatedAfter | string | – | Only keep posts created on or after this date (YYYY-MM-DD). Plain dates are normalized to the start of the day in UTC. |
postsCreatedBefore | string | – | Only keep posts created on or before this date (YYYY-MM-DD). Plain dates are normalized to the end of the day in UTC. |
onlyWithFlair | boolean | false | Only keep posts that have an assigned flair tag. |
flairContains | string | – | Match posts where the flair text contains this keyword. |
titleContains | string | – | Only keep posts whose title contains this keyword. |
textContains | string | – | Only keep posts whose body text contains this keyword. |
domainContains | string | – | Only keep posts linking to this domain (e.g. github.com, nytimes.com). |
excludeStickied | boolean | false | Exclude pinned/stickied moderator announcement posts from results. |
excludeKeywords | string[] | – | Exclude any posts containing these keywords in title or body. |
includeNsfw | boolean | true | Include posts marked as NSFW or 18+ in the output dataset. |
🤖 AI Analytics & Enrichment
Add AI-powered analysis to every extracted record — no external API keys needed.
| Parameter | Type | Default | Description |
|---|---|---|---|
sentiment_analysis | boolean | false | When enabled, adds sentiment_score, sentiment_confidence, and sentiment_label (positive, negative, neutral, mixed, uncertain) to posts and comments. |
content_analysis | boolean | false | When enabled, classifies each post against a content taxonomy and adds content_category_label and content_category_path to post records. |
Sentiment Labels
When sentiment_analysis is enabled, sentiment_label can have these values:
| Label | Meaning | Example |
|---|---|---|
positive | The text has a clear positive tone or favorable sentiment. | "This product is amazing, best purchase I've made!" |
negative | The text has a clear negative tone or unfavorable sentiment. | "Terrible experience, would not recommend to anyone." |
neutral | The text has little or no sentiment signal, or the sentiment score is close to zero. | "The package arrived on Tuesday with the latest invoice attached." |
mixed | The text has both positive and negative cues, and they remain unresolved or balanced. | "I love the cast and visuals. The story is boring and the ending is awful." |
uncertain | The analyzer does not have enough confidence to classify the tone, or the text is empty, bot boilerplate, or moderator boilerplate. | "I am a bot, and this action was performed automatically." |
Content Taxonomy Classification
When content_analysis is enabled, the actor assigns post-level topic labels from a stable content taxonomy and returns both a human-readable category and its taxonomy path. The labeling is designed for research pipelines, monitoring systems, and workflows that need a fast topical read of large Reddit collections before routing, clustering, summarizing, or enriching records downstream.
The classifier is intentionally conservative: when the record does not contain enough topic evidence, the actor leaves the category fields empty instead of forcing a low-confidence label. This keeps category fields useful as high-signal routing metadata for analysts, dashboards, warehouse models, and LLM agents.
content_category_label— Human-readable category name (e.g., "Artificial Intelligence")content_category_path— Full taxonomy path (e.g.,["Technology", "Artificial Intelligence"])
⚙️ Limits & Performance
| Parameter | Type | Default | Description |
|---|---|---|---|
maxPostsPerSubreddit | integer | 100 | Maximum number of posts to collect per subreddit or query source. |
maxTotalItems | integer | 500 | Safety ceiling for the total number of records across all inputs in a single run. |
maximize_coverage | boolean | false | Turn on high-coverage collection for broad or competitive topics. See Maximum Coverage Mode below. |
Maximum Coverage Mode
Reddit search and listing surfaces do not behave like unbounded database queries. For many seeds, Reddit exposes only a limited practical result window, commonly around 250 posts per query/ranking combination, even when the topic has far more historical discussion. A single relevance, top, hot, new, or comments traversal can therefore look complete while still missing major clusters of posts, subreddits, and time periods.
When maximize_coverage is enabled, the actor treats each seed as a research topic instead of a single result page. It expands the collection plan across complementary result views, date-aware traversals, and high-signal community follow-ups, then deduplicates overlapping posts into one stable dataset. This is useful for market research, incident monitoring, topic mapping, retrieval corpora, and workflows that need a broader evidence base before summarization, clustering, lead scoring, or trend analysis.
This mode can use more requests and take longer than a standard run, but it gives downstream systems a better chance of seeing the shape of the conversation rather than only the first capped slice of a single ranking surface. Use it when recall matters more than preserving a single ranking order.
Example Inputs
Scenario: Query-driven monitoring
{"mode": "search","queries": ["ai video generator", "synthetic media"],"sort": "new","timeframe": "week","postsCreatedAfter": "2026-03-01","postsCreatedBefore": "2026-03-31","maximize_coverage": true,"strictSearch": true,"strictTokenFilter": true,"maxTotalItems": 500}
Scenario: Direct URL collection with comments and AI
{"mode": "directUrls","urls": ["https://www.reddit.com/r/technology/","https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/"],"includeComments": true,"sentiment_analysis": true,"content_analysis": true,"maxCommentsPerPost": 200,"maxTotalItems": 150}
Scenario: Targeted subreddit feed
{"mode": "subreddits","subreddits": ["startups", "SaaS", "Entrepreneur"],"sort": "top","timeframe": "month","minScore": 50,"minComments": 10,"excludeStickied": true,"includeComments": true,"maxCommentsPerPost": 100,"commentsSort": "top","maxPostsPerSubreddit": 200,"maxTotalItems": 500}
Scenario: User profile research
{"mode": "users","usernames": ["spez", "GovSchwarzenegger"],"userContent": "all","maxTotalItems": 200}
Scenario: Advanced filtered search with date range
{"mode": "search","queries": ["product recall", "safety issue"],"sort": "new","postsCreatedAfter": "2026-01-01","postsCreatedBefore": "2026-06-30","minScore": 25,"postType": "text","excludeKeywords": ["meme", "shitpost"],"excludeStickied": true,"includeNsfw": false,"maximize_coverage": true,"sentiment_analysis": true,"content_analysis": true,"maxPostsPerSubreddit": 500,"maxTotalItems": 2000}
Scenario: Discover trending subreddits
{"mode": "discover","maxTotalItems": 100}
Output
Output Destination
The actor writes results to an Apify dataset as JSON records. The dataset is designed for direct consumption by analytics tools, ETL pipelines, and downstream APIs without post-processing. Download in JSON, CSV, Excel (XLSX), XML, HTML Table, or RSS formats directly from the Apify Console or via API.
Record Envelope (All Items)
Every record includes a stable category field, a Reddit identifier, and a canonical URL:
kind(string, required): Logical record type, such aspostorcomment.id(string, required): Stable Reddit identifier for the entity.url(string, required): Canonical Reddit URL for the record.
Recommended idempotency key: kind + ":" + id
Use this key for deduplication and upserts, especially when the same Reddit entity appears in overlapping queries, subreddit runs, or direct URL inputs.
Run Summary Artifact
At the end of each successful run, the actor writes RUN-SUMMARY to the key-value store with totals, query and subreddit breakdowns, date range, engagement highlights, optional sentiment/category counts, tier information, and a complete run overview.
The visual report is saved as RUN-MAP.html and renders the summary values as an interactive dashboard with stats cards, engagement breakdowns, mode indicators, and tier badges. Access both from the Output tab after your run completes.
Example Output: Post Record (kind = "post")
{"kind": "post","query": "cheesecake","id": "1hvoazn","title": "My best cheesecake so far","body": "Found my new favorite recipe (no water bath). Next time I will make a thicker crust. Added a raspberry compote.","sentiment_score": 2,"sentiment_label": "positive","sentiment_confidence": 0.78,"sentiment_score_normalized": 0.86,"content_category_label": "Desserts and Baking","content_category_path": ["Food & Drink", "Desserts and Baking"],"author": "ClearlyBulky","score": 3489,"upvote_ratio": 1,"num_comments": 43,"subreddit": "Baking","created_utc": "2025-01-07T10:09:56.000Z","url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/","permalink": "/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/","canonical_url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/","old_reddit_url": "https://old.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/","flair": "Recipe","post_hint": "link","over_18": false,"is_self": false,"spoiler": false,"locked": false,"is_video": false,"is_gallery": true,"hidden": false,"edited": false,"archived": false,"pinned": false,"domain": "old.reddit.com","thumbnail": "https://b.thumbs.redditmedia.com/example.jpg","url_overridden_by_dest": "https://www.reddit.com/gallery/1hvoazn","num_duplicates": 0,"subreddit_id": "t5_2qx1h","subreddit_name_prefixed": "r/Baking","subreddit_subscribers": 4322940,"media": null,"media_metadata": {},"gallery_data": {"items": [{ "is_deleted": false, "media_id": "kny1nmhlqjbe1", "id": 581711947 },{ "is_deleted": false, "media_id": "wjqc6mhlqjbe1", "id": 581711948 }]},"gallery_images": [{"media_id": "kny1nmhlqjbe1","caption": "","width": 3024,"height": 4032,"url": "https://preview.redd.it/kny1nmhlqjbe1.jpg?width=3024&format=pjpg&auto=webp","previews": ["https://preview.redd.it/kny1nmhlqjbe1.jpg?width=640&crop=smart&auto=webp"]}],"media_assets": [{"type": "Image","media_id": "kny1nmhlqjbe1","mime_type": "image/jpg","original_url": "https://preview.redd.it/kny1nmhlqjbe1.jpg?width=3024&format=pjpg&auto=webp","preview_urls": ["https://preview.redd.it/kny1nmhlqjbe1.jpg?width=640&crop=smart&auto=webp"]}],"age_hours": 10916.1333,"retrieved_at": "2026-04-07T00:00:00.000Z","media_type": "gallery","has_media": true,"gallery_count": 2,"outbound_url_host": "www.reddit.com","title_length": 26,"body_length": 112,"word_count": 25,"score_per_hour": 0.3196,"comments_per_hour": 0.0039,"is_deleted_or_removed": false,"engagement_total": 3532,"comment_to_score_ratio": 0.0123,"is_high_engagement": true,"content_flags": [],"stickied": false,"distinguished": null,"score_hidden": false,"total_awards_received": 0,"all_awardings": [],"gilded": 0,"num_crossposts": 0,"is_original_content": false,"author_fullname": "t2_dr3vyilor","author_flair_text": null,"author_premium": false,"body_html": "<div class=\"md\"><p>Found my new favorite recipe...</p></div>","preview": null,"secure_media": null,"secure_media_embed": {},"crosspost_parent_list": null}
Example Output: Comment Record (kind = "comment")
{"kind": "comment","query": "cheesecake","id": "m5un6bj","postId": "1hvoazn","postUrl": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/","parentId": "t3_1hvoazn","body": "This looks absolutely incredible! Can you share the full recipe?","sentiment_score": 3,"sentiment_label": "positive","sentiment_confidence": 0.91,"sentiment_score_normalized": 0.92,"author": "BakingFanatic","score": 76,"subreddit": "Baking","created_utc": "2025-01-07T10:13:48.000Z","url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/","permalink": "/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/","canonical_url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/","old_reddit_url": "https://old.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/","root_comment_id": "m5un6bj","parent_kind": "post","is_deleted_or_removed": false,"subreddit_id": "t5_2qx1h","subreddit_name_prefixed": "r/Baking","edited": false,"retrieved_at": "2026-04-07T00:00:00.000Z","age_hours": 10916.0678,"body_length": 63,"word_count": 10,"score_per_hour": 0.007,"stickied": false,"distinguished": null,"is_submitter": false,"score_hidden": false,"controversiality": 0,"depth": 0,"total_awards_received": 0,"all_awardings": [],"gilded": 0,"author_fullname": "t2_example123","author_flair_text": null,"author_premium": false,"body_html": "<div class=\"md\"><p>This looks absolutely incredible! Can you share the full recipe?</p></div>","collapsed": false,"collapsed_reason": null,"collapsed_because_crowd_control": false,"unrepliable_reason": null}
Field Reference
Post Fields (kind = "post")
| Field | Type | Required | Description |
|---|---|---|---|
kind | string | ✅ | Record category — always "post" for post records. |
query | string | – | Input query or source label that produced the record. |
id | string | ✅ | Stable Reddit post identifier. |
title | string | ✅ | Post title. |
body | string | – | Post body text. |
sentiment_score | number | – | Raw sentiment score computed from the post title and body when sentiment_analysis is enabled. |
sentiment_score_normalized | number | – | Length-normalized bounded sentiment score for the post when sentiment_analysis is enabled. |
sentiment_confidence | number | – | Heuristic confidence score for the sentiment result when sentiment_analysis is enabled. |
sentiment_label | string | – | Sentiment label derived from the post sentiment analysis. See Sentiment Labels for possible values. |
content_category_label | string | – | Readable content category name chosen for the post when content_analysis is enabled. |
content_category_path | array | – | Topic path from the top-level content category down to the matched post category when content_analysis is enabled. |
author | string | – | Username shown on the post. |
score | number | – | Post score (upvotes minus downvotes) at collection time. |
upvote_ratio | number | – | Upvote ratio (0.0–1.0) when available. |
num_comments | number | – | Comment count shown on the post. |
subreddit | string | – | Subreddit name (without r/). |
created_utc | string | – | Post creation time in ISO format. This is the timestamp used for exact post date filtering when postsCreatedAfter or postsCreatedBefore is provided. |
url | string | ✅ | Canonical Reddit URL for the post. |
permalink | string | – | Relative Reddit permalink. |
canonical_url | string | – | Canonical full URL. |
old_reddit_url | string | – | Alternate legacy Reddit URL (old.reddit.com). |
flair | string | – | Post flair text. |
post_hint | string | – | Reddit post hint, useful for downstream media classification. |
over_18 | boolean | – | Whether the post is marked NSFW. |
is_self | boolean | – | Whether the post is a self/text post. |
spoiler | boolean | – | Whether the post is marked as a spoiler. |
locked | boolean | – | Whether the post is locked. |
is_video | boolean | – | Whether the post is a video post. |
is_gallery | boolean | – | Whether the post is a Reddit image gallery. |
hidden | boolean | – | Whether the post is hidden for the viewing account. |
edited | boolean/number | – | false when untouched, otherwise Reddit's edited timestamp payload. |
archived | boolean | – | Whether the post is archived. |
pinned | boolean | – | Whether the post is pinned in the subreddit. |
domain | string | – | Source or linked domain. |
thumbnail | string | – | Thumbnail URL or Reddit thumbnail marker. |
url_overridden_by_dest | string | – | Final outbound destination URL when present. |
num_duplicates | number | – | Duplicate count reported by Reddit. |
subreddit_id | string | – | Internal Reddit subreddit reference. |
subreddit_name_prefixed | string | – | Prefixed subreddit label such as r/Baking. |
subreddit_subscribers | number | – | Subscriber count at collection time. |
media | object | – | Media object when available. |
media_metadata | object | – | Raw media metadata keyed by media ID. |
gallery_data | object | – | Reddit gallery metadata with item list. |
gallery_images | array | – | Normalized gallery image list with dimensions, URLs, and previews. |
media_assets | array | – | Normalized media asset list with type, MIME, original URL, and preview URLs. |
age_hours | number | – | Post age in hours at collection time. |
retrieved_at | string | – | Actor capture time in ISO format. |
media_type | string | – | Normalized media class: text, image, gallery, video, gif, or link. |
has_media | boolean | – | Convenience flag for image, gallery, GIF, or video posts. |
gallery_count | number | – | Number of normalized gallery images. |
outbound_url_host | string | – | Parsed host from url_overridden_by_dest when present. |
title_length | number | – | Character length of the title. |
body_length | number | – | Character length of the body text. |
word_count | number | – | Whitespace-based word count across title and body. |
score_per_hour | number | – | Score divided by post age with a minimum age floor. |
comments_per_hour | number | – | Comment count divided by post age with a minimum age floor. |
is_deleted_or_removed | boolean | – | Deletion/removal flag derived from visible placeholders and removal metadata. |
engagement_total | number | – | Combined engagement metric derived from score and comments. |
comment_to_score_ratio | number | – | Comments-to-score ratio. |
is_high_engagement | boolean | – | Convenience flag for high engagement. |
content_flags | array | – | Content classification flags when present. |
stickied | boolean | – | Whether the post is pinned/stickied. |
distinguished | string | – | Distinguishing label, such as moderator status. |
score_hidden | boolean | – | Whether Reddit reports the post score as hidden. |
total_awards_received | number | – | Total awards on the post. |
all_awardings | array | – | Raw awards list. |
gilded | number | – | Gilding count. |
num_crossposts | number | – | Number of crossposts. |
is_original_content | boolean | – | Whether the post is marked original content. |
author_fullname | string | – | Internal Reddit author reference when available. |
author_flair_text | string | – | Author flair text. |
author_premium | boolean | – | Whether the author has Reddit Premium status. |
body_html | string | – | HTML-formatted post body. |
preview | object | – | Preview object when available. |
secure_media | object | – | Secure media object when available. |
secure_media_embed | object | – | Secure media embed metadata. |
crosspost_parent_list | array | – | Crosspost parent data when available. |
Comment Fields (kind = "comment")
| Field | Type | Required | Description |
|---|---|---|---|
kind | string | ✅ | Record category — always "comment". |
query | string | – | Input query or source label that produced the record. |
id | string | ✅ | Stable Reddit comment identifier. |
postId | string | ✅ | Parent post identifier. |
postUrl | string | ✅ | Parent post URL. |
parentId | string | ✅ | Parent Reddit object identifier (post or parent comment). |
body | string | – | Comment body text. |
sentiment_score | number | – | Raw sentiment score computed from the comment body when sentiment_analysis is enabled. |
sentiment_score_normalized | number | – | Length-normalized bounded sentiment score for the comment when sentiment_analysis is enabled. |
sentiment_confidence | number | – | Heuristic confidence score for the comment sentiment result when sentiment_analysis is enabled. |
sentiment_label | string | – | Sentiment label derived from comment sentiment analysis. See Sentiment Labels for possible values. |
author | string | – | Username shown on the comment. |
score | number | – | Comment score at collection time. |
subreddit | string | – | Subreddit name when Reddit provides it on the comment payload. |
created_utc | string | – | Comment creation time in ISO format. |
url | string | ✅ | Canonical Reddit URL for the comment. |
permalink | string | – | Relative Reddit permalink. |
canonical_url | string | – | Canonical full URL. |
old_reddit_url | string | – | Alternate legacy Reddit URL. |
root_comment_id | string | – | Root comment ID for the thread. |
parent_kind | string | – | Parent record type: "post" or "comment". |
is_deleted_or_removed | boolean | – | Deletion/removal flag. |
subreddit_id | string | – | Internal Reddit subreddit reference. |
subreddit_name_prefixed | string | – | Prefixed subreddit label such as r/Baking. |
edited | boolean/number | – | false when untouched, otherwise Reddit's edited timestamp. |
retrieved_at | string | – | Actor capture time in ISO format. |
age_hours | number | – | Comment age in hours at collection time. |
body_length | number | – | Character length of the comment body. |
word_count | number | – | Whitespace-based word count for the comment body. |
score_per_hour | number | – | Score divided by comment age with a minimum age floor. |
stickied | boolean | – | Whether the comment is pinned. |
distinguished | string | – | Distinguishing label, such as moderator status. |
is_submitter | boolean | – | Whether the author is the original post creator. |
score_hidden | boolean | – | Whether the score is hidden. |
controversiality | number | – | Reddit controversiality indicator. |
depth | number | – | Nesting depth in the comment tree (0 = top-level reply). |
total_awards_received | number | – | Total awards on the comment. |
all_awardings | array | – | Raw awards list. |
gilded | number | – | Gilding count. |
author_fullname | string | – | Internal Reddit author reference when available. |
author_flair_text | string | – | Author flair text. |
author_premium | boolean | – | Whether the author has Reddit Premium. |
body_html | string | – | HTML-formatted comment body. |
collapsed | boolean | – | Whether Reddit reports the comment as collapsed. |
collapsed_reason | string | – | Reason Reddit provides for a collapsed comment. |
collapsed_because_crowd_control | boolean | – | Whether crowd control collapsed the comment. |
unrepliable_reason | string | – | Reason Reddit provides when the comment cannot be replied to. |
Data Guarantees & Handling
- Best-effort extraction: Fields may vary by region, session, availability, and Reddit surface changes or experiments.
- Optional fields: Always null-check in downstream code because many fields may be empty or unavailable.
- Time filtering:
timeframenarrows the Reddit source query, whilepostsCreatedAfter/postsCreatedBeforeapply exact record-level filtering in the actor output. - Maximum coverage mode: When
maximize_coverage=true, the actor prioritizes recall over a single Reddit ranking window and may use broader overlap-safe traversal before deduplicating records. - Deduplication: Recommend
kind + ":" + id. Stable identifiers make deduplication and upserts straightforward when the same entity is discovered through overlapping inputs.
How to Run on Apify
- Open the Actor in Apify Console.
- Select your Scraping Mode from the dropdown (Search, Subreddits, Posts, Users, Direct URLs, or Discover).
- Configure your inputs — keywords, subreddit names, usernames, or URLs depending on your chosen mode.
- Set optional filters — score thresholds, date ranges, flair keywords, post types, domain restrictions.
- Toggle AI Analytics — enable Sentiment Analysis and/or Content Taxonomy if needed.
- Set your limits — configure maximum items per source and total items.
- Click Start and wait for the run to finish.
- Download results in JSON, CSV, Excel, XML, or other supported formats.
Scheduling & Automation
Automated Data Collection
You can schedule recurring runs to keep your Reddit dataset current without manual work. This is useful for monitoring trends, tracking brand mentions, and maintaining fresh inputs for dashboards or data pipelines.
- Navigate to Schedules in Apify Console
- Create a new schedule (daily, weekly, or custom cron expression)
- Configure input parameters for your recurring run
- Enable notifications for run completion
- Optional: add webhooks for automated downstream processing
Integration Options
| Integration | Use Case |
|---|---|
| Webhooks | Trigger downstream actions when a run completes |
| Zapier | Connect to 5,000+ apps without coding |
| Make (Integromat) | Build multi-step automation workflows |
| Google Sheets | Auto-export results to a spreadsheet |
| Slack / Discord | Receive notifications and run summaries |
| Send automated reports via email | |
| API | Programmatic access to datasets and runs |
API Access
Start runs and retrieve results programmatically using the Apify API:
# Python SDK examplefrom apify_client import ApifyClientclient = ApifyClient("YOUR_API_TOKEN")run = client.actor("YOUR_ACTOR_ID").call(run_input={"mode": "subreddits","subreddits": ["technology", "startups"],"sort": "hot","maxPostsPerSubreddit": 50,"sentiment_analysis": True,})dataset_items = client.dataset(run["defaultDatasetId"]).list_items().itemsfor item in dataset_items:print(item["title"], item["score"], item.get("sentiment_label"))
// JavaScript SDK exampleimport { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });const run = await client.actor('YOUR_ACTOR_ID').call({mode: 'search',queries: ['machine learning', 'deep learning'],sentiment_analysis: true,content_analysis: true,maxTotalItems: 200,});const { items } = await client.dataset(run.defaultDatasetId).listItems();console.log(`Collected ${items.length} records`);
Performance
Estimated run times:
| Run Size | Estimated Time |
|---|---|
| Small (< 100 items) | ~30 seconds – 1 minute |
| Medium (100–500 items) | ~1–5 minutes |
| Large (500–2,000 items) | ~5–15 minutes |
| Very Large (2,000+ items) | ~15–30 minutes |
For planning purposes, many runs targeting around 100–500 outputs fall into the small-to-medium range, but execution time varies based on the number of items, comment depth, filter complexity, result volume, and whether AI analytics are enabled. Runs with maximize_coverage enabled may take longer due to multi-view traversal and deduplication.
Use Cases & Recipes
Brand Monitoring
Track mentions of your brand, product, or competitors across all of Reddit:
{"mode": "search","queries": ["YourBrand", "YourProduct", "CompetitorName"],"maximize_coverage": true,"sentiment_analysis": true,"maxTotalItems": 500}
Product Feedback Mining
Find user complaints, feature requests, and pain points:
{"mode": "subreddits","subreddits": ["YourProductSubreddit"],"sort": "new","timeframe": "month","includeComments": true,"maxCommentsPerPost": 100,"sentiment_analysis": true,"content_analysis": true,"maxPostsPerSubreddit": 200}
Trend Tracking & Rising Topics
Monitor emerging topics and viral content:
{"mode": "subreddits","subreddits": ["technology", "Futurology", "singularity"],"sort": "rising","minScore": 100,"maxPostsPerSubreddit": 50,"content_analysis": true}
Content Research for SEO
Mine Reddit for real user language, questions, and topics for content creation:
{"mode": "search","queries": ["how to", "best way to", "recommend"],"searchSubreddit": "YourNicheSubreddit","sort": "top","timeframe": "year","minComments": 20,"maxTotalItems": 300}
Academic Research Dataset
Collect structured discussion data for analysis:
{"mode": "search","queries": ["climate change discussion"],"postsCreatedAfter": "2025-01-01","postsCreatedBefore": "2025-12-31","includeComments": true,"flattenComments": true,"maxCommentsPerPost": 200,"sentiment_analysis": true,"content_analysis": true,"maximize_coverage": true,"maxTotalItems": 2000}
Compliance & Ethics
Responsible Data Collection
This actor collects publicly available Reddit posts, comments, and discussion metadata from https://www.reddit.com for legitimate business purposes, including:
- Consumer research and market analysis
- Brand monitoring and competitive tracking
- Product feedback discovery and trend analysis
- Academic research and discourse analysis
Users are responsible for making sure their use of the collected data complies with applicable laws, regulations, internal policies, and the target site's terms. This section is informational and not legal advice.
Best Practices
- Use collected data in accordance with applicable laws, regulations, and the target site's terms
- Respect individual privacy and personal information
- Use data responsibly and avoid disruptive or excessive collection
- Do not use this actor for spamming, harassment, or other harmful purposes
- Follow relevant data protection requirements where applicable, such as GDPR and CCPA
FAQ
Q: How many posts can I extract per run?
Free users can extract up to 20 items per run (4 runs/day). Paid subscribers have no limits — configure maxTotalItems as high as you need.
Q: Can I extract comments from posts?
Yes. Set includeComments to true and configure maxCommentsPerPost to control how many comments per thread. Use flattenComments to save each comment as a separate dataset row.
Q: What sorting options are available?
Posts can be sorted by hot, new, top, rising, controversial, or relevance. Comments can be sorted by confidence (Best), top, new, controversial, old, or qa.
Q: Can I filter posts by date range?
Yes. Use postsCreatedAfter and postsCreatedBefore with YYYY-MM-DD format to define exact date windows.
Q: What is "Maximize Coverage" mode? When enabled, the actor expands collection across multiple Reddit ranking surfaces (relevance, top, new, comments) and deduplicates overlapping results. Use it when you need broader topic coverage rather than a single ranked list.
Q: Does sentiment analysis require an external API key? No. AI Sentiment Analysis and Content Taxonomy Classification are built into the actor. No external API keys, subscriptions, or configuration needed — just toggle them on.
Q: Can I schedule recurring runs? Yes. Use Apify's built-in scheduling to run daily, weekly, or on custom cron expressions. Combine with webhooks, Zapier, or Make for automated workflows.
Q: What output formats are supported? JSON, CSV, Excel (XLSX), XML, HTML Table, and RSS. Download directly from Apify Console or access via API.
Q: How do I scrape a specific post or thread?
Use mode: "posts" or mode: "directUrls" and paste the full Reddit URL into the urls field. Enable includeComments to also extract the comment thread.
Q: Can I search within a specific subreddit?
Yes. Use mode: "search" with your keywords in queries and set searchSubreddit to the subreddit name (without r/).
Support
For help, use the Issues tab on the actor page in Apify Console. Include:
- The input configuration you used (with sensitive values redacted)
- The run ID from Apify Console
- Expected behavior vs. actual behavior
- A small output sample if helpful
Categories
Social Media · Data Extraction · Reddit · Reddit Scraper · Automation · Market Research · Sentiment Analysis · Web Scraping · AI Analytics · Brand Monitoring