Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s avatar

Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s

Pricing

from $1.60 / 1,000 results

Go to Apify Store
Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s

Reddit Scraper - Posts, Comments & Subreddits | 10 Results/1s

Lightning-fast Reddit scraping at 10 items in just 1 sec. Extract posts, comments, user profiles, subreddit feeds, and trending communities into structured JSON for research, monitoring, sentiment analysis, competitive intelligence, and lead generation

Pricing

from $1.60 / 1,000 results

Rating

0.0

(0)

Developer

Mikolabs

Mikolabs

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 days ago

Last modified

Categories

Share

Reddit Scraper — Complete Suite for Posts, Comments, Users & Subreddits

Extract Reddit posts, full comment threads, user profiles, subreddit feeds, and trending communities at scale. Collect structured JSON data from Reddit for market research, brand monitoring, sentiment analysis, competitive intelligence, lead generation, and content strategy — no coding required.

Overview

Reddit Scraper collects publicly available Reddit posts and, when enabled, comment records with the metadata teams usually need for analysis, enrichment, and automation. Each record can include core content, authorship, engagement metrics, subreddit context, timestamps, canonical URLs, flair data, and media-related fields when available. Reddit is one of the web's richest sources of real-world opinions, product feedback, community discussion, and trend signals, which makes it valuable for research, monitoring, and downstream decision-making.

This actor turns that information into consistent dataset records so teams can replace repetitive manual collection with repeatable runs. It supports six scraping modes — keyword search, subreddit feeds, post threads with nested comments, user profiles, direct URL extraction, and subreddit discovery — all configurable through a simple point-and-click interface. It is designed for production use with stable identifiers, overlap-friendly outputs for deduplication, and reliable scheduled collection workflows.

Whether you need 20 posts for a quick check or thousands of records for a research pipeline, this actor handles it with automatic pagination, deduplication, and reliable rate management.


Why Use This Actor

  • Market research and analytics teams: Track conversation volume, engagement, subreddit activity, and topic trends across keywords, communities, and time windows to inform strategic decisions.
  • Product and content teams: Discover user pain points, feature requests, language patterns, and high-performing discussion themes for product roadmap planning and editorial calendars.
  • Brand monitoring and competitive intelligence teams: Watch competitor mentions, product launch reactions, brand sentiment shifts, category-level discussion spikes, and recurring conversation patterns without manually checking Reddit every day.
  • Developers and data engineering teams: Feed Reddit data into ETL pipelines, data warehouses, dashboards, and APIs using structured JSON records that are easy to upsert and model.
  • Lead generation and enrichment teams: Identify relevant communities, active buyer discussions, context around buyer interests, brand mentions, and niche topics.
  • Academic and social researchers: Collect discussion datasets for discourse analysis, public opinion research, community dynamics studies, and longitudinal content tracking.
  • SEO and digital marketing teams: Mine Reddit for content ideas, keyword research, real user language, and high-engagement topics to improve search rankings, ad copy, and editorial strategy.
  • Operations and automation teams: Run recurring jobs on a schedule and use stable record keys to deduplicate overlapping results from queries, subreddit searches, and direct URLs.

Pricing & Plans (No Hidden Fees)

Transparent and predictable pricing with no extra proxy costs, no setup fees, and no hidden maintenance charges.

Tiered Pricing Structure

Tier / Discount LevelPrice per 1,000 ResultsEffective SavingsMinimum Scrape
No Discount (Standard / Pay-As-You-Go)$4.00 / 1,000 itemsStandard Rate1 item
🥉 Bronze Discount$2.00 / 1,000 items50% OFF20 items
🥈 Silver Discount$1.80 / 1,000 items55% OFF20 items
🥇 Gold Discount$1.60 / 1,000 items60% OFF20 items

Plan Comparison

FeatureFree TierSubscriber / Paid Tier
Free Daily Allowance20 items / run (4 runs / day free)Unlimited
Pricing$4.00 / 1,000 results (or free allowance)Down to $1.60 / 1,000 results
Additional Fees$0.00 (No extra fees)$0.00 (No extra fees)
Proxy / Bandwidth CostsIncluded ($0.00)Included ($0.00)
All 6 Scraping Modes
AI Sentiment Analysis
AI Content Taxonomy
Advanced Filters
Run Summary Dashboard
Scheduling & Automation

Free users can test every feature (search, subreddits, comments, user profiles, AI sentiment, content taxonomy) with up to 20 items per run and 4 runs per day completely free of charge. Upgrade for volume discounts down to $1.60 / 1,000 items with zero hidden fees.


How to Use Reddit Scraper — Step by Step

Step 1: Choose a Scraping Mode

Select one of six modes from the Scraping Mode dropdown at the top of the input form:

ModeWhat It DoesWhen to Use
🔍 SearchSearch Reddit by keywords and phrasesFinding posts about a topic, brand, or product across all of Reddit
🌐 SubredditsScrape subreddit feeds (Hot, Top, New, Rising)Monitoring specific communities and collecting feed-based datasets
💬 PostsExtract specific post threads with nested commentsDeep-diving into individual discussions and comment trees
👤 UsersCollect user profiles, karma, and submission historyResearching influencers, key accounts, or user activity patterns
🔗 Direct URLsScrape any Reddit URL you provideExtracting specific pages, threads, or search results you already found
🧭 DiscoverFind popular and trending subredditsDiscovering new communities, niches, and emerging topics

Step 2: Configure Your Mode

Fill in the settings section that corresponds to your chosen mode:

  • Search mode → Enter your keywords in "Search Keywords / Phrases"
  • Subreddits mode → Enter subreddit names in "Subreddits to Scrape"
  • Users mode → Enter Reddit usernames in "Reddit Usernames"
  • Direct URLs mode → Paste Reddit URLs in "Direct Reddit URLs"
  • Posts mode → Paste specific post URLs in "Direct Reddit URLs"
  • Discover mode → No additional input needed, the actor finds trending communities automatically

Step 3: Set Filters (Optional)

Use the Filters & Content Quality Bounds section to narrow your results by post type (text, image, video, gallery, link), minimum/maximum upvote score, date ranges, flair keywords, title or body keywords, domain restrictions, and NSFW inclusion/exclusion.

Step 4: Enable AI Analytics (Optional)

Toggle AI Sentiment Analysis and/or AI Content Taxonomy Classification to automatically enrich every record with sentiment scores and topic labels. No external API keys required.

Step 5: Set Limits & Run

Configure your item limits in the Limits & Performance section, then click Start to begin the extraction. Results are saved to your Apify dataset as clean JSON, ready for download in JSON, CSV, Excel, XML, or RSS formats.


Input Parameters

Provide any combination of modes, queries, URLs, and filters to control what the actor collects and how focused the results should be.

🎯 Step 1: Mode Selection

ParameterTypeDefaultDescription
modestringsearchScraping mode. Options: search, subreddits, posts, users, directUrls, discover

🔍 Step 2 — Search Settings (Mode = search)

Use when you want to discover relevant posts across all of Reddit matching your keywords.

ParameterTypeDefaultDescription
queriesstring[]Keywords or search phrases to look for across Reddit. Accepts multiple terms. Use this when you want the actor to discover relevant posts for you.
searchSubredditstringOptional: restrict search to a single subreddit (without r/). Leave empty to search all of Reddit.
strictSearchbooleanfalseMake Reddit search follow your keywords more closely, relying more on exact keywords and less on loose semantic matching. This usually returns fewer posts, but keeps results closer to your query.
strictTokenFilterbooleanfalseMake the actor scan each saved post's title, body, and URL and keep only posts that match all of your query keywords. This reduces output size, but keeps the most accurate results.

🌐 Step 2 — Subreddit & Feed Settings (Mode = subreddits)

Use when you want to collect posts from one or more subreddit feeds.

ParameterTypeDefaultDescription
subredditsstring[]["AskReddit", "technology"]List of subreddit names to scrape (without r/). Accepts multiple subreddits.
sortstringhotRanking order for subreddit feeds or search results. Options: hot (Trending), new (Latest), top (Most Upvoted), rising, controversial, relevance
timeframestringallTime window for top and controversial sorts. Options: hour, day, week, month, year, all. Use postsCreatedAfter and postsCreatedBefore for exact record-level date filtering.
fullSubredditModebooleanfalseDeep crawl mode — paginates as far back as the Reddit API permits for maximum historical coverage.

🔗 Step 2 — Direct URLs & Post Links (Mode = posts or directUrls)

Use when you have specific Reddit links to extract.

ParameterTypeDefaultDescription
urlsstring[]Direct Reddit URLs to scrape, such as post URLs, comment permalinks, subreddit pages, user pages, or Reddit search pages. When provided, URL input takes priority over search queries.

👤 Step 2 — User Profile Settings (Mode = users)

Use when you want to collect user profiles, karma scores, and submission history.

ParameterTypeDefaultDescription
usernamesstring[]["spez"]Reddit usernames to extract (without u/).
userContentstringallWhat activity to collect for user profiles. Options: all (Profile + Posts + Comments), submitted (Posts Only), comments (Comments Only)

💬 Comments & Discussion Settings

Configure comment extraction for any post-based mode.

ParameterTypeDefaultDescription
includeCommentsbooleanfalseWhen enabled, the actor also saves comments from each collected post. Useful for sentiment analysis, deeper discussion review, and thread-level context.
maxCommentsPerPostinteger25Maximum number of comments to collect per post when comment collection is enabled. Lower values help keep runs faster and datasets smaller.
commentsSortstringconfidenceRanking order for comments. Options: confidence (Best), top (Most Upvoted), new, controversial, old, qa (Q&A)
flattenCommentsbooleanfalseWhen true, saves each comment as a separate independent row in the dataset instead of nesting under the parent post.

⚡ Filters & Content Quality Bounds

Fine-tune results with granular post-level filters. All filters are optional and can be combined.

ParameterTypeDefaultDescription
postTypestringallFilter by media format. Options: all, text (Self Text Only), image, video, gallery, link (External Links Only)
minScoreintegerOnly keep posts with at least this upvote score.
maxScoreintegerOnly keep posts at or below this upvote score.
minCommentsintegerOnly keep posts with at least this many comments.
minUpvoteRationumberFilter by upvote ratio (e.g. 0.85 for ≥85% positive).
postsCreatedAfterstringOnly keep posts created on or after this date (YYYY-MM-DD). Plain dates are normalized to the start of the day in UTC.
postsCreatedBeforestringOnly keep posts created on or before this date (YYYY-MM-DD). Plain dates are normalized to the end of the day in UTC.
onlyWithFlairbooleanfalseOnly keep posts that have an assigned flair tag.
flairContainsstringMatch posts where the flair text contains this keyword.
titleContainsstringOnly keep posts whose title contains this keyword.
textContainsstringOnly keep posts whose body text contains this keyword.
domainContainsstringOnly keep posts linking to this domain (e.g. github.com, nytimes.com).
excludeStickiedbooleanfalseExclude pinned/stickied moderator announcement posts from results.
excludeKeywordsstring[]Exclude any posts containing these keywords in title or body.
includeNsfwbooleantrueInclude posts marked as NSFW or 18+ in the output dataset.

🤖 AI Analytics & Enrichment

Add AI-powered analysis to every extracted record — no external API keys needed.

ParameterTypeDefaultDescription
sentiment_analysisbooleanfalseWhen enabled, adds sentiment_score, sentiment_confidence, and sentiment_label (positive, negative, neutral, mixed, uncertain) to posts and comments.
content_analysisbooleanfalseWhen enabled, classifies each post against a content taxonomy and adds content_category_label and content_category_path to post records.

Sentiment Labels

When sentiment_analysis is enabled, sentiment_label can have these values:

LabelMeaningExample
positiveThe text has a clear positive tone or favorable sentiment."This product is amazing, best purchase I've made!"
negativeThe text has a clear negative tone or unfavorable sentiment."Terrible experience, would not recommend to anyone."
neutralThe text has little or no sentiment signal, or the sentiment score is close to zero."The package arrived on Tuesday with the latest invoice attached."
mixedThe text has both positive and negative cues, and they remain unresolved or balanced."I love the cast and visuals. The story is boring and the ending is awful."
uncertainThe analyzer does not have enough confidence to classify the tone, or the text is empty, bot boilerplate, or moderator boilerplate."I am a bot, and this action was performed automatically."

Content Taxonomy Classification

When content_analysis is enabled, the actor assigns post-level topic labels from a stable content taxonomy and returns both a human-readable category and its taxonomy path. The labeling is designed for research pipelines, monitoring systems, and workflows that need a fast topical read of large Reddit collections before routing, clustering, summarizing, or enriching records downstream.

The classifier is intentionally conservative: when the record does not contain enough topic evidence, the actor leaves the category fields empty instead of forcing a low-confidence label. This keeps category fields useful as high-signal routing metadata for analysts, dashboards, warehouse models, and LLM agents.

  • content_category_label — Human-readable category name (e.g., "Artificial Intelligence")
  • content_category_path — Full taxonomy path (e.g., ["Technology", "Artificial Intelligence"])

⚙️ Limits & Performance

ParameterTypeDefaultDescription
maxPostsPerSubredditinteger100Maximum number of posts to collect per subreddit or query source.
maxTotalItemsinteger500Safety ceiling for the total number of records across all inputs in a single run.
maximize_coveragebooleanfalseTurn on high-coverage collection for broad or competitive topics. See Maximum Coverage Mode below.

Maximum Coverage Mode

Reddit search and listing surfaces do not behave like unbounded database queries. For many seeds, Reddit exposes only a limited practical result window, commonly around 250 posts per query/ranking combination, even when the topic has far more historical discussion. A single relevance, top, hot, new, or comments traversal can therefore look complete while still missing major clusters of posts, subreddits, and time periods.

When maximize_coverage is enabled, the actor treats each seed as a research topic instead of a single result page. It expands the collection plan across complementary result views, date-aware traversals, and high-signal community follow-ups, then deduplicates overlapping posts into one stable dataset. This is useful for market research, incident monitoring, topic mapping, retrieval corpora, and workflows that need a broader evidence base before summarization, clustering, lead scoring, or trend analysis.

This mode can use more requests and take longer than a standard run, but it gives downstream systems a better chance of seeing the shape of the conversation rather than only the first capped slice of a single ranking surface. Use it when recall matters more than preserving a single ranking order.


Example Inputs

Scenario: Query-driven monitoring

{
"mode": "search",
"queries": ["ai video generator", "synthetic media"],
"sort": "new",
"timeframe": "week",
"postsCreatedAfter": "2026-03-01",
"postsCreatedBefore": "2026-03-31",
"maximize_coverage": true,
"strictSearch": true,
"strictTokenFilter": true,
"maxTotalItems": 500
}

Scenario: Direct URL collection with comments and AI

{
"mode": "directUrls",
"urls": [
"https://www.reddit.com/r/technology/",
"https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/"
],
"includeComments": true,
"sentiment_analysis": true,
"content_analysis": true,
"maxCommentsPerPost": 200,
"maxTotalItems": 150
}

Scenario: Targeted subreddit feed

{
"mode": "subreddits",
"subreddits": ["startups", "SaaS", "Entrepreneur"],
"sort": "top",
"timeframe": "month",
"minScore": 50,
"minComments": 10,
"excludeStickied": true,
"includeComments": true,
"maxCommentsPerPost": 100,
"commentsSort": "top",
"maxPostsPerSubreddit": 200,
"maxTotalItems": 500
}

Scenario: User profile research

{
"mode": "users",
"usernames": ["spez", "GovSchwarzenegger"],
"userContent": "all",
"maxTotalItems": 200
}

Scenario: Advanced filtered search with date range

{
"mode": "search",
"queries": ["product recall", "safety issue"],
"sort": "new",
"postsCreatedAfter": "2026-01-01",
"postsCreatedBefore": "2026-06-30",
"minScore": 25,
"postType": "text",
"excludeKeywords": ["meme", "shitpost"],
"excludeStickied": true,
"includeNsfw": false,
"maximize_coverage": true,
"sentiment_analysis": true,
"content_analysis": true,
"maxPostsPerSubreddit": 500,
"maxTotalItems": 2000
}
{
"mode": "discover",
"maxTotalItems": 100
}

Output

Output Destination

The actor writes results to an Apify dataset as JSON records. The dataset is designed for direct consumption by analytics tools, ETL pipelines, and downstream APIs without post-processing. Download in JSON, CSV, Excel (XLSX), XML, HTML Table, or RSS formats directly from the Apify Console or via API.

Record Envelope (All Items)

Every record includes a stable category field, a Reddit identifier, and a canonical URL:

  • kind (string, required): Logical record type, such as post or comment.
  • id (string, required): Stable Reddit identifier for the entity.
  • url (string, required): Canonical Reddit URL for the record.

Recommended idempotency key: kind + ":" + id

Use this key for deduplication and upserts, especially when the same Reddit entity appears in overlapping queries, subreddit runs, or direct URL inputs.

Run Summary Artifact

At the end of each successful run, the actor writes RUN-SUMMARY to the key-value store with totals, query and subreddit breakdowns, date range, engagement highlights, optional sentiment/category counts, tier information, and a complete run overview.

The visual report is saved as RUN-MAP.html and renders the summary values as an interactive dashboard with stats cards, engagement breakdowns, mode indicators, and tier badges. Access both from the Output tab after your run completes.


Example Output: Post Record (kind = "post")

{
"kind": "post",
"query": "cheesecake",
"id": "1hvoazn",
"title": "My best cheesecake so far",
"body": "Found my new favorite recipe (no water bath). Next time I will make a thicker crust. Added a raspberry compote.",
"sentiment_score": 2,
"sentiment_label": "positive",
"sentiment_confidence": 0.78,
"sentiment_score_normalized": 0.86,
"content_category_label": "Desserts and Baking",
"content_category_path": ["Food & Drink", "Desserts and Baking"],
"author": "ClearlyBulky",
"score": 3489,
"upvote_ratio": 1,
"num_comments": 43,
"subreddit": "Baking",
"created_utc": "2025-01-07T10:09:56.000Z",
"url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"permalink": "/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"canonical_url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"old_reddit_url": "https://old.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"flair": "Recipe",
"post_hint": "link",
"over_18": false,
"is_self": false,
"spoiler": false,
"locked": false,
"is_video": false,
"is_gallery": true,
"hidden": false,
"edited": false,
"archived": false,
"pinned": false,
"domain": "old.reddit.com",
"thumbnail": "https://b.thumbs.redditmedia.com/example.jpg",
"url_overridden_by_dest": "https://www.reddit.com/gallery/1hvoazn",
"num_duplicates": 0,
"subreddit_id": "t5_2qx1h",
"subreddit_name_prefixed": "r/Baking",
"subreddit_subscribers": 4322940,
"media": null,
"media_metadata": {},
"gallery_data": {
"items": [
{ "is_deleted": false, "media_id": "kny1nmhlqjbe1", "id": 581711947 },
{ "is_deleted": false, "media_id": "wjqc6mhlqjbe1", "id": 581711948 }
]
},
"gallery_images": [
{
"media_id": "kny1nmhlqjbe1",
"caption": "",
"width": 3024,
"height": 4032,
"url": "https://preview.redd.it/kny1nmhlqjbe1.jpg?width=3024&format=pjpg&auto=webp",
"previews": ["https://preview.redd.it/kny1nmhlqjbe1.jpg?width=640&crop=smart&auto=webp"]
}
],
"media_assets": [
{
"type": "Image",
"media_id": "kny1nmhlqjbe1",
"mime_type": "image/jpg",
"original_url": "https://preview.redd.it/kny1nmhlqjbe1.jpg?width=3024&format=pjpg&auto=webp",
"preview_urls": ["https://preview.redd.it/kny1nmhlqjbe1.jpg?width=640&crop=smart&auto=webp"]
}
],
"age_hours": 10916.1333,
"retrieved_at": "2026-04-07T00:00:00.000Z",
"media_type": "gallery",
"has_media": true,
"gallery_count": 2,
"outbound_url_host": "www.reddit.com",
"title_length": 26,
"body_length": 112,
"word_count": 25,
"score_per_hour": 0.3196,
"comments_per_hour": 0.0039,
"is_deleted_or_removed": false,
"engagement_total": 3532,
"comment_to_score_ratio": 0.0123,
"is_high_engagement": true,
"content_flags": [],
"stickied": false,
"distinguished": null,
"score_hidden": false,
"total_awards_received": 0,
"all_awardings": [],
"gilded": 0,
"num_crossposts": 0,
"is_original_content": false,
"author_fullname": "t2_dr3vyilor",
"author_flair_text": null,
"author_premium": false,
"body_html": "<div class=\"md\"><p>Found my new favorite recipe...</p></div>",
"preview": null,
"secure_media": null,
"secure_media_embed": {},
"crosspost_parent_list": null
}

Example Output: Comment Record (kind = "comment")

{
"kind": "comment",
"query": "cheesecake",
"id": "m5un6bj",
"postId": "1hvoazn",
"postUrl": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/",
"parentId": "t3_1hvoazn",
"body": "This looks absolutely incredible! Can you share the full recipe?",
"sentiment_score": 3,
"sentiment_label": "positive",
"sentiment_confidence": 0.91,
"sentiment_score_normalized": 0.92,
"author": "BakingFanatic",
"score": 76,
"subreddit": "Baking",
"created_utc": "2025-01-07T10:13:48.000Z",
"url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/",
"permalink": "/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/",
"canonical_url": "https://www.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/",
"old_reddit_url": "https://old.reddit.com/r/Baking/comments/1hvoazn/my_best_cheesecake_so_far/m5un6bj/",
"root_comment_id": "m5un6bj",
"parent_kind": "post",
"is_deleted_or_removed": false,
"subreddit_id": "t5_2qx1h",
"subreddit_name_prefixed": "r/Baking",
"edited": false,
"retrieved_at": "2026-04-07T00:00:00.000Z",
"age_hours": 10916.0678,
"body_length": 63,
"word_count": 10,
"score_per_hour": 0.007,
"stickied": false,
"distinguished": null,
"is_submitter": false,
"score_hidden": false,
"controversiality": 0,
"depth": 0,
"total_awards_received": 0,
"all_awardings": [],
"gilded": 0,
"author_fullname": "t2_example123",
"author_flair_text": null,
"author_premium": false,
"body_html": "<div class=\"md\"><p>This looks absolutely incredible! Can you share the full recipe?</p></div>",
"collapsed": false,
"collapsed_reason": null,
"collapsed_because_crowd_control": false,
"unrepliable_reason": null
}

Field Reference

Post Fields (kind = "post")

FieldTypeRequiredDescription
kindstringRecord category — always "post" for post records.
querystringInput query or source label that produced the record.
idstringStable Reddit post identifier.
titlestringPost title.
bodystringPost body text.
sentiment_scorenumberRaw sentiment score computed from the post title and body when sentiment_analysis is enabled.
sentiment_score_normalizednumberLength-normalized bounded sentiment score for the post when sentiment_analysis is enabled.
sentiment_confidencenumberHeuristic confidence score for the sentiment result when sentiment_analysis is enabled.
sentiment_labelstringSentiment label derived from the post sentiment analysis. See Sentiment Labels for possible values.
content_category_labelstringReadable content category name chosen for the post when content_analysis is enabled.
content_category_patharrayTopic path from the top-level content category down to the matched post category when content_analysis is enabled.
authorstringUsername shown on the post.
scorenumberPost score (upvotes minus downvotes) at collection time.
upvote_rationumberUpvote ratio (0.0–1.0) when available.
num_commentsnumberComment count shown on the post.
subredditstringSubreddit name (without r/).
created_utcstringPost creation time in ISO format. This is the timestamp used for exact post date filtering when postsCreatedAfter or postsCreatedBefore is provided.
urlstringCanonical Reddit URL for the post.
permalinkstringRelative Reddit permalink.
canonical_urlstringCanonical full URL.
old_reddit_urlstringAlternate legacy Reddit URL (old.reddit.com).
flairstringPost flair text.
post_hintstringReddit post hint, useful for downstream media classification.
over_18booleanWhether the post is marked NSFW.
is_selfbooleanWhether the post is a self/text post.
spoilerbooleanWhether the post is marked as a spoiler.
lockedbooleanWhether the post is locked.
is_videobooleanWhether the post is a video post.
is_gallerybooleanWhether the post is a Reddit image gallery.
hiddenbooleanWhether the post is hidden for the viewing account.
editedboolean/numberfalse when untouched, otherwise Reddit's edited timestamp payload.
archivedbooleanWhether the post is archived.
pinnedbooleanWhether the post is pinned in the subreddit.
domainstringSource or linked domain.
thumbnailstringThumbnail URL or Reddit thumbnail marker.
url_overridden_by_deststringFinal outbound destination URL when present.
num_duplicatesnumberDuplicate count reported by Reddit.
subreddit_idstringInternal Reddit subreddit reference.
subreddit_name_prefixedstringPrefixed subreddit label such as r/Baking.
subreddit_subscribersnumberSubscriber count at collection time.
mediaobjectMedia object when available.
media_metadataobjectRaw media metadata keyed by media ID.
gallery_dataobjectReddit gallery metadata with item list.
gallery_imagesarrayNormalized gallery image list with dimensions, URLs, and previews.
media_assetsarrayNormalized media asset list with type, MIME, original URL, and preview URLs.
age_hoursnumberPost age in hours at collection time.
retrieved_atstringActor capture time in ISO format.
media_typestringNormalized media class: text, image, gallery, video, gif, or link.
has_mediabooleanConvenience flag for image, gallery, GIF, or video posts.
gallery_countnumberNumber of normalized gallery images.
outbound_url_hoststringParsed host from url_overridden_by_dest when present.
title_lengthnumberCharacter length of the title.
body_lengthnumberCharacter length of the body text.
word_countnumberWhitespace-based word count across title and body.
score_per_hournumberScore divided by post age with a minimum age floor.
comments_per_hournumberComment count divided by post age with a minimum age floor.
is_deleted_or_removedbooleanDeletion/removal flag derived from visible placeholders and removal metadata.
engagement_totalnumberCombined engagement metric derived from score and comments.
comment_to_score_rationumberComments-to-score ratio.
is_high_engagementbooleanConvenience flag for high engagement.
content_flagsarrayContent classification flags when present.
stickiedbooleanWhether the post is pinned/stickied.
distinguishedstringDistinguishing label, such as moderator status.
score_hiddenbooleanWhether Reddit reports the post score as hidden.
total_awards_receivednumberTotal awards on the post.
all_awardingsarrayRaw awards list.
gildednumberGilding count.
num_crosspostsnumberNumber of crossposts.
is_original_contentbooleanWhether the post is marked original content.
author_fullnamestringInternal Reddit author reference when available.
author_flair_textstringAuthor flair text.
author_premiumbooleanWhether the author has Reddit Premium status.
body_htmlstringHTML-formatted post body.
previewobjectPreview object when available.
secure_mediaobjectSecure media object when available.
secure_media_embedobjectSecure media embed metadata.
crosspost_parent_listarrayCrosspost parent data when available.

Comment Fields (kind = "comment")

FieldTypeRequiredDescription
kindstringRecord category — always "comment".
querystringInput query or source label that produced the record.
idstringStable Reddit comment identifier.
postIdstringParent post identifier.
postUrlstringParent post URL.
parentIdstringParent Reddit object identifier (post or parent comment).
bodystringComment body text.
sentiment_scorenumberRaw sentiment score computed from the comment body when sentiment_analysis is enabled.
sentiment_score_normalizednumberLength-normalized bounded sentiment score for the comment when sentiment_analysis is enabled.
sentiment_confidencenumberHeuristic confidence score for the comment sentiment result when sentiment_analysis is enabled.
sentiment_labelstringSentiment label derived from comment sentiment analysis. See Sentiment Labels for possible values.
authorstringUsername shown on the comment.
scorenumberComment score at collection time.
subredditstringSubreddit name when Reddit provides it on the comment payload.
created_utcstringComment creation time in ISO format.
urlstringCanonical Reddit URL for the comment.
permalinkstringRelative Reddit permalink.
canonical_urlstringCanonical full URL.
old_reddit_urlstringAlternate legacy Reddit URL.
root_comment_idstringRoot comment ID for the thread.
parent_kindstringParent record type: "post" or "comment".
is_deleted_or_removedbooleanDeletion/removal flag.
subreddit_idstringInternal Reddit subreddit reference.
subreddit_name_prefixedstringPrefixed subreddit label such as r/Baking.
editedboolean/numberfalse when untouched, otherwise Reddit's edited timestamp.
retrieved_atstringActor capture time in ISO format.
age_hoursnumberComment age in hours at collection time.
body_lengthnumberCharacter length of the comment body.
word_countnumberWhitespace-based word count for the comment body.
score_per_hournumberScore divided by comment age with a minimum age floor.
stickiedbooleanWhether the comment is pinned.
distinguishedstringDistinguishing label, such as moderator status.
is_submitterbooleanWhether the author is the original post creator.
score_hiddenbooleanWhether the score is hidden.
controversialitynumberReddit controversiality indicator.
depthnumberNesting depth in the comment tree (0 = top-level reply).
total_awards_receivednumberTotal awards on the comment.
all_awardingsarrayRaw awards list.
gildednumberGilding count.
author_fullnamestringInternal Reddit author reference when available.
author_flair_textstringAuthor flair text.
author_premiumbooleanWhether the author has Reddit Premium.
body_htmlstringHTML-formatted comment body.
collapsedbooleanWhether Reddit reports the comment as collapsed.
collapsed_reasonstringReason Reddit provides for a collapsed comment.
collapsed_because_crowd_controlbooleanWhether crowd control collapsed the comment.
unrepliable_reasonstringReason Reddit provides when the comment cannot be replied to.

Data Guarantees & Handling

  • Best-effort extraction: Fields may vary by region, session, availability, and Reddit surface changes or experiments.
  • Optional fields: Always null-check in downstream code because many fields may be empty or unavailable.
  • Time filtering: timeframe narrows the Reddit source query, while postsCreatedAfter / postsCreatedBefore apply exact record-level filtering in the actor output.
  • Maximum coverage mode: When maximize_coverage=true, the actor prioritizes recall over a single Reddit ranking window and may use broader overlap-safe traversal before deduplicating records.
  • Deduplication: Recommend kind + ":" + id. Stable identifiers make deduplication and upserts straightforward when the same entity is discovered through overlapping inputs.

How to Run on Apify

  1. Open the Actor in Apify Console.
  2. Select your Scraping Mode from the dropdown (Search, Subreddits, Posts, Users, Direct URLs, or Discover).
  3. Configure your inputs — keywords, subreddit names, usernames, or URLs depending on your chosen mode.
  4. Set optional filters — score thresholds, date ranges, flair keywords, post types, domain restrictions.
  5. Toggle AI Analytics — enable Sentiment Analysis and/or Content Taxonomy if needed.
  6. Set your limits — configure maximum items per source and total items.
  7. Click Start and wait for the run to finish.
  8. Download results in JSON, CSV, Excel, XML, or other supported formats.

Scheduling & Automation

Automated Data Collection

You can schedule recurring runs to keep your Reddit dataset current without manual work. This is useful for monitoring trends, tracking brand mentions, and maintaining fresh inputs for dashboards or data pipelines.

  1. Navigate to Schedules in Apify Console
  2. Create a new schedule (daily, weekly, or custom cron expression)
  3. Configure input parameters for your recurring run
  4. Enable notifications for run completion
  5. Optional: add webhooks for automated downstream processing

Integration Options

IntegrationUse Case
WebhooksTrigger downstream actions when a run completes
ZapierConnect to 5,000+ apps without coding
Make (Integromat)Build multi-step automation workflows
Google SheetsAuto-export results to a spreadsheet
Slack / DiscordReceive notifications and run summaries
EmailSend automated reports via email
APIProgrammatic access to datasets and runs

API Access

Start runs and retrieve results programmatically using the Apify API:

# Python SDK example
from apify_client import ApifyClient
client = ApifyClient("YOUR_API_TOKEN")
run = client.actor("YOUR_ACTOR_ID").call(run_input={
"mode": "subreddits",
"subreddits": ["technology", "startups"],
"sort": "hot",
"maxPostsPerSubreddit": 50,
"sentiment_analysis": True,
})
dataset_items = client.dataset(run["defaultDatasetId"]).list_items().items
for item in dataset_items:
print(item["title"], item["score"], item.get("sentiment_label"))
// JavaScript SDK example
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_API_TOKEN' });
const run = await client.actor('YOUR_ACTOR_ID').call({
mode: 'search',
queries: ['machine learning', 'deep learning'],
sentiment_analysis: true,
content_analysis: true,
maxTotalItems: 200,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(`Collected ${items.length} records`);

Performance

Estimated run times:

Run SizeEstimated Time
Small (< 100 items)~30 seconds – 1 minute
Medium (100–500 items)~1–5 minutes
Large (500–2,000 items)~5–15 minutes
Very Large (2,000+ items)~15–30 minutes

For planning purposes, many runs targeting around 100–500 outputs fall into the small-to-medium range, but execution time varies based on the number of items, comment depth, filter complexity, result volume, and whether AI analytics are enabled. Runs with maximize_coverage enabled may take longer due to multi-view traversal and deduplication.


Use Cases & Recipes

Brand Monitoring

Track mentions of your brand, product, or competitors across all of Reddit:

{
"mode": "search",
"queries": ["YourBrand", "YourProduct", "CompetitorName"],
"maximize_coverage": true,
"sentiment_analysis": true,
"maxTotalItems": 500
}

Product Feedback Mining

Find user complaints, feature requests, and pain points:

{
"mode": "subreddits",
"subreddits": ["YourProductSubreddit"],
"sort": "new",
"timeframe": "month",
"includeComments": true,
"maxCommentsPerPost": 100,
"sentiment_analysis": true,
"content_analysis": true,
"maxPostsPerSubreddit": 200
}

Trend Tracking & Rising Topics

Monitor emerging topics and viral content:

{
"mode": "subreddits",
"subreddits": ["technology", "Futurology", "singularity"],
"sort": "rising",
"minScore": 100,
"maxPostsPerSubreddit": 50,
"content_analysis": true
}

Content Research for SEO

Mine Reddit for real user language, questions, and topics for content creation:

{
"mode": "search",
"queries": ["how to", "best way to", "recommend"],
"searchSubreddit": "YourNicheSubreddit",
"sort": "top",
"timeframe": "year",
"minComments": 20,
"maxTotalItems": 300
}

Academic Research Dataset

Collect structured discussion data for analysis:

{
"mode": "search",
"queries": ["climate change discussion"],
"postsCreatedAfter": "2025-01-01",
"postsCreatedBefore": "2025-12-31",
"includeComments": true,
"flattenComments": true,
"maxCommentsPerPost": 200,
"sentiment_analysis": true,
"content_analysis": true,
"maximize_coverage": true,
"maxTotalItems": 2000
}

Compliance & Ethics

Responsible Data Collection

This actor collects publicly available Reddit posts, comments, and discussion metadata from https://www.reddit.com for legitimate business purposes, including:

  • Consumer research and market analysis
  • Brand monitoring and competitive tracking
  • Product feedback discovery and trend analysis
  • Academic research and discourse analysis

Users are responsible for making sure their use of the collected data complies with applicable laws, regulations, internal policies, and the target site's terms. This section is informational and not legal advice.

Best Practices

  • Use collected data in accordance with applicable laws, regulations, and the target site's terms
  • Respect individual privacy and personal information
  • Use data responsibly and avoid disruptive or excessive collection
  • Do not use this actor for spamming, harassment, or other harmful purposes
  • Follow relevant data protection requirements where applicable, such as GDPR and CCPA

FAQ

Q: How many posts can I extract per run? Free users can extract up to 20 items per run (4 runs/day). Paid subscribers have no limits — configure maxTotalItems as high as you need.

Q: Can I extract comments from posts? Yes. Set includeComments to true and configure maxCommentsPerPost to control how many comments per thread. Use flattenComments to save each comment as a separate dataset row.

Q: What sorting options are available? Posts can be sorted by hot, new, top, rising, controversial, or relevance. Comments can be sorted by confidence (Best), top, new, controversial, old, or qa.

Q: Can I filter posts by date range? Yes. Use postsCreatedAfter and postsCreatedBefore with YYYY-MM-DD format to define exact date windows.

Q: What is "Maximize Coverage" mode? When enabled, the actor expands collection across multiple Reddit ranking surfaces (relevance, top, new, comments) and deduplicates overlapping results. Use it when you need broader topic coverage rather than a single ranked list.

Q: Does sentiment analysis require an external API key? No. AI Sentiment Analysis and Content Taxonomy Classification are built into the actor. No external API keys, subscriptions, or configuration needed — just toggle them on.

Q: Can I schedule recurring runs? Yes. Use Apify's built-in scheduling to run daily, weekly, or on custom cron expressions. Combine with webhooks, Zapier, or Make for automated workflows.

Q: What output formats are supported? JSON, CSV, Excel (XLSX), XML, HTML Table, and RSS. Download directly from Apify Console or access via API.

Q: How do I scrape a specific post or thread? Use mode: "posts" or mode: "directUrls" and paste the full Reddit URL into the urls field. Enable includeComments to also extract the comment thread.

Q: Can I search within a specific subreddit? Yes. Use mode: "search" with your keywords in queries and set searchSubreddit to the subreddit name (without r/).


Support

For help, use the Issues tab on the actor page in Apify Console. Include:

  • The input configuration you used (with sensitive values redacted)
  • The run ID from Apify Console
  • Expected behavior vs. actual behavior
  • A small output sample if helpful

Categories

Social Media · Data Extraction · Reddit · Reddit Scraper · Automation · Market Research · Sentiment Analysis · Web Scraping · AI Analytics · Brand Monitoring