Reddit Comments Scraper avatar

Reddit Comments Scraper

Pricing

$1.50 / 1,000 comment results

Go to Apify Store
Reddit Comments Scraper

Reddit Comments Scraper

Scrape complete public Reddit comment threads from one or many post or comment URLs. One flat row per comment with author, text, score, timestamps, parent, depth and path metadata, moderation flags and parent-post context. Control sort, reply depth and filters. No Reddit login or API key.

Pricing

$1.50 / 1,000 comment results

Rating

0.0

(0)

Developer

Delowar Munna

Delowar Munna

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Reddit Comments Scraper

Paste one or many Reddit post or comment URLs and get the discussion as a clean flat dataset — one row per comment, including nested replies, with parent, depth and path metadata, scores, timestamps and moderation flags. Control sort, depth, filters and hard cost caps without providing a Reddit account, cookies or API key.

Flatten the thread without losing the tree.

What it does

  • Takes post URLs, comment permalinks, redd.it short links, /r/<sub>/s/ share links or bare post IDs — any mix, in bulk.
  • Returns one row per comment: ID, author, text, score, created and edited timestamps, depth, parent ID, the full path from the top-level comment, OP and moderator flags, deleted/removed status, direct reply count and compact post context.
  • Follows nested replies and Reddit's collapsed "more comments" continuations.
  • Six sort modes, per-thread and run-wide caps, date/score/length/author/keyword filters, a skip list for incremental runs, and a target-comment mode that returns one comment with its ancestors and replies.
  • Pushes rows as it goes, so a late failure never loses what was already captured.
  • Reports thread completeness honestly in THREAD_SUMMARY rather than claiming a thread is whole when it is not.

Quick start

{
"postUrls": ["https://www.reddit.com/r/AskReddit/comments/1abc234/example/"],
"maxTotalComments": 1000,
"maxCommentsPerPost": 1000,
"includeNestedReplies": true,
"commentSort": "best",
"maxDepth": 10
}

Several threads at once, with filters:

{
"postUrls": ["https://www.reddit.com/r/investing/comments/1abc234/example/", "https://redd.it/1def456"],
"maxTotalComments": 5000,
"maxCommentsPerPost": 2500,
"commentSort": "top",
"maxDepth": 6,
"postedAfter": "2026-08-01",
"minScore": 3,
"excludeAuthors": ["AutoModerator"],
"excludeDeletedRemoved": true,
"excludeKeywords": ["I am a bot"],
"skipCommentIds": ["xyz987"]
}

One comment with its context — a comment permalink in target mode:

{
"postUrls": ["https://www.reddit.com/r/legaladvice/comments/1wjgckp/comment/pairqoe/"],
"focusOnTargetComment": true,
"commentContextDepth": 3,
"maxCommentsPerPost": 200
}

Full thread vs. target-comment context

By default every input — even a comment permalink — scrapes the whole thread. Switch on Focus on the linked comment to collect just that comment, up to Ancestor levels in focus mode parents above it, and every reply below it. Rows still carry absolute depth and the full path from the top-level comment, so the context reads correctly next to a full-thread export.

Sort modes

best, top, new, controversial, old, qa. The sort decides which comments a per-thread cap keeps: comments are collected breadth-first — the top-level comments Reddit ranks first, then their replies — so a cap on a huge thread returns the highest-ranked part of it.

Nested replies and "more comments"

Replies are followed to Max reply depth (default 10; 0 = top-level only). Reddit's collapsed continuations are resolved when Expand 'more comments' placeholders is on (default). Switching replies off is the cheapest mode — one request per page of top-level comments.

Input reference

FieldDefaultMeaning
postUrls—Post URLs, comment permalinks, redd.it links, share links or post IDs. Required.
maxTotalComments5000Hard cap on comments delivered by the whole run.
maxCommentsPerPost500Cap per thread.
includeNestedRepliestrueFollow reply branches.
maxDepth10Deepest reply level; 0 = top-level only.
commentSortbestOne of the six sorts above.
expandMoreCommentstrueResolve collapsed continuations.
minScore / maxScore—Score bounds. Hidden scores are kept.
excludeDeletedRemovedfalseDrop deleted and removed comments.
excludeStickiedfalseDrop moderator-pinned comments where the flag is observable (see Limitations).
postedAfter / postedBefore—ISO dates.
minLength / maxLength—Body length bounds.
includeAuthors / excludeAuthors—Username allow / deny lists.
includeKeywords / excludeKeywords—Case-insensitive substring filters on the text.
focusOnTargetCommentfalseTarget-comment mode for comment permalinks.
commentContextDepth3Ancestor levels in target mode.
includePostContexttruePost title, author and Reddit's comment count on every row.
skipCommentIds—IDs already collected; skipped, never charged, replies still explored.
maxConcurrency3Threads collected in parallel (1–5).
requestDelayMs—Extra pacing between requests.
proxyConfigurationApify DatacenterSee Proxy below.

Caps and cost

Two caps bound every run: maxTotalComments across the run and maxCommentsPerPost per thread. A comment beyond either cap is never fetched and never charged. Filters run before anything is saved, so filtered comments cost nothing either.

Output and tree reconstruction

Every row carries commentId, parentId (null for a top-level comment), parentFullId (the post's t3_ fullname for a top-level comment), depth and path — the list of comment IDs from the top-level comment down to this one. Group by postId, order by capturedSequence, and the tree rebuilds from the rows alone. directReplyCount is the number of direct replies known for the comment; on a thread cut short by a cap it is a lower bound.

Reddit Comments Scraper — Comments view, table (one row per comment with depth, parent and path)

The dataset has five views. One real record from each, taken from a run over nine r/personalfinance and r/legaladvice threads:

Comments — every field

A depth-2 reply, with its ancestors in path:

{
"recordType": "comment",
"commentId": "pairqoe",
"commentFullId": "t1_pairqoe",
"postId": "1wjgckp",
"postUrl": "https://www.reddit.com/r/legaladvice/comments/1wjgckp/pet_sitter_is_suing_for_lost_potential_income/",
"commentUrl": "https://www.reddit.com/r/legaladvice/comments/1wjgckp/pet_sitter_is_suing_for_lost_potential_income/pairqoe/",
"authorUsername": "example_replier",
"authorIsDeleted": false,
"text": "In addition, I would say print/record the text chain where he has confirmed the 60$ for 3 hours, info of your flight time showcasing he didn’t show up on time AND tried to re-negotiate. Get a lawyer and make the defense of it being under duress.",
"score": 125,
"isScoreHidden": false,
"createdAt": "2026-09-18T06:09:57.607Z",
"editedAt": null,
"isEdited": false,
"depth": 2,
"parentId": "paiisn8",
"parentFullId": "t1_paiisn8",
"path": [
"paiiokg",
"paiisn8",
"pairqoe"
],
"isTopLevel": false,
"isSubmitter": null,
"distinguished": null,
"isStickied": false,
"isLocked": false,
"isCollapsed": false,
"isArchived": false,
"isDeletedOrRemoved": false,
"removedReason": null,
"awardCount": null,
"directReplyCount": 1,
"postTitle": "Pet sitter is suing for lost potential income",
"postAuthorUsername": "example_op",
"postCommentCount": 86,
"subreddit": "legaladvice",
"sourceInputs": [
"https://www.reddit.com/r/legaladvice/comments/1wjgckp/"
],
"capturedSequence": 195,
"scrapedAt": "2026-09-24T07:25:53.399Z"
}

Top-level

Top-level comments only, with the reply count under each:

{
"commentId": "pailgv9",
"subreddit": "legaladvice",
"postTitle": "Pet sitter is suing for lost potential income",
"authorUsername": "example_user",
"text": "What happened to your dogs? He bait and switched you under duress, then didn’t follow through on his agreement.\n\nI wonder if he’s done this before.",
"score": 2143,
"createdAt": "2026-09-18T05:21:06.342Z",
"directReplyCount": 7,
"isSubmitter": null,
"distinguished": null,
"isDeletedOrRemoved": false,
"commentUrl": "https://www.reddit.com/r/legaladvice/comments/1wjgckp/pet_sitter_is_suing_for_lost_potential_income/pailgv9/"
}

Replies

Nested replies with their parent, depth and path:

{
"commentId": "pairqoe",
"parentId": "paiisn8",
"depth": 2,
"path": [
"paiiokg",
"paiisn8",
"pairqoe"
],
"authorUsername": "example_replier",
"text": "In addition, I would say print/record the text chain where he has confirmed the 60$ for 3 hours, info of your flight time showcasing he didn’t show up on time AND tried to re-negotiate. Get a lawyer and make the defense of it being under duress.",
"score": 125,
"createdAt": "2026-09-18T06:09:57.607Z",
"isSubmitter": null,
"isDeletedOrRemoved": false,
"directReplyCount": 1,
"commentUrl": "https://www.reddit.com/r/legaladvice/comments/1wjgckp/pet_sitter_is_suing_for_lost_potential_income/pairqoe/"
}

High score

Sorted for the highest-scoring comments across the run:

{
"score": 3480,
"authorUsername": "example_top_scorer",
"text": "I am actually wondering if the petsitter committed a crime by doing this - if they abandoned the dogs after undertaking their care it could be considered animal cruelty or negligence. You may consider calling a non emergency police line and ask/report.\n\nYou should seriously think about counter suing for breach of contract and damages caused by that breach. Not to mention, suing for negligence and any other animal related “damages” like emergency boarding, transportation costs to get back to your dogs, and sometimes even for emotional distress though that’s not as cut and dry.\nKeep every receipt related to any costs incurred bc of this person’s negligence and of course save any texts/emails etc and talk to an attorney.\n\nSo sorry you went through this, this person sounds like a scammer. I would be beyond furious and very emotionally distressed. I petsat for several years for many families and I always took the opportunity very seriously, knowing I’m in someone else’s home looking over a loved family member. I cannot fathom in a million years doing this. Truly awful",
"depth": 0,
"subreddit": "legaladvice",
"postTitle": "Pet sitter is suing for lost potential income",
"awardCount": null,
"createdAt": "2026-09-18T05:57:08.170Z",
"commentUrl": "https://www.reddit.com/r/legaladvice/comments/1wjgckp/pet_sitter_is_suing_for_lost_potential_income/paiq3so/"
}

Provenance

Which input produced each row, and in what order it was captured:

{
"capturedSequence": 195,
"commentId": "pairqoe",
"postId": "1wjgckp",
"sourceInputs": [
"https://www.reddit.com/r/legaladvice/comments/1wjgckp/"
],
"isTopLevel": false,
"depth": 2,
"scrapedAt": "2026-09-24T07:25:53.399Z"
}

isStickied, isLocked, isSubmitter and awardCount are null where the surface a thread was collected from does not publish them — see Limitations. text is plain text; a deleted or removed comment keeps its ID and position with an empty body and a removedReason.

Completeness flags

THREAD_SUMMARY in the run's key-value store lists every thread with commentsSeen, commentsEmitted, Reddit's own knownCommentCount, a coverageRatio, isThreadComplete and a truncationReason — one of maxTotalComments, maxCommentsPerPost, maxDepth, pagination-unavailable, moreRequestsCap, continuation-failed, provider-blocked or charge-limit. RUN_SUMMARY holds the run-wide counts. Reddit's own count includes removed and collapsed comments it never renders, so a complete thread can sit slightly below 1.0.

Incremental growing-thread runs

Pass the comment IDs you already hold as skipCommentIds (with or without the t1_ prefix). Known comments are neither saved nor charged, but their replies are still explored, so a scheduled run picks up only what is new. Combine with postedAfter for a cheap time window.

Pricing

Pay per result: one comment-result event per unique comment delivered to your dataset. Never charged: duplicates across overlapping inputs, comments removed by your filters, comments on your skip list, comments cut off by a cap, the parent-post context, invalid or unavailable post URLs, empty threads, and the run summaries. This README deliberately quotes no figures — the live price is on the Pricing tab and in the Console.

Proxy and reliability

Use Apify Datacenter proxy (the default) or no proxy. Requests are paced to Reddit's own public rate budget per exit address, and an address that draws Reddit's "prove your humanity" page is rested and replaced.

Apify Residential proxy is not supported. The run fails at startup if apifyProxyGroups includes RESIDENTIAL: comment pages are large and residential bandwidth is billed to the Actor rather than to your run, while Reddit serves the same pages to datacenter addresses. If you need residential routing, supply your own provider under Custom proxy URLs — that traffic goes through your account and is honoured.

Limitations

  • Very large threads. Logged out, Reddit does not page top-level comments beyond the first page, nor a comment's replies beyond its first 24. The Actor sweeps all six sorts and every reachable reply subtree to recover as much as it can, then marks the thread pagination-unavailable in THREAD_SUMMARY rather than calling it complete. Large threads are collected in full through a metered provider where the Actor is configured for it; the row shape is identical.
  • isStickied and isLocked are not published on Reddit's public comment markup and are null there; they are booleans on rows collected through the metered provider. isSubmitter and awardCount are the reverse: published on the public markup, null from the metered provider.
  • No controversiality flag. Neither surface publishes one per comment; the controversial sort orders by it.
  • Private, quarantined and login-only threads are reported as unavailable, not scraped.
  • Deleted and removed comments keep their IDs and tree position with an empty body and isDeletedOrRemoved: true; their original text is not recovered.
  • Reddit rate-limits by exit address; a run with many threads takes minutes, not seconds.

API and integrations

Feed post URLs straight from a Reddit post-search Actor and send authorUsername values on to a Reddit user Actor. The dataset is flat JSON, CSV-friendly, and each row is self-describing.

Responsible use

Public comments can contain personal data. Comply with applicable law and Reddit's terms; do not use this Actor to reconstruct deleted or private content or to target individuals.

Changelog

  • 1.0 — Initial release: flat comment rows with tree metadata, six sorts, nested reply expansion, target-comment mode, filters, skip lists, incremental pushing, honest completeness reporting.