Reddit Comment Scraper — Post Comments & Subreddit Monitoring avatar

Reddit Comment Scraper — Post Comments & Subreddit Monitoring

Pricing

$4.50 / 1,000 comment produceds

Go to Apify Store
Reddit Comment Scraper — Post Comments & Subreddit Monitoring

Reddit Comment Scraper — Post Comments & Subreddit Monitoring

Extract comments from specific Reddit posts or from the top posts of any subreddit. Supports all Reddit comment sort modes.

Pricing

$4.50 / 1,000 comment produceds

Rating

0.0

(0)

Developer

Automly

Automly

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

4 days ago

Last modified

Categories

Share

Extract the comments of any public Reddit post, or of the posts a subreddit is showing right now, as structured JSON with the full reply structure: parent, depth, score, author, text and HTML. No Reddit account or API key needed.

What it does

  • Post mode: give post URLs (or bare post IDs) and get up to maxComments comments from each post, including the replies hidden behind "load more comments" on large threads.
  • Subreddit mode: give subreddits and the scraper picks their first maxPosts posts (hot, new, top or rising) and collects comments from each. Used only when no post URL is given.
  • Sort comments the way Reddit does: best, top, new, controversial, old or Q&A.
  • Deleted and removed comments are left out; the replies under them are kept.

Input

FieldTypeDescription
postUrlsarray of stringsPost URLs or base-36 post IDs.
subredditsarray of stringsSubreddits whose posts to read (ignored when post URLs are given).
sortBystringPost order in subreddit mode: hot (default), new, top, rising.
commentSortstringconfidence (best, default), top, new, controversial, old, qa.
maxCommentsintegerComments from each post URL; in subreddit mode, the total shared between subreddits and posts (1–10000). Default 500.
maxPostsintegerPosts per subreddit in subreddit mode (1–100). Default 25.

Example input

{
"postUrls": ["https://www.reddit.com/r/python/comments/1abc234/some_title/"],
"commentSort": "top",
"maxComments": 1000
}

Output

One record per comment:

{
"commentId": "kx9y8z7",
"postId": "1abc234",
"parentId": "kx9a1b2",
"author": "some_user",
"body": "Great write-up, thanks!",
"bodyHtml": "<div class=\"md\"><p>Great write-up, thanks!</p></div>",
"score": 42,
"createdUtc": 1790000000,
"edited": false,
"depth": 1,
"subreddit": "Python",
"permalink": "/r/Python/comments/1abc234/some_title/kx9y8z7/",
"isSubmitter": false,
"stickied": false,
"distinguished": null,
"controversiality": 0
}

Each record also carries gilded, gildings, totalAwardsReceived, collapsed and collapsedReason. depth is 0 for a top-level comment, 1 for a reply to it, and so on; parentId is the parent comment, or the post for a top-level comment.

Pricing

Pay per event: you are charged for each comment saved to the dataset, and never for duplicates.

Tips

  • Set commentSort: new and run on a schedule to follow a live discussion.
  • Large threads are read in full: raise maxComments to get every comment.