Reddit Scraper - Subreddit Posts and Comments, No Login avatar

Reddit Scraper - Subreddit Posts and Comments, No Login

Pricing

from $0.49 / 1,000 reddit posts

Go to Apify Store
Reddit Scraper - Subreddit Posts and Comments, No Login

Reddit Scraper - Subreddit Posts and Comments, No Login

Scrape Reddit posts from any subreddit or user profile, no login. Each row holds the title, author, full body text, permalink and timestamp, plus narration-ready text split into sentences. Vote counts and comments need your own free Reddit app. $0.50 per 1,000 posts.

Pricing

from $0.49 / 1,000 reddit posts

Rating

5.0

(1)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

12

Total users

7

Monthly active users

2 days ago

Last modified

Share

Give it subreddits or Reddit user profiles and get their posts back with the full body text, one row per post: title, author, permalink, timestamp, and a cleaned narration version of the body with the sentences already split.

One thing to know first, because it changes what the rows look like. The default method, rss, needs no login and carries no vote data. score and numComments arrive as 0, upvoteRatio and flair as null, and over18 as false, on every row. That is the keyless feed staying quiet, not a post nobody voted on. For the real numbers, add your own free Reddit app and set method to oauth.

InputSubreddit names, user profiles, or reddit.com links
OutputOne row per post: title, author, full body text, permalink, timestamp and narration-ready text
Ceiling100 posts per source, per run
Account neededNone by default. Your own Reddit app only if you want vote counts and comments
Price$0.50 per 1,000 posts, flat on every plan. The free plan's $5 a month covers about 10,000

🔍 What Reddit Scraper does

You hand it a mixed list. askreddit, r/tifu, u/GallowBoob, a full reddit.com link: all four shapes work, in the same list. It reads one page per source, in the sort order you picked, and writes one row per post.

Every row carries the full selftext, not a cut-down preview. With cleanText left on you also get narration, the body with markdown, links and edit stamps taken out, and ttsSegments, the same text split into sentences, which is the shape a text-to-speech step wants.

dedupeAcrossRuns is on by default. It remembers the last 5,000 post ids it gave you, so a daily run on the same subreddit hands you what is new rather than the same ten posts again.

📋 What data you get from each Reddit post

What you getField
The post id, where it came from and the sort it was read inid, subreddit, sourceType, sourceName, sort, timeRange
Title, author, permalink and timestamptitle, author, url, createdUtc
The whole body, and whether it is a text or a link postselftext, postType
Text ready for a voice: title plus body, the body alone, one sentence per entryscriptText, narration, ttsSegments
Length, a read-aloud estimate and whether it fits your word rangewordCount, readTimeSeconds, fitsShort
How strong the opening line is, 0 to 100hookScore
Votes, comment count, upvote ratio, flair and the over-18 flag, real only with your own appscore, numComments, upvoteRatio, flair, over18
Top comments, when you ask for themcomments
When the row was readfetchedAt

▶️ How to scrape Reddit posts from a subreddit or profile

  1. Open Reddit Scraper and click Try for free.
  2. Put your subreddits and profiles into Sources (subreddits or user profiles), one per line.
  3. Pick a Sort and, for top or controversial, a Time range.
  4. Set Max stories per subreddit. Start at 10 while you look at the row shape.
  5. Click Start, then download the dataset as JSON, CSV or Excel, or read it from the API.

💰 How much does it cost to scrape Reddit posts?

$0.50 per 1,000 posts. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 10,000 posts.

You pay per post delivered. Posts your filters drop are never written, so they are never billed. A source that comes back empty produces no rows, and a post that dedupeAcrossRuns already gave you is not delivered again. The most a run can bill for is maxPostsPerSubreddit times the number of sources.

📥 What you give it

{
"sources": ["r/tifu", "r/AskReddit", "u/GallowBoob"],
"sort": "top",
"time": "week",
"maxPostsPerSubreddit": 25
}
FieldDefaultWhat it is
sourcesnoneSubreddits and user profiles, mixed freely. Give it at least one. The form opens with r/AskReddit.
subreddits, usersnoneOlder aliases. Anything here is merged into sources.
methodrssrss needs no login. oauth uses your app credentials and fills in votes and comments. json is an older route that rarely returns anything now, so pick one of the other two.
sorttoptop, hot, new, rising or controversial.
timedayThe window for top and controversial. Ignored by the other sorts.
maxPostsPerSubreddit10Posts per source, 1 to 100. This is your spend cap.
cleanTexttrueCleans the text in scriptText and narration. Turn it off to keep the raw body.
requireStoryfalseText posts only, with the word-count range applied. Off means every post type.
minWords, maxWords30, 800The word-count range. Only filters when requireStory is on.
minHookScore0Drops posts whose opening line scores below this. 0 keeps everything.
minScore0Minimum upvotes. Only does anything on oauth.
includeNsfwfalseOnly does anything on oauth. See the limits below.
commentLimit0Top comments per post. Needs oauth to work in practice.
dedupeAcrossRunstrueSkips post ids this actor already delivered to you.
rawModefalseWrites each post as it was read instead of the shaped row: Reddit's own object on oauth, a much thinner one on the keyless path.
redditClientId, redditClientSecretnoneYour own free script app from reddit.com/prefs/apps. Used only when method is oauth.
proxyConfigurationApify defaultLeave it alone for a normal run. Servers of your own are used exactly as given.

Give it at least one source. Leave sources empty and the run does not stop: it falls back to r/amitheasshole, fetches those posts and bills you for them.

📤 What you get back

A real row from a real run on the keyless path, with the long text fields cut short here so the block stays readable:

{
"id": "1wnf1c6",
"subreddit": "tifu",
"sourceType": "subreddit",
"sourceName": "tifu",
"sort": "top",
"timeRange": "month",
"title": "TIFU by getting AI generated wallpaper installed in my bedroom",
"url": "https://www.reddit.com/r/tifu/comments/1wnf1c6/tifu_by_getting_ai_generated_wallpaper_installed/",
"author": "pixelperfect728",
"score": 0,
"upvoteRatio": null,
"numComments": 0,
"createdUtc": 1790096173,
"over18": false,
"isStory": true,
"flair": null,
"selftext": "I bought mural wallpaper from Etsy (~$500USD). Didn’t notice anything suspicious in the product photos. When it arrived, I looked at the sample...",
"postType": "text",
"scriptText": "TIFU by getting AI generated wallpaper installed in my bedroom\n\nI bought mural wallpaper from Etsy (~$500USD). Didn’t notice anything suspicious...",
"narration": "I bought mural wallpaper from Etsy (~$500USD). Didn’t notice anything suspicious in the product photos...",
"ttsSegments": [
"I bought mural wallpaper from Etsy (~$500USD).",
"Didn’t notice anything suspicious in the product photos."
],
"wordCount": 186,
"readTimeSeconds": 74.4,
"hookScore": 44,
"fitsShort": true,
"fetchedAt": "2026-09-30T23:04:04.740Z"
}
FieldHow to read it
idReddit's post id without the t3_ prefix. Use it as your dedupe key.
urlAlways the Reddit permalink. On a link post it is not the external address.
postTypetext or link. selftext is empty on a link post.
scriptTextTitle plus body, cleaned when cleanText is on. narration is the body alone.
hookScore0 to 100, worked out from the opening line.
fitsShortWhether wordCount sits between minWords and maxWords.
readTimeSecondsWord count divided by 2.5, so a rough read-aloud estimate.
commentsOnly present when commentLimit is above 0. Each entry has author, body, score and depth.

🧾 Reading the output

Every row in the dataset is a post, and every post row is billed. This actor writes no sample rows and no diagnostic rows, so there is nothing to filter out.

Which fields actually carry a value depends on the method you ran:

FieldKeyless (rss)With your app (oauth)
title, author, url, selftext, createdUtcfilledfilled
score, numCommentsalways 0the real counts
upvoteRatio, flairalways nullfilled when Reddit has them
over18always false, whatever the post isthe real flag
commentsusually emptyfilled up to commentLimit

A source that is private, banned or misspelled writes a warning into the run log and produces no row. If every source fails that way the run still finishes green with an empty dataset, so check the row count rather than the run status.

💡 What people use it for

  • Pulling story posts for short-form video, which is what narration, ttsSegments and fitsShort exist for.
  • Watching a handful of subreddits daily and reading only what is new.
  • Building a text corpus where the full body matters and a preview snippet would be useless.
  • Following what one prolific poster puts out, by pointing it at u/<name> instead of a sub.

From a subreddit to narration a voice can read, in three steps:

  1. Run this actor on r/tifu with requireStory on and maxWords set to 250.
  2. Keep the rows where fitsShort is true and sort them by hookScore.
  3. Put the ones you picked into Reddit Text Cleaner as texts. It reads the text off each row as it is, spells out shorthand like AITA and tones down the swearing with profanityMode set to soft.

🚧 What it does not do

  • No keyword search. It reads the sources you name. Searching Reddit is a separate actor, linked below.
  • No vote data without your own app. On the keyless path score and numComments are written as 0 rather than left empty, which is easy to misread as a post nobody touched.
  • minScore and includeNsfw do nothing on the keyless path. There is no score to filter on, and over-18 posts come through labelled over18: false. Pinned posts are not told apart there either. Run oauth if any of that matters.
  • Credentials alone are not enough. Set method to oauth as well, or the run stays keyless and ignores them.
  • Profiles always stay keyless. u/ sources are read without your app even on oauth, so their rows never carry vote data. A profile asked for top posts of the last day often comes back empty.
  • One page per source. 100 posts is the hard ceiling for one source in one run, and filters work inside that page, so strict ones return fewer rows rather than reading further.
  • Link posts hide the destination. url is the Reddit permalink, never the article it points at.
  • rawMode skips comments. You get the post as it was read and no comments array.
  • Comments are unreliable while keyless. Set commentLimit without oauth and they usually come back empty.

🧭 Which Reddit actor do you need?

If you wantUse
Posts and their full text from subreddits or profiles you nameThis one
Posts matching a keyword, across Reddit or inside one subredditReddit Search Scraper
Reddit videos as one MP4 with the sound in itReddit Video Scraper
Post text you already have, cleaned up for a voice to readReddit Text Cleaner
A post rewritten into a hook and a short-form scriptStory to Script Rewriter
A finished faceless short video from a subreddit or a storyAI Faceless Video Generator

❓ Questions people ask

Do I need a Reddit account?

Not for the default run. The credential fields are optional and only buy you vote counts, comment counts and comment bodies, and only with method set to oauth.

Why is every score 0?

The keyless method does not carry vote counts, and the row writes 0 rather than leaving the field empty. Add your own app credentials and set method to oauth.

My second run returned nothing. Is it broken?

Probably not. dedupeAcrossRuns is on, so a run that finds only posts you already have delivers nothing, and says nothing about it. Turn it off to see everything again.

Can I get more than 100 posts from one subreddit?

Not in one run. Split the work across sorts and time windows, or run it every day and let dedupe collect new posts as they arrive.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/reddit-scraper. Either way the run happens on your Apify account at the same price.

This reads public posts only, never private messages or accounts. The rows can still hold personal data, which GDPR and similar laws cover, so have a reason for keeping it. Apify's write-up on the legality of web scraping is a sensible starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Include the sources you asked for and the run ID, and the log tells us the rest.