Reddit Scraper avatar

Reddit Scraper

Pricing

from $1.20 / 1,000 results

Go to Apify Store
Reddit Scraper

Reddit Scraper

Extract Reddit posts, comments, search results, and user history via the Arctic Shift archive with RSS fallback. No OAuth, no API key, no login.

Pricing

from $1.20 / 1,000 results

Rating

0.0

(0)

Developer

Aurenic

Aurenic

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Extract Reddit posts, comments, search results, and user history via the Arctic Shift archive with RSS fallback. No OAuth, no API key, no login.

What does Reddit Scraper do?

Scrape Reddit in four modes:

  • Subreddit posts — pass one or more subreddits, get posts with title, author, score, upvote ratio, comment count, flair, body, and permalink. Sort by hot, new, top, rising, or controversial, with a time filter.
  • Keyword search — search across all of Reddit or within one subreddit.
  • Post comments — pass a post URL or ID, get the comment tree.
  • User history — pass a username, get their post and comment history.

Every record includes the fields that matter for real analysis — score, upvote ratio, comment count, awards — not just the bare title and URL that RSS-only scrapers return.

Why this works in 2026. Reddit deprecated unauthenticated .json access in May 2026 — every request to reddit.com/...json now returns HTTP 403, from datacenter and residential IPs. Self-service API app creation was also shut down under Reddit's Responsible Builder Policy. The actor sidesteps all of it by using two openly accessible backends:

  1. Arctic Shift — a free, keyless, community-run Reddit archive with an HTTP API. Full post and comment data including scores and vote ratios. Covers May 2025 to present with a ~48-hour indexing lag.
  2. Reddit RSS feeds — still open in 2026, used as a fallback for the freshest posts that haven't been indexed by Arctic Shift yet.

Neither requires OAuth, an API key, or account approval.

Output fields

Posts

FieldDescription
idReddit post ID
subredditSubreddit name
titlePost title
authorUsername
bodyPost body / selftext (truncated to 20,000 chars)
scoreNet upvotes
upvote_ratioUpvote ratio (0–1)
num_commentsComment count
created_utcISO 8601 timestamp
permalinkReddit comment page URL
urlExternal URL for link posts, else permalink
domainLink domain
isSelfText post flag
over18NSFW flag
linkFlairPost flair
thumbnailThumbnail URL
awardsTotal awards received
sourcearctic-shift or rss-fallback

Comments

FieldDescription
idComment ID
postIdParent post ID
parentIdParent comment ID (or post ID for top-level)
subreddit / authorContext
bodyComment text (truncated to 10,000 chars)
scoreNet upvotes
created_utcISO 8601 timestamp
permalinkDirect comment URL
isSubmitterWhether commenter is the OP
depthComment tree depth

Who is it for?

  • Brand monitoring teams tracking product mentions across thousands of subreddits
  • Lead generation agencies finding high-intent discussions ("what tool should I use for…")
  • Market researchers measuring sentiment on any topic, product, or event
  • AI and RAG builders ingesting Reddit corpora for training and grounding
  • Data journalists pulling community discussions by keyword and time window
  • Growth teams monitoring competitor subreddits for sentiment shifts

Pricing

$1.10 per 1,000 results. No subscription.

ResultsCost
100$0.11
1,000$1.10
10,000$11.00

How to use it

  1. Pick a Mode.
  2. Enter the mode's inputs: subreddits, search query, post URLs/IDs, or usernames.
  3. Set Sort and Time Filter for posts and search.
  4. Set Max Items (default 200).
  5. Click Start.

Output example

{
"recordType": "post",
"id": "1dxyz42",
"subreddit": "python",
"title": "What's the best way to handle async retries in 2026?",
"author": "some_user",
"body": "I've been using tenacity but wondering if there's something better...",
"score": 342,
"upvote_ratio": 0.94,
"num_comments": 87,
"created_utc": "2026-09-15T14:32:10.000Z",
"permalink": "https://www.reddit.com/r/python/comments/1dxyz42/whats_the_best_way/",
"url": "https://www.reddit.com/r/python/comments/1dxyz42/whats_the_best_way/",
"domain": "self.python",
"isSelf": true,
"over18": false,
"linkFlair": "Discussion",
"thumbnail": "",
"awards": 3,
"source": "arctic-shift",
"scrapedAt": "2026-09-21T12:00:00.000Z"
}

Technical details

  • No OAuth, no API key, no login. Reddit's .json endpoints are dead as of May 2026 — the actor uses openly accessible backends instead.
  • Primary source: Arctic Shift (arctic-shift.photon-reddit.com) — a free, keyless, community archive with full post and comment data including scores.
  • Fallback source: Reddit RSS — still returns 200 in 2026 for subreddit feeds and search, used when Arctic Shift has not yet indexed very recent posts (~48-hour lag).
  • Self-throttled at ~700ms between Arctic Shift requests to respect the free service.
  • No proxy required. Arctic Shift accepts datacenter IPs.
  • Two extraction paths — if Arctic Shift returns no data, RSS takes over transparently and records are tagged source: rss-fallback.

Known limits

  • Arctic Shift has ~48-hour indexing lag. Very recent posts may not yet be indexed. Enable Use RSS Fallback to get them (scores will be null on RSS records).
  • Arctic Shift is a free volunteer-run service. No uptime SLA. The actor retries on 5xx with exponential backoff and falls back to RSS.
  • RSS records lack scores and upvote ratios. RSS carries no vote data. Records tagged source: rss-fallback will have score: null, upvote_ratio: null.
  • Post body is truncated at 20,000 characters. Comment body at 10,000.
  • Arctic Shift does not index private or quarantined subreddits. Only public data is available.

FAQ

Do I need a Reddit account? No. Neither backend requires authentication.

Do I need a proxy? No. Arctic Shift accepts datacenter IPs.

Why is score null on some records? RSS fallback records carry no vote data — the RSS feed doesn't include scores. Arctic Shift records always have scores.

Why are recent posts missing? Arctic Shift has ~48-hour indexing lag. Enable RSS fallback for fresh posts, or wait and re-run.

Is this legal / allowed? The actor uses two publicly accessible data sources: a community-run archive API that requires no authentication, and Reddit's own open RSS feeds. No login, no terms-of-service circumvention.

How do I export data? After a run, go to Storage → Export as JSON, CSV, Excel.

Support

Open an issue on the Actor's page for bugs or feature requests.