Reddit Scraper — Subreddits, Posts, Comments & Search avatar

Reddit Scraper — Subreddits, Posts, Comments & Search

Under maintenance

Pricing

from $2.00 / 1,000 results

Go to Apify Store
Reddit Scraper — Subreddits, Posts, Comments & Search

Reddit Scraper — Subreddits, Posts, Comments & Search

Under maintenance

Scrape Reddit subreddits, posts with comments, user profiles, and search results by parsing server-rendered HTML. Works after the May 2026 .json API shutdown.

Pricing

from $2.00 / 1,000 results

Rating

0.0

(0)

Developer

Joren Maurissen

Joren Maurissen

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 days ago

Last modified

Share

Reddit Scraper

Scrape Reddit subreddits, posts with comments, user profiles, and search results by parsing server-rendered HTML. Works after the May 2026 .json API shutdown.

Why This Actor Exists

Reddit shut down its public .json API in May 2026, breaking many existing scrapers that relied on appending .json to Reddit URLs. This actor parses Reddit's server-rendered HTML directly — no API key, no .json endpoint, no authentication required.

Features

  • Subreddit scraping — Get posts from any subreddit sorted by hot, new, top, or rising
  • Post + comments — Scrape a single post with all its comments and replies
  • User profiles — Get a user's posts and comments
  • Search — Search Reddit for keywords, optionally restricted to a subreddit
  • Pagination — Uses Reddit's after cursor for unlimited scrolling
  • Time filtering — Filter top posts by hour, day, week, month, year, or all time
  • Proxy support — Apify proxy integration for avoiding rate limits

Input

FieldTypeRequiredDefaultDescription
modeselectYessubredditsubreddit, post, user, or search
subredditstringNoSubreddit name (without r/) for subreddit and search modes
usernamestringNoReddit username for user mode
postUrlstringNoFull Reddit post URL for post mode
searchQuerystringNoSearch query for search mode
sortselectNohothot, new, top, rising
timeFilterselectNodayhour, day, week, month, year, all (top sort only)
maxResultsintegerNo100Max results (1-1000)
proxyConfigurationproxyNoonApify proxy recommended

Output

Each result contains: title, author, score, upvote_ratio, num_comments, created_utc, url, selftext, subreddit, flair, permalink, and (for post mode) nested comments.

Technical Approach

  • Fetches HTML from old.reddit.com (simpler HTML, fewer anti-bot measures)
  • Falls back to www.reddit.com if old.reddit.com fails
  • Parses embedded JSON data from <script> tags when available
  • Falls back to HTML parsing with BeautifulSoup for structured extraction
  • Uses browser-like User-Agent to avoid bot detection
  • Handles 429 rate limiting with exponential backoff
  • No external APIs, no API keys, no paid services