# Changelog of Reddit Scraper (`prodiger/reddit-scraper`) Actor

- **URL**: https://apify.com/prodiger/reddit-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/prodiger/reddit-scraper.md

## Changelog

All notable changes to this project are documented here. Format follows [Keep a Changelog](https://keepachangelog.com/en/1.1.0/).

### \[0.5] — 2026-09-29

#### Fixed

- **The actor returned 0 results while reporting success.** Reddit put old.reddit.com behind a login wall: every logged-out request is now answered with a `302` to `/login/?reason=lor2&dest=…`. The crawler followed the redirect, received the login page with HTTP 200, found no posts in it, and finished `SUCCEEDED` with an empty dataset — for every source type (subreddit, user, search, post).

#### Changed

- **Data source is now www.reddit.com's server-rendered partials**, the same HTML fragments Reddit's logged-out frontend loads (`/svc/shreddit/…` feeds, search and comment trees), plus the public post embed page for directly scraped post URLs. Still no Reddit account or API key.
- **`upvoteRatio` and `totalAwards` are back** for subreddit and user listings (the old.reddit HTML did not expose them).
- **Listing posts now include `selfText`.** Feeds carry the full post body, so `includeComments` and the AI output formats no longer need an extra request per post just to read the body.
- **Comments are fetched page by page.** The first page renders ~25 top-level comments with their visible replies; "more replies" and "view more comments" pages are followed (replies first) until `maxCommentsPerPost` is reached, bounded to 40 pages per post.
- **A post is never lost to a failed comment page**: in the default format the post is written before its comments; in the AI formats it is written with the comments collected so far.
- **A bare `403` is retried on a fresh IP** instead of being skipped as "private". On www.reddit.com inaccessible sources answer `200`, so a `403` without a recognizable private/banned page is a per-IP rejection. Skipped sources now log the reason.
- **Login-wall redirects and JS challenge pages are treated as blocks** (retry on a fresh residential IP; fail the run if nothing could be scraped) instead of being parsed as an empty page. A source whose first page has no posts is reported as skipped, with a warning.

#### Limitations

- `subredditSubscribers` is still not available (`null`).
- Search results and directly scraped post URLs have `upvoteRatio` / `totalAwards` of `0`; search results have no `selfText` unless `includeComments` or an AI output format is used.
- Reddit renders the same empty feed for private, banned, misspelled and genuinely empty subreddits, so these cannot be told apart; all are reported as skipped.

### \[0.4] — 2026-06-02

#### Changed

- **Data source is now old.reddit.com server-rendered HTML (Cheerio), not Reddit's API.** The 0.3 OAuth approach proved unusable: Reddit's Responsible Builder Policy (Nov 2025) ended self-service app creation and does not approve commercial scraping on the free tier, so requiring API credentials made the actor unusable for its audience. This release restores the **no-API-key** experience by scraping old.reddit HTML behind residential proxies — the same credential-free technique the popular Apify Reddit actors use. Switched `HttpCrawler` → `CheerioCrawler` with a generated Chrome browser fingerprint.
- **Removed the OAuth credential requirement** (`redditClientId` / `redditClientSecret` and the `REDDIT_CLIENT_*` env fallback are gone). No Reddit account or API key is needed.
- **Block handling tuned for HTML scraping:** a 403 "blocked by network security" page retires the session so Crawlee retries on a fresh residential IP (how rate-limit blocks clear); 404/private are benign per-source skips; 429 honors `Retry-After`. The fail-loud guard from 0.3 is kept — a run blocked on every request fails loudly instead of reporting an empty success; a run that reaches Reddit (even if all posts are filtered) succeeds.

#### Limitations (vs the old JSON API)

- **`upvoteRatio`, `totalAwards`, and `subredditSubscribers` are no longer available** from HTML (set to `0` / `0` / `null`); gallery `imageUrls` is best-effort.
- **Comment depth is bounded by what old.reddit renders** (fetched with `?limit=500`). "Load more comments" / deeply collapsed threads use an AJAX endpoint Reddit blocks, so the deepest tails are not retrieved.
- **Listing/search posts carry metadata only (no `selfText`).** Self-text + comments are populated when scraping a post URL directly, with `includeComments=true`, or for the AI output formats (all fetch the post page). Keyword filtering on listings matches the title.
- **RESIDENTIAL proxies are now required**, not just recommended — datacenter IPs are blocked and per-IP rate-limits are cleared by rotation.

### \[0.3] — 2026-06-02 (superseded by 0.4)

#### Fixed

- **The actor returned 0 results while reporting success.** Reddit shut down its unauthenticated public `.json` API — every `www.reddit.com/*.json` request (and `old.reddit.com/*.json`) now returns an HTTP 403 "blocked by network security" HTML page, regardless of User-Agent, browser fingerprint, cookies, or proxy IP (the block is endpoint-level, so residential proxies don't help). The actor treated those 403s as benign skips (`ignoreHttpErrorStatusCodes: [403]` + a "403 = private subreddit" assumption + non-JSON bodies producing only a warning + no zero-result guard), so every run drained its queue and exited successfully with an empty dataset.

#### Changed

- **Data now comes from Reddit's official OAuth API (`oauth.reddit.com`)**, which returns the identical Listing/Thing JSON the parsers already consume — so post/comment mapping is unchanged. Requests carry a bearer token obtained via the application-only (`client_credentials`) grant.
- **Credentials** are resolved from input (`redditClientId` / `redditClientSecret`) or, as a fallback, from `REDDIT_CLIENT_ID` / `REDDIT_CLIENT_SECRET` environment variables (so a maintainer can set one shared app via Apify secrets and keep end users key-free). Credentials are validated up front; a missing or invalid pair fails the run immediately with setup instructions (https://www.reddit.com/prefs/apps).
- **Error handling fails loud.** 401 refreshes the OAuth token and retries; 403/404 are benign per-source skips; 429 honors `Retry-After`. If a run produces zero items *because* requests were blocked or rejected, it now calls `Actor.fail()` with a diagnostic instead of reporting an empty success. A legitimately empty source (valid but no matching posts) still succeeds.
- User-Agent is now a stable, descriptive identifier per Reddit's API terms (the previous Chrome-fingerprint rotation was anti-bot theater for the now-dead public endpoints).

### \[0.2] — 2026-04-19

#### Fixed

- **Posts and comments now land in the run's default dataset.** Previously the actor wrote to account-level named datasets (`Actor.openDataset('posts')` / `'comments'`), which made the Apify Console "Storage" tab, `run.defaultDatasetId` API access, and standard SDK smoke tests all see an empty dataset even though items were silently accumulating on the user's account-wide named datasets across runs. Both record types now share the per-run default dataset and are discriminated by the existing `type: 'post' | 'comment'` field on each record.

#### Changed

- `dataset_schema.json` now declares both post and comment field shapes and adds a "Comments" view alongside the existing "Posts" view. The `type` enum now includes `comment`.

#### Migration note

- Existing data on the account-level named "posts" and "comments" datasets is unaffected (still readable via the Apify dataset API by name). New runs from 0.2 onward write to the per-run default dataset only.

### \[0.1.2] — 2026-04-18

#### Fixed

- Removed redundant explicit `Actor.charge('actor_start')` call in `src/main.ts` — Apify Console uses the synthetic `apify-actor-start` event which fires automatically. The explicit call was logging an unknown-event warning per run without doing anything useful.

### \[0.1] — 2026-04-18

#### Added

- Initial release. Reddit scraper covering subreddits, posts, comments, users, and search.
- Three output formats: `default` (standard JSON), `jsonl-finetune` (OpenAI chat-format SFT records), `rag-markdown` (vector-DB-ready markdown documents with stable `chunkId`).
- Pure HTTP via Crawlee `HttpCrawler` (no headless browser).
- Pay-per-event pricing matching `automation-lab/reddit-scraper` rates as of 2026-04-18.
- RESIDENTIAL proxy default with explicit DATACENTER override option (datacenter is no longer reliable against Reddit in 2026).
- Hard cap `maxCommentsPerPost ≤ 1000` to bound per-run proxy/compute cost.
- Pre-flight DATACENTER + large-run guard with WARN at run start.
- Reactive 429 backoff honoring `Retry-After` headers; session retirement on persistent blocks.
- Charge-event ordering: `actor_start` after input validation (so failed-validation runs don't bill); `post` and `comment` charges only after successful dataset writes; `comment` charges suppressed for AI formats (comments are bundled into the post record).
