# Changelog of Threads Scraper (`yasaslive/threads-scraper`) Actor

- **URL**: https://apify.com/yasaslive/threads-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/yasaslive/threads-scraper.md

## Changelog

All notable changes to this Actor are documented here. Format: [Keep a Changelog](https://keepachangelog.com/en/1.1.0/); versioning: [SemVer](https://semver.org/).

### \[1.0.0] - 2026-09-27

#### Added

- Scrape public Threads **profiles**, a profile's **posts** (multi-part threads split and linked), **post + replies** (nested reply chains linked via `parentPostId`) and **keyword search**, without login.
- HTTP-first route (`HttpCrawler` + got-scraping): parses the server-rendered `data-sjs` JSON, then paginates via `/api/graphql` using the page's own `lsd` token and preloaded `doc_id` + variables.
- Token-refresh routine: when GraphQL rejects `lsd`/`doc_id` (HTML, 401/403, FB error envelope), the page is re-fetched on the same session for fresh tokens (max 2 per request) before failing over to session rotation.
- Resumable pagination: the GraphQL cursor and per-input counters are stored on the request, so retries on a new session continue instead of restarting.
- Browser fallback (`PlaywrightCrawler`): activated only after **3 consecutive login walls** on the HTTP route; reuses the same parsers and runs the same GraphQL pagination via `fetch()` inside the page. Each item records `scrapedVia: "http" | "browser"`.
- Pay-per-event charging (`actor-start` $0.01, `post-scraped` $0.001, `profile-scraped` $0.003) with budget pre-checks, charge-after-push, serialized pushes, and a clean stop with `budgetReached: true` in `STATS`.
- Deduplication by stable ID, persisted across migrations (`STATE` record).
- Categorized error counters (blocked / rateLimited / proxy / network / parse / loginWall / notFound / other) in `STATS`, logged every 30 s.
- `FAILED_INPUTS` record explaining every input that produced nothing.
- Soft-failure detection: HTTP 200 login pages, data-less app shells, checkpoints/challenges and rate-limit pages are treated as blocks and retried with exponential backoff (jittered, capped at 30 s) and session rotation.

#### Decisions (made without asking, per project conventions)

- **Node 22, not 20.** Node 20 reached end-of-life in April 2026 and Apify no longer publishes Node 20 Playwright images; Vitest 5 also requires Node ≥ 22.12.
- **Base image `apify/actor-node-playwright-chrome:22-1.63.0`** instead of `apify/actor-node`, because the required browser fallback needs Chrome in the image. Cost: larger image and slower cold start; the HTTP route is still used first for everything.
- **`/api/graphql`, not `/graphql/query`.** `/graphql/query` now returns 403 "Page Not Found" to logged-out clients; `/api/graphql` with `lsd` + `doc_id` works.
- **Full browser-like headers are required.** Threads serves an empty app shell (no data) to clients without `Sec-Fetch-*` / client-hint headers; got-scraping's header generator provides them. The empty shell counts as a login wall.
- **Search depth.** Logged-out search returns one page of top results (typically 10–25) with no cursor. Pagination code is in place in case Threads exposes a cursor.
- **Nonexistent vs. private vs. blocked profiles** all redirect to `/login`, so they are indistinguishable. They are retried, can trigger the browser fallback, and are reported in `FAILED_INPUTS` with an explanation. Deleted posts (`?error=invalid_post`) are detected and never retried.
- **`maxResults`** caps total items (profiles + posts) per run; default 50 so the default input finishes fast and cheaply. `0` = unlimited.
- **Added `maxPostsPerSearch`, `startUrls`, `enableBrowserFallback`, `maxConcurrency`** inputs beyond the brief, for per-query limits, bulk URL import, and cost/speed control.
- **`hashtags`** contains inline `#tags` plus the post's Threads topic tag (also exposed alone as `topicTag`), since most Threads posts use topic tags instead of inline hashtags.
- **Outbound links** are unwrapped from `l.threads.com/?u=` redirects.
- **Dependency audit:** 13 moderate advisories all stem from `stream-json@1.9.1` (via Crawlee). The advisory affects its `pick/filter/replace` streamers on attacker-supplied JSON; Crawlee only uses `StreamArray` on its own state. Accepted; CI fails on `high`+ only. Revisit when Crawlee upgrades.
