Scrape public Threads profiles, posts, replies and search results without login. Get text, likes, replies, reposts, media, hashtags and follower counts. Export to JSON, CSV, Excel or API.
All notable changes to this Actor are documented here. Format: Keep a Changelog ; versioning: SemVer .
[1.0.0] - 2026-09-27
Added
Scrape public Threads profiles, a profile's posts (multi-part threads split and linked), post + replies (nested reply chains linked via parentPostId) and keyword search, without login.
HTTP-first route (HttpCrawler + got-scraping): parses the server-rendered data-sjs JSON, then paginates via /api/graphql using the page's own lsd token and preloaded doc_id + variables.
Token-refresh routine: when GraphQL rejects lsd/doc_id (HTML, 401/403, FB error envelope), the page is re-fetched on the same session for fresh tokens (max 2 per request) before failing over to session rotation.
Resumable pagination: the GraphQL cursor and per-input counters are stored on the request, so retries on a new session continue instead of restarting.
Browser fallback (PlaywrightCrawler): activated only after 3 consecutive login walls on the HTTP route; reuses the same parsers and runs the same GraphQL pagination via fetch() inside the page. Each item records scrapedVia: "http" | "browser".
Pay-per-event charging (actor-start $0.01, post-scraped $0.001, profile-scraped $0.003) with budget pre-checks, charge-after-push, serialized pushes, and a clean stop with budgetReached: true in STATS.
Deduplication by stable ID, persisted across migrations (STATE record).
Categorized error counters (blocked / rateLimited / proxy / network / parse / loginWall / notFound / other) in STATS, logged every 30 s.
FAILED_INPUTS record explaining every input that produced nothing.
Soft-failure detection: HTTP 200 login pages, data-less app shells, checkpoints/challenges and rate-limit pages are treated as blocks and retried with exponential backoff (jittered, capped at 30 s) and session rotation.
Decisions (made without asking, per project conventions)
Node 22, not 20. Node 20 reached end-of-life in April 2026 and Apify no longer publishes Node 20 Playwright images; Vitest 5 also requires Node ≥ 22.12.
Base image apify/actor-node-playwright-chrome:22-1.63.0 instead of apify/actor-node, because the required browser fallback needs Chrome in the image. Cost: larger image and slower cold start; the HTTP route is still used first for everything.
/api/graphql, not /graphql/query./graphql/query now returns 403 "Page Not Found" to logged-out clients; /api/graphql with lsd + doc_id works.
Full browser-like headers are required. Threads serves an empty app shell (no data) to clients without Sec-Fetch-* / client-hint headers; got-scraping's header generator provides them. The empty shell counts as a login wall.
Search depth. Logged-out search returns one page of top results (typically 10–25) with no cursor. Pagination code is in place in case Threads exposes a cursor.
Nonexistent vs. private vs. blocked profiles all redirect to /login, so they are indistinguishable. They are retried, can trigger the browser fallback, and are reported in FAILED_INPUTS with an explanation. Deleted posts (?error=invalid_post) are detected and never retried.
maxResults caps total items (profiles + posts) per run; default 50 so the default input finishes fast and cheaply. 0 = unlimited.
Added maxPostsPerSearch, startUrls, enableBrowserFallback, maxConcurrency inputs beyond the brief, for per-query limits, bulk URL import, and cost/speed control.
hashtags contains inline #tags plus the post's Threads topic tag (also exposed alone as topicTag), since most Threads posts use topic tags instead of inline hashtags.
Outbound links are unwrapped from l.threads.com/?u= redirects.
Dependency audit: 13 moderate advisories all stem from stream-json@1.9.1 (via Crawlee). The advisory affects its pick/filter/replace streamers on attacker-supplied JSON; Crawlee only uses StreamArray on its own state. Accepted; CI fails on high+ only. Revisit when Crawlee upgrades.