# Changelog of Bluesky Historical Archive Scraper - Old Posts by Date & Author (`logiover/bluesky-historical-archive-scraper`) Actor

- **URL**: https://apify.com/logiover/bluesky-historical-archive-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/logiover/bluesky-historical-archive-scraper.md

## Changelog

### 2026-09-23

- Fleet-wide quality audit. Verified end to end against live data and re-checked the input schema, the output columns and the run configuration.
- Output verified on a live run: 31 columns returned, 79.9% of cells populated.
- Run reliability reviewed: 100.0% of public runs succeeded in the last 30 days.
- Input schema, output schema and pricing configuration reviewed.
- Noted that 2 column(s) came back empty in this sample (`replyToUri`, `rootUri`); these are under review.

### 2026-09-01

- Fleet-wide health check. Verified against this Actor's real run history: 30-day success rate, output row counts, per-field fill rates, and peak memory against the configured memory limit.
- Reviewed for the failure patterns that have cost this fleet runs — unguarded proxy setup, retry loops that can outlast the run's own time budget, and full-page HTML parsing that can exhaust a small container.
- No change to input, output fields or scraping logic.

### 2026-08-11

- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.

### 2026-08-01

- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.
- Declared 31 dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder.
- Declared 31 nullable dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder or brittle to sparse modes.

### 2026-07-30

- Quality-test fix: added a 3.5-minute wall-clock safety deadline that always exits SUCCEEDED with everything collected so far. Author-profile enrichment on large repost-heavy feeds (e.g. `nytimes.com`) can no longer push a run past the platform's 5-minute automated test window — near the deadline, enrichment is skipped so the feed walk itself finishes. Lowered the default prefilled **Max posts per account** to 150 and trimmed the per-request timeout to 25 s. Raise the run timeout to reach a full archive.

### 2026-07-22

- Initial release of the Bluesky Historical Archive Scraper.
- Full public account history via Bluesky's keyless AT-Protocol API (`app.bsky.feed.getAuthorFeed`), paginated by cursor to the account's first post.
- Optional thread / reply-tree expansion (`app.bsky.feed.getPostThread`).
- Curated Bluesky list feeds (`app.bsky.feed.getListFeed`).
- Author profile enrichment — follower, following and post counts (`app.bsky.actor.getProfile`), cached once per author.
- Client-side `afterDate` / `beforeDate` filtering with early-stop pagination (feeds are newest-first).
- 24 flat fields per row, deduplicated by URI.
- Fresh-proxy retry with exponential backoff on transient 429/403/5xx/network errors.
