# Changelog of Internet Archive Metadata Scraper — Bulk archive.org Export (`logiover/internet-archive-metadata-scraper`) Actor

- **URL**: https://apify.com/logiover/internet-archive-metadata-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/logiover/internet-archive-metadata-scraper.md

## Changelog

### 2026-09-01

- Fleet-wide health check. Verified against this Actor's real run history: 30-day success rate, output row counts, per-field fill rates, and peak memory against the configured memory limit.
- Reviewed for the failure patterns that have cost this fleet runs — unguarded proxy setup, retry loops that can outlast the run's own time budget, and full-page HTML parsing that can exhaust a small container.
- No change to input, output fields or scraping logic.

### 2026-08-11

- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.

### 2026-08-01

- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.
- Replaced the field-level output schema with a canonical `results` dataset link so run results open correctly in the Apify Console.
- Declared 16 dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder.
- Declared 18 nullable dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder or brittle to sparse modes.

### 2026-07-22

- Initial release of the Internet Archive Metadata Scraper.
- Bulk item-metadata export from archive.org via the keyless **Scrape API** (cursor pagination, millions of items) with the **Advanced Search API** used automatically whenever a `sort` is requested (capped ~10,000 results).
- Query builder: pass a raw Lucene `query` OR combine structured filters — `collection`, `mediaType`, `creator`, `subject` and a `dateFrom`/`dateTo` range — joined with AND.
- Selectable metadata `fields`, optional per-item file-list enrichment (`fetchFiles`) from the `/metadata/` endpoint, `maxItems` cap honored exactly across pagination, and dedupe by `identifier`.
- Safe normalization of fields that archive.org returns as either a string or an array (`collection`, `subject`, `format`, `language`).
- Datacenter proxy by default with fresh IP + exponential backoff on 429/5xx, graceful empty exit, and per-item error isolation.
