# Changelog of Wikipedia Revision History Scraper — Page Edit History by Date (`logiover/wikipedia-revision-history-scraper`) Actor

- **URL**: https://apify.com/logiover/wikipedia-revision-history-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/logiover/wikipedia-revision-history-scraper.md

## Changelog

### 2026-09-23

- Fleet-wide quality audit. Verified end to end against live data and re-checked the input schema, the output columns and the run configuration.
- Output verified on a live run: 17 columns returned, 99.4% of cells populated.
- Run reliability reviewed: 100.0% of public runs succeeded in the last 30 days.
- Input schema, output schema and pricing configuration reviewed.

### 2026-09-01

- Fleet-wide health check. Verified against this Actor's real run history: 30-day success rate, output row counts, per-field fill rates, and peak memory against the configured memory limit.
- Reviewed for the failure patterns that have cost this fleet runs — unguarded proxy setup, retry loops that can outlast the run's own time budget, and full-page HTML parsing that can exhaust a small container.
- No change to input, output fields or scraping logic.

### 2026-08-11

- Maintenance release: refreshed the build and dependencies.
- Re-verified live execution, non-empty structured output and dataset field/type integrity.
- Reviewed reliability (retries, pagination) and output quality as part of a full-fleet QA pass.

### 2026-08-01

- Completed the August 2026 full health check: verified empty/programmatic default, Console UI default, and two source-informed alternative inputs on Apify.
- Confirmed successful live execution, non-empty structured output, dataset-field/type integrity, and logical sample quality within the 5-minute quality window.
- Replaced the field-level output schema with a canonical `results` dataset link so run results open correctly in the Apify Console.
- Added a working default Wikipedia target so an empty-input run produces revision records.
- Aligned page, revision, parent and user IDs with the numeric types documented in the README.
- Corrected the deployed dataset contract regression that still declared numeric MediaWiki IDs as strings, and made all-target request failure fail closed instead of succeeding empty.
- Declared 17 dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder.
- Declared 17 nullable dataset fields from typed live cloud samples so the output contract is no longer an empty placeholder or brittle to sparse modes.

### 1.0 — 2026-07-22

#### Initial release

- Full page **revision (edit) history** extraction via the keyless **MediaWiki Action API** (`action=query&prop=revisions`).
- **18 fields per revision**: page, pageId, language, project, revId, parentId, timestamp, user, userId, isAnonymous, isMinor, comment, size, **sizeDelta**, tags, sha1, revisionUrl, and optional content.
- **Byte-level `sizeDelta`** computed against the immediately preceding version, with an extra anchor-revision fetch so the oldest emitted row is accurate at the cap boundary.
- **Full `rvcontinue` pagination** — pages through a page's entire history in batches of up to 500 revisions.
- **Category expansion** via `list=categorymembers` to fan a single category out into hundreds of member pages.
- **Filters**: `afterDate` / `beforeDate` (mapped to `rvend` / `rvstart` with `rvdir=older`), `userFilter` (`rvuser`), `maxRevisionsPerPage`, and global `maxItems`.
- **Any language edition and any MediaWiki project** (Wikipedia, Wiktionary, Wikibooks, Wikiquote, Wikinews, Wikisource, Wikiversity, Wikivoyage).
- Optional **full wikitext** per revision via `includeContent`.
- Politeness & reliability: descriptive **User-Agent** on every request, `maxlag=5` with exponential backoff, retries on 5xx/429/timeout with a fresh proxy IP per attempt, per-page try/catch, missing-page skip, and clean exit on empty input.
- Dedupe by `revId`; datacenter proxy default; 512 MB memory.
