All notable changes to this Actor are documented here. Format: Keep a Changelog , versioning: SemVer .
[1.0.2] - 2026-09-30
Changed
Lower price for text pages: page-text $0.001 → $0.0004 (GOLD and above $0.0003). A 100-page report now costs about $0.04. OCR and extraction prices are unchanged.
[1.0.1] - 2026-09-29
First public release on Apify Store.
Added
Output schema: Documents, Tables, Markdown and the key-value store files.
Changed
Pay-per-event prices set for all six tiers: page-text $0.001 → $0.0007, page-ocr $0.008 → $0.0065, page-extract $0.005 → $0.0035 (FREE → GOLD+). No per-document charge.
Actor-start event description corrected to the 1 GB default memory.
[1.0.0] - 2026-09-29
First release.
Added
PDF to Markdown with column-aware reading order, headings, lists and de-hyphenation, per document and per page.
Table detection:
ruled tables from drawn cell borders, and borderless tables from column alignment;
header mapping for multi-line and spanning headers, leader-dot removal, superscript footnote markers and wrapped row labels;
output as rows keyed by column name, a raw cell grid and RFC 4180 CSV.
OCR with Tesseract 5 for pages without a usable text layer (ocrMode: auto, always, never):
pages are rendered at the scan's native resolution (200–300 dpi);
grid lines are detected and erased before OCR, and their positions rebuild scanned tables;
any Tesseract language is supported, with English built in and others downloaded on demand.
Inputs: document URLs, an Apify dataset plus URL field (chaining), and key-value store records.
Optional schema-guided extraction with your own Anthropic or OpenAI key: JSON Schema validation and page citations.
Metadata (title, author, dates, page count, language, PDF version, encryption) and provenance (url, sha256, scrapedAt).
Per-document error records: NOT_FOUND, ACCESS_DENIED, HTTP_ERROR, NETWORK_ERROR, NOT_PDF, EMPTY_FILE, FILE_TOO_LARGE, ENCRYPTED, CORRUPT_PDF. A partial success still ends as SUCCEEDED, and a SUMMARY record is saved.
Downloads are retried with exponential backoff on network errors, 429 and 5xx (honouring Retry-After); 404 and other client errors are not retried.
each page is charged as soon as it is processed, and processing stops at the run's spending limit;
state is migration-safe, so no page is charged twice.
Oversized results move their Markdown and tables to the key-value store to stay under the 9 MB dataset item limit.
Run-timeout awareness: when the run nears ACTOR_TIMEOUT_AT, processing stops and finished (charged) pages are delivered with truncatedReason: "runTimeout".
Dataset schema with Overview, Tables and Markdown views; input schema with examples on every field.
Accuracy benchmark corpus and regression tests (unit, OCR, live end-to-end).