# Changelog of PDF to Markdown, Tables & OCR Extractor (`cylindrical_lighthouse/document-intelligence`) Actor

- **URL**: https://apify.com/cylindrical\_lighthouse/document-intelligence/changelog.md
- **Full Actor documentation**: https://apify.com/cylindrical\_lighthouse/document-intelligence.md

## Changelog

All notable changes to this Actor are documented here. Format: [Keep a Changelog](https://keepachangelog.com/), versioning: [SemVer](https://semver.org/).

### \[1.0.2] - 2026-09-30

#### Changed

- Lower price for text pages: `page-text` $0.001 → **$0.0004** (GOLD and above $0.0003). A 100-page report now costs about $0.04. OCR and extraction prices are unchanged.

### \[1.0.1] - 2026-09-29

First public release on Apify Store.

#### Added

- Output schema: Documents, Tables, Markdown and the key-value store files.

#### Changed

- Pay-per-event prices set for all six tiers: `page-text` $0.001 → $0.0007, `page-ocr` $0.008 → $0.0065, `page-extract` $0.005 → $0.0035 (FREE → GOLD+). No per-document charge.
- Actor-start event description corrected to the 1 GB default memory.

### \[1.0.0] - 2026-09-29

First release.

#### Added

- PDF to Markdown with column-aware reading order, headings, lists and de-hyphenation, per document and per page.
- Table detection:
  - ruled tables from drawn cell borders, and borderless tables from column alignment;
  - header mapping for multi-line and spanning headers, leader-dot removal, superscript footnote markers and wrapped row labels;
  - output as rows keyed by column name, a raw cell grid and RFC 4180 CSV.
- OCR with Tesseract 5 for pages without a usable text layer (`ocrMode`: `auto`, `always`, `never`):
  - pages are rendered at the scan's native resolution (200–300 dpi);
  - grid lines are detected and erased before OCR, and their positions rebuild scanned tables;
  - any Tesseract language is supported, with English built in and others downloaded on demand.
- Inputs: document URLs, an Apify dataset plus URL field (chaining), and key-value store records.
- Optional schema-guided extraction with your own Anthropic or OpenAI key: JSON Schema validation and page citations.
- Metadata (title, author, dates, page count, language, PDF version, encryption) and provenance (`url`, `sha256`, `scrapedAt`).
- Per-document error records: `NOT_FOUND`, `ACCESS_DENIED`, `HTTP_ERROR`, `NETWORK_ERROR`, `NOT_PDF`, `EMPTY_FILE`, `FILE_TOO_LARGE`, `ENCRYPTED`, `CORRUPT_PDF`. A partial success still ends as SUCCEEDED, and a `SUMMARY` record is saved.
- Downloads are retried with exponential backoff on network errors, 429 and 5xx (honouring `Retry-After`); 404 and other client errors are not retried.
- Pay-per-event pricing (`page-text`, `page-ocr`, `page-extract`):
  - each page is charged as soon as it is processed, and processing stops at the run's spending limit;
  - state is migration-safe, so no page is charged twice.
- Oversized results move their Markdown and tables to the key-value store to stay under the 9 MB dataset item limit.
- Run-timeout awareness: when the run nears `ACTOR_TIMEOUT_AT`, processing stops and finished (charged) pages are delivered with `truncatedReason: "runTimeout"`.
- Dataset schema with Overview, Tables and Markdown views; input schema with examples on every field.
- Accuracy benchmark corpus and regression tests (unit, OCR, live end-to-end).
