PDF to Markdown, Tables & OCR Extractor
Pricing
from $0.30 / 1,000 page (text layer)s
PDF to Markdown, Tables & OCR Extractor
Convert PDFs and scanned documents to Markdown, header-mapped tables (JSON + CSV) and schema JSON, with OCR and page-level provenance.
Pricing
from $0.30 / 1,000 page (text layer)s
Rating
0.0
(0)
Developer
Lighthouse Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Turn PDFs into clean Markdown, header-mapped tables (JSON + CSV) and schema-guided JSON, including scanned documents (OCR) and borderless financial and statistical tables that plain text extractors flatten into word soup. Every result carries page numbers, SHA-256 and source URL, so RAG pipelines, data teams and AI agents can cite exactly where a value came from.
- Tables that survive: ruled grids and borderless tables with multi-line and spanning headers, leader dots and footnote markers come out as rows keyed by column name, plus CSV.
- Scans that read: pages without a text layer are OCR-ed automatically (Tesseract 5); grid lines are detected and removed first, so scanned tables keep their structure.
- Page provenance: per-page Markdown,
pageNumberon every table, page citations for extracted fields. - Chains with your crawler: feed it the dataset of Website Content Crawler (or any Actor) or files in a key-value store.
- Pay per page: $0.0004 per text page, $0.008 per OCR page. No subscription.
Try it: the default input converts the Federal Reserve's 3-page G.19 Consumer Credit release in about 5 seconds, for about $0.001.
What does it do?
For each PDF it downloads (or reads from Apify storage), the Actor:
- Reads the text layer with exact positions (pdf.js) and rebuilds the reading order, including multi-column layouts.
- Detects tables: from drawn cell borders (ruled tables) and from column alignment (borderless tables), then maps the header rows onto column names.
- OCRs pages that have no usable text layer (
ocrMode: auto), or every page (always). - Writes one dataset item per document: full Markdown, per-page Markdown, tables (JSON rows + CSV), metadata (title, author, dates, page count, language) and provenance (
url,sha256,scrapedAt). - Optionally sends the Markdown to your own Anthropic or OpenAI key to fill a JSON Schema you provide, with page citations per field.
Failed documents (404, not a PDF, password-protected, corrupt, too large) get a clear error record, and the rest of the run carries on.
Why use it?
| This Actor | Plain PDF-to-text tools | |
|---|---|---|
| Borderless tables (financial statements, statistics) | Rows keyed by header, CSV | Numbers run together on one line |
| Scanned pages | OCR with grid-line removal | Empty text |
| Multi-column layouts | Column-aware reading order | Lines interleaved across columns |
| Provenance | Page numbers, SHA-256, source URL | Usually none |
| Structured JSON | Your JSON Schema + page citations (BYO key) | No |
| Price | $0.0004 per text page, $0.008 per OCR page | Varies |
Measured accuracy (open benchmark, see Accuracy): 99–100% table cell accuracy on synthetic, Federal Reserve and Census tables, 99.6% character accuracy on a noisy 3-column scan.
Use cases
- RAG and AI agents: convert PDFs found by a crawler into Markdown chunks with page numbers for citations.
- Data extraction from reports: statistical releases, annual reports, price lists and regulatory filings, straight into rows and CSV.
- Scanned archives: OCR old reports, letters and forms into searchable text.
- Invoices and forms at scale: schema-guided JSON (invoice number, totals, dates) with your own LLM key.
- Monitoring: schedule a run on a URL that publishes a new PDF each month and diff the tables.
How to use it
- Click Try for free (or open the Actor in Apify Console).
- Paste one or more PDF URLs into Document URLs, or set Dataset ID to chain from another Actor.
- Keep OCR mode on
auto. Set Max pages per document to cap cost on very long files. - Click Start. Download results as JSON, CSV or Excel, or use the Tables and Markdown views.
Example input
{"documentUrls": [{ "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf" }],"ocrMode": "auto","maxPagesPerDocument": 50}
Chain it after Website Content Crawler
Website Content Crawler can save the PDFs it finds (saveFiles) or list their URLs in its dataset. Pass that dataset to this Actor:
{ "datasetId": "<WCC run's default dataset ID>", "datasetUrlField": "url", "maxDocuments": 500 }
In Apify Console you can do this automatically: on the crawler's task, add an integration → Run Actor with this Actor and the input above (use {{resource.defaultDatasetId}} as datasetId). Combined with a schedule, you get a weekly "crawl site → convert new PDFs" pipeline. Unchanged files can be skipped downstream by sha256.
Files already in a key-value store (for example WCC's saveFiles output) work too:
{ "keyValueStoreId": "<store ID>", "keyValueStoreKeys": ["report-2025.pdf", "report-2026.pdf"] }
Call it from code or an AI agent
- API:
POST https://api.apify.com/v2/acts/<username>~document-intelligence/run-sync-get-dataset-items?token=<TOKEN>with the input as JSON. The response is the dataset items. - MCP: add the Actor to the Apify MCP server and let the agent call it with
documentUrls. Field descriptions are written for tool use.
Input
| Field | Type | Description | Example |
|---|---|---|---|
documentUrls | array | Public PDF URLs (text or scanned). One dataset item per URL. | [{"url": "https://…/report.pdf"}] |
ocrMode | string | auto (OCR only pages without text), always, never | auto |
ocrLanguage | string | Tesseract language codes joined with +. English is built in; others download on demand. | eng, deu, eng+fra, chi_sim |
maxPagesPerDocument | integer | Pages processed per document, from page 1 (default 200) | 50 |
maxDocuments | integer | Documents per run, including failed ones (default 100) | 100 |
datasetId | string | Read URLs from an Apify dataset (chaining) | aBcD1234EfGh5678 |
datasetUrlField | string | Field with the URL; dot paths and arrays allowed (default url) | url |
keyValueStoreId | string | Key-value store holding PDF files | username~my-store |
keyValueStoreKeys | array | Record keys of the PDFs in that store | ["invoice-001.pdf"] |
extractionSchema | object | Optional JSON Schema of data to extract with your LLM key | see below |
extractionInstructions | string | Extra guidance for the model | Amounts in USD |
llmProvider | string | anthropic or openai | anthropic |
llmModel | string | Model ID exactly as in your provider's docs (required with a schema) | |
llmApiKey | secret | Your API key; stored encrypted by Apify, never logged or output | |
maxFileSizeMb | integer | Skip larger files with an error record (default 100) | 100 |
proxyConfiguration | object | Only if a server blocks direct downloads | {"useApifyProxy": true} |
Schema extraction example
{"documentUrls": [{ "url": "https://example.com/invoice-0042.pdf" }],"extractionSchema": {"type": "object","properties": {"invoiceNumber": { "type": "string" },"invoiceDate": { "type": "string", "description": "ISO date" },"total": { "type": "number" }},"required": ["invoiceNumber", "total"]},"llmProvider": "anthropic","llmModel": "<model ID from your provider>","llmApiKey": "<your key>"}
Output
One item per document. Real output for a one-page invoice, shortened:
{"url": "https://example.com/invoice-0042.pdf","status": "succeeded","error": null,"fileName": "invoice-0042.pdf","sha256": "5b884246dca561b0f49300e95d84dc3a1acde5552eb292ade884ec503a18b259","metadata": { "title": "Invoice INV-2026-0042", "author": null, "pageCount": 1, "language": null, "creationDate": "2026-09-29T20:25:15.000Z" },"pageCount": 1,"pagesProcessed": 1,"pagesOcr": 0,"truncated": false,"markdown": "# INVOICE\n\nInvoice number: INV-2026-0042\n\n…\n\n| Description | Qty | Unit price | Amount |\n| --- | --- | --- | --- |\n| Widget, standard | 10 | 4.50 | 45.00 |\n…","pages": [{ "pageNumber": 1, "markdown": "# INVOICE …", "tableIds": ["p1-t1"], "ocrUsed": false, "ocrConfidence": null, "error": null }],"tables": [{"tableId": "p1-t1","pageNumber": 1,"detection": "aligned","columns": ["Description", "Qty", "Unit price", "Amount"],"rows": [{ "Description": "Widget, standard", "Qty": "10", "Unit price": "4.50", "Amount": "45.00" },{ "Description": "Total", "Qty": null, "Unit price": null, "Amount": "76.95" }],"csv": "Description,Qty,Unit price,Amount\r\n\"Widget, standard\",10,4.50,45.00\r\n…"}],"extracted": null,"scrapedAt": "2026-09-29T21:33:46.868Z"}
- Missing values are
null, never empty strings. Table cells are strings exactly as printed ("1,234.5"), so nothing is lost. Footnote markers are kept as superscripts (Nonrevolving³). statusissucceeded,partial(some pages failed, or the spending limit or run timeout was reached) orfailed, with anerror.codesuch asNOT_FOUND,NOT_PDF,ENCRYPTED,CORRUPT_PDForFILE_TOO_LARGE.- A run summary is saved in the key-value store as
SUMMARY. - Dataset views: Overview (one row per document), Tables (one row per table with CSV), Markdown.
Pricing
Pay per event: you pay only for pages processed.
| Event | Price | When |
|---|---|---|
| Page (text layer) | $0.0004 | Each page with a text layer, converted with its tables |
| Page (OCR) | $0.008 | Each scanned or image-only page read with OCR |
| Page (schema extraction) | $0.005 | Each page sent to your LLM for extractionSchema. LLM tokens are billed to your key by your provider |
| Actor start | $0.00005 per GB of memory | Apify's standard start event ($0.00005 per run at the default 1 GB) |
Failed documents and failed pages are not charged. Set a maximum cost per run in Console and the Actor stops cleanly when it is reached.
Cost examples
| Job | Pages | Cost |
|---|---|---|
| Default input (G.19, 3 text pages) | 3 text | about $0.001 |
| 100-page annual report | 100 text | $0.04 |
| 20-page scanned contract | 20 OCR | $0.16 |
| 1,000 mixed pages (80% text, 20% scans) | 800 text + 200 OCR | $2.40 |
| 50 one-page invoices with schema extraction | 50 text + 50 extraction | $0.27 + your LLM tokens |
Platform usage (compute) is included in these prices.
New to Apify? A free Apify account includes monthly platform credit.
Accuracy benchmark
Measured on an open corpus of public-domain and synthetic PDFs with hand-checked ground truth The same metrics run as regression tests on every change.
| Test | Metric | Result |
|---|---|---|
| Synthetic ruled grid tables (98 cells) | cell accuracy | 100% |
| Synthetic borderless + booktabs tables (94 cells) | cell accuracy | 100% |
| Federal Reserve G.19, rotated page, footnotes (126 cells) | cell accuracy | 99.2% |
| Census P60-280 Table A-3, leader dots, 13 columns (156 cells) | cell accuracy | 99.4% |
| Federal Reserve H.8 p.3, held out from tuning (91 cells) | cell accuracy | 92.3% strict; every numeric cell correct, line numbers get their own column |
| Two-column article | text accuracy (reading order) | 100% |
| Scan: 3-column IRS page, 200 dpi | character accuracy | 99.6% |
| Scan: 1949 NACA report | word recall | 98.9% |
| Scan: ruled table with 0.6° skew and noise | cell accuracy | 91–99% |
In the same benchmark, a pdfplumber-based pipeline averaged 59.8% cell accuracy on the tables (it missed borderless tables entirely), and Docling averaged 80.9% while needing 2+ GB RAM and minutes per page on CPU.
FAQ
Which files are supported? PDF, both text PDFs and scans or image-only PDFs. DOCX, PPTX, XLSX and HTML are planned. Other files get a NOT_PDF error record.
Do I need to choose OCR? No. auto OCRs only pages without a usable text layer. Use always for scans that carry a poor embedded text layer (common with old OCR software).
Which OCR languages? Any of Tesseract's ~100 languages: eng is built in, others (deu, fra, spa, chi_sim, jpn, ara…) are downloaded once per run. Combine them with +.
Are numbers converted to numbers? Table cells are kept as printed strings, so formats like (1,234), n.a. and footnote markers are not lost. Convert downstream as needed.
How are password-protected PDFs handled? They are reported with ENCRYPTED. The Actor never tries to break or bypass passwords. PDFs that are only "owner-locked" (printing or copying restricted) are processed.
What about my LLM key? It is an encrypted secret input. It is never logged or written to the output, and it is only sent to the provider you chose. Without an extractionSchema, no LLM is called.
Is it legal? You decide which documents to process; only use documents you are allowed to access. This Actor does not log in, solve CAPTCHAs or bypass paywalls. Documents may contain personal data, which is protected by the GDPR and other regulations: process personal data only with a legitimate reason, and consult your lawyers if unsure.
Privacy and retention: documents are processed in your own Apify run. Nothing is stored anywhere else, and results live only in your run's storage under your account's data retention settings. When schema extraction is on, page text is sent to your chosen LLM provider under your own key.
Limitations
- PDF only in this version.
- Very complex tables (nested tables, cells with several lines of text in borderless tables, heavily merged cells) may need clean-up. Spanning group headers are repeated into each column name they cover, which is sometimes approximate.
- Charts and images are not described; text inside images is read only when the page is OCR-ed.
- OCR is slower than text extraction: about 15–20 s per page at the default 1 GB memory (about 12 s at 2 GB). Text pages take a fraction of a second. For scans longer than about 150 pages, raise the run timeout (default 1 hour) or the memory. If the run gets close to its timeout, the Actor stops early and still delivers the finished pages (
truncatedReason: "runTimeout"). Handwriting is not supported. - Text in a page's text layer is used as-is: scans with a bad embedded OCR layer need
ocrMode: always. - Schema extraction sends at most about 400,000 characters (about 100k tokens) per document; later pages are left out and listed in
extracted.pagesUsed. - Dataset items are limited to 9 MB. For huge documents the Markdown (and, if needed, the tables) move to the key-value store; see
markdownStoreKey.
Example run
Ready-made examples are on the Example tasks tab, for instance "Extract tables from a PDF report to JSON and CSV" and "OCR a scanned PDF to Markdown".
Changelog
See ./CHANGELOG.md.
Support
Found a bug or need a feature, such as another document format? Open an issue on the Issues tab. We reply within 48 hours. Please include the run ID and, if possible, a public link to the PDF.