PDF Text, Tables & OCR Extractor (Markdown output) avatar

PDF Text, Tables & OCR Extractor (Markdown output)

Pricing

from $3.00 / 1,000 pdf processeds

Go to Apify Store
PDF Text, Tables & OCR Extractor (Markdown output)

PDF Text, Tables & OCR Extractor (Markdown output)

Turn PDFs into usable data: per-page text, tables as JSON, document metadata and a Markdown version with headings, plus OCR only where a page has no text layer. Handles page ranges, size caps and encrypted files gracefully. Built for RAG pipelines and document processing.

Pricing

from $3.00 / 1,000 pdf processeds

Rating

0.0

(0)

Developer

Paul Vasquez

Paul Vasquez

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

PDF Text and Tables Extractor with OCR

Turn public PDFs or uploaded files into document rows, full text, Markdown, and structured table JSON. This Python 3.12 actor uses PyMuPDF for text, metadata and rendering, pdfplumber for tables, and local Tesseract through pytesseract for OCR. It needs no API keys, external OCR service, or separate website. Documents are processed sequentially in the actor container. Use it for report indexing, document search, research collections, and preparing material for downstream analysis. Extraction preserves source wording rather than summarizing documents.

Quick start

Use the included INPUT.json to process three public PDFs: W3C's sample, an arXiv paper, and a Census retail ecommerce report. The local validation completed well below two minutes; network availability and hosted OCR can change timings. For local development, create a Python 3.12 virtual environment, install requirements.txt, and run:

.venv/Scripts/python.exe -m unittest discover -s tests -v
apify validate-schema .actor/input_schema.json
powershell -NoProfile -ExecutionPolicy Bypass -File validation/run_live.ps1

The validation script copies INPUT into a fresh local store, runs the actual Apify SDK entry point, saves counts and timings, and checks persisted output. For normal execution use python -m src with INPUT in the default KVS. Uploaded PDFs must be binary records with application/pdf content type in that same run's default key-value store. Set keyValueStoreKeys to their record names; no separate storage ID or account token belongs in actor input.

Input options

pdfUrls accepts HTTP or HTTPS URLs; keyValueStoreKeys is the optional upload alternative. Either list can be omitted. Exact duplicates within each list are processed once. URLs must not contain embedded username/password credentials. pages accepts one-based selections such as 1-5,8; duplicate selections are merged and results use ascending document order. Out-of-range pages are ignored. A selection containing no actual pages produces an uncharged summary row.

extractText and extractTables default to true. Disabling text also disables OCR. ocrMode defaults to auto, which attempts OCR only when a selected page has fewer than 20 stripped native characters. always replaces each selected page's native text with recognized text; never disables recognition. ocrLanguage defaults to eng. The Docker image includes English; other Tesseract language packs require a custom image. Multiple installed languages can be requested with codes such as eng+deu.

outputMarkdown defaults to true. maxPages, default 200, rejects an entire document above the limit, even when only a subset was requested. maxFileMb, default 50, means MiB and applies to downloaded or uploaded bytes. Downloads check both declared length and streamed content. timeoutSecs defaults to 30 and bounds each download attempt and each OCR subprocess, not total extraction CPU time. The example uses 15 seconds. HTTP 429, server failures and transport failures receive two retries after one and two seconds; other HTTP errors do not.

Results and files

One dataset row represents each document. It includes source, fileName, fileSizeBytes, total pageCount, pagesProcessed, metadata, textPages, ocrPages, tables, warnings, and nullable error. Metadata includes title, author, subject, creator, producer, createdAt and modifiedAt. Dates retain their original PDF representation; missing values are null.

fullText contains at most 50,000 characters. The complete extracted text is saved as TEXT-<hash>.txt, identified by the additional textKey field. markdownKey points to MARKDOWN-<hash>.md; tablesKey points to TABLES-<hash>.json. Hashes incorporate source, document bytes and extraction settings. Tables contain {page, index, rows} with one-based page/table numbers and nested cell arrays, including null cells when applicable. Each pages entry contains page number, complete character count, OCR status, and at most 2,000 characters of text. SUMMARY stores document totals, eligible event requests and per-document timings. Empty source lists produce one free summary.

Pricing and extraction limits

The custom events are pdf-processed at $0.003 per successfully opened, accepted document with selected pages; text-page at $0.0005 per page with nonempty native extracted text; and ocr-page at $0.008 per completed OCR page, including recognition yielding empty text. A page with native text that also undergoes OCR incurs both page events. Failed downloads, malformed or encrypted PDFs, and over-limit documents produce uncharged error rows. Failed OCR attempts are uncharged and retain native text with a warning. Missing Tesseract also produces warnings. Local SDK runs request events but do not bill.

Artifacts are stored before charging, then the dataset row is written. Budget checks reject insufficient document budgets before charges start. Billing and storage are not transactional: a later charge or dataset failure cannot undo earlier charges. Configure all three custom events and disable synthetic charges before publication. Docker and real hosted billing remain unverified locally.

Markdown headings are font-size guesses relative to the dominant character font; OCR text has page headings only. Multi-column reading order, mathematical notation, merged cells, and borderless tables can be imperfect. Table detection uses native PDF geometry and does not reconstruct scanned tables from OCR. OCR accuracy depends on scan quality. Windows tests skip real OCR when Tesseract is absent; no system software is installed by validation. See VALIDATION.md for measured coverage and the exact remaining limitations.

Example output

Recorded local validation output from storage/live-20260926-064724/datasets/default/000000001.json, dataset row 1. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.

{
"source": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
"fileName": "dummy.pdf",
"fileSizeBytes": 13264,
"pageCount": 1,
"pagesProcessed": 1,
"textPages": 1,
"ocrPages": 0,
"tables": 0,
"fullText": "Dummy PDF file",
"warnings": [
"Page 1: OCR unavailable; Tesseract not installed"
],
"error": null
}

Use cases

  • A research librarian extracts native text from public papers and uses the full-text artifacts to populate a document index.
  • A retail analyst extracts tables from government reports, then checks cell alignment against the original PDFs before spreadsheet analysis.
  • A records digitization contractor processes scanned pages with installed OCR languages and reviews warnings and recognized text for quality.
  • A knowledge-base administrator converts selected report pages into Markdown and retains source links for later editorial verification.

Pricing example

100 accepted PDFs, 1,000 nonempty native-text pages, and 50 completed OCR pages cost (100 x $0.003) + (1,000 x $0.0005) + (50 x $0.008) = $1.20 in declared events. These are event counts: a page can contribute to both page categories. Rates come from .actor/pay_per_event.json. This calculation is an event subtotal, not a measured invoice; local validation does not bill.

Limitations

This saved row retains its actual missing-Tesseract warning; it demonstrates native extraction, not successful OCR. Extraction does not establish factual accuracy of the source document. Review page order and table cells before downstream use, and retrieve the text artifact when the dataset text cap would omit content.