PDF Text, Tables & OCR Extractor (Markdown output)
Pricing
from $3.00 / 1,000 pdf processeds
PDF Text, Tables & OCR Extractor (Markdown output)
Turn PDFs into usable data: per-page text, tables as JSON, document metadata and a Markdown version with headings, plus OCR only where a page has no text layer. Handles page ranges, size caps and encrypted files gracefully. Built for RAG pipelines and document processing.
Pricing
from $3.00 / 1,000 pdf processeds
Rating
0.0
(0)
Developer
Paul Vasquez
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Categories
Share
PDF Text and Tables Extractor with OCR
Turn public PDFs or uploaded files into document rows, full text, Markdown, and structured table JSON. This Python 3.12 actor uses PyMuPDF for text, metadata and rendering, pdfplumber for tables, and local Tesseract through pytesseract for OCR. It needs no API keys, external OCR service, or separate website. Documents are processed sequentially in the actor container. Use it for report indexing, document search, research collections, and preparing material for downstream analysis. Extraction preserves source wording rather than summarizing documents.
Quick start
Use the included INPUT.json to process three public PDFs: W3C's sample, an
arXiv paper, and a Census retail ecommerce report. The local validation completed
well below two minutes; network availability and hosted OCR can change timings.
For local development, create a Python 3.12 virtual environment, install
requirements.txt, and run:
.venv/Scripts/python.exe -m unittest discover -s tests -vapify validate-schema .actor/input_schema.jsonpowershell -NoProfile -ExecutionPolicy Bypass -File validation/run_live.ps1
The validation script copies INPUT into a fresh local store, runs the actual
Apify SDK entry point, saves counts and timings, and checks persisted output.
For normal execution use python -m src with INPUT in the default KVS. Uploaded
PDFs must be binary records with application/pdf content type in that same
run's default key-value store. Set keyValueStoreKeys to their record names;
no separate storage ID or account token belongs in actor input.
Input options
pdfUrls accepts HTTP or HTTPS URLs; keyValueStoreKeys is the optional upload
alternative. Either list can be omitted. Exact duplicates within each list are
processed once. URLs must not contain embedded username/password credentials.
pages accepts one-based selections such as 1-5,8; duplicate selections are
merged and results use ascending document order. Out-of-range pages are ignored.
A selection containing no actual pages produces an uncharged summary row.
extractText and extractTables default to true. Disabling text also disables
OCR. ocrMode defaults to auto, which attempts OCR only when a selected page
has fewer than 20 stripped native characters. always replaces each selected
page's native text with recognized text; never disables recognition.
ocrLanguage defaults to eng. The Docker image includes English; other
Tesseract language packs require a custom image. Multiple installed languages
can be requested with codes such as eng+deu.
outputMarkdown defaults to true. maxPages, default 200, rejects an entire
document above the limit, even when only a subset was requested. maxFileMb,
default 50, means MiB and applies to downloaded or uploaded bytes. Downloads
check both declared length and streamed content. timeoutSecs defaults to 30
and bounds each download attempt and each OCR subprocess, not total extraction
CPU time. The example uses 15 seconds. HTTP 429, server failures and transport
failures receive two retries after one and two seconds; other HTTP errors do not.
Results and files
One dataset row represents each document. It includes source, fileName,
fileSizeBytes, total pageCount, pagesProcessed, metadata, textPages,
ocrPages, tables, warnings, and nullable error. Metadata includes title,
author, subject, creator, producer, createdAt and modifiedAt. Dates retain their
original PDF representation; missing values are null.
fullText contains at most 50,000 characters. The complete extracted text is
saved as TEXT-<hash>.txt, identified by the additional textKey field.
markdownKey points to MARKDOWN-<hash>.md; tablesKey points to
TABLES-<hash>.json. Hashes incorporate source, document bytes and extraction
settings. Tables contain {page, index, rows} with one-based page/table numbers
and nested cell arrays, including null cells when applicable. Each pages
entry contains page number, complete character count, OCR status, and at most
2,000 characters of text. SUMMARY stores document totals, eligible event
requests and per-document timings. Empty source lists produce one free summary.
Pricing and extraction limits
The custom events are pdf-processed at $0.003 per successfully opened,
accepted document with selected pages; text-page at $0.0005 per page with
nonempty native extracted text; and ocr-page at $0.008 per completed OCR page,
including recognition yielding empty text. A page with native text that also
undergoes OCR incurs both page events. Failed downloads, malformed or encrypted
PDFs, and over-limit documents produce uncharged error rows. Failed OCR attempts
are uncharged and retain native text with a warning. Missing Tesseract also
produces warnings. Local SDK runs request events but do not bill.
Artifacts are stored before charging, then the dataset row is written. Budget checks reject insufficient document budgets before charges start. Billing and storage are not transactional: a later charge or dataset failure cannot undo earlier charges. Configure all three custom events and disable synthetic charges before publication. Docker and real hosted billing remain unverified locally.
Markdown headings are font-size guesses relative to the dominant character
font; OCR text has page headings only. Multi-column reading order, mathematical
notation, merged cells, and borderless tables can be imperfect. Table detection
uses native PDF geometry and does not reconstruct scanned tables from OCR.
OCR accuracy depends on scan quality. Windows tests skip real OCR when
Tesseract is absent; no system software is installed by validation. See
VALIDATION.md for measured coverage and the exact remaining limitations.
Example output
Recorded local validation output from storage/live-20260926-064724/datasets/default/000000001.json, dataset row 1. Fields are omitted for brevity; retained values are unchanged. This is a historical example, not a current-source claim.
{"source": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf","fileName": "dummy.pdf","fileSizeBytes": 13264,"pageCount": 1,"pagesProcessed": 1,"textPages": 1,"ocrPages": 0,"tables": 0,"fullText": "Dummy PDF file","warnings": ["Page 1: OCR unavailable; Tesseract not installed"],"error": null}
Use cases
- A research librarian extracts native text from public papers and uses the full-text artifacts to populate a document index.
- A retail analyst extracts tables from government reports, then checks cell alignment against the original PDFs before spreadsheet analysis.
- A records digitization contractor processes scanned pages with installed OCR languages and reviews warnings and recognized text for quality.
- A knowledge-base administrator converts selected report pages into Markdown and retains source links for later editorial verification.
Pricing example
100 accepted PDFs, 1,000 nonempty native-text pages, and 50 completed OCR pages cost (100 x $0.003) + (1,000 x $0.0005) + (50 x $0.008) = $1.20 in declared events. These are event counts: a page can contribute to both page categories. Rates come from .actor/pay_per_event.json. This calculation is an event subtotal, not a measured invoice; local validation does not bill.
Limitations
This saved row retains its actual missing-Tesseract warning; it demonstrates native extraction, not successful OCR. Extraction does not establish factual accuracy of the source document. Review page order and table cells before downstream use, and retrieve the text artifact when the dataset text cap would omit content.