PDF to Markdown & RAG Chunks: Tables to CSV, OCR
Pricing
from $3.00 / 1,000 pdf processeds
PDF to Markdown & RAG Chunks: Tables to CSV, OCR
Convert PDFs into clean Markdown and RAG-ready chunks with page numbers and heading paths. Tables export to CSV and JSON, and offline OCR reads scanned pages. Failed files are never charged.
Pricing
from $3.00 / 1,000 pdf processeds
Rating
0.0
(0)
Developer
Bongo Seakhoa
Maintained by CommunityActor stats
1
Bookmarked
2
Total users
1
Monthly active users
10 minutes ago
Last modified
Categories
Share
Turn PDFs into clean Markdown and retrieval-ready chunks you can cite. Every chunk carries its page span, its heading path and exact character offsets into the Markdown. It also works as a PDF to JSON converter and extracts tables from PDF files three ways: inline in the Markdown, as JSON rows and as a downloadable CSV file. Scanned pages are read with offline OCR. Nothing is sent to a language model or a third-party API.
What the PDF to Markdown converter outputs
- Markdown with headings, paragraphs joined across line breaks, de-hyphenated words, bulleted and numbered lists, and tables as GitHub-flavoured Markdown. Multi-column pages are read in column order, and running headers, footers and page numbers are removed.
- RAG chunks sized in characters, split at headings, paragraphs and sentences, with
page_start/page_end,heading_path,char_start/char_endand a stablechunk_id. The same PDF with the same settings on the same Actor build always gives the same chunk IDs, so re-indexing a vector store is idempotent. - Tables as JSON rows (header plus body) and as CSV and JSON files. A table that continues on the next page is merged into one when its columns line up, and a repeated header row is dropped. A cell merged down several rows is repeated in each of them, so each CSV row stands on its own.
- Filled-in PDF forms: values typed into form fields appear next to their labels, and checkboxes show as
[x]or[ ]. - OCR for scans with Tesseract, offline, in English, German, French, Spanish, Italian, Portuguese, Dutch and Hungarian. In the default
automode only pages without a usable text layer are OCR'd (including scans whose text layer is just a scanner stamp or a Bates number), and only pages whose text came from OCR are charged at the OCR price. - An estimate mode that counts pages, tells you which pages would need OCR and shows the projected price before you convert anything. It charges none of this Actor's events.
- Honest failures. Broken, encrypted, oversized or unreachable files get a status row with an error code and are never charged.
Typical uses: a PDF chunker that feeds a vector database for retrieval-augmented generation, LangChain or LlamaIndex PDF pipelines that need page citations, extracting tables from reports and invoices into spreadsheets, and turning scanned archives into searchable text.
How to convert PDFs to Markdown and RAG chunks
- Paste one or more PDF links into PDF URLs. Links must point at the PDF file itself; redirects are followed and Dropbox
?dl=0links are converted to downloads. - Or pick a key-value store that holds your PDFs (upload them in Apify Console under Storage), optionally naming the record keys to use.
- Leave the defaults or adjust chunk size, OCR and limits, then start the run.
- Open the Output tab. The Documents, Chunks and Tables views each list every row of the run, so rows of the other two types show empty cells; sort or filter on
record_type, or use the.md,.csvand.jsonfiles in the run's key-value store.
Starting the Actor with no sources at all converts a small bundled synthetic sample PDF so you can see the output format first. The sample is charged none of this Actor's events; Apify's standard Actor start fee still applies, as it does to every run.
Input: PDF URLs or files in a key-value store
{"sources": ["https://example.com/reports/annual-report-2025.pdf","https://www.dropbox.com/s/abc123/scan.pdf?dl=0"],"mode": "convert","outputs": ["markdown", "chunks", "tables"],"chunkSize": 1500,"chunkOverlap": 150,"ocrMode": "auto","ocrLanguages": ["eng"],"pageRange": "1-50","maxPages": 200,"maxFileSizeMb": 50,"perDocumentTimeoutSecs": 900,"includePageMarkers": false,"csvFormulaGuard": true}
| Field | Default | What it does |
|---|---|---|
sources | [] | PDF URLs (http or https). |
keyValueStoreId, keyValueStoreKeys | none | Read PDFs from a key-value store. Without keys, every record whose key ends in .pdf is used (up to 1,000). |
useBundledSample | false | Also convert the 3-page synthetic sample (no event charges). |
mode | convert | estimate only counts pages and projects the price; it charges none of this Actor's events. |
outputs | all three | Any of markdown, chunks, tables. |
chunkSize, chunkOverlap | 1500, 150 | Characters (about 4 per token). Overlap is at most half the chunk size. |
ocrMode | auto | auto (pages without usable text only), force (every page, each charged as ocr-page) or off. |
ocrLanguages | ["eng"] | Any of eng deu fra spa ita por nld hun. |
pageRange | all pages | For example 1-5,8,10-. |
maxPages | 200 | Pages per document; later pages are not converted and not charged. |
maxFileSizeMb | 50 | Larger files are rejected before download completes. Also capped at a quarter of the run's memory: 256 MB at the default 1 GB. |
perDocumentTimeoutSecs | 900 | A slow document returns the pages finished so far as partial. |
includePageMarkers | false | Adds a <!-- page N --> line before the first block that starts on each page. Blank pages, and pages whose text only continues a paragraph or table from the page before, get no marker. |
csvFormulaGuard | true | In the table CSV files, a cell that starts with =, +, -, @, a tab or a carriage return (also after leading spaces or invisible characters, or in full-width form) gets a leading ' so spreadsheets show it as text instead of running it as a formula. Signed numbers such as -12.5, +3 or -45%, and cells holding only dashes or plus signs (-, --, often a nil placeholder), are left alone. JSON rows, the Markdown and the dataset always keep the raw text. Set false for raw CSV cells. |
password | none | Password for encrypted PDFs (secret field, never logged). |
Output: Markdown, RAG chunks and tables as JSON and CSV
The dataset holds three kinds of rows, told apart by record_type. Filter on it when you read the dataset through the API. The examples below are the rows of the bundled sample, trimmed where marked with ....
Document row (one per PDF):
{"record_type": "document","document_id": "f6a2ab3234b42a02","document_index": 0,"source": "bundled-sample:sample.pdf","source_type": "sample","mode": "convert","status": "ok","error_code": null,"pages_total": 3,"pages_processed": 3,"text_pages": 2,"ocr_pages": 1,"empty_pages": 0,"table_count": 1,"chunk_count": 3,"markdown": "# Synthetic Sample Report\n\nThis is a synthetic sample PDF bundled with the PDF Markdown Actor. ...","markdown_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02.md?signature=<signature>","warnings": [],"charged_events": {}}
The sample is not charged, so its charged_events is empty. For a PDF from your own URL it lists what was charged, for example {"pdf-processed": 1, "page-processed": 2, "ocr-page": 1}.
Chunk row:
{"record_type": "chunk","document_id": "f6a2ab3234b42a02","chunk_id": "a730736c675c2a6b183d6ea1","chunk_index": 2,"text": "## Quarterly figures\n\nThe invented quarterly figures below are for demonstration.\n\n| Quarter | Orders | Revenue (EUR) | Returns |\n| --- | --- | --- | --- |\n| Q1 | 1,204 | 18,930.00 | 31 |\n... \n\nThe Actor reads it with offline OCR.","char_start": 325,"char_end": 750,"page_start": 2,"page_end": 3,"heading_path": ["Synthetic Sample Report", "Quarterly figures"],"table_ids": ["t1"],"table_header": null,"approx_tokens": 107}
Table row:
{"record_type": "table","document_id": "f6a2ab3234b42a02","table_id": "t1","page_start": 2,"page_end": 2,"n_rows": 5,"n_cols": 4,"header": ["Quarter", "Orders", "Revenue (EUR)", "Returns"],"rows": [["Q1", "1,204", "18,930.00", "31"], ["Q2", "1,388", "21,115.50", "27"], ["Q3", "1,512", "23,480.25", "40"], ["Q4", "1,690", "26,002.75", "35"]],"csv_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02-t1.csv?signature=<signature>","json_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02-t1.json?signature=<signature>"}
markdown[char_start:char_end] is exactly the chunk's text, so you can always point back to the source passage. A chunk that starts partway down a table also carries the table's header row and separator in table_header (outside text), so you can embed the two together and the columns keep their names. Very long Markdown is left out of the document row (markdown is null) and is always available from markdown_url. A very large table's rows are left out of its table row (rows is null, rows_omitted is true) and stay in its CSV and JSON files. The run's key-value store also gets a SUMMARY record with the counts for the whole run.
Document status is ok, partial (some content may be missing: the warnings say why, for example a page limit, a timeout or a page OCR could not read), failed (nothing extracted; error_code says why, for example BLOCKED_DESTINATION for a link to a private or local network address) or skipped (not started: CHARGE_LIMIT_REACHED when your maximum charge was reached, RUN_TIME_LIMIT when the run was about to time out). If Apify restarts a run (for example when it moves the run to another server) while a document is being charged, that document is not converted again: its row carries a BILLING_UNCERTAIN_AFTER_RESTART warning, and its charged_events may not list every charge made before the restart.
Pricing per PDF and per page
This Actor is priced per event: you pay for what was converted.
| Event | Price | When it is charged |
|---|---|---|
pdf-processed | $0.003 per document | Once per document that produced content (status ok or partial). |
page-processed | $0.0003 per page | Each page whose text came from the PDF's text layer. |
ocr-page | $0.006 per page | Each page whose text came from OCR, instead of page-processed. |
With ocrMode set to force, every page is charged as ocr-page, even pages that have a good text layer. Apify adds the Actor start fee to every run ($0.00005 per run at up to 1 GB of memory, once more for each extra GB). Chunks, tables, files and dataset rows are not charged separately.
Worked examples:
- A 10-page born-digital report: $0.003 + 10 × $0.0003 = $0.006.
- A 20-page scanned contract: $0.003 + 20 × $0.006 = $0.123.
- 1,000 single-page invoices with a text layer: 1,000 × ($0.003 + $0.0003) = $3.30.
- A 30-page document of which 4 pages are scans: $0.003 + 26 × $0.0003 + 4 × $0.006 = $0.0348.
Never charged by this Actor: documents that fail (broken, encrypted without the right password, not a PDF, too large, unreachable), blank pages and pages that produce no text, pages beyond maxPages or outside pageRange, the bundled sample, and estimate mode. Only Apify's start fee applies to those runs.
Your maximum charge is respected. Before each document the Actor checks what is left of your run's maximum total charge and converts only the pages it can charge for. It always keeps one text page's price ($0.0003) unspent below your limit, so a document whose exact price would just fit can still be cut short or skipped. How pages are budgeted depends on ocrMode:
off: every page at $0.0003.force: every page at $0.006.auto: every page at $0.006 while that fits. When it does not, a quick pre-check of the document predicts which pages need OCR, and the other pages are budgeted at $0.0003. If the pre-check misses a page that needs OCR, Apify stops charging at your limit, so you never pay more than your maximum.
A document that does not fit in full is converted up to the last page that fits and marked partial with a CHARGE_LIMIT_PAGES warning. A document that cannot pay for $0.003 plus its first page, and every document after it, is listed as skipped with CHARGE_LIMIT_REACHED and is not charged; the run normally finishes as Succeeded. If Apify itself ends the run at the limit, documents with no row were not processed and were not charged. Rerun the skipped documents with a higher limit; for born-digital PDFs, ocrMode off also makes the most of a small limit.
Limits: read before you buy
- Layout is heuristic. Reading order handles one to four columns. Sidebars next to a column, text wrapped around figures and right-to-left scripts are not reconstructed correctly. Headings are inferred from font size and bold text and go down to H3 only.
- Tables. Ruled tables (with lines) are detected well. Borderless tables are detected conservatively, so some are missed and stay as text. A cell merged down several rows is repeated in each row; a cell merged across columns (a section row or a note) keeps its text in its first column. A simple two-row grouped header (one spanning cell over its sub-columns) becomes one row ("Readings Min", "Readings Max"). A table continued on the next page is merged only when its columns line up with the first part. Tables on rotated pages are not detected. In the CSV files, cells that a spreadsheet would run as a formula start with
'(seecsvFormulaGuard). - Forms. Values of fillable form fields are read on upright pages only, and not on pages that are OCR'd.
- OCR output is plain paragraphs. Scanned pages give text, not headings or tables. Handwriting is not supported. Accuracy depends on scan quality; low-confidence pages get an
OCR_LOW_CONFIDENCEwarning. Only the eight languages listed above are installed. With OCR off, a scan whose text layer is only a stamp gets aSCANNED_PAGE_NOT_OCREDwarning. - OCR is slow at the default memory. Apify gives a run a quarter of a CPU core at 1 GB of memory. In our tests a scanned page took 8 to 17 seconds at 1 GB, depending on how dense the scan is. At 2 GB, the maximum for this Actor, a sparse scan took about 3 seconds per page; dense scans take longer. With the defaults, a scan of 50 to 100 pages can hit the 900-second limit per document and come back
partial. For large scans use 2 GB of memory, raiseperDocumentTimeoutSecs, or split the work withpageRange. Text pages are fast (well under a second each). - Downloads. Links must return the PDF itself. Pages that show a viewer, a cookie wall or a login (Google Drive previews, SharePoint pages and the like) are reported as
NOT_A_PDF. Servers that block automated downloads returnHTTP_ERRORwith the status code. Links must point to the public internet: see "Security". - Sideways text in a minority orientation on a page (a margin note or stamp) is left out, with a warning.
- Damaged files that lost their cross-reference table are reported as
CORRUPT; they are not repaired. - Token counts in
approx_tokensare characters divided by four, not a tokenizer count. - Per run: at most 1,000 listed sources plus up to 1,000
.pdfrecords read from a key-value store,maxPagesup to 5,000, and files up to 500 MB but never more than a quarter of the run's memory (128 MB at 512 MB, 256 MB at 1 GB, 500 MB at 2 GB). Defaults are 200 pages and 50 MB.
FAQ
Is my document sent to an AI service? No. Text extraction uses PDFium and pdfminer, and OCR uses Tesseract, all inside the Actor's container. There is no language model and no external OCR API.
Why is a document partial? Some content may be missing. The warnings list gives the reason codes, for example PAGES_TRUNCATED (page limit), CHARGE_LIMIT_PAGES (cut short to fit your maximum charge), DEADLINE_REACHED (time limit), OCR_FAILED or NO_TEXT_LAYER (a scanned page with OCR off). You pay only for pages that produced text.
How do I check the price before converting? Run with "mode": "estimate". Each document row then shows pages_total, text_pages, ocr_pages and projected_cost_usd, and none of this Actor's events are charged (Apify's start fee still applies). The estimate reads character counts only, so a page whose text layer is unreadable may be counted as a text page, and a scan with a stamp is counted as an OCR page even if the conversion ends up billing it as a text page.
Can I use password-protected PDFs? Yes: set password. It applies to every document in the run, is stored as a secret input and is never logged.
Are the chunk IDs stable? Yes, within one Actor build. A chunk ID is a hash of the document content, the chunk's position and its text. The same file with the same settings on the same build always produces the same IDs. A new build that improves extraction (or updates Tesseract) can change the text or the chunk boundaries, and with them the IDs, so re-index after an update if you rely on them.
Can I read the output from code? Yes. Use the dataset items endpoint and filter on record_type, or download the .md, .csv and .json files from the run's key-value store.
Something does not work. Open an issue on the Actor's Issues tab with the run link and, if you can share it, a PDF that shows the problem. Issues are read on a best-effort basis; there is no guaranteed response time.
Security
- Only public internet hosts are fetched. Before each request, and again for every redirect, the Actor looks up the link's host and refuses it if any of its addresses is private, local or reserved:
localhostand other loopback addresses, private networks (10.x, 172.16-31.x, 192.168.x, fc00::/7), link-local addresses including the cloud metadata address 169.254.169.254, carrier-grade NAT (100.64.0.0/10), multicast and reserved ranges, and IPv6 forms that wrap such an IPv4 address. Numeric hosts such ashttp://2130706433/count as the address they stand for. Links with a user name or password in them (https://user:pass@host/...) are refused too. A refused link becomes afaileddocument witherror_codeBLOCKED_DESTINATIONand is never charged. To convert a file from a private network, upload it to a key-value store and usekeyValueStoreId. - CSV files are safe to open in a spreadsheet by default (
csvFormulaGuard, above).
Privacy and data retention
The Actor reads only the URLs and key-value store records you give it. Your PDFs are processed inside your run and are not sent anywhere else. The log contains counts, status and error codes, never document text, URLs or passwords. The Actor keeps nothing after the run ends: the outputs live in your run's dataset and key-value store under your Apify account, and Apify's storage retention applies to them (unnamed storages are deleted automatically after a retention period).
About the sample and the tests
The bundled sample and every test file used to develop this Actor are synthetic PDFs generated for the purpose and labelled "synthetic test fixture" in their metadata. No third-party documents were used.
Licences
The Actor is proprietary software by Bongo Seakhoa. It is built on open-source components under permissive licences (PDFium via pypdfium2, pdfminer.six, pdfplumber, Pillow, Tesseract, httpx and the Apify SDK, among others) and uses no AGPL or GPL Python library such as PyMuPDF. certifi (MPL-2.0) is included unmodified. The Debian base image contains GPL and LGPL system tools and libraries, used unmodified as separate programs. The full list and licence texts are in the THIRD_PARTY_NOTICES file shipped with the Actor image.