PDF & Document to Markdown for LLMs (OCR, Tables, RAG Chunks) avatar

PDF & Document to Markdown for LLMs (OCR, Tables, RAG Chunks)

Pricing

from $2.00 / 1,000 document converteds

Go to Apify Store
PDF & Document to Markdown for LLMs (OCR, Tables, RAG Chunks)

PDF & Document to Markdown for LLMs (OCR, Tables, RAG Chunks)

Convert PDFs, Word, PowerPoint, Excel, HTML and scanned images into clean, LLM-ready Markdown. Keeps headings, tables and two-column reading order, OCRs scanned pages, and outputs heading-aware chunks for RAG. Pay only for documents that convert.

Pricing

from $2.00 / 1,000 document converteds

Rating

0.0

(0)

Developer

Qiwei He

Qiwei He

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

PDF & Document to Markdown for LLMs

Turn PDFs, Word, PowerPoint, Excel, HTML and scanned images into clean Markdown that LLMs and RAG pipelines can use. Headings, tables and lists are kept, multi-column layouts (papers, government documents, newsletters) are read in the right order, and scanned pages are OCR'd automatically. You can also get heading-aware chunks ready for embeddings.

You pay per document and per page. Documents that fail or contain no text are free, and OCR is billed only for pages where it finds text.

What makes the output clean

Problem with naive PDF text extractionWhat this Actor does
Multi-column pages come out as interleaved half-linesDetects 2–4 columns and reads each column in order, keeping full-width titles, mastheads and headings in their place
Tables turn into a soup of numbersRuled tables become Markdown tables with a header row, one row per record even when the table is only ruled every few rows
Every page repeats "Journal of X · Page 12"Running headers, footers and page numbers are removed
Headings are lost, so chunks have no contextFont sizes and bold lines become #, ##, ### headings
Scanned PDFs return nothingPages without a text layer (or with garbled fonts) are OCR'd with Tesseract
Words split across lines ("exam- ple")Line-end hyphenation is repaired; compounds like "state-of-the-art" are kept
Chunks cut sections in halfChunks start at section boundaries and carry their section path and page numbers

Supported formats

  • PDF: text PDFs, scanned PDFs (OCR), password-protected PDFs (with pdfPassword)
  • Microsoft Office: Word .docx, PowerPoint .pptx, Excel .xlsx and .xls
  • Web and data: HTML, CSV, JSON, Jupyter notebooks, Markdown, plain text, EPUB
  • Images (OCR): PNG, JPG, TIFF (multi-page, up to 200 frames), BMP, WebP, GIF

Share links work as-is: Google Drive files and Google Docs/Slides/Sheets shared as "Anyone with the link" (exported as Office files), Dropbox and GitHub.

How to use it

  1. Paste one or more document URLs, or upload files.
  2. Optionally set a chunk size (for example 500 tokens) if you're loading a vector database.
  3. Run. Each document becomes one row in the dataset with its Markdown and metadata.

Input example

{
"urls": [
{ "url": "https://arxiv.org/pdf/1706.03762" },
{ "url": "https://example.com/handbook.docx" }
],
"ocrMode": "auto",
"chunkSizeTokens": 500,
"chunkOverlapTokens": 50
}

Output example

{
"url": "https://example.com/q2-sales-report.pdf",
"status": "success",
"fileType": "pdf",
"title": "Quarterly Sales Report",
"pageCount": 1,
"pagesProcessed": 1,
"ocrPageCount": 0,
"tableCount": 1,
"tokenEstimate": 240,
"markdown": "# Quarterly Sales Report\n\n# Summary\n\nRevenue grew in every region...\n\n| Region | Q1 revenue | Q2 revenue | Growth |\n| --- | --- | --- | --- |\n| Ontario | $1.2M | $1.5M | 25% |",
"chunks": [
{
"chunkIndex": 1,
"text": "### Regional breakdown\n\n| Region | Q1 revenue | ...",
"section": "Summary > Regional breakdown",
"pageStart": 1,
"pageEnd": 1,
"tokenEstimate": 50
}
],
"warnings": []
}

Every input gets a row, and its status is one of:

  • success: converted and charged.
  • failed: not charged. The error explains why in plain English, for example "The PDF is password-protected. Provide the password in the pdfPassword input."
  • no_text: the file had no extractable text. Not charged.
  • skipped: not processed because the run reached the maximum cost you set.

Very large documents keep their Markdown (and chunks) in the run's key-value store, linked from markdownFileUrl / chunksFileUrl, because a dataset row is limited to about 9 MB.

Use it from code, automations or AI agents

HTTP API: one request returns the results:

curl -X POST "https://api.apify.com/v2/acts/YOUR_USERNAME~document-to-markdown/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"urls":[{"url":"https://arxiv.org/pdf/1706.03762"}]}'

Python

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("YOUR_USERNAME/document-to-markdown").call(run_input={
"urls": [{"url": "https://arxiv.org/pdf/1706.03762"}],
"chunkSizeTokens": 500,
})
for doc in client.dataset(run["defaultDatasetId"]).iterate_items():
if doc["status"] == "success":
print(doc["title"], doc["tokenEstimate"])
else:
print(doc["url"], doc["status"], doc.get("error"))

AI agents (MCP): add this Actor to Claude, Cursor or any MCP client through the Apify MCP server (https://mcp.apify.com), and your agent can read any PDF or Office file it finds.

Vector databases: set Chunk output to One dataset item per chunk and connect the dataset to the Pinecone, Qdrant or other vector database integrations.

Zapier, Make, n8n: use the Apify integration and map the markdown field.

Pricing

Pay per event: you only pay for what converts.

EventPrice
Document converted$0.002
Page processed$0.0001
Page OCR (scanned pages only)$0.004

Each run also has Apify's standard start fee of $0.00005 per GB of memory (1 GB by default).

Examples:

DocumentCost
10-page text PDF$0.003
100-page report$0.012
5-page scanned PDF$0.0225
Word document of about 6,000 characters$0.0022

For formats without real pages, one "page" is 3,000 characters of output (PowerPoint counts slides). Empty spreadsheet cells are not counted. Documents that fail or contain no text are free.

Maximum cost per run is always respected. Near the limit, the Actor converts one document at a time so it can stop exactly. Documents it can't afford are listed with status skipped, and a document cut short says how many pages were included.

Benchmark against other PDF Actors in Apify Store (21 September 2026)

Five public PDFs were run through this Actor and five other PDF-to-text Actors from Apify Store, using the same URLs for every tool, and scored against an answer key written from the page images. Each check is an exact text match on the tool's Markdown (or its plain text, if it has no Markdown) after dropping case, accents, spacing and punctuation. Formatting differences cost nothing; only missing, garbled or out-of-order text fails. The French check keeps accents, because that is what it tests.

Test (what it measures)This ActorActor AActor BActor CActor DActor E
25-column statistics table: rows kept intact (5 rows)5/55/55/55/50/50/5
3-column Federal Register: sentences in reading order, including across column and page breaks (8)8/80/80/85/80/83/8
Scanned page: sentences recovered by OCR (8)8/88/85/80/88/80/8
19-page report: sentences (2) and table rows (2)2/2 and 2/22/2 and 2/22/2 and 2/22/2 and 2/22/2 and 2/22/2 and 2/2
French accents preserved (4 sentences)4/44/44/44/43/44/4
Running headers and footers left in the text (46 on these pages; lower is better)04646464622
Cost for all five documents (38 pages)$0.018$0.033$0.045$0.019$0.013$0.120

The typical failures: on the three-column document, most tools mix lines from neighbouring columns into the same sentence. On the statistics table, some tools list every row label first and all the numbers after, or cut rows short. Some can't read scanned pages at all. And every other tool leaves some or all of the running headers and footers in the text, so a sentence that crosses a page break gets a footer line in the middle of it.

Test files (pinned commits, so anyone can repeat the test with any tool): NICS background checks · Federal Register 2020-17221 · scanned LinnSequencer page · National Hydro Network data model · French meeting minutes

The other Actors ran with OCR, tables and header removal switched on wherever they offer those options. Their costs are what each run was charged on 20–21 September 2026, and their prices and results may have changed since. This Actor's cost is at its current price.

Input options

OptionDefaultWhat it does
urls / filesDocuments to convert
ocrModeautoauto OCRs only pages without usable text, force OCRs every page, off never OCRs
ocrLanguagesengTesseract languages, e.g. eng+fra. Installed: eng, fra, deu, spa, ita, por, nld, pol, chi_sim, jpn
pageRangeallPDF pages to convert, e.g. 1-5, 8, 10-
maxPages0 (no limit)Cap pages per PDF
extractTablestrueMarkdown tables from ruled PDF tables
removeHeadersFooterstrueDrop running headers, footers and page numbers
includePageBreaksfalseInsert <!-- page N --> markers for citations
pdfPasswordPassword for protected PDFs
chunkSizeTokens0 (off)Heading-aware chunks of about this many tokens
chunkOverlapTokens50Overlap between chunks in the same section
chunkOutputnestednested in each document, or separateItems (one row per chunk)
saveMarkdownFilesfalseAlso save a downloadable .md file per document
httpHeadersHeaders for private documents, e.g. Authorization
maxFileSizeMb100Skip larger files (up to 200)
maxConcurrency3Documents converted in parallel

FAQ

Does it handle multi-column documents? Yes. It detects the gaps between 2–4 columns and reads each column in order, while titles, abstracts and headings that span the page stay in place. It was tested on academic papers and the three-column Federal Register.

How accurate is OCR? It uses Tesseract 5, which does well on clean scans at typical office resolution. Photos of documents, handwriting and very low-resolution scans will be less accurate.

What if a scan has no readable text? The document is returned with status no_text and isn't charged, and OCR pages without text are never billed.

Are images, charts or equations described? No. The Actor extracts text. Text inside images is read only when a page is OCR'd.

Tables without borders? Tables need ruling lines to become Markdown tables. The text of borderless tables is still extracted, as plain lines.

Is my data stored? Documents are processed in memory during your run. Results are stored in your own Apify dataset under your account's retention settings.

Legacy .doc or .ppt files? Save them as .docx, .pptx or PDF first.

Changelog

  • 1.0.1 (21 September 2026): tables ruled only every few rows now give one Markdown row per record; title pages that open with a licence notice or strapline now keep the real title as the document title. The page price is halved.
  • 1.0: first release, covering PDF (layout-aware, tables, OCR), Office, HTML, images, RAG chunking and share-link support.