Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML)
Pricing
Pay per event
Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML)
Convert PDF, Word, Excel, PowerPoint, HTML and scanned documents to clean Markdown with OCR, metadata and RAG ready chunks for LLM ingestion and AI agents.
Pricing
Pay per event
Rating
0.0
(0)
Developer
Rod Services
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does Document to Markdown for LLMs do?
Document to Markdown for LLMs converts PDF, Word (DOCX), Excel (XLSX, XLS), PowerPoint (PPTX), HTML, EPUB, CSV and scanned images into clean Markdown that large language models read well. It keeps headings, lists and tables, removes running headers and page numbers, and runs OCR on scanned pages automatically.
Every document also comes back as RAG ready chunks with page numbers and section headings. Feed them straight into a vector database, an embeddings pipeline or an AI agent.
Paste a file URL and click Start. The example PDF converts in about 5 seconds. As an Apify Actor you also get an API, scheduling, webhooks, integrations with Make, Zapier, n8n and LangChain, and run monitoring.
Why use this PDF to Markdown converter?
- LLM ingestion and RAG. Turn document libraries into chunks for pgvector, Pinecone, Qdrant, Weaviate or Chroma.
- AI agents and MCP tools. Give an agent one call that reads any PDF, DOCX or web page as Markdown.
- Scanned PDFs and images. Tesseract OCR in 35 languages runs only on pages that have no text layer.
- Tables survive. PDF tables, Word tables and Excel sheets become Markdown tables.
- Two column PDFs. Columns are read in order, not line by line across the page.
- Metadata. Title, author, dates, page count, word count and HTML meta tags.
- Fair pricing. You pay per converted document and per OCR page. Failed files are free.
How to convert PDF to Markdown
- Open the Input tab.
- Add one or more links in File URLs, or upload a file with Upload a file.
- Pick an OCR mode. Keep
autounless you know the files are scans. - Set Chunk size for your embedding model. 500 to 1,000 tokens works for most.
- Click Start. Results appear in the Output tab.
- Download JSON, CSV or Excel, or call the API from your pipeline.
Google Drive, Google Docs, Sheets and Slides, Dropbox and GitHub share links are turned into direct downloads for you. The file must be shared with "Anyone with the link". Google Docs are exported as DOCX, Sheets as XLSX and Slides as PPTX.
Input
All fields are on the Input tab. Example:
{"fileUrls": ["https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf","https://example.com/report.docx"],"ocr": "auto","languages": ["eng", "deu"],"chunkSize": 800,"chunkOverlap": 80,"includeImages": false,"maxPages": 0}
| Field | Default | What it does |
|---|---|---|
fileUrls | Direct links to documents. Apify key-value store record URLs work too. | |
uploadedFile | One file uploaded from your computer in Console. | |
keyValueStoreRecords | [] | Record keys in this run's key-value store, or storeId/key. Useful when another Actor calls this one. |
ocr | auto | auto OCRs only scanned pages. force OCRs every page. off never OCRs. |
languages | ["eng"] | Tesseract codes, for example eng, deu, fra, spa, lit, pol, rus, jpn, chi_sim. |
chunkSize | 1000 | Approximate tokens per chunk. 0 turns chunking off. |
chunkOverlap | 100 | Approximate tokens repeated between chunks. |
includeImages | false | Save embedded images to the key-value store and link them in the Markdown. |
maxPages | 0 | Only the first N pages or slides. 0 is all, but text PDFs stop at 300 pages. |
saveMarkdownFiles | false | Also save each document as a .md file. |
maxFileSizeMb | 100 | Skip bigger downloads. |
maxConcurrency | 2 | Documents converted in parallel. |
Output
One dataset item per document. Shortened example:
{"sourceUrl": "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf","fileName": "USCODE-2011-title17-chap1-sec107.pdf","mimeType": "application/pdf","pages": 5,"title": "§107. Limitations on exclusive rights: Fair use","markdown": "by two or more authors, a waiver of rights under this paragraph made by one such author waives such rights for all such authors.\n\n(2) Ownership of the rights conferred by subsection (a)...","chunks": [{"index": 0,"text": "by two or more authors, a waiver of rights under this paragraph...","tokensApprox": 845,"page": 1,"heading": null}],"chunksCount": 13,"metadata": {"creator": "Federal Digital System, U. S. Government Publishing Office","createdAt": "2019-10-14T08:51:38+00:00","pdfVersion": "1.5"},"ocrPagesCount": 0,"warnings": [],"error": null}
You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: Overview, Markdown and RAG chunks with one row per chunk.
Documents longer than 200,000 characters keep a truncated copy in the dataset. The full text is saved as a .md file in the key-value store, linked in markdownUrl.
Data fields
| Field | Description |
|---|---|
sourceUrl | Original link, or the key-value store record URL |
fileName | File name from the URL or the download headers |
mimeType | Detected type, checked against the file content |
pages | Pages for PDF, slides for PPTX, frames for images |
title | Document title from metadata or the first heading |
markdown | Clean Markdown of the whole document |
chunks | index, text, tokensApprox, page, heading |
metadata | Author, subject, keywords, created and modified dates, word count, HTML meta |
ocrPagesCount | Pages that went through OCR |
images | Links to extracted images when includeImages is on |
markdownUrl | Full .md file for long documents |
warnings | Notes such as pages skipped by maxPages |
error | Why a document failed. Failed documents are not charged |
How much does it cost to convert PDF to Markdown?
The Actor uses pay per event pricing:
- $0.003 per converted document
- $0.01 per OCR page, only when a page is really scanned or when you force OCR
- $0.001 per run start (per GB of memory)
1,000 normal PDFs cost about $3. A 100 page scanned book costs about $1. Failed downloads, broken files and password protected PDFs cost nothing. Apify's free plan credit covers hundreds of documents each month.
Set Maximum cost per run in the run options. The Actor never goes over it. Documents that do not fit are skipped, and scanned pages that do not fit are left out with a warning on the document.
Tips for better results
- Keep OCR on
auto. Text PDFs are fast and cheap. OCR costs time and money. - List only the OCR languages you need. Each extra language slows OCR down.
- Use
maxPagesto preview long files before a full run. - Text PDFs are fast: about 30 pages per second at 1 GB, table heavy reports about 10.
- For big OCR jobs, give the run more memory. 1 GB is the cheapest per page. 4 GB is about 2.7 times faster.
- Tune chunk size to your embedding model. Chunks break at headings and paragraphs, and table chunks repeat the table header.
- Call it from LangChain, LlamaIndex or n8n through the Apify integration, then embed
chunks[].text.
FAQ, limits and support
Which formats are supported? PDF, DOCX, XLSX, XLS, PPTX, HTML, EPUB, CSV, JSON, XML, TXT, Markdown, Outlook MSG, PNG, JPG, TIFF, WEBP, BMP and GIF. Legacy DOC and PPT are not supported. Save them as DOCX or PPTX first.
Does it work with scanned PDFs? Yes. Pages without a text layer are rendered and read with Tesseract OCR. Handwriting and very low quality scans give weak results.
Are password protected PDFs supported? No. They return an error and are not charged. PDFs that only restrict printing or copying open normally.
What if a link opens a login or error page? The document fails with a clear error and is not charged. This happens with private Google Drive or Dropbox files and expired links.
Is there a page limit? Text PDFs over 300 pages stop at page 300 when maxPages is 0. The document gets a warning. To convert more, set maxPages to the page count you need, for example 1000. You can also split the file. Scanned PDFs have no such limit, since you pay for each OCR page.
How accurate are PDF headings and tables? Headings come from font sizes, tables from ruled lines. Complex forms and borderless tables can come out as plain text.
Is my data safe? Files are processed inside your own Apify run. Nothing is sent to third party AI services.
Is it legal? You must have the right to process the documents you submit.
Found a bug or need a feature? Open an issue in the Issues tab. Custom pipelines, other formats and private deployments are available on request.