Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML) avatar

Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML)

Pricing

Pay per event

Go to Apify Store
Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML)

Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML)

Convert PDF, Word, Excel, PowerPoint, HTML and scanned documents to clean Markdown with OCR, metadata and RAG ready chunks for LLM ingestion and AI agents.

Pricing

Pay per event

Rating

0.0

(0)

Developer

Rod Services

Rod Services

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Document to Markdown for LLMs do?

Document to Markdown for LLMs converts PDF, Word (DOCX), Excel (XLSX, XLS), PowerPoint (PPTX), HTML, EPUB, CSV and scanned images into clean Markdown that large language models read well. It keeps headings, lists and tables, removes running headers and page numbers, and runs OCR on scanned pages automatically.

Every document also comes back as RAG ready chunks with page numbers and section headings. Feed them straight into a vector database, an embeddings pipeline or an AI agent.

Paste a file URL and click Start. The example PDF converts in about 5 seconds. As an Apify Actor you also get an API, scheduling, webhooks, integrations with Make, Zapier, n8n and LangChain, and run monitoring.

Why use this PDF to Markdown converter?

  • LLM ingestion and RAG. Turn document libraries into chunks for pgvector, Pinecone, Qdrant, Weaviate or Chroma.
  • AI agents and MCP tools. Give an agent one call that reads any PDF, DOCX or web page as Markdown.
  • Scanned PDFs and images. Tesseract OCR in 35 languages runs only on pages that have no text layer.
  • Tables survive. PDF tables, Word tables and Excel sheets become Markdown tables.
  • Two column PDFs. Columns are read in order, not line by line across the page.
  • Metadata. Title, author, dates, page count, word count and HTML meta tags.
  • Fair pricing. You pay per converted document and per OCR page. Failed files are free.

How to convert PDF to Markdown

  1. Open the Input tab.
  2. Add one or more links in File URLs, or upload a file with Upload a file.
  3. Pick an OCR mode. Keep auto unless you know the files are scans.
  4. Set Chunk size for your embedding model. 500 to 1,000 tokens works for most.
  5. Click Start. Results appear in the Output tab.
  6. Download JSON, CSV or Excel, or call the API from your pipeline.

Google Drive, Google Docs, Sheets and Slides, Dropbox and GitHub share links are turned into direct downloads for you. The file must be shared with "Anyone with the link". Google Docs are exported as DOCX, Sheets as XLSX and Slides as PPTX.

Input

All fields are on the Input tab. Example:

{
"fileUrls": [
"https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf",
"https://example.com/report.docx"
],
"ocr": "auto",
"languages": ["eng", "deu"],
"chunkSize": 800,
"chunkOverlap": 80,
"includeImages": false,
"maxPages": 0
}
FieldDefaultWhat it does
fileUrlsDirect links to documents. Apify key-value store record URLs work too.
uploadedFileOne file uploaded from your computer in Console.
keyValueStoreRecords[]Record keys in this run's key-value store, or storeId/key. Useful when another Actor calls this one.
ocrautoauto OCRs only scanned pages. force OCRs every page. off never OCRs.
languages["eng"]Tesseract codes, for example eng, deu, fra, spa, lit, pol, rus, jpn, chi_sim.
chunkSize1000Approximate tokens per chunk. 0 turns chunking off.
chunkOverlap100Approximate tokens repeated between chunks.
includeImagesfalseSave embedded images to the key-value store and link them in the Markdown.
maxPages0Only the first N pages or slides. 0 is all, but text PDFs stop at 300 pages.
saveMarkdownFilesfalseAlso save each document as a .md file.
maxFileSizeMb100Skip bigger downloads.
maxConcurrency2Documents converted in parallel.

Output

One dataset item per document. Shortened example:

{
"sourceUrl": "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf",
"fileName": "USCODE-2011-title17-chap1-sec107.pdf",
"mimeType": "application/pdf",
"pages": 5,
"title": "§107. Limitations on exclusive rights: Fair use",
"markdown": "by two or more authors, a waiver of rights under this paragraph made by one such author waives such rights for all such authors.\n\n(2) Ownership of the rights conferred by subsection (a)...",
"chunks": [
{
"index": 0,
"text": "by two or more authors, a waiver of rights under this paragraph...",
"tokensApprox": 845,
"page": 1,
"heading": null
}
],
"chunksCount": 13,
"metadata": {
"creator": "Federal Digital System, U. S. Government Publishing Office",
"createdAt": "2019-10-14T08:51:38+00:00",
"pdfVersion": "1.5"
},
"ocrPagesCount": 0,
"warnings": [],
"error": null
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: Overview, Markdown and RAG chunks with one row per chunk.

Documents longer than 200,000 characters keep a truncated copy in the dataset. The full text is saved as a .md file in the key-value store, linked in markdownUrl.

Data fields

FieldDescription
sourceUrlOriginal link, or the key-value store record URL
fileNameFile name from the URL or the download headers
mimeTypeDetected type, checked against the file content
pagesPages for PDF, slides for PPTX, frames for images
titleDocument title from metadata or the first heading
markdownClean Markdown of the whole document
chunksindex, text, tokensApprox, page, heading
metadataAuthor, subject, keywords, created and modified dates, word count, HTML meta
ocrPagesCountPages that went through OCR
imagesLinks to extracted images when includeImages is on
markdownUrlFull .md file for long documents
warningsNotes such as pages skipped by maxPages
errorWhy a document failed. Failed documents are not charged

How much does it cost to convert PDF to Markdown?

The Actor uses pay per event pricing:

  • $0.003 per converted document
  • $0.01 per OCR page, only when a page is really scanned or when you force OCR
  • $0.001 per run start (per GB of memory)

1,000 normal PDFs cost about $3. A 100 page scanned book costs about $1. Failed downloads, broken files and password protected PDFs cost nothing. Apify's free plan credit covers hundreds of documents each month.

Set Maximum cost per run in the run options. The Actor never goes over it. Documents that do not fit are skipped, and scanned pages that do not fit are left out with a warning on the document.

Tips for better results

  • Keep OCR on auto. Text PDFs are fast and cheap. OCR costs time and money.
  • List only the OCR languages you need. Each extra language slows OCR down.
  • Use maxPages to preview long files before a full run.
  • Text PDFs are fast: about 30 pages per second at 1 GB, table heavy reports about 10.
  • For big OCR jobs, give the run more memory. 1 GB is the cheapest per page. 4 GB is about 2.7 times faster.
  • Tune chunk size to your embedding model. Chunks break at headings and paragraphs, and table chunks repeat the table header.
  • Call it from LangChain, LlamaIndex or n8n through the Apify integration, then embed chunks[].text.

FAQ, limits and support

Which formats are supported? PDF, DOCX, XLSX, XLS, PPTX, HTML, EPUB, CSV, JSON, XML, TXT, Markdown, Outlook MSG, PNG, JPG, TIFF, WEBP, BMP and GIF. Legacy DOC and PPT are not supported. Save them as DOCX or PPTX first.

Does it work with scanned PDFs? Yes. Pages without a text layer are rendered and read with Tesseract OCR. Handwriting and very low quality scans give weak results.

Are password protected PDFs supported? No. They return an error and are not charged. PDFs that only restrict printing or copying open normally.

What if a link opens a login or error page? The document fails with a clear error and is not charged. This happens with private Google Drive or Dropbox files and expired links.

Is there a page limit? Text PDFs over 300 pages stop at page 300 when maxPages is 0. The document gets a warning. To convert more, set maxPages to the page count you need, for example 1000. You can also split the file. Scanned PDFs have no such limit, since you pay for each OCR page.

How accurate are PDF headings and tables? Headings come from font sizes, tables from ruled lines. Complex forms and borderless tables can come out as plain text.

Is my data safe? Files are processed inside your own Apify run. Nothing is sent to third party AI services.

Is it legal? You must have the right to process the documents you submit.

Found a bug or need a feature? Open an issue in the Issues tab. Custom pipelines, other formats and private deployments are available on request.