Document & PDF to Markdown for LLMs: Word, Excel, OCR avatar

Document & PDF to Markdown for LLMs: Word, Excel, OCR

Pricing

from $3.00 / 1,000 page extracteds

Go to Apify Store
Document & PDF to Markdown for LLMs: Word, Excel, OCR

Document & PDF to Markdown for LLMs: Word, Excel, OCR

Convert PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX), CSV and HTML files into clean, LLM-ready Markdown and JSON. Keeps headings, lists and tables, outputs per-page chunks for RAG and vector databases, and OCRs scanned PDFs and images. Failed files are free. $3 per 1,000 pages.

Pricing

from $3.00 / 1,000 page extracteds

Rating

0.0

(0)

Developer

Huss

Huss

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Share

Document & PDF to Markdown is a PDF to Markdown converter and document parser API for AI. It converts PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX), CSV and HTML files into clean, LLM-ready Markdown and JSON. Give it document URLs and get back structured Markdown with headings, lists and tables preserved, per-page chunks ready for RAG, document metadata, and OCR for scanned PDFs and images.

It's made for AI agents, RAG pipelines, vector databases and anyone who needs document text in a form an LLM can read. You can call it from your own code, Zapier, Make, n8n, LangChain, LlamaIndex or straight from AI agents through the Apify MCP server.

Why use this PDF to Markdown converter?

Document & PDF to MarkdownPDF-only text extractorsSelf-hosted parsers (PyMuPDF, MarkItDown, Docling)
PDF, Word, PowerPoint, Excel, CSV and HTMLAll in one ActorPDF onlyOne library per format
Tables kept as Markdown tablesYesOften flattened to textDepends on the library
OCR for scanned pagesAutomatic, only where neededSometimesExtra setup (Tesseract)
RAG-ready per-page chunks and metadataYesPartlyYou write the code
Servers, memory and scalingHandled by ApifyHandledYour job
Failed or empty documentsNot chargedVariesStill cost compute
Price$3 per 1,000 pages, pay as you goPer file or per pageYour infrastructure

Who uses it?

  • AI and RAG developers loading PDFs, contracts, manuals and reports into vector databases.
  • AI agents that need to read a document link a user shares, via the Apify MCP server.
  • Legal, finance and research teams turning filings, papers and scanned contracts into searchable text.
  • Data teams extracting tables from PDFs and spreadsheets into a consistent format.
  • No-code builders adding document reading to Make, Zapier or n8n workflows.

What can this document converter do?

  • ๐Ÿ“„ PDF to Markdown with headings, paragraphs, bullet and numbered lists, and tables as Markdown tables, including multi-column layouts read in the right order.
  • ๐Ÿ”Ž OCR for scanned PDFs and images (PNG, JPEG, TIFF). In Auto mode OCR runs only on pages that have no usable text layer, so you pay the OCR price only where it's needed. English, German, French, Spanish, Portuguese, Italian and Dutch.
  • ๐Ÿ“ Word to Markdown: headings, lists, bold and italic text, links and tables from DOCX files.
  • ๐Ÿ“Š PowerPoint to Markdown: one section per slide with the slide title, text, tables, chart data and speaker notes.
  • ๐Ÿ“ˆ Excel and CSV to Markdown tables, one heading per sheet, split into chunks that repeat the header row.
  • ๐ŸŒ HTML to Markdown: keeps the main article content and drops menus, sidebars, language lists and footers (or converts the whole page if you prefer).
  • โœ‚๏ธ Page ranges like 1-5, 8, 20- and a per-document page cap, so you only pay for what you need.
  • ๐Ÿงฉ RAG-ready chunks: every document comes with a pages array, or switch to one dataset item per page.
  • ๐Ÿท๏ธ Metadata: title, author, dates, page count, word count, table count, file size and a SHA-256 hash for de-duplication.
  • ๐Ÿงน Removes repeating headers, footers and page numbers from PDFs.
  • ๐Ÿ”— Google Drive, Google Docs and Dropbox share links are turned into download links automatically.
  • ๐Ÿšฆ Failed documents are free. Broken links, unsupported files and empty documents are reported with a clear reason and never charged.

Supported formats

FormatExtensionsWhat counts as one page
PDF (text).pdfOne PDF page
PDF (scanned) and images.pdf, .png, .jpg, .tiff, .webp, .bmp, .gifOne OCR page per page or image frame
PowerPoint.pptxOne slide
Word.docxOne section of up to 3,000 characters of output
Excel and CSV.xlsx, .csv, .tsvOne section of up to 3,000 characters of output
HTML, text and Markdown.html, .txt, .mdOne section of up to 3,000 characters of output

Legacy .doc, .ppt and .xls files and OpenDocument files aren't supported. Save them as DOCX, PPTX, XLSX or PDF first.

How to convert documents to Markdown

  1. Click Try for free.
  2. Add your Document URLs: type them in, paste a list, or upload a text file with one link per line.
  3. Optionally set a Page range, the OCR mode, or One item per page output.
  4. Click Start. Each document becomes a dataset item with its Markdown, metadata and per-page chunks.
  5. Download the results as JSON, CSV or Excel, or read them through the API.

Input example

{
"startUrls": [
{ "url": "https://arxiv.org/pdf/1706.03762" },
{ "url": "https://example.com/files/quarterly-report.docx" }
],
"pageRange": "",
"maxPagesPerDocument": 1000,
"ocrMode": "auto",
"ocrLanguages": ["eng"],
"outputMode": "document",
"includePages": true
}

Output example

One item per document (default). Fields are the same for every format, so your pipeline can rely on a stable schema:

{
"url": "https://arxiv.org/pdf/1706.03762",
"status": "success",
"error": null,
"fileName": "1706.03762v7.pdf",
"format": "pdf",
"contentType": "application/pdf",
"fileSizeBytes": 2215244,
"sha256": "bdfaa68d8984f0dc02beaca527b76f207d99b666d31d1da728ee0728182df697",
"title": "Attention Is All You Need",
"author": null,
"subject": null,
"createdAt": "2024-04-10T21:11:43",
"modifiedAt": "2024-04-10T21:11:43",
"pageUnit": "page",
"pageCount": 15,
"pagesExtracted": 15,
"ocrPages": 0,
"tableCount": 8,
"wordCount": 6358,
"charCount": 40032,
"truncated": false,
"warnings": [],
"markdownFileUrl": null,
"markdown": "## Provided proper attribution is provided, Google hereby grants permission ...\n\n# Attention Is All You Need\n\n| Ashish Vaswani โˆ— | Noam Shazeer โˆ— | Niki Parmar โˆ— | Jakob Uszkoreit โˆ— |\n| --- | --- | --- | --- |\n| Google Brain | Google Brain | Google Research | Google Research |\n\n...\n\n## Abstract\n\nThe dominant sequence transduction models are based on ...",
"pages": [
{ "pageNumber": 1, "markdown": "## Provided proper attribution ...", "wordCount": 420, "isOcr": false, "tableCount": 2 }
],
"chargedPages": 15,
"chargedOcrPages": 0,
"extractedAt": "2026-10-08T15:44:47Z"
}

With Output mode: One item per page, each page is its own dataset item with pageNumber, isOcr, pageWordCount and that page's markdown, plus the document fields. This is the easiest way to load chunks into Pinecone, Qdrant, Weaviate, pgvector or any other vector store.

Documents that fail have "status": "error" and an error such as Download failed with HTTP 404. or Legacy Excel .xls files are not supported. A RUN_SUMMARY record in the key-value store lists all failures and charged pages.

How much does it cost to convert PDF to Markdown?

This Actor uses pay-per-event pricing. You pay only for pages that were actually extracted:

EventPrice
Page extracted (text PDF page, slide, or 3,000-character section)$3.00 per 1,000 pages ($0.003 each)
OCR page (scanned PDF page or image)$8.00 per 1,000 OCR pages ($0.008 each)
Actor start$0.00005 per run

Examples:

  • A 15-page research paper costs $0.045.
  • 1,000 text PDF pages cost $3.
  • A 20-slide PowerPoint costs $0.06.
  • A 10-page Word document (about 30,000 characters) costs about $0.03.
  • A 50-page scanned contract costs $0.40.

OCR pages are charged only at the OCR price, not both. Blank pages, failed downloads, unsupported files and documents with no extractable text are free. Use Page range and Max pages per document to control costs, and set Maximum cost per run in the run options. The Actor stops cleanly when the limit is reached.

Tips

  • Large documents: the default 1 GB of memory handles most files. For PDFs with hundreds of pages or long OCR jobs, use 2โ€“4 GB. It runs faster, and the price per page stays the same.
  • Scanned PDFs: keep OCR mode on Auto. Use Always if a PDF has a broken or garbled text layer, and Never if you only want embedded text.
  • Fewer tokens: turn off Keep links and Include per-page chunks if you only need the full Markdown.
  • Spreadsheets are converted to Markdown tables with the first row as the header. Huge sheets produce many sections, so use Max pages per document to cap them.
  • Files in your own storage: any URL that downloads the file works, including signed S3 or GCS URLs and Apify key-value store record URLs.

Use it with AI agents (MCP) and the API

Add this Actor as a tool in Claude, ChatGPT, Cursor, VS Code or any other MCP client through the Apify MCP server:

https://mcp.apify.com?tools=magenta_waterwheel/document-to-markdown

Then just ask, for example: "Convert this PDF to Markdown and summarize the tables: https://arxiv.org/pdf/1706.03762" The agent fills in the input, runs the Actor and reads the results. The Actor runs with limited permissions, so it can only access its own run storage.

To get results in a single HTTP request, call the synchronous endpoint with your Apify API token:

curl -X POST "https://api.apify.com/v2/acts/magenta_waterwheel~document-to-markdown/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://arxiv.org/pdf/1706.03762"}],"pageRange":"1-3"}'

You can also use the official Python and JavaScript clients, or the ready-made code on the API tab.

LangChain, LlamaIndex and vector databases

Use the Apify dataset loader in LangChain or LlamaIndex and map markdown to the document text and the other fields to metadata. With Output mode: One item per page, each dataset item is a ready-made chunk for Pinecone, Qdrant, Weaviate, pgvector or Chroma.

Integrations and scheduling

  • โฐ Schedule runs hourly, daily or weekly in Apify Console, with no server to maintain.
  • ๐Ÿ”— Send results anywhere: Google Sheets, Slack, Google Drive, Airbyte, webhooks, or no-code tools such as Make, Zapier and n8n.
  • ๐Ÿ“ค Export the dataset as JSON, CSV, Excel, XML, RSS or HTML.
  • ๐Ÿ“ˆ Monitoring: get notified if a run fails, and see every run's log and cost in Console.

FAQ

Is there a free trial?

Yes. Apify's Free plan includes $5 of usage credit every month, enough for about 1,600 text pages (or about 600 OCR pages) with this Actor. No credit card is needed to start.

How do I convert a PDF to Markdown?

Paste the PDF's URL into Document URLs and click Start. The markdown field holds the whole document, and pages holds one chunk per page. To convert a file from your computer, upload it to any storage that gives you a download link (Google Drive, Dropbox, S3) and paste that link.

Is the PDF table extraction accurate?

Tables with ruling lines are extracted cell by cell. Borderless tables are rebuilt from column alignment, which works well for regular tables but can merge cells in very irregular layouts. Word, PowerPoint, Excel and HTML tables are always exact.

How good is the OCR?

OCR uses Tesseract, the most widely used open-source OCR engine, at 200 DPI. It works well on clean scans and printed text. Handwriting, very low-resolution scans and complex forms give weaker results.

Can it read password-protected or DRM-locked PDFs?

No. Encrypted PDFs that need a password are reported as failed and are not charged.

Does it crawl websites?

No. This Actor converts the documents you link to. To crawl a website, use a web crawler and pass the document links it finds to this Actor.

Is there a file size limit?

The default limit is 50 MB per file and you can raise it to 200 MB. Very large files may need more memory.

Do you store my documents?

Documents are processed in your own Actor run. The output stays in your run's storage under your Apify data retention settings, and nothing is kept elsewhere.

I found a bug or need a feature

Open an issue in the Issues tab with a link to your run. We aim to reply within one business day.

More tools from the same developer

ActorWhat it doesPrice
Career Site Jobs APIJobs from Greenhouse, Lever, Ashby and SmartRecruiters career sites, with salaries$2 / 1,000 jobs
Website Screenshot & PDF APIFull-page PNG/JPEG screenshots and web page to PDF in bulk$2.50 / 1,000 screenshots
App Store Reviews ScraperApple App Store reviews and app details in any country$0.25 / 1,000 reviews