Document to Markdown - PDF, Word, Excel, PowerPoint to Text avatar

Document to Markdown - PDF, Word, Excel, PowerPoint to Text

Pricing

from $4.00 / 1,000 documents

Go to Apify Store
Document to Markdown - PDF, Word, Excel, PowerPoint to Text

Document to Markdown - PDF, Word, Excel, PowerPoint to Text

Convert PDF, DOCX, XLSX, PPTX, HTML and CSV files from URLs into clean Markdown or plain text for AI, RAG and LLM pipelines. Tables kept, pages split, pay per document.

Pricing

from $4.00 / 1,000 documents

Rating

0.0

(0)

Developer

Björn Ólafur

Björn Ólafur

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

What does Document to Markdown do?

Document to Markdown converts PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML, CSV and text files from any URL into clean Markdown or plain text. Headings, lists and tables are kept, documents can be split into pages, slides or sheets, and you pay only for documents that convert successfully.

It's built for AI agents, RAG pipelines and LLM workflows: give it links, get back text that's ready to chunk, embed or summarise. Run it in Apify Console, call it from the API, schedule it, or use it from your AI agent through the Apify MCP server. Connect it to Make, Zapier, n8n, LangChain or LlamaIndex with Apify integrations.

Why use Document to Markdown?

  • 📄 One tool for every office format: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT, Markdown and JSON.
  • 🔎 OCR built in: scanned PDFs and images (PNG, JPEG, WebP, TIFF) become searchable text, in 100+ languages. In auto mode OCR runs only on pages that need it.
  • 🧠 LLM-ready Markdown: headings, bullet lists and GitHub-style tables, so models understand the structure.
  • ✂️ Page-level output: optional pages array (one per PDF page, slide or sheet) for citations and chunking.
  • 🔍 Automatic file-type detection from the file content, so links without an extension (download links, ?id= links) work.
  • 🗒️ Speaker notes from PowerPoint slides and all sheets from Excel workbooks.
  • 💸 Fair pricing: broken links, oversized files and unsupported formats are free.
  • ⚡ Fast and cheap: no browser, 10 documents in parallel by default.

Typical uses: feeding reports, contracts, manuals and research papers to ChatGPT/Claude, building RAG knowledge bases, extracting tables from Excel files on websites, indexing slide decks, and monitoring published PDFs such as price lists, tenders or annual reports.

How to convert PDF, Word, Excel or PowerPoint to Markdown

  1. Click Try for free.
  2. Paste document links into Document URLs, one per line.
  3. Choose Markdown or Plain text.
  4. Click Start. Results appear in the Output tab within seconds.
  5. Download them as JSON, CSV or Excel, or fetch them via the API.

Input

Only Document URLs is required. See the Input tab for all options:

OptionWhat it doesDefault
urlsLinks to documents–
outputFormatmarkdown or textmarkdown
includePagesAdd a pages array (per page, slide or sheet)false
ocrauto (scanned pages and images), always or offauto
ocrLanguagese.g. eng, deu, isl, eng+fraeng
maxOcrPagesOCR page limit per document50
maxPagesPage/slide limit per document1000
maxRowsPerSheetRow limit per Excel sheet or CSV5000
maxFileSizeMbSkip larger files50 MB
{
"urls": ["https://arxiv.org/pdf/1706.03762", "https://calibre-ebook.com/downloads/demos/demo.docx"],
"outputFormat": "markdown",
"includePages": true
}

Output

One item per document:

{
"url": "https://arxiv.org/pdf/1706.03762",
"success": true,
"fileName": "1706.03762",
"fileType": "pdf",
"title": null,
"pageCount": 15,
"wordCount": 6477,
"characterCount": 39790,
"format": "markdown",
"content": "## Page 1\n\nAttention Is All You Need\n\nAshish Vaswani ...",
"contentTruncated": false,
"fullContentUrl": null,
"pages": ["Attention Is All You Need ...", "..."],
"metadata": { "creator": "LaTeX with hyperref", "createdAt": "2024-04-10" },
"fileSizeBytes": 2215244,
"convertedAt": "2026-09-30T00:40:00.000Z",
"error": null
}

Failed documents are listed too, with success: false and the reason in error, and are not charged. Very long documents (over 2 million characters) are cut in the dataset item, and the full text is saved as a file linked in fullContentUrl. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

How much does it cost to convert documents?

$4 per 1,000 documents ($0.004 each) plus $0.002 per run, whatever the size or format. OCR costs $0.01 per recognised page, charged only for scanned pages and images that actually need it (the ocrPages field shows how many). Failed documents are free and platform usage is included. With Apify's free plan you can convert a few hundred documents a month at no cost. You can set a maximum cost per run; the Actor stops cleanly when it's reached.

Tips

  • Use Split into pages when you need page numbers for citations or want to chunk by page.
  • Use Plain text for search indexing and Markdown for LLMs.
  • Scanned PDFs (images of text) have no text layer. With OCR on auto they are recognised automatically; with OCR off they are marked metadata.scannedOrImageOnly: true.
  • Set OCR languages to the document's language for best accuracy (metadata.ocrConfidence shows 0–100).
  • Use Max OCR pages to cap cost on long scanned books.
  • Old binary formats (.doc, .xls, .ppt) aren't supported. Save them as .docx, .xlsx or .pptx first.
  • If a server blocks data-centre downloads, enable Proxy.

FAQ

Does it work with files behind a login? No, the links must be publicly downloadable.

Can I upload files instead of links? Put the file anywhere with a public link (for example an Apify key-value store, S3 or Google Drive direct-download link) and pass that link.

Is it legal? Converting documents you're allowed to access is fine. You're responsible for respecting copyright and the source's terms.

Something not working? Open an issue in the Issues tab with the document link and I'll fix it quickly. Custom versions are available on request.

  • Website Screenshot: full-page screenshots and PDFs of any website, cookie banners removed
  • Sitemap Extractor: every URL of a website from its sitemaps, plus a broken-link check
  • RSS Feed Reader: read and monitor any RSS/Atom feed, only new items, full article text