Document to Markdown - PDF, Word, Excel, PowerPoint to Text
Pricing
from $4.00 / 1,000 documents
Document to Markdown - PDF, Word, Excel, PowerPoint to Text
Convert PDF, DOCX, XLSX, PPTX, HTML and CSV files from URLs into clean Markdown or plain text for AI, RAG and LLM pipelines. Tables kept, pages split, pay per document.
Pricing
from $4.00 / 1,000 documents
Rating
0.0
(0)
Developer
Björn Ólafur
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
What does Document to Markdown do?
Document to Markdown converts PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML, CSV and text files from any URL into clean Markdown or plain text. Headings, lists and tables are kept, documents can be split into pages, slides or sheets, and you pay only for documents that convert successfully.
It's built for AI agents, RAG pipelines and LLM workflows: give it links, get back text that's ready to chunk, embed or summarise. Run it in Apify Console, call it from the API, schedule it, or use it from your AI agent through the Apify MCP server. Connect it to Make, Zapier, n8n, LangChain or LlamaIndex with Apify integrations.
Why use Document to Markdown?
- 📄 One tool for every office format: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT, Markdown and JSON.
- 🔎 OCR built in: scanned PDFs and images (PNG, JPEG, WebP, TIFF) become searchable text, in 100+ languages. In auto mode OCR runs only on pages that need it.
- 🧠 LLM-ready Markdown: headings, bullet lists and GitHub-style tables, so models understand the structure.
- ✂️ Page-level output: optional
pagesarray (one per PDF page, slide or sheet) for citations and chunking. - 🔍 Automatic file-type detection from the file content, so links without an extension (download links,
?id=links) work. - 🗒️ Speaker notes from PowerPoint slides and all sheets from Excel workbooks.
- 💸 Fair pricing: broken links, oversized files and unsupported formats are free.
- ⚡ Fast and cheap: no browser, 10 documents in parallel by default.
Typical uses: feeding reports, contracts, manuals and research papers to ChatGPT/Claude, building RAG knowledge bases, extracting tables from Excel files on websites, indexing slide decks, and monitoring published PDFs such as price lists, tenders or annual reports.
How to convert PDF, Word, Excel or PowerPoint to Markdown
- Click Try for free.
- Paste document links into Document URLs, one per line.
- Choose Markdown or Plain text.
- Click Start. Results appear in the Output tab within seconds.
- Download them as JSON, CSV or Excel, or fetch them via the API.
Input
Only Document URLs is required. See the Input tab for all options:
| Option | What it does | Default |
|---|---|---|
urls | Links to documents | – |
outputFormat | markdown or text | markdown |
includePages | Add a pages array (per page, slide or sheet) | false |
ocr | auto (scanned pages and images), always or off | auto |
ocrLanguages | e.g. eng, deu, isl, eng+fra | eng |
maxOcrPages | OCR page limit per document | 50 |
maxPages | Page/slide limit per document | 1000 |
maxRowsPerSheet | Row limit per Excel sheet or CSV | 5000 |
maxFileSizeMb | Skip larger files | 50 MB |
{"urls": ["https://arxiv.org/pdf/1706.03762", "https://calibre-ebook.com/downloads/demos/demo.docx"],"outputFormat": "markdown","includePages": true}
Output
One item per document:
{"url": "https://arxiv.org/pdf/1706.03762","success": true,"fileName": "1706.03762","fileType": "pdf","title": null,"pageCount": 15,"wordCount": 6477,"characterCount": 39790,"format": "markdown","content": "## Page 1\n\nAttention Is All You Need\n\nAshish Vaswani ...","contentTruncated": false,"fullContentUrl": null,"pages": ["Attention Is All You Need ...", "..."],"metadata": { "creator": "LaTeX with hyperref", "createdAt": "2024-04-10" },"fileSizeBytes": 2215244,"convertedAt": "2026-09-30T00:40:00.000Z","error": null}
Failed documents are listed too, with success: false and the reason in error, and are not charged. Very long documents (over 2 million characters) are cut in the dataset item, and the full text is saved as a file linked in fullContentUrl. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.
How much does it cost to convert documents?
$4 per 1,000 documents ($0.004 each) plus $0.002 per run, whatever the size or format. OCR costs $0.01 per recognised page, charged only for scanned pages and images that actually need it (the ocrPages field shows how many). Failed documents are free and platform usage is included. With Apify's free plan you can convert a few hundred documents a month at no cost. You can set a maximum cost per run; the Actor stops cleanly when it's reached.
Tips
- Use Split into pages when you need page numbers for citations or want to chunk by page.
- Use Plain text for search indexing and Markdown for LLMs.
- Scanned PDFs (images of text) have no text layer. With OCR on
autothey are recognised automatically; with OCRoffthey are markedmetadata.scannedOrImageOnly: true. - Set OCR languages to the document's language for best accuracy (
metadata.ocrConfidenceshows 0–100). - Use Max OCR pages to cap cost on long scanned books.
- Old binary formats (.doc, .xls, .ppt) aren't supported. Save them as .docx, .xlsx or .pptx first.
- If a server blocks data-centre downloads, enable Proxy.
FAQ
Does it work with files behind a login? No, the links must be publicly downloadable.
Can I upload files instead of links? Put the file anywhere with a public link (for example an Apify key-value store, S3 or Google Drive direct-download link) and pass that link.
Is it legal? Converting documents you're allowed to access is fine. You're responsible for respecting copyright and the source's terms.
Something not working? Open an issue in the Issues tab with the document link and I'll fix it quickly. Custom versions are available on request.
Related tools
- Website Screenshot: full-page screenshots and PDFs of any website, cookie banners removed
- Sitemap Extractor: every URL of a website from its sitemaps, plus a broken-link check
- RSS Feed Reader: read and monitor any RSS/Atom feed, only new items, full article text