PDF to Markdown Converter - Extract & Format Text avatar

PDF to Markdown Converter - Extract & Format Text

Pricing

from $26.00 / 1,000 pdf converteds

Go to Apify Store
PDF to Markdown Converter - Extract & Format Text

PDF to Markdown Converter - Extract & Format Text

Convert PDF documents into clean, readable Markdown with the document structure preserved — headings, lists and paragraph flow. Process files in bulk and get output that drops straight into LLM pipelines, RAG systems, knowledge bases and documentation sites.

Pricing

from $26.00 / 1,000 pdf converteds

Rating

0.0

(0)

Developer

daehwan kim

daehwan kim

Maintained by Community

Actor stats

0

Bookmarked

1

Total users

0

Monthly active users

13 days ago

Last modified

Share

Office to Markdown — RAG-Ready Document Extractor

Convert PDF, DOCX, PPTX, XLSX, HTML, images, audio, and 20+ other formats into clean LLM-ready Markdown in one API call. Powered by Microsoft MarkItDown — the highest-fidelity open-source document-to-Markdown engine available.

Optimized for RAG pipelines, embedding ingestion, AI knowledge bases, and document understanding workflows where the quality of upstream chunking determines the quality of downstream retrieval.

2026-09-16 — Format detection follows the bytes that arrived, not the URL ending, so a PDF served at /report.html no longer comes back as replacement characters. Transient 429/5xx answers are retried instead of costing the document.

2026-09-15 — Batch input: fileUrls takes an array and returns one Markdown row per document. fileUrl and pdfUrl still work as single-URL aliases.

v2.0 (2026-05-25) — Upgraded from PDF-only to multi-format.

Why This Actor

pdf-parse and pdfplumber extract raw text but lose structure: headings collapse, tables stringify, lists flatten. For RAG, that means smaller retrieval precision and more hallucination downstream.

Microsoft MarkItDown preserves:

  • Heading hierarchy → #, ##, ### mapped from document outline
  • Tables → real Markdown tables, not pipe-broken text
  • Lists → bullet and numbered list integrity
  • Code blocks → fenced code fences preserved
  • Image alt-text → embedded into the flow for context

For DOCX and PPTX, semantic structure (slide titles, footnotes, comments) is preserved. For images and audio, OCR / transcription fallback runs automatically.

Supported Formats

CategoryFormats
DocumentsPDF, DOCX, PPTX, XLSX, ODT, RTF
Web / MarkupHTML, HTM, XML, MHTML
DataCSV, JSON, TSV
PlainTXT, MD
Images (with OCR)PNG, JPG, JPEG, GIF, BMP, WEBP
Audio (with transcription)MP3, WAV, M4A
ArchivesZIP (recursive), EPUB
OthersYouTube URLs (transcript), Outlook MSG

Max file size: 100 MB per document, up to 200 documents per run.

Use Cases

  • RAG ingestion — Convert document libraries into Markdown chunks before embedding with OpenAI / Voyage / Cohere
  • AI knowledge bases — Bulk import company wikis, training material, manuals into vector DBs
  • Document Q&A — Pre-process source documents for Claude / GPT structured extraction
  • Compliance archival — Normalize multi-format historical records to searchable Markdown
  • Migration projects — Move from SharePoint / Confluence to modern docs-as-code platforms
  • LLM fine-tuning data prep — Clean Markdown corpus from heterogeneous source files

Input

FieldTypeRequiredDescription
fileUrlsarray of string✅Direct HTTPS URLs to supported documents (max 100 MB each, 200 per run)
fileUrl / pdfUrlstring—Legacy single-URL aliases (v1/v2.0 compatibility)
includePageBreaksboolean—For PDFs, extract page by page and insert --- between pages. Uses a plain-text extractor instead of MarkItDown's table-aware PDF path, so tables come out as text. Default false
truncateCharsinteger—Cap each document's Markdown at N characters. 0 = full text (default)

URLs without a scheme (example.com/report.pdf) are fetched over HTTPS. Only http:// and https:// are fetched; any other scheme returns a row explaining why. When a run approaches its timeout, the remaining URLs are skipped and reported in a notice row so the rows already converted are kept.

{
"fileUrls": [
"https://arxiv.org/pdf/2305.10601",
"https://example.com/report.docx"
],
"includePageBreaks": true,
"truncateChars": 0
}

Output

One dataset item per document:

FieldTypeDescription
fileUrlstringSource URL
fileFormatstringDetected file extension (pdf, docx, ...), resolved from the content type when the URL disagrees
contentTypestringContent type the server declared
byteSizeintegerBytes downloaded
charCountintegerCharacter length of resulting Markdown
wordCountintegerWhitespace-tokenized word count
pageBreaksAppliedbooleanWhether per-page extraction was actually used
truncatedbooleanWhether truncateChars cut the text
downloadAttemptsintegerHow many requests the download took (transient 429/5xx answers are retried)
extractionEmptybooleantrue when the source carried no extractable text, e.g. an image-only scan
markdownstringFinal cleaned Markdown
disclaimerstringConversion accuracy notice
errorstringPopulated only on failure
noticestringStatus message row (free-plan, charging limit, run-limit or time-budget stop); not a result
{
"fileUrl": "https://arxiv.org/pdf/2305.10601",
"fileFormat": "pdf",
"byteSize": 1043820,
"charCount": 48230,
"wordCount": 7821,
"pageBreaksApplied": true,
"truncated": false,
"markdown": "# Tree of Thoughts: Deliberate Problem Solving with Large Language Models\n\n## Abstract\n\nLanguage models are increasingly being deployed for general problem solving..."
}

Pricing

Free plan: each run processes up to 25 documents (download or conversion error rows count toward the 25). Paid Apify plans receive the full result set.

  • $0.05 per document converted (event: pdf-converted)
  • Charged per document delivered to the dataset
  • Download / conversion error rows and status notices are not charged. A document that converts successfully but contains no extractable text (for example an image-only scan, marked extractionEmpty: true) is still charged as one converted document
  • Apify platform compute usage is billed separately to users (passOnCosts enabled)

Quick Start

curl

curl -X POST "https://api.apify.com/v2/acts/ntriqpro~pdf-to-markdown/runs?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"fileUrls": ["https://arxiv.org/pdf/2305.10601"],
"includePageBreaks": true
}'

Python (Apify Client)

from apify_client import ApifyClient
client = ApifyClient("YOUR_TOKEN")
run = client.actor("ntriqpro/pdf-to-markdown").call(run_input={
"fileUrls": ["https://example.com/report.docx", "https://example.com/manual.pdf"]
})
items = list(client.dataset(run["defaultDatasetId"]).iterate_items())
print(items[0]["markdown"])

JavaScript (Apify Client)

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('ntriqpro/pdf-to-markdown').call({
fileUrls: ['https://example.com/slides.pptx']
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].markdown);

Limitations

LimitationDetail
Scanned image-only PDFsOCR is applied but accuracy depends on scan quality
Encrypted / password-protected filesNot supported
Files > 100 MBHard-rejected to protect compute cost
Non-Latin scripts (CJK, Arabic, etc.)Supported but proofreading recommended for production
Streaming sources (S3 signed URLs)Supported as long as URL is HTTPS-reachable

Always validate critical extractions against the source.

Technology Stack

Disclaimer

This Actor is an unofficial open-source wrapper around Microsoft MarkItDown. It is not affiliated with, sponsored by, or endorsed by Microsoft Corporation. Conversion fidelity depends on source-document structure; results are provided for informational and AI ingestion purposes only and are not a substitute for human review of critical or regulated documents.

Changelog

  • 2.0 (2026-05-25) — Migrated to Python + Microsoft MarkItDown. Multi-format support (DOCX, PPTX, XLSX, HTML, images, audio, etc.). Output schema enriched with fileFormat / byteSize / charCount. pdfUrl retained as alias of fileUrl.
  • 1.0 (2026-04-14) — Initial release with pdf-parse JavaScript backend (PDF only).

⭐ Rate this Actor

If this saves you time, please leave a review — it helps other teams discover it.