PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG avatar

PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG

Pricing

from $0.35 / 1,000 pages

Go to Apify Store
PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG

PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG

Convert PDF, Word, Excel and PowerPoint files to clean Markdown for AI and RAG: tables as real Markdown tables and JSON rows, scanned pages via OCR, one row per document, page or chunk. Fills in the files the Website Content Crawler downloads but leaves empty.

Pricing

from $0.35 / 1,000 pages

Rating

0.0

(0)

Developer

Perceptron Data

Perceptron Data

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

PDF & DOCX to Markdown Converter — Excel, PowerPoint, Tables & OCR for RAG

Convert PDF, Word, Excel and PowerPoint files into clean Markdown and plain text for AI, RAG pipelines and vector databases. Tables come out as real Markdown tables and as JSON rows; scanned pages are read with OCR automatically. Get one row per document, per page or per ready-made chunk.

No LLM, no third-party API: your documents are converted inside your own Apify run and never sent to an outside AI service.

Fills the gap in Website Content Crawler

The Website Content Crawler can download the PDFs and Office files it finds ("Save files") — but it does not read them. Measured on a real crawl: every downloaded file comes back as a record with an empty text and an empty markdown, plus a fileUrl.

Give this actor the crawler's run ID or dataset ID and it fills them in — with the same url, crawl and metadata fields, so the results merge straight into your crawl data:

{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C" }

To run it automatically after every crawl, add it in the crawler's Integrations tab ("Run succeeded" → this actor) with the input:

{ "crawler_dataset": "{{resource.defaultDatasetId}}", "output": "chunks" }

Formats

FormatHow it is read
PDF with a text layerText in reading order, headings detected from font size, tables as Markdown + JSON rows
PDF scans (pages without text)OCR with Tesseract — English, German, French, Spanish, Italian, Dutch, Polish, Portuguese
Word (.docx)Headings, lists, tables
Excel (.xlsx, .xls)Every sheet as a table, also as JSON rows
PowerPoint (.pptx)Slide titles and text, slide by slide
HTML, CSV, JSON, XML, TXTConverted to Markdown

Old binary formats (.doc, .ppt) are not supported — the run says so instead of returning garbage. A link that claims to be a PDF but returns a web page (a login wall, an error page) is rejected and not charged.

Measured speed

On one CPU core, 5 October 2026:

DocumentTimeResult
IRS Form W-9, 6 pages0.7 stitle, headings, 1 table
IRS Publication 15-T, 14 pages of tax tables2.8 s8 tables as JSON rows
4 PDFs from a crawl, 116 pages12.4 s—
Scanned PDF, 2 pages5.3 sread by OCR
Word, Excel, PowerPoint file0.4–0.5 s eachtables included

Runs use 1 GB of memory by default. More memory gives the run more CPU and converts proportionally faster — at the same cost per page.

Output

One row per document (default) — shortened:

{
"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"status": "ok",
"fileName": "fw9.pdf",
"format": "pdf",
"title": "Form W-9 (Rev. March 2024)",
"pageCount": 6,
"pagesConverted": 6,
"ocrPages": 0,
"markdown": "<!-- Page 1 -->\n\n## Form W-9\n\nRequest for Taxpayer Identification Number and Certification …",
"text": "Form W-9\n\nRequest for Taxpayer Identification Number and Certification …",
"tables": [{ "page": 3, "rows": [["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."], ["• Corporation", "Corporation."], "…"] }],
"warnings": []
}

Tables from an Excel sheet, as they come out:

"tables": [
{ "page": 1, "rows": [["Artikel", "Preis", "Bestand"], ["Schrauben M6", "0.12", "5400"], ["Dübel 8mm", "0.08", "12000"]] }
]

One row per page ("output": "pages") adds page and ocr — useful when your answers must cite a page. RAG chunks ("output": "chunks") split on paragraph boundaries into pieces of chunk_size characters with chunk_overlap, and every chunk carries pageStart and pageEnd. Form W-9 with 1,500-character chunks: 33 chunks.

Files that cannot be converted come back as a row with "status": "failed" and an error explaining why — and are not charged.

Input examples

Convert a list of files:

{
"urls": [
"https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"https://example.com/report.docx",
"https://example.com/prices.xlsx"
]
}

RAG chunks from a crawl, scanned pages in German and English:

{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C", "output": "chunks", "chunk_size": 1500, "ocr_languages": ["deu", "eng"] }

A PDF whose text layer is garbage (copy-paste gives nonsense): force OCR.

{ "urls": ["https://example.com/broken-text.pdf"], "ocr": "force" }

What does a run cost?

$0.005 per run start, then per result — cheaper on bigger Apify plans:

Your Apify planDocumentPagePage read by OCR
Free$0.001$0.0005$0.003
Starter$0.0009$0.00045$0.0027
Scale$0.0008$0.0004$0.0024
Business & Enterprise$0.0007$0.00035$0.0021

1,000 ten-page PDFs with a text layer: $6.01 on the Free plan. Word, Excel and PowerPoint files count one page per 3,000 characters of output. Failed files are free.

Honest limits

  • Tables without lines (columns aligned only by spacing) are often returned as text, not as a table. Ruled tables — the common case in reports, invoices and statistics — are detected reliably.
  • Multi-column layouts (newspapers, magazines) are read column by column where the layout is clear; very complex layouts can mix lines.
  • Charts and images are not described — there is no vision model involved.
  • OCR quality follows scan quality. Handwriting is not supported.
  • Password-protected PDFs fail with a clear message.
  • Long documents stop at max_pages (default 500) and say so in warnings.

Data protection

Files are downloaded, converted and written to the dataset of your own run — nothing is sent to an AI provider or any other third party, and nothing is kept outside your Apify storage. Only public http(s) addresses are fetched; addresses inside private networks are refused.

FAQ

Why not just use an LLM to read PDFs? Cost and privacy. This actor converts 1,000 pages for $0.50 without sending a single page to an outside model. Use an LLM afterwards, on clean Markdown.

Does it work with scanned documents? Yes. Pages without a text layer are detected automatically and read with OCR; ocrPages tells you how many. Blank pages and chapter dividers are not sent to OCR — they are not scans.

Can I get only the tables? Every row has tables with the rows of each table. Read that field and ignore the rest; Excel sheets always appear there.

How large can files be? 50 MB by default, up to 500 MB with max_file_mb, and 500 pages per document by default.

Does it keep the crawler's data? Yes — with crawler_dataset, every record keeps the crawler's url and crawl fields, so you can join it with the crawl's web pages.

Auf Deutsch: PDF, Word und Excel in Markdown umwandeln

Dieser Actor wandelt PDF-, Word-, Excel- und PowerPoint-Dateien in sauberes Markdown und reinen Text um — für KI-Anwendungen, RAG und Vektordatenbanken. Tabellen kommen als Markdown-Tabellen und als JSON-Zeilen heraus, eingescannte Seiten werden automatisch per Texterkennung (OCR) gelesen, auch auf Deutsch.

Er ergänzt den Website Content Crawler: Dessen heruntergeladene Dateien bleiben dort leer — hier werden sie gefüllt. Ohne KI-Dienst eines Drittanbieters: Die Dokumente verlassen den eigenen Apify-Lauf nicht. Ausgabe wahlweise je Dokument, je Seite oder als fertige Abschnitte für RAG.

ActorWhat it does
Website Content CrawlerCrawls websites and downloads their files — this actor reads them
Imprint & Impressum ScraperCompany data from the Impressum of any DACH website
Google Ads Transparency ScraperWhich companies advertise on Google, how many ads, who runs them
TED Tenders APIEU public tenders — their documents convert here
EUDAMED ScraperEU medical devices, manufacturers and certificates
DACH & EU Data Source FinderFree: tells you which ready-made scraper covers your data source

Support

Bug reports are welcome in the Issues tab.