PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG
Pricing
from $0.35 / 1,000 pages
PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG
Convert PDF, Word, Excel and PowerPoint files to clean Markdown for AI and RAG: tables as real Markdown tables and JSON rows, scanned pages via OCR, one row per document, page or chunk. Fills in the files the Website Content Crawler downloads but leaves empty.
Pricing
from $0.35 / 1,000 pages
Rating
0.0
(0)
Developer
Perceptron Data
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
PDF & DOCX to Markdown Converter — Excel, PowerPoint, Tables & OCR for RAG
Convert PDF, Word, Excel and PowerPoint files into clean Markdown and plain text for AI, RAG pipelines and vector databases. Tables come out as real Markdown tables and as JSON rows; scanned pages are read with OCR automatically. Get one row per document, per page or per ready-made chunk.
No LLM, no third-party API: your documents are converted inside your own Apify run and never sent to an outside AI service.
Fills the gap in Website Content Crawler
The Website Content Crawler
can download the PDFs and Office files it finds ("Save files") — but it does not
read them. Measured on a real crawl: every downloaded file comes back as a record
with an empty text and an empty markdown, plus a fileUrl.
Give this actor the crawler's run ID or dataset ID and it fills them in —
with the same url, crawl and metadata fields, so the results merge
straight into your crawl data:
{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C" }
To run it automatically after every crawl, add it in the crawler's Integrations tab ("Run succeeded" → this actor) with the input:
{ "crawler_dataset": "{{resource.defaultDatasetId}}", "output": "chunks" }
Formats
| Format | How it is read |
|---|---|
| PDF with a text layer | Text in reading order, headings detected from font size, tables as Markdown + JSON rows |
| PDF scans (pages without text) | OCR with Tesseract — English, German, French, Spanish, Italian, Dutch, Polish, Portuguese |
| Word (.docx) | Headings, lists, tables |
| Excel (.xlsx, .xls) | Every sheet as a table, also as JSON rows |
| PowerPoint (.pptx) | Slide titles and text, slide by slide |
| HTML, CSV, JSON, XML, TXT | Converted to Markdown |
Old binary formats (.doc, .ppt) are not supported — the run says so instead of returning garbage. A link that claims to be a PDF but returns a web page (a login wall, an error page) is rejected and not charged.
Measured speed
On one CPU core, 5 October 2026:
| Document | Time | Result |
|---|---|---|
| IRS Form W-9, 6 pages | 0.7 s | title, headings, 1 table |
| IRS Publication 15-T, 14 pages of tax tables | 2.8 s | 8 tables as JSON rows |
| 4 PDFs from a crawl, 116 pages | 12.4 s | — |
| Scanned PDF, 2 pages | 5.3 s | read by OCR |
| Word, Excel, PowerPoint file | 0.4–0.5 s each | tables included |
Runs use 1 GB of memory by default. More memory gives the run more CPU and converts proportionally faster — at the same cost per page.
Output
One row per document (default) — shortened:
{"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf","status": "ok","fileName": "fw9.pdf","format": "pdf","title": "Form W-9 (Rev. March 2024)","pageCount": 6,"pagesConverted": 6,"ocrPages": 0,"markdown": "<!-- Page 1 -->\n\n## Form W-9\n\nRequest for Taxpayer Identification Number and Certification …","text": "Form W-9\n\nRequest for Taxpayer Identification Number and Certification …","tables": [{ "page": 3, "rows": [["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."], ["• Corporation", "Corporation."], "…"] }],"warnings": []}
Tables from an Excel sheet, as they come out:
"tables": [{ "page": 1, "rows": [["Artikel", "Preis", "Bestand"], ["Schrauben M6", "0.12", "5400"], ["Dübel 8mm", "0.08", "12000"]] }]
One row per page ("output": "pages") adds page and ocr — useful when
your answers must cite a page. RAG chunks ("output": "chunks") split on
paragraph boundaries into pieces of chunk_size characters with chunk_overlap,
and every chunk carries pageStart and pageEnd. Form W-9 with 1,500-character
chunks: 33 chunks.
Files that cannot be converted come back as a row with "status": "failed" and
an error explaining why — and are not charged.
Input examples
Convert a list of files:
{"urls": ["https://www.irs.gov/pub/irs-pdf/fw9.pdf","https://example.com/report.docx","https://example.com/prices.xlsx"]}
RAG chunks from a crawl, scanned pages in German and English:
{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C", "output": "chunks", "chunk_size": 1500, "ocr_languages": ["deu", "eng"] }
A PDF whose text layer is garbage (copy-paste gives nonsense): force OCR.
{ "urls": ["https://example.com/broken-text.pdf"], "ocr": "force" }
What does a run cost?
$0.005 per run start, then per result — cheaper on bigger Apify plans:
| Your Apify plan | Document | Page | Page read by OCR |
|---|---|---|---|
| Free | $0.001 | $0.0005 | $0.003 |
| Starter | $0.0009 | $0.00045 | $0.0027 |
| Scale | $0.0008 | $0.0004 | $0.0024 |
| Business & Enterprise | $0.0007 | $0.00035 | $0.0021 |
1,000 ten-page PDFs with a text layer: $6.01 on the Free plan. Word, Excel and PowerPoint files count one page per 3,000 characters of output. Failed files are free.
Honest limits
- Tables without lines (columns aligned only by spacing) are often returned as text, not as a table. Ruled tables — the common case in reports, invoices and statistics — are detected reliably.
- Multi-column layouts (newspapers, magazines) are read column by column where the layout is clear; very complex layouts can mix lines.
- Charts and images are not described — there is no vision model involved.
- OCR quality follows scan quality. Handwriting is not supported.
- Password-protected PDFs fail with a clear message.
- Long documents stop at
max_pages(default 500) and say so inwarnings.
Data protection
Files are downloaded, converted and written to the dataset of your own run — nothing is sent to an AI provider or any other third party, and nothing is kept outside your Apify storage. Only public http(s) addresses are fetched; addresses inside private networks are refused.
FAQ
Why not just use an LLM to read PDFs? Cost and privacy. This actor converts 1,000 pages for $0.50 without sending a single page to an outside model. Use an LLM afterwards, on clean Markdown.
Does it work with scanned documents?
Yes. Pages without a text layer are detected automatically and read with OCR;
ocrPages tells you how many. Blank pages and chapter dividers are not sent to
OCR — they are not scans.
Can I get only the tables?
Every row has tables with the rows of each table. Read that field and ignore
the rest; Excel sheets always appear there.
How large can files be?
50 MB by default, up to 500 MB with max_file_mb, and 500 pages per document by
default.
Does it keep the crawler's data?
Yes — with crawler_dataset, every record keeps the crawler's url and crawl
fields, so you can join it with the crawl's web pages.
Auf Deutsch: PDF, Word und Excel in Markdown umwandeln
Dieser Actor wandelt PDF-, Word-, Excel- und PowerPoint-Dateien in sauberes Markdown und reinen Text um — für KI-Anwendungen, RAG und Vektordatenbanken. Tabellen kommen als Markdown-Tabellen und als JSON-Zeilen heraus, eingescannte Seiten werden automatisch per Texterkennung (OCR) gelesen, auch auf Deutsch.
Er ergänzt den Website Content Crawler: Dessen heruntergeladene Dateien bleiben dort leer — hier werden sie gefüllt. Ohne KI-Dienst eines Drittanbieters: Die Dokumente verlassen den eigenen Apify-Lauf nicht. Ausgabe wahlweise je Dokument, je Seite oder als fertige Abschnitte für RAG.
Related actors
| Actor | What it does |
|---|---|
| Website Content Crawler | Crawls websites and downloads their files — this actor reads them |
| Imprint & Impressum Scraper | Company data from the Impressum of any DACH website |
| Google Ads Transparency Scraper | Which companies advertise on Google, how many ads, who runs them |
| TED Tenders API | EU public tenders — their documents convert here |
| EUDAMED Scraper | EU medical devices, manufacturers and certificates |
| DACH & EU Data Source Finder | Free: tells you which ready-made scraper covers your data source |
Support
Bug reports are welcome in the Issues tab.