PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT
Pricing
from $3.00 / 1,000 document converteds
PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT
Extract text and tables from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV into clean LLM-ready Markdown plus structured JSON tables, with optional RAG chunks. Document parser for AI agents (MCP) and data pipelines. $0.003 per document; failed files are free.
Pricing
from $3.00 / 1,000 document converteds
Rating
0.0
(0)
Developer
Ventura WorkAlong
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
PDF to Markdown & Text Extractor: Tables, Word, Excel, PowerPoint
PDF to Markdown & Text Extractor turns documents into clean, LLM-ready Markdown and structured JSON tables in one call. It extracts text and tables from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV files. It's a bulk document parser for AI agents (via the Apify MCP server), RAG pipelines and data teams. $0.003 per document, and failed files are free.
- One Actor, many formats: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT/Markdown. The file type is detected from the file itself, not just the extension.
- Real tables, not flattened text: every table comes back as
columns+rows(and optionallyrecords), and as a GFM table inside the Markdown. Spreadsheet title rows become acaption, and two-row headers ("Population Estimate" over "2021 / 2022 / 2023") are merged into proper column names. - RAG-ready chunks: set
chunkSizeto get chunks split on heading, paragraph and table boundaries, each tagged with its section heading. - Pay only for what converts: failed downloads and unreadable files are free.
How to convert PDF to Markdown (and extract tables)
- Paste one or more document URLs into Document URLs: PDFs, Word, Excel or PowerPoint files, web pages or CSVs.
- Keep Extract tables on to get every table as JSON. Set Chunk size if you want RAG chunks.
- Click Start. Each document becomes one dataset item. Export it as JSON or CSV, or read it through the API or MCP.
Supported formats: PDF, Word, Excel, PowerPoint, HTML, CSV
| Input | Markdown | Tables |
|---|---|---|
Text per page (<!-- page N --> markers), rotated margin text removed | Ruled tables detected per page, with page location | |
| Word (DOCX) | Headings, lists and paragraphs in document order | Every Word table |
| Excel (XLSX) | One section per sheet | One table per sheet, with title rows as caption and merged headers |
| PowerPoint (PPTX) | One section per slide: title, nested bullets, speaker notes | Slide tables |
| HTML | Main content only (nav/footer/scripts dropped), links made absolute | Data tables; layout tables rendered as text |
| CSV | GFM table | Delimiter auto-detected (, ; tab |) |
Input example
{"documentUrls": ["https://www.irs.gov/pub/irs-pdf/fw9.pdf","https://www2.census.gov/programs-surveys/popest/tables/2020-2023/state/totals/NST-EST2023-POP.xlsx"],"extractTables": true,"includeTableRecords": true,"chunkSize": 2000}
Output examples (one dataset item per document, shortened)
PDF text extraction with a table: IRS Form W-9 (real output)
{"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf","status": "ok","fileType": "pdf","title": "Form W-9 (Rev. March 2024)","pageCount": 6,"wordCount": 6317,"tableCount": 1,"markdown": "<!-- page 1 -->\n\nW-9\nRequest for Taxpayer\nForm Give form to the\n(Rev. March 2024) Identification Number and Certification ...","tables": [{"tableIndex": 0,"location": "page 3","columns": ["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."],"rows": [["• Corporation", "Corporation."], ["• Individual or • Sole proprietorship", "Individual/sole proprietor."]],"rowCount": 5}],"metadata": {"subject": "Request for Taxpayer Identification Number and Certification", "createdAt": "2024-03-06"},"warnings": [],"error": null}
PDF text keeps the layout's reading order, but form-style PDFs can still interleave labels, as above.
Excel to Markdown and JSON table: US Census workbook (real output)
{"url": "https://www2.census.gov/.../NST-EST2023-POP.xlsx","status": "ok","fileType": "xlsx","pageCount": 1,"tableCount": 1,"markdown": "## Sheet: NST-EST2023-POP\n\n**Annual Estimates of the Resident Population ...**\n\n| Geographic Area | April 1, 2020 Estimates Base | Population Estimate (as of July 1) 2020 | ...","tables": [{"tableIndex": 0,"location": "sheet NST-EST2023-POP","caption": "Annual Estimates of the Resident Population ...","columns": ["Geographic Area", "April 1, 2020 Estimates Base", "Population Estimate (as of July 1) 2020", "..."],"rows": [["United States", "331464948", "331526933", "..."]],"rowCount": 57}],"chunks": [{"chunkIndex": 0, "heading": "Sheet: NST-EST2023-POP", "charCount": 1987, "text": "..."}],"metadata": {"sheetCount": 1},"warnings": [],"error": null}
Failed documents still produce an item with status: "failed" and an error message, so batches never silently lose files.
Pricing (pay per event)
| Event | Price |
|---|---|
| Document converted (includes the first 20 pages/slides) | $0.003 |
| Each additional page/slide beyond 20 | $0.0002 |
Examples: 1,000 one-page invoices cost $3. A 100-page PDF costs $0.003 + 80 × $0.0002 = $0.019. Use maxPagesPerDocument to cap spend per file. The run stops cleanly when it reaches your maximum total charge.
Use with AI agents (MCP)
This Actor is callable as a tool through the Apify MCP server. An agent passes documentUrls and gets Markdown plus tables back, ready to reason over. Tip for agents: set includeMarkdown: false and extractTables: true when you only need the numbers.
Limitations (honest list)
- No OCR. Scanned or image-only PDF pages produce no text; the item's
warningsnames those pages. Use a dedicated OCR Actor for scans. - PDF table detection works on tables with ruling lines. Borderless "whitespace" tables come through as text.
- Old binary formats (.doc, .xls, .ppt) are not supported; save them as .docx/.xlsx/.pptx.
- Charts, images and embedded objects aren't extracted (image alt text is kept for HTML).
- Files must be reachable by a direct URL. Pages behind logins or paywalls aren't supported, by design.
Responsible use
- Fetches only the URLs you provide, one request per document, and honors each site's
robots.txtby default. - Refuses private, internal and cloud-metadata addresses.
- Documents are processed in memory for your run only; the Actor does not store or reuse your files. You're responsible for having the right to process the documents you submit, especially any containing personal data.
FAQ
Can I upload a file instead of a URL?
Upload it to an Apify key-value store (or any storage with a direct link) and pass that URL.
Does it extract tables from PDF?
Yes, for tables drawn with ruling lines. Each table comes back as columns + rows JSON and as a Markdown table. Borderless tables come through as text.
Can I convert Word (DOCX) or Excel (XLSX) to Markdown?
Yes. Word headings, lists and tables keep their structure. Each Excel sheet becomes a Markdown table plus a JSON table with merged headers resolved.
Password-protected PDFs?
Set pdfPassword.
Huge spreadsheets?
maxRowsPerSheet (default 10,000) keeps output manageable. Set it to 0 for all rows.
Related Actors
- Sitemap URL Extractor: list every page or PDF a site publishes, then convert them here.
- Tech Stack Detector and Bulk WHOIS Domain Lookup for company and website research.
Found a problem or need a format? Open an issue on the Actor's Issues tab. We read and fix them quickly.