PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT avatar

PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT

Pricing

from $3.00 / 1,000 document converteds

Go to Apify Store
PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT

PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT

Extract text and tables from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV into clean LLM-ready Markdown plus structured JSON tables, with optional RAG chunks. Document parser for AI agents (MCP) and data pipelines. $0.003 per document; failed files are free.

Pricing

from $3.00 / 1,000 document converteds

Rating

0.0

(0)

Developer

Ventura WorkAlong

Ventura WorkAlong

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

5 hours ago

Last modified

Share

PDF to Markdown & Text Extractor: Tables, Word, Excel, PowerPoint

PDF to Markdown & Text Extractor turns documents into clean, LLM-ready Markdown and structured JSON tables in one call. It extracts text and tables from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV files. It's a bulk document parser for AI agents (via the Apify MCP server), RAG pipelines and data teams. $0.003 per document, and failed files are free.

  • One Actor, many formats: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT/Markdown. The file type is detected from the file itself, not just the extension.
  • Real tables, not flattened text: every table comes back as columns + rows (and optionally records), and as a GFM table inside the Markdown. Spreadsheet title rows become a caption, and two-row headers ("Population Estimate" over "2021 / 2022 / 2023") are merged into proper column names.
  • RAG-ready chunks: set chunkSize to get chunks split on heading, paragraph and table boundaries, each tagged with its section heading.
  • Pay only for what converts: failed downloads and unreadable files are free.

How to convert PDF to Markdown (and extract tables)

  1. Paste one or more document URLs into Document URLs: PDFs, Word, Excel or PowerPoint files, web pages or CSVs.
  2. Keep Extract tables on to get every table as JSON. Set Chunk size if you want RAG chunks.
  3. Click Start. Each document becomes one dataset item. Export it as JSON or CSV, or read it through the API or MCP.

Supported formats: PDF, Word, Excel, PowerPoint, HTML, CSV

InputMarkdownTables
PDFText per page (<!-- page N --> markers), rotated margin text removedRuled tables detected per page, with page location
Word (DOCX)Headings, lists and paragraphs in document orderEvery Word table
Excel (XLSX)One section per sheetOne table per sheet, with title rows as caption and merged headers
PowerPoint (PPTX)One section per slide: title, nested bullets, speaker notesSlide tables
HTMLMain content only (nav/footer/scripts dropped), links made absoluteData tables; layout tables rendered as text
CSVGFM tableDelimiter auto-detected (, ; tab |)

Input example

{
"documentUrls": [
"https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"https://www2.census.gov/programs-surveys/popest/tables/2020-2023/state/totals/NST-EST2023-POP.xlsx"
],
"extractTables": true,
"includeTableRecords": true,
"chunkSize": 2000
}

Output examples (one dataset item per document, shortened)

PDF text extraction with a table: IRS Form W-9 (real output)

{
"url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
"status": "ok",
"fileType": "pdf",
"title": "Form W-9 (Rev. March 2024)",
"pageCount": 6,
"wordCount": 6317,
"tableCount": 1,
"markdown": "<!-- page 1 -->\n\nW-9\nRequest for Taxpayer\nForm Give form to the\n(Rev. March 2024) Identification Number and Certification ...",
"tables": [{
"tableIndex": 0,
"location": "page 3",
"columns": ["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."],
"rows": [["• Corporation", "Corporation."], ["• Individual or • Sole proprietorship", "Individual/sole proprietor."]],
"rowCount": 5
}],
"metadata": {"subject": "Request for Taxpayer Identification Number and Certification", "createdAt": "2024-03-06"},
"warnings": [],
"error": null
}

PDF text keeps the layout's reading order, but form-style PDFs can still interleave labels, as above.

Excel to Markdown and JSON table: US Census workbook (real output)

{
"url": "https://www2.census.gov/.../NST-EST2023-POP.xlsx",
"status": "ok",
"fileType": "xlsx",
"pageCount": 1,
"tableCount": 1,
"markdown": "## Sheet: NST-EST2023-POP\n\n**Annual Estimates of the Resident Population ...**\n\n| Geographic Area | April 1, 2020 Estimates Base | Population Estimate (as of July 1) 2020 | ...",
"tables": [{
"tableIndex": 0,
"location": "sheet NST-EST2023-POP",
"caption": "Annual Estimates of the Resident Population ...",
"columns": ["Geographic Area", "April 1, 2020 Estimates Base", "Population Estimate (as of July 1) 2020", "..."],
"rows": [["United States", "331464948", "331526933", "..."]],
"rowCount": 57
}],
"chunks": [{"chunkIndex": 0, "heading": "Sheet: NST-EST2023-POP", "charCount": 1987, "text": "..."}],
"metadata": {"sheetCount": 1},
"warnings": [],
"error": null
}

Failed documents still produce an item with status: "failed" and an error message, so batches never silently lose files.

Pricing (pay per event)

EventPrice
Document converted (includes the first 20 pages/slides)$0.003
Each additional page/slide beyond 20$0.0002

Examples: 1,000 one-page invoices cost $3. A 100-page PDF costs $0.003 + 80 × $0.0002 = $0.019. Use maxPagesPerDocument to cap spend per file. The run stops cleanly when it reaches your maximum total charge.

Use with AI agents (MCP)

This Actor is callable as a tool through the Apify MCP server. An agent passes documentUrls and gets Markdown plus tables back, ready to reason over. Tip for agents: set includeMarkdown: false and extractTables: true when you only need the numbers.

Limitations (honest list)

  • No OCR. Scanned or image-only PDF pages produce no text; the item's warnings names those pages. Use a dedicated OCR Actor for scans.
  • PDF table detection works on tables with ruling lines. Borderless "whitespace" tables come through as text.
  • Old binary formats (.doc, .xls, .ppt) are not supported; save them as .docx/.xlsx/.pptx.
  • Charts, images and embedded objects aren't extracted (image alt text is kept for HTML).
  • Files must be reachable by a direct URL. Pages behind logins or paywalls aren't supported, by design.

Responsible use

  • Fetches only the URLs you provide, one request per document, and honors each site's robots.txt by default.
  • Refuses private, internal and cloud-metadata addresses.
  • Documents are processed in memory for your run only; the Actor does not store or reuse your files. You're responsible for having the right to process the documents you submit, especially any containing personal data.

FAQ

Can I upload a file instead of a URL?

Upload it to an Apify key-value store (or any storage with a direct link) and pass that URL.

Does it extract tables from PDF?

Yes, for tables drawn with ruling lines. Each table comes back as columns + rows JSON and as a Markdown table. Borderless tables come through as text.

Can I convert Word (DOCX) or Excel (XLSX) to Markdown?

Yes. Word headings, lists and tables keep their structure. Each Excel sheet becomes a Markdown table plus a JSON table with merged headers resolved.

Password-protected PDFs?

Set pdfPassword.

Huge spreadsheets?

maxRowsPerSheet (default 10,000) keeps output manageable. Set it to 0 for all rows.

Found a problem or need a format? Open an issue on the Actor's Issues tab. We read and fix them quickly.