# PDF to Markdown & Text Extractor: Tables, Word, Excel, PPT (`ventura_workalong/doc-to-markdown-tables`) Actor

Convert PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV docs to clean LLM-ready Markdown: extract text and structured JSON tables, with optional RAG chunks. Document parser for AI agents (MCP) and data pipelines. $0.003 per document; failed files are free.

- **URL**: https://apify.com/ventura_workalong/doc-to-markdown-tables.md
- **Developed by:** [Ventura WorkAlong](https://apify.com/ventura_workalong) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 document converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Markdown & Text Extractor: Tables, Word, Excel, PowerPoint

**PDF to Markdown & Text Extractor** turns documents into **clean, LLM-ready Markdown** and **structured JSON tables** in one call. It extracts text and tables from PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML and CSV files. It's a bulk document parser for AI agents (via the Apify MCP server), RAG pipelines and data teams. **$0.003 per document**, and failed files are free.

- **One Actor, many formats:** PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT/Markdown. The file type is detected from the file itself, not just the extension.
- **Real tables, not flattened text:** every table comes back as `columns` + `rows` (and optionally `records`), *and* as a GFM table inside the Markdown. Spreadsheet title rows become a `caption`, and two-row headers ("Population Estimate" over "2021 / 2022 / 2023") are merged into proper column names.
- **RAG-ready chunks:** set `chunkSize` to get chunks split on heading, paragraph and table boundaries, each tagged with its section heading.
- **Pay only for what converts:** failed downloads and unreadable files are free.

### How to convert PDF to Markdown (and extract tables)

1. Paste one or more document URLs into **Document URLs**: PDFs, Word, Excel or PowerPoint files, web pages or CSVs.
2. Keep **Extract tables** on to get every table as JSON. Set **Chunk size** if you want RAG chunks.
3. Click **Start**. Each document becomes one dataset item. Export it as JSON or CSV, or read it through the API or MCP.

### Supported formats: PDF, Word, Excel, PowerPoint, HTML, CSV

| Input | Markdown | Tables |
|---|---|---|
| PDF | Text per page (`<!-- page N -->` markers), rotated margin text removed | Ruled tables detected per page, with page location |
| Word (DOCX) | Headings, lists and paragraphs in document order | Every Word table |
| Excel (XLSX) | One section per sheet | One table per sheet, with title rows as caption and merged headers |
| PowerPoint (PPTX) | One section per slide: title, nested bullets, speaker notes | Slide tables |
| HTML | Main content only (nav/footer/scripts dropped), links made absolute | Data tables; layout tables rendered as text |
| CSV | GFM table | Delimiter auto-detected (`,` `;` tab `\|`) |

### Input example

```json
{
  "documentUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    "https://www2.census.gov/programs-surveys/popest/tables/2020-2023/state/totals/NST-EST2023-POP.xlsx"
  ],
  "extractTables": true,
  "includeTableRecords": true,
  "chunkSize": 2000
}
```

### Output examples (one dataset item per document, shortened)

#### PDF text extraction with a table: IRS Form W-9 (real output)

```json
{
  "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
  "status": "ok",
  "fileType": "pdf",
  "title": "Form W-9 (Rev. March 2024)",
  "pageCount": 6,
  "wordCount": 6317,
  "tableCount": 1,
  "markdown": "<!-- page 1 -->\n\nW-9\nRequest for Taxpayer\nForm Give form to the\n(Rev. March 2024) Identification Number and Certification ...",
  "tables": [{
    "tableIndex": 0,
    "location": "page 3",
    "columns": ["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."],
    "rows": [["• Corporation", "Corporation."], ["• Individual or • Sole proprietorship", "Individual/sole proprietor."]],
    "rowCount": 5
  }],
  "metadata": {"subject": "Request for Taxpayer Identification Number and Certification", "createdAt": "2024-03-06"},
  "warnings": [],
  "error": null
}
```

PDF text keeps the layout's reading order, but form-style PDFs can still interleave labels, as above.

#### Excel to Markdown and JSON table: US Census workbook (real output)

```json
{
  "url": "https://www2.census.gov/.../NST-EST2023-POP.xlsx",
  "status": "ok",
  "fileType": "xlsx",
  "pageCount": 1,
  "tableCount": 1,
  "markdown": "## Sheet: NST-EST2023-POP\n\n**Annual Estimates of the Resident Population ...**\n\n| Geographic Area | April 1, 2020 Estimates Base | Population Estimate (as of July 1) 2020 | ...",
  "tables": [{
    "tableIndex": 0,
    "location": "sheet NST-EST2023-POP",
    "caption": "Annual Estimates of the Resident Population ...",
    "columns": ["Geographic Area", "April 1, 2020 Estimates Base", "Population Estimate (as of July 1) 2020", "..."],
    "rows": [["United States", "331464948", "331526933", "..."]],
    "rowCount": 57
  }],
  "chunks": [{"chunkIndex": 0, "heading": "Sheet: NST-EST2023-POP", "charCount": 1987, "text": "..."}],
  "metadata": {"sheetCount": 1},
  "warnings": [],
  "error": null
}
```

Failed documents still produce an item with `status: "failed"` and an `error` message, so batches never silently lose files.

### Pricing (pay per event)

| Event | Price |
|---|---|
| Document converted (includes the first 20 pages/slides) | $0.003 |
| Each additional page/slide beyond 20 | $0.0002 |

Examples: 1,000 one-page invoices cost $3. A 100-page PDF costs $0.003 + 80 × $0.0002 = $0.019. Use `maxPagesPerDocument` to cap spend per file. The run stops cleanly when it reaches your maximum total charge.

### Use with AI agents (MCP)

This Actor is callable as a tool through the [Apify MCP server](https://mcp.apify.com). An agent passes `documentUrls` and gets Markdown plus tables back, ready to reason over. Tip for agents: set `includeMarkdown: false` and `extractTables: true` when you only need the numbers.

### Limitations (honest list)

- **No OCR.** Scanned or image-only PDF pages produce no text; the item's `warnings` names those pages. Use a dedicated OCR Actor for scans.
- PDF table detection works on tables with ruling lines. Borderless "whitespace" tables come through as text.
- Old binary formats (.doc, .xls, .ppt) are not supported; save them as .docx/.xlsx/.pptx.
- Charts, images and embedded objects aren't extracted (image alt text is kept for HTML).
- Files must be reachable by a direct URL. Pages behind logins or paywalls aren't supported, by design.

### Responsible use

- Fetches only the URLs you provide, one request per document, and honors each site's `robots.txt` by default.
- Refuses private, internal and cloud-metadata addresses.
- Documents are processed in memory for your run only; the Actor does not store or reuse your files. You're responsible for having the right to process the documents you submit, especially any containing personal data.

### FAQ

#### Can I upload a file instead of a URL?

Upload it to an Apify key-value store (or any storage with a direct link) and pass that URL.

#### Does it extract tables from PDF?

Yes, for tables drawn with ruling lines. Each table comes back as `columns` + `rows` JSON and as a Markdown table. Borderless tables come through as text.

#### Can I convert Word (DOCX) or Excel (XLSX) to Markdown?

Yes. Word headings, lists and tables keep their structure. Each Excel sheet becomes a Markdown table plus a JSON table with merged headers resolved.

#### Password-protected PDFs?

Set `pdfPassword`.

#### Huge spreadsheets?

`maxRowsPerSheet` (default 10,000) keeps output manageable. Set it to 0 for all rows.

### Related Actors

- [Sitemap URL Extractor](https://apify.com/ventura_workalong/sitemap-url-extractor): list every page or PDF a site publishes, then convert them here.
- [Tech Stack Detector](https://apify.com/ventura_workalong/tech-stack-detector) and [Bulk WHOIS Domain Lookup](https://apify.com/ventura_workalong/domain-lookup-bundle) for company and website research.

Found a problem or need a format? Open an issue on the Actor's Issues tab. We read and fix them quickly.

# Actor input Schema

## `documentUrls` (type: `array`):

Direct links to the files to convert: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT/MD. Public URLs or Apify key-value store URLs. The file type is detected from the file itself.

## `extractTables` (type: `boolean`):

Return every detected table as structured columns + rows, in addition to the Markdown version.

## `includeTableRecords` (type: `boolean`):

Add a `records` array (one object per row, keyed by column name) to each table. Handy for spreadsheets and databases.

## `chunkSize` (type: `integer`):

Split the Markdown into chunks of at most this many characters on heading/paragraph/table boundaries. 0 disables chunking.

## `includeMarkdown` (type: `boolean`):

Turn off to return only tables/chunks and save output size.

## `maxPagesPerDocument` (type: `integer`):

Process at most this many PDF pages or PPTX slides per document (0 = all). Pages beyond 20 are billed per page, so this also caps cost.

## `maxRowsPerSheet` (type: `integer`):

Truncate each spreadsheet or CSV to this many data rows (0 = all).

## `maxFileSizeMb` (type: `integer`):

Skip files larger than this.

## `pdfPassword` (type: `string`):

Password for protected PDFs (applies to all PDFs in this run).

## `keepHtmlNavigation` (type: `boolean`):

By default, nav, header, footer and aside elements are dropped from HTML pages.

## `respectRobotsTxt` (type: `boolean`):

Skip URLs that the site's robots.txt disallows.

## `maxConcurrency` (type: `integer`):

Documents downloaded and converted in parallel.

## Actor input object example

```json
{
  "documentUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "extractTables": true,
  "includeTableRecords": false,
  "chunkSize": 0,
  "includeMarkdown": true,
  "maxPagesPerDocument": 200,
  "maxRowsPerSheet": 10000,
  "maxFileSizeMb": 50,
  "keepHtmlNavigation": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 4
}
```

# Actor output Schema

## `results` (type: `string`):

All results from this run, one item per input, in the default dataset.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documentUrls": [
        "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("ventura_workalong/doc-to-markdown-tables").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "documentUrls": ["https://www.irs.gov/pub/irs-pdf/fw9.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("ventura_workalong/doc-to-markdown-tables").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documentUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ]
}' |
apify call ventura_workalong/doc-to-markdown-tables --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ventura_workalong/doc-to-markdown-tables"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DTu5S6hGNmBiRDGgX/builds/nEDboV2bRl9o72kkQ/openapi.json
