# PDF & DOCX to Markdown – Excel, PowerPoint, Tables, OCR for RAG (`perceptr0n/pdf-to-markdown-converter`) Actor

Convert PDF, Word, Excel and PowerPoint files to clean Markdown for AI and RAG: tables as real Markdown tables and JSON rows, scanned pages via OCR, one row per document, page or chunk. Fills in the files the Website Content Crawler downloads but leaves empty.

- **URL**: https://apify.com/perceptr0n/pdf-to-markdown-converter.md
- **Developed by:** [Perceptron Data](https://apify.com/perceptr0n) (community)
- **Categories:** AI, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.35 / 1,000 pages

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF & DOCX to Markdown Converter — Excel, PowerPoint, Tables & OCR for RAG

Convert **PDF, Word, Excel and PowerPoint** files into clean **Markdown** and
plain text for AI, RAG pipelines and vector databases. Tables come out as real
Markdown tables **and** as JSON rows; scanned pages are read with **OCR**
automatically. Get one row per document, per page or per ready-made chunk.

**No LLM, no third-party API:** your documents are converted inside your own
Apify run and never sent to an outside AI service.

### Fills the gap in Website Content Crawler

The [Website Content Crawler](https://apify.com/apify/website-content-crawler)
can download the PDFs and Office files it finds ("Save files") — but it does not
read them. Measured on a real crawl: every downloaded file comes back as a record
with an **empty `text` and an empty `markdown`**, plus a `fileUrl`.

Give this actor the crawler's **run ID or dataset ID** and it fills them in —
with the same `url`, `crawl` and `metadata` fields, so the results merge
straight into your crawl data:

```json
{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C" }
```

To run it automatically after every crawl, add it in the crawler's
**Integrations** tab ("Run succeeded" → this actor) with the input:

```json
{ "crawler_dataset": "{{resource.defaultDatasetId}}", "output": "chunks" }
```

### Formats

| Format | How it is read |
|---|---|
| **PDF** with a text layer | Text in reading order, headings detected from font size, tables as Markdown + JSON rows |
| **PDF scans** (pages without text) | OCR with Tesseract — English, German, French, Spanish, Italian, Dutch, Polish, Portuguese |
| **Word** (.docx) | Headings, lists, tables |
| **Excel** (.xlsx, .xls) | Every sheet as a table, also as JSON rows |
| **PowerPoint** (.pptx) | Slide titles and text, slide by slide |
| HTML, CSV, JSON, XML, TXT | Converted to Markdown |

Old binary formats (.doc, .ppt) are not supported — the run says so instead of
returning garbage. A link that claims to be a PDF but returns a web page (a login
wall, an error page) is rejected and not charged.

### Measured speed

On **one CPU core**, 5 October 2026:

| Document | Time | Result |
|---|---|---|
| IRS Form W-9, 6 pages | 0.7 s | title, headings, 1 table |
| IRS Publication 15-T, 14 pages of tax tables | 2.8 s | 8 tables as JSON rows |
| 4 PDFs from a crawl, 116 pages | 12.4 s | — |
| Scanned PDF, 2 pages | 5.3 s | read by OCR |
| Word, Excel, PowerPoint file | 0.4–0.5 s each | tables included |

Runs use 1 GB of memory by default. More memory gives the run more CPU and
converts proportionally faster — at the same cost per page.

### Output

**One row per document** (default) — shortened:

```json
{
    "url": "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
    "status": "ok",
    "fileName": "fw9.pdf",
    "format": "pdf",
    "title": "Form W-9 (Rev. March 2024)",
    "pageCount": 6,
    "pagesConverted": 6,
    "ocrPages": 0,
    "markdown": "<!-- Page 1 -->\n\n## Form  W-9\n\nRequest for Taxpayer Identification Number and Certification …",
    "text": "Form  W-9\n\nRequest for Taxpayer Identification Number and Certification …",
    "tables": [{ "page": 3, "rows": [["IF the entity/individual on line 1 is a(n) . . .", "THEN check the box for . . ."], ["• Corporation", "Corporation."], "…"] }],
    "warnings": []
}
```

Tables from an Excel sheet, as they come out:

```json
"tables": [
    { "page": 1, "rows": [["Artikel", "Preis", "Bestand"], ["Schrauben M6", "0.12", "5400"], ["Dübel 8mm", "0.08", "12000"]] }
]
```

**One row per page** (`"output": "pages"`) adds `page` and `ocr` — useful when
your answers must cite a page. **RAG chunks** (`"output": "chunks"`) split on
paragraph boundaries into pieces of `chunk_size` characters with `chunk_overlap`,
and every chunk carries `pageStart` and `pageEnd`. Form W-9 with 1,500-character
chunks: 33 chunks.

Files that cannot be converted come back as a row with `"status": "failed"` and
an `error` explaining why — and are **not charged**.

### Input examples

Convert a list of files:

```json
{
    "urls": [
        "https://www.irs.gov/pub/irs-pdf/fw9.pdf",
        "https://example.com/report.docx",
        "https://example.com/prices.xlsx"
    ]
}
```

RAG chunks from a crawl, scanned pages in German and English:

```json
{ "crawler_dataset": "Myhr0xrkHRkf0Ub4C", "output": "chunks", "chunk_size": 1500, "ocr_languages": ["deu", "eng"] }
```

A PDF whose text layer is garbage (copy-paste gives nonsense): force OCR.

```json
{ "urls": ["https://example.com/broken-text.pdf"], "ocr": "force" }
```

### What does a run cost?

$0.005 per run start, then per result — cheaper on bigger Apify plans:

| Your Apify plan | Document | Page | Page read by OCR |
|---|---|---|---|
| Free | $0.001 | $0.0005 | $0.003 |
| Starter | $0.0009 | $0.00045 | $0.0027 |
| Scale | $0.0008 | $0.0004 | $0.0024 |
| Business & Enterprise | $0.0007 | $0.00035 | $0.0021 |

1,000 ten-page PDFs with a text layer: **$6.01** on the Free plan. Word, Excel and
PowerPoint files count one page per 3,000 characters of output. Failed files are
free.

### Honest limits

- **Tables without lines** (columns aligned only by spacing) are often returned
  as text, not as a table. Ruled tables — the common case in reports, invoices and
  statistics — are detected reliably.
- **Multi-column layouts** (newspapers, magazines) are read column by column
  where the layout is clear; very complex layouts can mix lines.
- **Charts and images** are not described — there is no vision model involved.
- **OCR quality follows scan quality.** Handwriting is not supported.
- **Password-protected PDFs** fail with a clear message.
- Long documents stop at `max_pages` (default 500) and say so in `warnings`.

### Data protection

Files are downloaded, converted and written to the dataset of **your own run** —
nothing is sent to an AI provider or any other third party, and nothing is kept
outside your Apify storage. Only public http(s) addresses are fetched; addresses
inside private networks are refused.

### FAQ

**Why not just use an LLM to read PDFs?**
Cost and privacy. This actor converts 1,000 pages for $0.50 without sending a
single page to an outside model. Use an LLM afterwards, on clean Markdown.

**Does it work with scanned documents?**
Yes. Pages without a text layer are detected automatically and read with OCR;
`ocrPages` tells you how many. Blank pages and chapter dividers are not sent to
OCR — they are not scans.

**Can I get only the tables?**
Every row has `tables` with the rows of each table. Read that field and ignore
the rest; Excel sheets always appear there.

**How large can files be?**
50 MB by default, up to 500 MB with `max_file_mb`, and 500 pages per document by
default.

**Does it keep the crawler's data?**
Yes — with `crawler_dataset`, every record keeps the crawler's `url` and `crawl`
fields, so you can join it with the crawl's web pages.

### Auf Deutsch: PDF, Word und Excel in Markdown umwandeln

Dieser Actor wandelt **PDF-, Word-, Excel- und PowerPoint-Dateien** in sauberes
**Markdown** und reinen Text um — für KI-Anwendungen, RAG und Vektordatenbanken.
Tabellen kommen als Markdown-Tabellen und als JSON-Zeilen heraus, eingescannte
Seiten werden automatisch per **Texterkennung (OCR)** gelesen, auch auf Deutsch.

Er ergänzt den Website Content Crawler: Dessen heruntergeladene Dateien bleiben
dort leer — hier werden sie gefüllt. **Ohne KI-Dienst eines Drittanbieters:** Die
Dokumente verlassen den eigenen Apify-Lauf nicht. Ausgabe wahlweise je Dokument,
je Seite oder als fertige Abschnitte für RAG.

### Related actors

| Actor | What it does |
|---|---|
| [Website Content Crawler](https://apify.com/apify/website-content-crawler) | Crawls websites and downloads their files — this actor reads them |
| [Imprint & Impressum Scraper](https://apify.com/perceptr0n/imprint-impressum-website-contact-scraper) | Company data from the Impressum of any DACH website |
| [Google Ads Transparency Scraper](https://apify.com/perceptr0n/google-ads-transparency-scraper) | Which companies advertise on Google, how many ads, who runs them |
| [TED Tenders API](https://apify.com/perceptr0n/eu-tenders-ted-scraper) | EU public tenders — their documents convert here |
| [EUDAMED Scraper](https://apify.com/perceptr0n/eudamed-medical-device-scraper) | EU medical devices, manufacturers and certificates |
| [DACH & EU Data Source Finder](https://apify.com/perceptr0n/dach-eu-data-source-finder) | Free: tells you which ready-made scraper covers your data source |

### Support

Bug reports are welcome in the **Issues** tab.

# Actor input Schema

## `urls` (type: `array`):

Links to PDF, Word (.docx), Excel (.xlsx, .xls), PowerPoint (.pptx), HTML, CSV, JSON, XML or text files — one per line.

## `crawler_dataset` (type: `string`):

Convert the files a **Website Content Crawler** run downloaded (option 'Save files'). The crawler stores them with an empty `text` and `markdown`; this fills them in, with the same `url`, `crawl` and `metadata` fields, so the results merge straight into your crawl. Paste the run ID or its dataset ID.

## `output` (type: `string`):

**Documents**: one row per file with the full Markdown. **Pages**: one row per page — handy for citations. **Chunks**: ready-made pieces for vector databases and RAG, with the pages each chunk covers.

## `ocr` (type: `string`):

**Auto**: pages without a text layer (scans, photos of paper) are read with OCR and charged as OCR pages; all others are read directly. **Off**: never OCR. **Force**: OCR every page, for PDFs whose text layer is broken.

## `ocr_languages` (type: `array`):

Languages of the scanned documents. Fewer languages recognise faster and more accurately.

## `extract_tables` (type: `boolean`):

Detects tables in PDFs and returns them as Markdown tables in place and as structured rows in `tables`. Excel and Word tables are always included.

## `chunk_size` (type: `integer`):

Only for output 'RAG chunks'. Chunks break at paragraph boundaries; a table stays in one piece where it fits.

## `chunk_overlap` (type: `integer`):

Only for output 'RAG chunks'. Whole paragraphs from the end of one chunk are repeated at the start of the next, up to this length.

## `max_pages` (type: `integer`):

Longer PDFs are cut off here and say so in `warnings`. Protects your budget from a surprise 3,000-page file.

## `max_file_mb` (type: `integer`):

Larger files are skipped with a message, and not charged.

## Actor input object example

```json
{
  "urls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "output": "documents",
  "ocr": "auto",
  "ocr_languages": [
    "eng",
    "deu"
  ],
  "extract_tables": true,
  "chunk_size": 2000,
  "chunk_overlap": 200,
  "max_pages": 500,
  "max_file_mb": 50
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `resultsCsv` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("perceptr0n/pdf-to-markdown-converter").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://www.irs.gov/pub/irs-pdf/fw9.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("perceptr0n/pdf-to-markdown-converter").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ]
}' |
apify call perceptr0n/pdf-to-markdown-converter --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,perceptr0n/pdf-to-markdown-converter"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/2C6gJTOxp2XzkRhx0/builds/KkkG2rhCROLSTVdwI/openapi.json
