# Document & PDF to Markdown for LLMs: Word, Excel, OCR (`magenta_waterwheel/document-to-markdown`) Actor

Convert PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX), CSV and HTML files into clean, LLM-ready Markdown and JSON. Keeps headings, lists and tables, outputs per-page chunks for RAG and vector databases, and OCRs scanned PDFs and images. Failed files are free. $3 per 1,000 pages.

- **URL**: https://apify.com/magenta_waterwheel/document-to-markdown.md
- **Developed by:** [Huss](https://apify.com/magenta_waterwheel) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 page extracteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Document & PDF to Markdown for LLMs: Word, Excel, OCR

**Document & PDF to Markdown** is a **PDF to Markdown converter and document parser API** for AI. It converts **PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX), CSV and HTML** files into clean, **LLM-ready Markdown and JSON**. Give it document URLs and get back structured Markdown with headings, lists and **tables preserved**, **per-page chunks** ready for RAG, document metadata, and **OCR for scanned PDFs and images**.

It's made for AI agents, RAG pipelines, vector databases and anyone who needs document text in a form an LLM can read. You can call it from your own code, Zapier, Make, n8n, LangChain, LlamaIndex or straight from AI agents through the **Apify MCP server**.

### Why use this PDF to Markdown converter?

| | Document & PDF to Markdown | PDF-only text extractors | Self-hosted parsers (PyMuPDF, MarkItDown, Docling) |
| --- | --- | --- | --- |
| PDF, Word, PowerPoint, Excel, CSV and HTML | **All in one Actor** | PDF only | One library per format |
| Tables kept as Markdown tables | Yes | Often flattened to text | Depends on the library |
| OCR for scanned pages | Automatic, only where needed | Sometimes | Extra setup (Tesseract) |
| RAG-ready per-page chunks and metadata | Yes | Partly | You write the code |
| Servers, memory and scaling | Handled by Apify | Handled | Your job |
| Failed or empty documents | **Not charged** | Varies | Still cost compute |
| Price | **$3 per 1,000 pages**, pay as you go | Per file or per page | Your infrastructure |

### Who uses it?

- **AI and RAG developers** loading PDFs, contracts, manuals and reports into vector databases.
- **AI agents** that need to read a document link a user shares, via the Apify MCP server.
- **Legal, finance and research teams** turning filings, papers and scanned contracts into searchable text.
- **Data teams** extracting tables from PDFs and spreadsheets into a consistent format.
- **No-code builders** adding document reading to Make, Zapier or n8n workflows.

### What can this document converter do?

- 📄 **PDF to Markdown** with headings, paragraphs, bullet and numbered lists, and **tables as Markdown tables**, including multi-column layouts read in the right order.
- 🔎 **OCR for scanned PDFs and images** (PNG, JPEG, TIFF). In Auto mode OCR runs only on pages that have no usable text layer, so you pay the OCR price only where it's needed. English, German, French, Spanish, Portuguese, Italian and Dutch.
- 📝 **Word to Markdown**: headings, lists, bold and italic text, links and tables from DOCX files.
- 📊 **PowerPoint to Markdown**: one section per slide with the slide title, text, tables, chart data and speaker notes.
- 📈 **Excel and CSV to Markdown tables**, one heading per sheet, split into chunks that repeat the header row.
- 🌐 **HTML to Markdown**: keeps the main article content and drops menus, sidebars, language lists and footers (or converts the whole page if you prefer).
- ✂️ **Page ranges** like `1-5, 8, 20-` and a per-document page cap, so you only pay for what you need.
- 🧩 **RAG-ready chunks**: every document comes with a `pages` array, or switch to one dataset item per page.
- 🏷️ **Metadata**: title, author, dates, page count, word count, table count, file size and a SHA-256 hash for de-duplication.
- 🧹 **Removes repeating headers, footers and page numbers** from PDFs.
- 🔗 **Google Drive, Google Docs and Dropbox share links** are turned into download links automatically.
- 🚦 **Failed documents are free.** Broken links, unsupported files and empty documents are reported with a clear reason and never charged.

### Supported formats

| Format | Extensions | What counts as one page |
| --- | --- | --- |
| PDF (text) | `.pdf` | One PDF page |
| PDF (scanned) and images | `.pdf`, `.png`, `.jpg`, `.tiff`, `.webp`, `.bmp`, `.gif` | One OCR page per page or image frame |
| PowerPoint | `.pptx` | One slide |
| Word | `.docx` | One section of up to 3,000 characters of output |
| Excel and CSV | `.xlsx`, `.csv`, `.tsv` | One section of up to 3,000 characters of output |
| HTML, text and Markdown | `.html`, `.txt`, `.md` | One section of up to 3,000 characters of output |

Legacy `.doc`, `.ppt` and `.xls` files and OpenDocument files aren't supported. Save them as DOCX, PPTX, XLSX or PDF first.

### How to convert documents to Markdown

1. Click **Try for free**.
2. Add your **Document URLs**: type them in, paste a list, or upload a text file with one link per line.
3. Optionally set a **Page range**, the **OCR mode**, or **One item per page** output.
4. Click **Start**. Each document becomes a dataset item with its Markdown, metadata and per-page chunks.
5. Download the results as JSON, CSV or Excel, or read them through the API.

### Input example

```json
{
    "startUrls": [
        { "url": "https://arxiv.org/pdf/1706.03762" },
        { "url": "https://example.com/files/quarterly-report.docx" }
    ],
    "pageRange": "",
    "maxPagesPerDocument": 1000,
    "ocrMode": "auto",
    "ocrLanguages": ["eng"],
    "outputMode": "document",
    "includePages": true
}
```

### Output example

One item per document (default). Fields are the same for every format, so your pipeline can rely on a stable schema:

```json
{
    "url": "https://arxiv.org/pdf/1706.03762",
    "status": "success",
    "error": null,
    "fileName": "1706.03762v7.pdf",
    "format": "pdf",
    "contentType": "application/pdf",
    "fileSizeBytes": 2215244,
    "sha256": "bdfaa68d8984f0dc02beaca527b76f207d99b666d31d1da728ee0728182df697",
    "title": "Attention Is All You Need",
    "author": null,
    "subject": null,
    "createdAt": "2024-04-10T21:11:43",
    "modifiedAt": "2024-04-10T21:11:43",
    "pageUnit": "page",
    "pageCount": 15,
    "pagesExtracted": 15,
    "ocrPages": 0,
    "tableCount": 8,
    "wordCount": 6358,
    "charCount": 40032,
    "truncated": false,
    "warnings": [],
    "markdownFileUrl": null,
    "markdown": "## Provided proper attribution is provided, Google hereby grants permission ...\n\n# Attention Is All You Need\n\n| Ashish Vaswani ∗ | Noam Shazeer ∗ | Niki Parmar ∗ | Jakob Uszkoreit ∗ |\n| --- | --- | --- | --- |\n| Google Brain | Google Brain | Google Research | Google Research |\n\n...\n\n## Abstract\n\nThe dominant sequence transduction models are based on ...",
    "pages": [
        { "pageNumber": 1, "markdown": "## Provided proper attribution ...", "wordCount": 420, "isOcr": false, "tableCount": 2 }
    ],
    "chargedPages": 15,
    "chargedOcrPages": 0,
    "extractedAt": "2026-10-08T15:44:47Z"
}
```

With **Output mode: One item per page**, each page is its own dataset item with `pageNumber`, `isOcr`, `pageWordCount` and that page's `markdown`, plus the document fields. This is the easiest way to load chunks into Pinecone, Qdrant, Weaviate, pgvector or any other vector store.

Documents that fail have `"status": "error"` and an `error` such as `Download failed with HTTP 404.` or `Legacy Excel .xls files are not supported.` A `RUN_SUMMARY` record in the key-value store lists all failures and charged pages.

### How much does it cost to convert PDF to Markdown?

This Actor uses **pay-per-event** pricing. You pay only for pages that were actually extracted:

| Event | Price |
| --- | --- |
| Page extracted (text PDF page, slide, or 3,000-character section) | **$3.00 per 1,000 pages** ($0.003 each) |
| OCR page (scanned PDF page or image) | **$8.00 per 1,000 OCR pages** ($0.008 each) |
| Actor start | $0.00005 per run |

Examples:

- A 15-page research paper costs **$0.045**.
- 1,000 text PDF pages cost **$3**.
- A 20-slide PowerPoint costs **$0.06**.
- A 10-page Word document (about 30,000 characters) costs about **$0.03**.
- A 50-page scanned contract costs **$0.40**.

OCR pages are charged only at the OCR price, not both. Blank pages, failed downloads, unsupported files and documents with no extractable text are **free**. Use **Page range** and **Max pages per document** to control costs, and set **Maximum cost per run** in the run options. The Actor stops cleanly when the limit is reached.

### Tips

- **Large documents**: the default 1 GB of memory handles most files. For PDFs with hundreds of pages or long OCR jobs, use 2–4 GB. It runs faster, and the price per page stays the same.
- **Scanned PDFs**: keep **OCR mode** on Auto. Use Always if a PDF has a broken or garbled text layer, and Never if you only want embedded text.
- **Fewer tokens**: turn off **Keep links** and **Include per-page chunks** if you only need the full Markdown.
- **Spreadsheets** are converted to Markdown tables with the first row as the header. Huge sheets produce many sections, so use **Max pages per document** to cap them.
- **Files in your own storage**: any URL that downloads the file works, including signed S3 or GCS URLs and Apify key-value store record URLs.

### Use it with AI agents (MCP) and the API

Add this Actor as a tool in Claude, ChatGPT, Cursor, VS Code or any other MCP client through the [Apify MCP server](https://docs.apify.com/mcp):

```text
https://mcp.apify.com?tools=magenta_waterwheel/document-to-markdown
```

Then just ask, for example: *"Convert this PDF to Markdown and summarize the tables: https://arxiv.org/pdf/1706.03762"* The agent fills in the input, runs the Actor and reads the results. The Actor runs with **limited permissions**, so it can only access its own run storage.

To get results in a single HTTP request, call the synchronous endpoint with your [Apify API token](https://console.apify.com/settings/integrations):

```bash
curl -X POST "https://api.apify.com/v2/acts/magenta_waterwheel~document-to-markdown/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"startUrls":[{"url":"https://arxiv.org/pdf/1706.03762"}],"pageRange":"1-3"}'
```

You can also use the official [Python](https://docs.apify.com/api/client/python) and [JavaScript](https://docs.apify.com/api/client/js) clients, or the ready-made code on the **API** tab.

#### LangChain, LlamaIndex and vector databases

Use the Apify dataset loader in LangChain or LlamaIndex and map `markdown` to the document text and the other fields to metadata. With **Output mode: One item per page**, each dataset item is a ready-made chunk for Pinecone, Qdrant, Weaviate, pgvector or Chroma.

### Integrations and scheduling

- ⏰ **Schedule runs** hourly, daily or weekly in Apify Console, with no server to maintain.
- 🔗 **Send results anywhere**: Google Sheets, Slack, Google Drive, Airbyte, webhooks, or no-code tools such as **Make, Zapier and n8n**.
- 📤 **Export** the dataset as JSON, CSV, Excel, XML, RSS or HTML.
- 📈 **Monitoring**: get notified if a run fails, and see every run's log and cost in Console.

### FAQ

#### Is there a free trial?

Yes. Apify's Free plan includes **$5 of usage credit every month**, enough for about **1,600 text pages** (or about 600 OCR pages) with this Actor. No credit card is needed to start.

#### How do I convert a PDF to Markdown?

Paste the PDF's URL into **Document URLs** and click **Start**. The `markdown` field holds the whole document, and `pages` holds one chunk per page. To convert a file from your computer, upload it to any storage that gives you a download link (Google Drive, Dropbox, S3) and paste that link.

#### Is the PDF table extraction accurate?

Tables with ruling lines are extracted cell by cell. Borderless tables are rebuilt from column alignment, which works well for regular tables but can merge cells in very irregular layouts. Word, PowerPoint, Excel and HTML tables are always exact.

#### How good is the OCR?

OCR uses Tesseract, the most widely used open-source OCR engine, at 200 DPI. It works well on clean scans and printed text. Handwriting, very low-resolution scans and complex forms give weaker results.

#### Can it read password-protected or DRM-locked PDFs?

No. Encrypted PDFs that need a password are reported as failed and are not charged.

#### Does it crawl websites?

No. This Actor converts the documents you link to. To crawl a website, use a web crawler and pass the document links it finds to this Actor.

#### Is there a file size limit?

The default limit is 50 MB per file and you can raise it to 200 MB. Very large files may need more memory.

#### Do you store my documents?

Documents are processed in your own Actor run. The output stays in your run's storage under your Apify data retention settings, and nothing is kept elsewhere.

#### I found a bug or need a feature

Open an issue in the **Issues** tab with a link to your run. We aim to reply within one business day.

### More tools from the same developer

| Actor | What it does | Price |
| --- | --- | --- |
| [Career Site Jobs API](https://apify.com/magenta_waterwheel/career-site-jobs-api) | Jobs from Greenhouse, Lever, Ashby and SmartRecruiters career sites, with salaries | $2 / 1,000 jobs |
| [Website Screenshot & PDF API](https://apify.com/magenta_waterwheel/website-screenshot-pdf) | Full-page PNG/JPEG screenshots and web page to PDF in bulk | $2.50 / 1,000 screenshots |
| [App Store Reviews Scraper](https://apify.com/magenta_waterwheel/app-store-reviews-details) | Apple App Store reviews and app details in any country | $0.25 / 1,000 reviews |

# Actor input Schema

## `startUrls` (type: `array`):

Direct links to documents: PDF, DOCX, PPTX, XLSX, CSV, HTML, TXT, Markdown, or images (PNG, JPEG, TIFF) for OCR. Google Drive, Google Docs and Dropbox share links are converted to download links automatically. You can paste many links at once or upload a text file.

## `pageRange` (type: `string`):

Pages to extract, for example "1-5, 8, 20-". Applies to PDF pages, PowerPoint slides and the sections of other formats. Leave empty for all pages.

## `maxPagesPerDocument` (type: `integer`):

Stop after this many pages of each document. Protects you from paying for unexpectedly huge files.

## `ocrMode` (type: `string`):

Auto runs OCR only on pages with no usable text layer (scanned pages, broken fonts). Always runs OCR on every page. Never skips OCR. OCR pages are billed at the OCR page price.

## `ocrLanguages` (type: `array`):

Languages in the scanned documents. Adding more languages makes OCR slower; pick only those you need.

## `includeTables` (type: `boolean`):

Convert tables in PDFs to Markdown tables. Word, PowerPoint, Excel and HTML tables are always kept.

## `removeHeadersFooters` (type: `boolean`):

Drop text that repeats at the top or bottom of most PDF pages, such as running titles and page numbers.

## `includeSpeakerNotes` (type: `boolean`):

Add PowerPoint speaker notes under each slide.

## `includeLinks` (type: `boolean`):

Keep hyperlinks as Markdown links in Word and HTML documents. Turn off for plain text with fewer tokens.

## `htmlMainContentOnly` (type: `boolean`):

For HTML pages, keep only the main article or content area and drop menus, sidebars, language lists and footers. Turn off to convert the whole page.

## `outputMode` (type: `string`):

One dataset item per document (full Markdown plus a per-page array), or one item per page, which suits vector databases and RAG pipelines.

## `includePages` (type: `boolean`):

In document mode, add a pages array with the Markdown of each page. Turn off to get only the full Markdown.

## `saveMarkdownFiles` (type: `boolean`):

Also save each document as a .md file in the key-value store and add a download link (markdownFileUrl).

## `maxFileSizeMb` (type: `integer`):

Skip files larger than this. Files above 50 MB may need more memory.

## `maxConcurrency` (type: `integer`):

How many documents to download and process at the same time.

## `downloadTimeoutSecs` (type: `integer`):

Give up on a download after this many seconds without progress.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://arxiv.org/pdf/1706.03762"
    },
    {
      "url": "https://raw.githubusercontent.com/microsoft/markitdown/main/packages/markitdown/tests/test_files/test.docx"
    }
  ],
  "pageRange": "",
  "maxPagesPerDocument": 1000,
  "ocrMode": "auto",
  "ocrLanguages": [
    "eng"
  ],
  "includeTables": true,
  "removeHeadersFooters": true,
  "includeSpeakerNotes": true,
  "includeLinks": true,
  "htmlMainContentOnly": true,
  "outputMode": "document",
  "includePages": true,
  "saveMarkdownFiles": false,
  "maxFileSizeMb": 50,
  "maxConcurrency": 2,
  "downloadTimeoutSecs": 60
}
```

# Actor output Schema

## `documents` (type: `string`):

No description

## `files` (type: `string`):

No description

## `runSummary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://arxiv.org/pdf/1706.03762"
        },
        {
            "url": "https://raw.githubusercontent.com/microsoft/markitdown/main/packages/markitdown/tests/test_files/test.docx"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("magenta_waterwheel/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        { "url": "https://arxiv.org/pdf/1706.03762" },
        { "url": "https://raw.githubusercontent.com/microsoft/markitdown/main/packages/markitdown/tests/test_files/test.docx" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("magenta_waterwheel/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://arxiv.org/pdf/1706.03762"
    },
    {
      "url": "https://raw.githubusercontent.com/microsoft/markitdown/main/packages/markitdown/tests/test_files/test.docx"
    }
  ]
}' |
apify call magenta_waterwheel/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,magenta_waterwheel/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/syhcVYROTsyRdot5v/builds/jnqUXOadzcgEFXFsn/openapi.json
