# Document to Markdown - PDF, Word, Excel, PowerPoint to Text (`adaptive_arbor_rtz/document-to-markdown`) Actor

Convert PDF, DOCX, XLSX, PPTX, HTML and CSV files from URLs into clean Markdown or plain text for AI, RAG and LLM pipelines. Tables kept, pages split, pay per document.

- **URL**: https://apify.com/adaptive\_arbor\_rtz/document-to-markdown.md
- **Developed by:** [Björn Ólafur](https://apify.com/adaptive_arbor_rtz) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.00 / 1,000 documents

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Document to Markdown do?

**Document to Markdown** converts **PDF, Word (DOCX), Excel (XLSX), PowerPoint (PPTX), HTML, CSV and text files** from any URL into **clean Markdown or plain text**. Headings, lists and **tables are kept**, documents can be **split into pages, slides or sheets**, and you **pay only for documents that convert successfully**.

It's built for **AI agents, RAG pipelines and LLM workflows**: give it links, get back text that's ready to chunk, embed or summarise. Run it in Apify Console, call it from the API, schedule it, or use it from your AI agent through the Apify MCP server. Connect it to Make, Zapier, n8n, LangChain or LlamaIndex with Apify integrations.

### Why use Document to Markdown?

- 📄 **One tool for every office format**: PDF, DOCX, XLSX, PPTX, HTML, CSV, TXT, Markdown and JSON.
- 🔎 **OCR built in**: **scanned PDFs and images** (PNG, JPEG, WebP, TIFF) become searchable text, in 100+ languages. In auto mode OCR runs only on pages that need it.
- 🧠 **LLM-ready Markdown**: headings, bullet lists and GitHub-style tables, so models understand the structure.
- ✂️ **Page-level output**: optional `pages` array (one per PDF page, slide or sheet) for citations and chunking.
- 🔍 **Automatic file-type detection** from the file content, so links without an extension (download links, `?id=` links) work.
- 🗒️ **Speaker notes** from PowerPoint slides and **all sheets** from Excel workbooks.
- 💸 **Fair pricing**: broken links, oversized files and unsupported formats are free.
- ⚡ **Fast and cheap**: no browser, 10 documents in parallel by default.

**Typical uses:** feeding reports, contracts, manuals and research papers to ChatGPT/Claude, building RAG knowledge bases, extracting tables from Excel files on websites, indexing slide decks, and monitoring published PDFs such as price lists, tenders or annual reports.

### How to convert PDF, Word, Excel or PowerPoint to Markdown

1. Click **Try for free**.
2. Paste document links into **Document URLs**, one per line.
3. Choose **Markdown** or **Plain text**.
4. Click **Start**. Results appear in the **Output** tab within seconds.
5. Download them as JSON, CSV or Excel, or fetch them via the API.

### Input

Only **Document URLs** is required. See the **Input** tab for all options:

| Option            | What it does                                         | Default    |
| ----------------- | ---------------------------------------------------- | ---------- |
| `urls`            | Links to documents                                   | –          |
| `outputFormat`    | `markdown` or `text`                                 | `markdown` |
| `includePages`    | Add a `pages` array (per page, slide or sheet)       | `false`    |
| `ocr`             | `auto` (scanned pages and images), `always` or `off` | `auto`     |
| `ocrLanguages`    | e.g. `eng`, `deu`, `isl`, `eng+fra`                  | `eng`      |
| `maxOcrPages`     | OCR page limit per document                          | 50         |
| `maxPages`        | Page/slide limit per document                        | 1000       |
| `maxRowsPerSheet` | Row limit per Excel sheet or CSV                     | 5000       |
| `maxFileSizeMb`   | Skip larger files                                    | 50 MB      |

```json
{
    "urls": ["https://arxiv.org/pdf/1706.03762", "https://calibre-ebook.com/downloads/demos/demo.docx"],
    "outputFormat": "markdown",
    "includePages": true
}
```

### Output

One item per document:

```json
{
    "url": "https://arxiv.org/pdf/1706.03762",
    "success": true,
    "fileName": "1706.03762",
    "fileType": "pdf",
    "title": null,
    "pageCount": 15,
    "wordCount": 6477,
    "characterCount": 39790,
    "format": "markdown",
    "content": "## Page 1\n\nAttention Is All You Need\n\nAshish Vaswani ...",
    "contentTruncated": false,
    "fullContentUrl": null,
    "pages": ["Attention Is All You Need ...", "..."],
    "metadata": { "creator": "LaTeX with hyperref", "createdAt": "2024-04-10" },
    "fileSizeBytes": 2215244,
    "convertedAt": "2026-09-30T00:40:00.000Z",
    "error": null
}
```

Failed documents are listed too, with `success: false` and the reason in `error`, and are not charged. Very long documents (over 2 million characters) are cut in the dataset item, and the full text is saved as a file linked in `fullContentUrl`. You can download the dataset in various formats such as JSON, HTML, CSV or Excel.

### How much does it cost to convert documents?

**$4 per 1,000 documents** ($0.004 each) plus **$0.002 per run**, whatever the size or format. **OCR costs $0.01 per recognised page**, charged only for scanned pages and images that actually need it (the `ocrPages` field shows how many). Failed documents are free and platform usage is included. With Apify's free plan you can convert a few hundred documents a month at no cost. You can set a **maximum cost per run**; the Actor stops cleanly when it's reached.

### Tips

- Use **Split into pages** when you need page numbers for citations or want to chunk by page.
- Use **Plain text** for search indexing and **Markdown** for LLMs.
- Scanned PDFs (images of text) have no text layer. With OCR on `auto` they are recognised automatically; with OCR `off` they are marked `metadata.scannedOrImageOnly: true`.
- Set **OCR languages** to the document's language for best accuracy (`metadata.ocrConfidence` shows 0–100).
- Use **Max OCR pages** to cap cost on long scanned books.
- Old binary formats (.doc, .xls, .ppt) aren't supported. Save them as .docx, .xlsx or .pptx first.
- If a server blocks data-centre downloads, enable **Proxy**.

### FAQ

**Does it work with files behind a login?** No, the links must be publicly downloadable.

**Can I upload files instead of links?** Put the file anywhere with a public link (for example an Apify key-value store, S3 or Google Drive direct-download link) and pass that link.

**Is it legal?** Converting documents you're allowed to access is fine. You're responsible for respecting copyright and the source's terms.

**Something not working?** Open an issue in the **Issues** tab with the document link and I'll fix it quickly. Custom versions are available on request.

### Related tools

- [Website Screenshot](https://apify.com/adaptive_arbor_rtz/website-screenshot): full-page screenshots and PDFs of any website, cookie banners removed
- [Sitemap Extractor](https://apify.com/adaptive_arbor_rtz/sitemap-extractor): every URL of a website from its sitemaps, plus a broken-link check
- [RSS Feed Reader](https://apify.com/adaptive_arbor_rtz/rss-feed-reader): read and monitor any RSS/Atom feed, only new items, full article text

# Actor input Schema

## `urls` (type: `array`):

Links to documents, one per line. Supported: PDF, Word (.docx), Excel (.xlsx), PowerPoint (.pptx), HTML pages, CSV, TXT, Markdown and JSON, plus scanned PDFs and images (PNG, JPEG, WebP, TIFF) via OCR. The file type is detected from the content, so links without a file extension work too.

## `outputFormat` (type: `string`):

markdown keeps headings, lists and tables (best for AI/LLM and RAG pipelines). text is plain text without formatting.

## `includePages` (type: `boolean`):

Also return a 'pages' array with one entry per PDF page, slide or spreadsheet sheet. Useful for citations and chunking.

## `ocr` (type: `string`):

auto = recognise text only on PDF pages without a text layer (scanned pages) and in image files (PNG, JPEG, WebP, TIFF, GIF, BMP). always = OCR every PDF page. off = never. OCR is charged per recognised page.

## `ocrLanguages` (type: `string`):

Language(s) of the scanned text, e.g. 'eng', 'deu', 'fra', 'isl', or several like 'eng+deu'. English names such as 'German' work too. Default English.

## `maxOcrPages` (type: `integer`):

Stop OCR after this many pages per document, to cap cost on long scans.

## `maxPages` (type: `integer`):

Stop after this many PDF pages or slides per document.

## `maxRowsPerSheet` (type: `integer`):

Stop after this many rows per Excel sheet or CSV file.

## `maxFileSizeMb` (type: `integer`):

Skip larger files (not charged).

## `timeoutSecs` (type: `integer`):

Give up on a download after this many seconds (not charged).

## `maxConcurrency` (type: `integer`):

How many documents to process at the same time.

## `proxyConfiguration` (type: `object`):

Optional. Use a proxy if a server blocks data-centre downloads.

## Actor input object example

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://calibre-ebook.com/downloads/demos/demo.docx"
  ],
  "outputFormat": "markdown",
  "includePages": false,
  "ocr": "auto",
  "ocrLanguages": "eng",
  "maxOcrPages": 50,
  "maxPages": 1000,
  "maxRowsPerSheet": 5000,
  "maxFileSizeMb": 50,
  "timeoutSecs": 120,
  "maxConcurrency": 10,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://calibre-ebook.com/downloads/demos/demo.docx"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("adaptive_arbor_rtz/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://calibre-ebook.com/downloads/demos/demo.docx",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("adaptive_arbor_rtz/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://calibre-ebook.com/downloads/demos/demo.docx"
  ]
}' |
apify call adaptive_arbor_rtz/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,adaptive_arbor_rtz/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Dr2HDQQkAkXhfic0P/builds/wrqWziLO79tu67rcS/openapi.json
