# Document to Markdown for LLMs (PDF, DOCX, XLSX, PPTX, HTML) (`rod_analytics/doc-to-markdown`) Actor

Convert PDF, Word, Excel, PowerPoint, HTML and scanned documents to clean Markdown with OCR, metadata and RAG ready chunks for LLM ingestion and AI agents.

- **URL**: https://apify.com/rod\_analytics/doc-to-markdown.md
- **Developed by:** [Rod Services](https://apify.com/rod_analytics) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Document to Markdown for LLMs do?

**Document to Markdown for LLMs** converts **PDF, Word (DOCX), Excel (XLSX, XLS), PowerPoint (PPTX), HTML, EPUB, CSV and scanned images** into clean **Markdown** that large language models read well. It keeps headings, lists and tables, removes running headers and page numbers, and runs **OCR on scanned pages** automatically.

Every document also comes back as **RAG ready chunks** with page numbers and section headings. Feed them straight into a vector database, an embeddings pipeline or an **AI agent**.

Paste a file URL and click Start. The example PDF converts in about 5 seconds. As an Apify Actor you also get an API, scheduling, webhooks, integrations with Make, Zapier, n8n and LangChain, and run monitoring.

### Why use this PDF to Markdown converter?

- **LLM ingestion and RAG.** Turn document libraries into chunks for pgvector, Pinecone, Qdrant, Weaviate or Chroma.
- **AI agents and MCP tools.** Give an agent one call that reads any PDF, DOCX or web page as Markdown.
- **Scanned PDFs and images.** Tesseract OCR in 35 languages runs only on pages that have no text layer.
- **Tables survive.** PDF tables, Word tables and Excel sheets become Markdown tables.
- **Two column PDFs.** Columns are read in order, not line by line across the page.
- **Metadata.** Title, author, dates, page count, word count and HTML meta tags.
- **Fair pricing.** You pay per converted document and per OCR page. Failed files are free.

### How to convert PDF to Markdown

1. Open the **Input** tab.
2. Add one or more links in **File URLs**, or upload a file with **Upload a file**.
3. Pick an **OCR mode**. Keep `auto` unless you know the files are scans.
4. Set **Chunk size** for your embedding model. 500 to 1,000 tokens works for most.
5. Click **Start**. Results appear in the **Output** tab.
6. Download JSON, CSV or Excel, or call the API from your pipeline.

Google Drive, Google Docs, Sheets and Slides, Dropbox and GitHub share links are turned into direct downloads for you. The file must be shared with "Anyone with the link". Google Docs are exported as DOCX, Sheets as XLSX and Slides as PPTX.

### Input

All fields are on the Input tab. Example:

```json
{
    "fileUrls": [
        "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf",
        "https://example.com/report.docx"
    ],
    "ocr": "auto",
    "languages": ["eng", "deu"],
    "chunkSize": 800,
    "chunkOverlap": 80,
    "includeImages": false,
    "maxPages": 0
}
```

| Field | Default | What it does |
| --- | --- | --- |
| `fileUrls` | | Direct links to documents. Apify key-value store record URLs work too. |
| `uploadedFile` | | One file uploaded from your computer in Console. |
| `keyValueStoreRecords` | `[]` | Record keys in this run's key-value store, or `storeId/key`. Useful when another Actor calls this one. |
| `ocr` | `auto` | `auto` OCRs only scanned pages. `force` OCRs every page. `off` never OCRs. |
| `languages` | `["eng"]` | Tesseract codes, for example `eng`, `deu`, `fra`, `spa`, `lit`, `pol`, `rus`, `jpn`, `chi_sim`. |
| `chunkSize` | `1000` | Approximate tokens per chunk. `0` turns chunking off. |
| `chunkOverlap` | `100` | Approximate tokens repeated between chunks. |
| `includeImages` | `false` | Save embedded images to the key-value store and link them in the Markdown. |
| `maxPages` | `0` | Only the first N pages or slides. `0` is all, but text PDFs stop at 300 pages. |
| `saveMarkdownFiles` | `false` | Also save each document as a `.md` file. |
| `maxFileSizeMb` | `100` | Skip bigger downloads. |
| `maxConcurrency` | `2` | Documents converted in parallel. |

### Output

One dataset item per document. Shortened example:

```json
{
    "sourceUrl": "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf",
    "fileName": "USCODE-2011-title17-chap1-sec107.pdf",
    "mimeType": "application/pdf",
    "pages": 5,
    "title": "§107. Limitations on exclusive rights: Fair use",
    "markdown": "by two or more authors, a waiver of rights under this paragraph made by one such author waives such rights for all such authors.\n\n(2) Ownership of the rights conferred by subsection (a)...",
    "chunks": [
        {
            "index": 0,
            "text": "by two or more authors, a waiver of rights under this paragraph...",
            "tokensApprox": 845,
            "page": 1,
            "heading": null
        }
    ],
    "chunksCount": 13,
    "metadata": {
        "creator": "Federal Digital System, U. S. Government Publishing Office",
        "createdAt": "2019-10-14T08:51:38+00:00",
        "pdfVersion": "1.5"
    },
    "ocrPagesCount": 0,
    "warnings": [],
    "error": null
}
```

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel. The Output tab has three views: **Overview**, **Markdown** and **RAG chunks** with one row per chunk.

Documents longer than 200,000 characters keep a truncated copy in the dataset. The full text is saved as a `.md` file in the key-value store, linked in `markdownUrl`.

### Data fields

| Field | Description |
| --- | --- |
| `sourceUrl` | Original link, or the key-value store record URL |
| `fileName` | File name from the URL or the download headers |
| `mimeType` | Detected type, checked against the file content |
| `pages` | Pages for PDF, slides for PPTX, frames for images |
| `title` | Document title from metadata or the first heading |
| `markdown` | Clean Markdown of the whole document |
| `chunks` | `index`, `text`, `tokensApprox`, `page`, `heading` |
| `metadata` | Author, subject, keywords, created and modified dates, word count, HTML meta |
| `ocrPagesCount` | Pages that went through OCR |
| `images` | Links to extracted images when `includeImages` is on |
| `markdownUrl` | Full `.md` file for long documents |
| `warnings` | Notes such as pages skipped by `maxPages` |
| `error` | Why a document failed. Failed documents are not charged |

### How much does it cost to convert PDF to Markdown?

The Actor uses pay per event pricing:

- **$0.003 per converted document**
- **$0.01 per OCR page**, only when a page is really scanned or when you force OCR
- **$0.001 per run start** (per GB of memory)

1,000 normal PDFs cost about **$3**. A 100 page scanned book costs about **$1**. Failed downloads, broken files and password protected PDFs cost nothing. Apify's free plan credit covers hundreds of documents each month.

Set **Maximum cost per run** in the run options. The Actor never goes over it. Documents that do not fit are skipped, and scanned pages that do not fit are left out with a warning on the document.

### Tips for better results

- Keep **OCR on `auto`**. Text PDFs are fast and cheap. OCR costs time and money.
- List only the **OCR languages** you need. Each extra language slows OCR down.
- Use **`maxPages`** to preview long files before a full run.
- Text PDFs are fast: about 30 pages per second at 1 GB, table heavy reports about 10.
- For big OCR jobs, give the run **more memory**. 1 GB is the cheapest per page. 4 GB is about 2.7 times faster.
- Tune **chunk size** to your embedding model. Chunks break at headings and paragraphs, and table chunks repeat the table header.
- Call it from **LangChain, LlamaIndex or n8n** through the Apify integration, then embed `chunks[].text`.

### FAQ, limits and support

**Which formats are supported?** PDF, DOCX, XLSX, XLS, PPTX, HTML, EPUB, CSV, JSON, XML, TXT, Markdown, Outlook MSG, PNG, JPG, TIFF, WEBP, BMP and GIF. Legacy DOC and PPT are not supported. Save them as DOCX or PPTX first.

**Does it work with scanned PDFs?** Yes. Pages without a text layer are rendered and read with Tesseract OCR. Handwriting and very low quality scans give weak results.

**Are password protected PDFs supported?** No. They return an error and are not charged. PDFs that only restrict printing or copying open normally.

**What if a link opens a login or error page?** The document fails with a clear error and is not charged. This happens with private Google Drive or Dropbox files and expired links.

**Is there a page limit?** Text PDFs over 300 pages stop at page 300 when `maxPages` is `0`. The document gets a warning. To convert more, set `maxPages` to the page count you need, for example `1000`. You can also split the file. Scanned PDFs have no such limit, since you pay for each OCR page.

**How accurate are PDF headings and tables?** Headings come from font sizes, tables from ruled lines. Complex forms and borderless tables can come out as plain text.

**Is my data safe?** Files are processed inside your own Apify run. Nothing is sent to third party AI services.

**Is it legal?** You must have the right to process the documents you submit.

Found a bug or need a feature? Open an issue in the **Issues** tab. Custom pipelines, other formats and private deployments are available on request.

# Actor input Schema

## `fileUrls` (type: `array`):

Direct links to documents: PDF, DOCX, XLSX, XLS, PPTX, HTML, EPUB, CSV, JSON, XML, TXT, images (PNG, JPG, TIFF, WEBP). Google Drive, Google Docs/Sheets/Slides, Dropbox and GitHub share links are converted to direct downloads. Shared files must be open to anyone with the link. Apify key-value store record URLs work when they are signed (recordPublicUrl).

## `uploadedFile` (type: `string`):

Upload one document from your computer. Apify stores it in a key-value store and passes its URL to the Actor. For many files, use File URLs.

## `keyValueStoreRecords` (type: `array`):

Keys of records that hold files in this run's default key-value store, for example when another Actor calls this one. Use `storeId/key` or `username~store-name/key` for other stores. The Actor runs with limited permissions, so for your own stores prefer the record URL with its signature in File URLs.

## `ocr` (type: `string`):

`auto` runs OCR only on scanned pages that have no text layer. `force` runs OCR on every page. `off` never runs OCR. Each OCR page is billed as an extra event.

## `languages` (type: `array`):

Tesseract language codes for OCR, for example `eng`, `deu`, `fra`, `spa`, `lit`. Several languages slow OCR down. Installed: eng, deu, fra, spa, ita, por, nld, pol, ces, slk, slv, hrv, hun, ron, bul, ell, lit, lav, est, fin, swe, dan, nor, rus, ukr, tur, ara, heb, hin, jpn, kor, chi\_sim, chi\_tra, vie, ind.

## `chunkSize` (type: `integer`):

Approximate chunk size for RAG in tokens (1 token is about 4 characters). Chunks respect headings, paragraphs and tables. Set 0 to skip chunking.

## `chunkOverlap` (type: `integer`):

Approximate overlap between neighbouring chunks in tokens. Capped at half of the chunk size.

## `includeImages` (type: `boolean`):

Extract embedded images from PDF, DOCX and PPTX into the key-value store and link them from the Markdown. HTML image links are kept as absolute URLs. Off by default to keep the text clean for LLMs.

## `maxPages` (type: `integer`):

Process only the first N pages of each PDF, slides of each PPTX or frames of each TIFF. 0 means all pages, except that text PDFs stop at 300 pages. Set a higher number to convert longer text PDFs.

## `saveMarkdownFiles` (type: `boolean`):

Also store every document as a .md file in the key-value store. Documents longer than 200,000 characters are always stored there, and the dataset keeps a truncated copy.

## `maxFileSizeMb` (type: `integer`):

Skip downloads larger than this.

## `maxConcurrency` (type: `integer`):

How many documents are converted in parallel. OCR is CPU heavy, so give the run more memory when you raise this.

## Actor input object example

```json
{
  "fileUrls": [
    "https://www.irs.gov/pub/irs-pdf/fw9.pdf"
  ],
  "keyValueStoreRecords": [],
  "ocr": "auto",
  "languages": [
    "eng"
  ],
  "chunkSize": 1000,
  "chunkOverlap": 100,
  "includeImages": false,
  "maxPages": 0,
  "saveMarkdownFiles": false,
  "maxFileSizeMb": 100,
  "maxConcurrency": 2
}
```

# Actor output Schema

## `documents` (type: `string`):

No description

## `overview` (type: `string`):

No description

## `chunks` (type: `string`):

No description

## `files` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "fileUrls": [
        "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("rod_analytics/doc-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "fileUrls": ["https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf"] }

# Run the Actor and wait for it to finish
run = client.actor("rod_analytics/doc-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "fileUrls": [
    "https://www.govinfo.gov/content/pkg/USCODE-2011-title17/pdf/USCODE-2011-title17-chap1-sec107.pdf"
  ]
}' |
apify call rod_analytics/doc-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rod_analytics/doc-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/k9w4WTTpMWdGZn7BC/builds/0V9uVV6HdhIEfz56h/openapi.json
