# Document to Markdown for AI & RAG (`egra_van/document-to-markdown`) Actor

Convert PDF, Word, PowerPoint, Excel, CSV, HTML and TXT files into clean Markdown or plain text for LLMs and RAG. Keeps headings, lists and tables, splits into token-sized chunks with page and heading metadata, and OCRs scanned PDFs.

- **URL**: https://apify.com/egra\_van/document-to-markdown.md
- **Developed by:** [Argentin Vazdautan](https://apify.com/egra_van) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.70 / 1,000 document page converted to texts

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Document to Markdown for AI & RAG

Convert **PDF, Word (DOCX), PowerPoint (PPTX), Excel (XLSX), CSV, HTML and TXT** files into **clean Markdown or plain text** for **LLMs, ChatGPT, Claude, LangChain, LlamaIndex and RAG pipelines**. Headings, lists and tables are preserved, documents are split into **token-sized chunks with page numbers and section headings**, and **scanned PDFs can be OCR'd**.

Give it a list of links (Google Drive, Dropbox, OneDrive and GitHub links work too) and get back one dataset item per chunk, ready to embed into Pinecone, Qdrant, Weaviate, Chroma, pgvector or any vector database.

### What it does

- 📄 **PDF to Markdown**: paragraphs rebuilt from the text layer, headings detected from font sizes, bullet lists, simple tables as Markdown tables, running headers/footers and page numbers removed, hyphenated words re-joined, multi-column layouts kept in reading order
- 📝 **Word (DOCX) to Markdown**: headings, bold/italic, links, nested lists and tables
- 📊 **PowerPoint (PPTX)**: every slide with its title, bullet levels, tables and **speaker notes**
- 📈 **Excel (XLSX), CSV, TSV**: every sheet as a Markdown table, dates normalised, formula results used
- 🌐 **HTML pages**: main content only (navigation, footers, scripts removed), tables converted
- ✂️ **RAG chunking**: by heading (one section per chunk) or fixed size with overlap, measured in **tokens** (cl100k\_base, the tokenizer of OpenAI embeddings) or characters. Tables that are too big are split by rows with the **header row repeated** in every chunk
- 🏷️ **Metadata per chunk**: source URL, file name, page range, heading path (e.g. `2. Methodology > 2.1 Chunking strategy`), chunk index, token count
- 🔍 **OCR for scanned PDFs and images** (English, Romanian, Russian). Pages that already have text are never OCR'd. Without OCR, scanned pages are detected and reported
- 🔗 **Share links just work**: Google Drive, Google Docs/Sheets/Slides (exported automatically), Dropbox, OneDrive, SharePoint, Box, GitHub
- 🧯 **One bad file doesn't stop the run**: failed documents get a dataset item with a clear error message

### How to use it

1. Paste links to your documents into **Document URLs** (or reference files in a key-value store, or a dataset of URLs from a crawler).
2. Pick **Markdown** or **Plain text**.
3. For RAG, choose **Chunking → By heading** or **Fixed size**, e.g. 800 tokens with 100 tokens overlap.
4. Turn on **OCR** if some PDFs are scans.
5. Click **Start** and download the results as JSON, CSV or Excel, or read them through the API.

### Input example

```json
{
  "urls": [
    { "url": "https://arxiv.org/pdf/1706.03762" },
    { "url": "https://drive.google.com/file/d/FILE_ID/view?usp=sharing" },
    { "url": "https://example.com/handbook.docx" }
  ],
  "outputFormat": "markdown",
  "chunking": "heading",
  "chunkSize": 800,
  "chunkSizeUnit": "tokens",
  "chunkOverlap": 100,
  "includePageNumbers": false,
  "ocr": true,
  "ocrLanguages": ["eng"]
}
```

Other sources:

- `keyValueStoreKeys`: files you uploaded to an Apify key-value store, as `"storeName/key"` or just `"key"` with `keyValueStoreId`
- `datasetId` + `datasetUrlField`: take URLs from another Actor's output (e.g. a website crawler that collected PDF links)

### Output

One dataset item per chunk (or per document when chunking is off):

```json
{
  "sourceUrl": "https://example.com/annual-report.pdf",
  "fileName": "annual-report.pdf",
  "fileType": "pdf",
  "pages": 3,
  "chunkIndex": 2,
  "chunkCount": 9,
  "text": "## 1.1 Key numbers\n\n| Quarter | Revenue | Students |\n| --- | --- | --- |\n| Q1 | 12,500 | 140 |\n| Q2 | 15,200 | 171 |",
  "tokenCount": 71,
  "charCount": 187,
  "documentUrl": "https://api.apify.com/v2/key-value-stores/.../records/0001-annual-report.md",
  "metadata": {
    "title": "Annual Report 2026",
    "totalPages": 3,
    "pagesEstimated": false,
    "pageStart": 1,
    "pageEnd": 1,
    "headings": ["1. Introduction", "1.1 Key numbers"],
    "headingPath": "1. Introduction > 1.1 Key numbers"
  },
  "status": "ok"
}
```

- `documentUrl`: the **full Markdown file** of the document in the run's key-value store (plus `textUrl` in plain-text mode)
- `metadata.ocrPages`: pages that were read with OCR, e.g. `"1-5"`
- `metadata.scannedPagesWithoutText`: scanned pages that were **not** read because OCR was off
- `metadata.warnings`: anything worth knowing (truncated sheets, formulas without saved values, page limits)
- Failed documents: `{ "sourceUrl": "...", "status": "failed", "error": "Download failed: HTTP 403 Forbidden. The file is private: share it as \"Anyone with the link can view\"." }`

The key-value store record `OUTPUT` has a run summary: documents succeeded/failed, pages, OCR pages, chunks and all errors.

### Pricing

**Pay per page.** You pay only for pages that are converted:

| Event | Price | When |
|---|---|---|
| Page processed | $0.0015 ($1.50 per 1,000 pages) | Each PDF page or PowerPoint slide. For formats without pages (DOCX, XLSX, CSV, HTML, TXT) every 3,000 characters of output count as one page |
| OCR page | + $0.005 ($5 per 1,000 pages) | Each scanned page or image read with OCR (only when OCR is on) |

Failed downloads and unsupported files are free. Set a **maximum cost per run** and the Actor converts only what fits in it, then stops cleanly. Use **Max pages per document** to cap long files.

### Supported formats

| Format | Notes |
|---|---|
| PDF | Text-layer PDFs; scanned pages with OCR. Password-protected PDFs are reported as errors |
| DOCX / DOCM | Word 2007+. Save legacy `.doc` as `.docx` |
| PPTX / PPTM | Slides in presentation order, speaker notes optional |
| XLSX / XLSM | All visible sheets, up to 100,000 rows each. Save legacy `.xls` as `.xlsx` |
| CSV / TSV | Delimiter auto-detected; UTF-8 (other encodings best-effort) |
| HTML | Main content extraction |
| TXT / Markdown / JSON / XML | Passed through as text |
| PNG / JPG / TIFF / WebP / BMP | OCR only |

### Use with LangChain, LlamaIndex, n8n, Make

Each item's `text` + `metadata` maps directly to a LangChain `Document(page_content, metadata)` or a LlamaIndex `TextNode`. Use the Apify integration for LangChain/LlamaIndex, or call the Actor from **Make, Zapier or n8n** and send the chunks to your vector database. The `pageStart`, `pageEnd` and `headingPath` fields let your chatbot **cite the exact page and section** of its answer.

### FAQ

**Are tables in PDFs supported?** Tables whose columns are clearly separated are rebuilt as Markdown tables. Complex tables (merged cells, no spacing, tables drawn as images) may come out as plain lines. DOCX, PPTX, XLSX and HTML tables are converted exactly.

**What about scanned PDFs?** Turn on **OCR**. Only pages without a text layer are OCR'd (≈2–3 seconds per page). Without OCR, the Actor tells you which pages are scanned.

**How big can files be?** Up to 100 MB by default (`maxFileSizeMb`, max 500 MB). A 300-page PDF converts in a few seconds.

**My Google Drive link fails.** Share the file as *Anyone with the link can view*. Google Docs, Sheets and Slides are exported to DOCX, XLSX and PPTX automatically.

**Why tokens?** Embedding models and LLM context windows are limited in tokens, not characters. Token counts use `cl100k_base` (OpenAI `text-embedding-3-*`, GPT-4); other tokenizers are usually within ±15%.

**Is my data stored?** Only in your own Apify storage for this run, under your account's retention settings.

# Actor input Schema

## `urls` (type: `array`):

Links to PDF, DOCX, PPTX, XLSX, CSV, TSV, HTML, TXT or Markdown files. Google Drive, Google Docs/Sheets/Slides, Dropbox, OneDrive/SharePoint, Box and GitHub share links are converted to direct downloads automatically (files must be shared as 'Anyone with the link').

## `keyValueStoreKeys` (type: `array`):

Keys of files uploaded to an Apify key-value store. Use 'storeIdOrName/key', or just 'key' together with the store below (default: this run's store).

## `keyValueStoreId` (type: `string`):

Store that holds the keys above.

## `datasetId` (type: `string`):

Apify dataset (e.g. output of a crawler) whose items contain document URLs.

## `datasetUrlField` (type: `string`):

Field of the dataset items that holds the URL.

## `outputFormat` (type: `string`):

Output format.

## `chunking` (type: `string`):

How to split documents into dataset items for embeddings / vector databases.

## `chunkSize` (type: `integer`):

Maximum chunk size, in the unit below. Typical for embeddings: 300–1000 tokens.

## `chunkSizeUnit` (type: `string`):

Chunk size unit.

## `chunkOverlap` (type: `integer`):

How much text from the end of a chunk is repeated at the start of the next one (same unit). Max half of the chunk size.

## `includePageNumbers` (type: `boolean`):

Adds  (Markdown) or \[Page N] (text) at every page/slide boundary. Page ranges are always in each item's metadata.

## `includeSpeakerNotes` (type: `boolean`):

Include PowerPoint speaker notes.

## `ocr` (type: `boolean`):

Read image-only (scanned) PDF pages and image files (PNG, JPG, TIFF, WebP) with OCR. Only pages without a text layer are OCR'd, and only those are charged as OCR pages.

## `ocrLanguages` (type: `array`):

Languages present in the scanned documents. Fewer languages = faster and more accurate.

## `maxPagesPerDocument` (type: `integer`):

Process only the first N pages/slides of each document (0 = all). For formats without pages the text is cut at N × 3,000 characters.

## `maxDocuments` (type: `integer`):

Max documents.

## `maxFileSizeMb` (type: `integer`):

Max file size (MB).

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://arxiv.org/pdf/1706.03762"
    }
  ],
  "datasetUrlField": "url",
  "outputFormat": "markdown",
  "chunking": "none",
  "chunkSize": 800,
  "chunkSizeUnit": "tokens",
  "chunkOverlap": 100,
  "includePageNumbers": false,
  "includeSpeakerNotes": true,
  "ocr": false,
  "ocrLanguages": [
    "eng"
  ],
  "maxPagesPerDocument": 0,
  "maxDocuments": 1000,
  "maxFileSizeMb": 100
}
```

# Actor output Schema

## `results` (type: `string`):

All result items of this run.

## `summary` (type: `string`):

Summary of the run with counts and download links.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://arxiv.org/pdf/1706.03762"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("egra_van/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [{ "url": "https://arxiv.org/pdf/1706.03762" }] }

# Run the Actor and wait for it to finish
run = client.actor("egra_van/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://arxiv.org/pdf/1706.03762"
    }
  ]
}' |
apify call egra_van/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,egra_van/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/BVKXqMPmMQwiYiKg5/builds/XP4CnTuywfXqFNtZL/openapi.json
