# PDF to Markdown & Document to Text for LLM/RAG, with OCR (`tinlark/document-to-markdown`) Actor

Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and images to clean Markdown or text, with tables, OCR for scans and RAG-ready chunks. Links or uploaded files, one row per document.

- **URL**: https://apify.com/tinlark/document-to-markdown.md
- **Developed by:** [Tinlark](https://apify.com/tinlark) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Markdown & Document to Text for LLM/RAG, with OCR

Convert PDF, Word, PowerPoint, Excel, CSV, HTML, EPUB and image files into clean Markdown or plain text. One Actor, one dataset row per document, or one row per chunk when you want ready-to-embed pieces for a RAG pipeline.

It keeps the structure that language models need: headings, lists and tables. Pages that are only a picture (scans, photos, screenshots) are read with OCR, and only those pages, so text PDFs stay fast.

### Who it is for

- Teams that feed documents to an LLM, a vector database or an AI agent and need Markdown instead of raw PDF text.
- People who have a pile of reports, contracts, slide decks, spreadsheets and scans and want searchable text.
- Anyone who wants one tool for many file types instead of a separate PDF, Word and OCR service.

### What it does

1. Downloads each link you give it (or reads files you upload), checks the real file type from the file's own bytes, and converts it.
2. Writes Markdown and/or plain text, plus title, metadata, page count, language guess, word count, tables found, pages read with OCR and warnings.
3. Optionally splits the text into chunks by headings or by size, with the heading path and page numbers of every chunk.
4. A document that cannot be converted (broken link, password-protected PDF, unsupported file) becomes an error row with the reason. It does not stop the run and it is not charged.

### Input

| Field | What it does | Default |
|---|---|---|
| Document links (`sources`) | http(s) links to documents. The type is detected from the file, so links without an extension work. | |
| Uploaded files (`files`) | Upload files in the console; they are stored in your account and read from there. | |
| Output format (`outputFormat`) | `markdown`, `text` or `both` | `markdown` |
| OCR (`ocr`) | `auto`: only pages with no text layer; `off`; `force`: every page | `auto` |
| OCR languages (`ocrLanguages`) | English, German, French, Spanish, Italian, Portuguese, Dutch, Russian, Chinese (simplified), Japanese | English |
| Max pages per document (`maxPagesPerDocument`) | PDF pages, slides, sheets or EPUB chapters converted per document | 50 |
| Page range (`pageRange`) | For example `1-5, 8, 12-` | all |
| Tables as Markdown tables (`includeTables`) | Detect tables in PDFs | on |
| Include document metadata (`includeMetadata`) | Author, dates and other properties stored in the file | on |
| Chunking (`chunking`) | `off`, `by-headings` or `by-size` | `off` |
| Chunk size / overlap | Characters. 2,000 characters are about 500 tokens of English | 2000 / 200 |
| Max file size (`maxFileSizeMb`) | Larger files get an error row | 50 |

Example input:

```json
{
    "sources": [
        "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
        "https://example.com/slides.pptx"
    ],
    "outputFormat": "markdown",
    "chunking": "by-headings",
    "chunkSize": 2000
}
```

### Output

One row per document, or one row per chunk with chunking on. A shortened row:

```json
{
    "recordType": "document",
    "status": "ok",
    "url": "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
    "fileName": "NIST.CSWP.04162018.pdf",
    "fileType": "pdf",
    "pages": 55,
    "pagesConverted": 5,
    "pagesOcr": 0,
    "title": "Framework for Improving Critical Infrastructure Cybersecurity, Version 1.1",
    "metadata": {"author": "National Institute of Standards and Technology", "created": "2018-04-17T13:45:53Z"},
    "language": "en",
    "tables": 1,
    "wordCount": 1189,
    "billablePages": 5,
    "markdown": "# Framework for Improving Critical Infrastructure Cybersecurity\n\nVersion 1.1\n\n...",
    "text": null,
    "warnings": ["The document has 55 pages; only the first 5 were converted (limit: Max pages per document)."],
    "error": null
}
```

A chunk row adds `chunkIndex`, `chunkCount`, `headingPath` (for example `["Introduction", "Scope"]`), `pageStart`, `pageEnd` and `tokensEstimate`. An error row has `status: "error"`, an `errorCode` (`not-found`, `access-denied`, `too-large`, `encrypted`, `corrupt`, `unsupported-type`, `blocked-address`, `timeout` and others) and an `error` text that says what to do. Rows are written as documents finish; `sourceIndex` is the position of the link in your input. A summary with totals is stored under the `SUMMARY` key.

### Supported formats

| Format | How it is read | Notes |
|---|---|---|
| PDF | Text layer with headings (by font size), lists, tables, two-column pages, repeated headers and footers removed; OCR for pages without text | Password-protected PDFs give a clear error |
| DOCX | Headings, lists, tables, bold and italic, links | Old `.doc` files are not supported: save as DOCX |
| PPTX | One section per slide: title, bullet levels, tables, speaker notes | Old `.ppt` is not supported |
| XLSX | One Markdown table per sheet, values as last saved | At most 5,000 rows per sheet; old `.xls` is not supported |
| CSV / TSV | One Markdown table | Delimiter detected |
| HTML | Main content as Markdown; navigation, footers, forms and scripts dropped | Pages that build their text with JavaScript come back empty with a warning |
| EPUB | One section per chapter | |
| TXT / Markdown | Passed through | |
| PNG, JPEG, TIFF, GIF, BMP, WebP | OCR | Multi-page TIFF is read frame by frame; needs OCR on |

### OCR

With `auto` (the default) a PDF page is read with OCR only when it has no usable text layer. Text PDFs never touch the OCR engine. Images are always read with OCR (with OCR set to `off` they give an error row). Choose the languages that appear in your documents; each page is read with all languages you select. OCR is Tesseract. Each row reports `pagesOcr` and `ocrConfidence` (0 to 100), and a warning appears when the confidence is low. OCR is slower than text extraction: expect a few seconds per page on a full CPU core. Use 4096 MB of memory for OCR-heavy runs: it costs about the same per page and is several times faster.

### Chunking for RAG

- `by-headings`: one chunk per section under a Markdown heading. A section longer than the chunk size is cut further at paragraph ends.
- `by-size`: chunks of about the chunk size, cut at paragraph boundaries, with overlap. Long tables are split by rows and every part repeats the header row.

Every chunk has its heading path and the pages it comes from, so answers can cite where they were found.

### Limits and safety

- Files up to 50 MB by default (200 MB at most). Downloads time out after 3 minutes and are retried with backoff on server errors.
- Only public http(s) links. Links that point to private or internal network addresses are refused.
- The file type comes from the file's bytes, not from the link or the Content-Type header.
- A single document is converted in its own process and stopped when it takes too long, so one bad file cannot stop the others.
- Default 50 pages per document; raise it up to 2,000.
- Not supported: password-protected files, scanned handwriting, text inside images embedded in Word or PowerPoint files, JavaScript-rendered web pages, legacy `.doc`/`.xls`/`.ppt`, DRM-protected files.
- Reading order of complicated layouts (magazines, posters) can be wrong, and tables without ruling lines are found only when their columns line up clearly. Check the output of documents where the layout matters.

### Pricing

**Free during launch (until 31 October 2026).** You pay only Apify's own platform usage for your runs.

From 1 November 2026: pay per event, **$1 per 1,000 converted pages, plus $4 per 1,000 pages read with OCR**. A page is a PDF page, a slide, a sheet, an image, or for Word, HTML, text, CSV and EPUB files one page per started 500 words. Failed documents are never charged. Every row shows `billablePages`.

Apify platform usage in our test runs (Free plan compute price): about $0.04 per 1,000 text pages and about $1.30 to $1.60 per 1,000 OCR pages. Higher memory finishes sooner at about the same cost per page.

### FAQ

**Does it keep tables?** Tables in Word, PowerPoint, Excel and HTML are kept as Markdown tables. In PDFs, ruled tables and tables with clearly aligned columns are found; merged header cells are collapsed.

**Can I upload files instead of using links?** Yes. Use the Uploaded files field in the console; the files are stored in your own account.

**Does it send my documents anywhere?** Documents are downloaded into the run, converted inside the run and deleted when the run ends. The only output is your dataset. No third-party AI or OCR service is called.

**Why did a PDF come back empty?** It probably has no text layer and OCR was set to `off`, or it is a scan of low quality. Check `warnings` and `pagesOcr`.

**Can I call it from code or an AI agent?** Yes, through the Apify API, the client libraries or Apify's MCP server: start the Actor with the input above and read the dataset items.

**Disclaimers and legality.** Process only documents you have the right to use. You are responsible for the lawful use of the documents you convert and of the output, including copyright, confidentiality and personal data in them. The Actor fetches exactly the links you give it, as a browser would; it does not search, crawl or follow links, and it does not log in to anything. Do not use it to get around access controls or terms of a site. Conversion is automatic and can contain errors: do not rely on it alone for legal, medical or financial decisions.

### Related Tinlark Actors

- [Audio & Podcast Transcriber](https://apify.com/tinlark/audio-podcast-transcriber): turns audio and video links and podcast RSS feeds into text and subtitles, so recordings and documents can go into the same RAG or search pipeline.

### Changelog

- 2026-10-03: added the Related Tinlark Actors section.

### Support

Something wrong or missing? Open an issue on this Actor's Issues tab with your input (the document link or file type) and what you expected.

# Actor input Schema

## `sources` (type: `array`):

Links (http or https) to PDF, Word (.docx), PowerPoint (.pptx), Excel (.xlsx), CSV, HTML, text, Markdown, EPUB or image files (PNG, JPEG, TIFF). The Actor detects the type from the file itself, so links without a file extension work. Links must be public and reachable from the internet.

## `files` (type: `array`):

Upload files from your computer. They are stored in a key-value store of your account and read from there. Use this together with, or instead of, document links.

## `outputFormat` (type: `string`):

markdown keeps headings, lists and tables. text is plain text without markup. both adds both fields to every row.

## `ocr` (type: `string`):

auto: read only the pages that have no text layer (scans, photos) with OCR. off: never run OCR (scanned pages stay empty and images fail). force: OCR every page, even when the PDF has text. OCR pages are slower and are priced higher from 1 November 2026.

## `ocrLanguages` (type: `array`):

Languages of the scanned text. Pick all that appear in the documents: more languages make OCR a little slower.

## `maxPagesPerDocument` (type: `integer`):

Converts at most this many pages (PDF pages, slides, sheets or EPUB chapters) of each document. Longer documents get a warning in their row.

## `pageRange` (type: `string`):

Only these pages, for example 1-5, 8, 12- (12 to the end). Applies to PDF pages, PowerPoint slides, Excel sheets and TIFF frames. Leave empty for all pages up to the limit above.

## `includeTables` (type: `boolean`):

Detect tables in PDFs and write them as Markdown tables. Word, PowerPoint, Excel and HTML tables are always kept; when off they become plain lines.

## `includeMetadata` (type: `boolean`):

Add the metadata object (author, creation date, subject and similar properties stored in the file).

## `chunking` (type: `string`):

off: one row per document. by-headings: one row per section, split under each Markdown heading. by-size: rows of about the chunk size, cut at paragraph ends. Each chunk row has its heading path and page numbers. Sections longer than the chunk size are cut further.

## `chunkSize` (type: `integer`):

Target maximum length of a chunk in characters. 2,000 characters are about 500 tokens of English text.

## `chunkOverlap` (type: `integer`):

With by-size chunking, each chunk starts with the end of the previous one, up to this many characters. At most half of the chunk size.

## `maxFileSizeMb` (type: `integer`):

Larger files are skipped with an error row.

## Actor input object example

```json
{
  "sources": [
    "https://example.com/report.pdf",
    "https://example.com/slides.pptx"
  ],
  "outputFormat": "markdown",
  "ocr": "auto",
  "ocrLanguages": [
    "eng"
  ],
  "maxPagesPerDocument": 5,
  "includeTables": true,
  "includeMetadata": true,
  "chunking": "off",
  "chunkSize": 2000,
  "chunkOverlap": 200,
  "maxFileSizeMb": 50
}
```

# Actor output Schema

## `documents` (type: `string`):

Overview of converted documents or chunks.

## `content` (type: `string`):

Converted content per document or chunk.

## `errors` (type: `string`):

Documents that could not be converted, with the reason.

## `summary` (type: `string`):

Counts of documents, pages and OCR pages, and the list of errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "sources": [
        "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
        "https://raw.githubusercontent.com/mwilliamson/python-mammoth/f3b7b9fc73fdbffe6ebac77e6b3ac6c418ee33d9/tests/test-data/tables.docx",
        "https://raw.githubusercontent.com/tesseract-ocr/test/232ff181c66516116ec0e84c4963f70de15050fd/testing/phototest.tif"
    ],
    "ocrLanguages": [
        "eng"
    ],
    "maxPagesPerDocument": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("tinlark/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "sources": [
        "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
        "https://raw.githubusercontent.com/mwilliamson/python-mammoth/f3b7b9fc73fdbffe6ebac77e6b3ac6c418ee33d9/tests/test-data/tables.docx",
        "https://raw.githubusercontent.com/tesseract-ocr/test/232ff181c66516116ec0e84c4963f70de15050fd/testing/phototest.tif",
    ],
    "ocrLanguages": ["eng"],
    "maxPagesPerDocument": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("tinlark/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "sources": [
    "https://nvlpubs.nist.gov/nistpubs/CSWP/NIST.CSWP.04162018.pdf",
    "https://raw.githubusercontent.com/mwilliamson/python-mammoth/f3b7b9fc73fdbffe6ebac77e6b3ac6c418ee33d9/tests/test-data/tables.docx",
    "https://raw.githubusercontent.com/tesseract-ocr/test/232ff181c66516116ec0e84c4963f70de15050fd/testing/phototest.tif"
  ],
  "ocrLanguages": [
    "eng"
  ],
  "maxPagesPerDocument": 5
}' |
apify call tinlark/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tinlark/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/eBpofdUtYh4tMbUtr/builds/536XhmcFGarAzwpzw/openapi.json
