# PDF Text Extractor & PDF to Markdown - OCR, Word, Excel (`tidytools/document-to-markdown`) Actor

Document parser and PDF parser for LLMs: convert PDF, Word (DOCX/DOC), PowerPoint, Excel and CSV to clean Markdown text. PDF OCR for scanned files, RAG chunks. URLs, uploads, base64. $2/1,000 docs.

- **URL**: https://apify.com/tidytools/document-to-markdown.md
- **Developed by:** [Yukai Lin](https://apify.com/tidytools) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does PDF Text Extractor & PDF to Markdown do?

Give it documents (links, an uploaded file, or base64 from your code) and get back **clean Markdown text** for each one: ready for ChatGPT / Claude prompts, vector databases, RAG pipelines, search indexes or content migration.

Supported formats:

| Format | Extensions |
|---|---|
| PDF | `.pdf` (text-based PDFs; scanned PDFs with the optional OCR) |
| Microsoft Word | `.docx`, `.doc`, `.rtf` |
| Microsoft PowerPoint | `.pptx`, `.ppt`, `.ppsx`, `.pps` (slide text, via PDF) |
| Microsoft Excel | `.xlsx`, `.xls` (every sheet becomes a Markdown table) |
| OpenDocument | `.odt`, `.ods`, `.odp` |
| Data | `.csv` (as a table), `.xml` (text content only) |
| Web pages | HTML pages work too |

DOC, RTF, PPT/PPTX/PPS/PPSX and ODP files are first converted to PDF with LibreOffice on our own server, then to Markdown (rows show `convertedFrom` and `via: "vm-office"`). DOCX, XLSX/XLS and ODT/ODS are converted directly, and fall back to the same LibreOffice path if the direct conversion fails. Same price for every format.

- 📥 **Three ways in**: public URLs, a **file upload**, or **base64** in the API input (for AI agents, n8n, Make, Zapier)
- 🔗 **Cloud share links work**: Google Drive, Google Docs/Sheets/Slides, Dropbox and OneDrive/SharePoint links are turned into direct downloads automatically
- 📄 **Keeps structure**: headings, paragraphs, lists and tables
- 🔍 **OCR for scanned PDFs (optional)**: turn on **OCR for scanned PDFs** and PDFs without a text layer are transcribed page by page by an AI vision model, tables included, in most languages (tested with English, French, Chinese and Japanese)
- 🚦 **Clear status per document**: `status` is `success`, `no_text`, `download_blocked`, `not_found`, `too_large`, `unsupported` or `failed`; scanned PDFs are flagged with `needsOcr: true` when OCR is off
- 🧹 **Cleaner PDF text**: words hyphenated at line ends are joined ("repre-sentation" → "representation"), spaces lost between text blocks are put back ("AbstractWe" → "Abstract We") and simple number tables are rebuilt as Markdown tables; each item reports the changes in `pdfCleanup` (turn off with **Clean up PDF text**)
- 🧾 **PDF facts**: `pageCount`, `author`, `createdAt`, plus `wordCount`, `bytes` and `fileName` for every document
- ✂️ **RAG chunks built in**: optional `chunks` array split at paragraph boundaries with overlap, each with a `headingPath`; or **one row per chunk** (`outputMode: "chunks"`) for direct import into a vector database or CSV
- ⚡ **Fast**: no browser needed; most documents are converted in about a second
- 💸 **Simple pricing**: **$2 per 1,000 documents**, no start fee, no compute charges
- 🛡️ **Pay only for success**: broken links, blocked downloads and empty documents are **free**

### How much does it cost?

| Event | Price |
|---|---|
| Converted document | **$2.00 / 1,000 documents** |
| OCR page (optional, scanned PDFs only) | **$5.00 / 1,000 pages** |

**No start fee.** You pay per document, **not per page**: a 100-page PDF costs the same $0.002 as a one-page file. With per-page pricing, the page fee is multiplied by 100. Failed, blocked and empty (scanned) documents are free. Your maximum charge limit is always respected. On Apify's Scale plan prices are 10% lower, on Business and higher 20% lower.

**OCR** is charged only when you turn it on and a PDF has no text layer: $0.005 per transcribed page, on top of the $0.002 document price, up to **OCR: max pages per document** (default 20). Blank pages and failed pages are not charged. A 3-page scan costs $0.002 + 3 × $0.005 = **$0.017**. For comparison (checked September 2026), memo23 PDF Text Extractor charges $15 and yabanana99 PDF, Word & Excel to Markdown $12 per 1,000 OCR pages, on top of their per-document and per-run fees.

For AI agents that convert one file per call, there is no start fee to add: one 10-page text PDF costs $0.002. Tools that charge a $0.005 start fee per run cost about $0.0085–$0.01 for the same call.

### Control your cost

- **What is charged:** each document converted successfully ($0.002, whatever its page count), and OCR pages when you turn OCR on.
- **What is free:** broken links (`not_found`), blocked downloads, empty or scanned documents with OCR off, unsupported files and invalid input lines. Every row has `charged: true` or `false`.
- At the start, the run logs its worst case, e.g. `Plan: 10 documents × $0.002 = at most $0.02` (plus the OCR page price when OCR is on), and warns when that is more than your **maximum charge per run**.
- When the maximum charge per run is reached, the run stops and keeps everything converted so far. The status message says so, and the `SUMMARY` record has `status: "LIMIT_REACHED"` and `notProcessed` (how many documents were not started, and up to 100 of them).
- **If Apify restarts the run** (server migration or Resurrect), items already finished are skipped and not charged again (`SUMMARY.resumedSkipped`).

### How to use it

1. Paste the **Document URLs**, one per line (or upload a text/CSV file with one link per line in the request-list field under the options), or use **Or upload a file** for a document on your computer.
2. Optional: set a **RAG chunk size** such as 2000 characters, and **Output rows** = one row per chunk.
3. Optional: turn on **OCR for scanned PDFs** (and set a language hint for non-Latin scripts).
4. Click **Start** and download the Markdown as JSON, CSV or Excel, or fetch it via API.

#### Input example

```json
{
    "urls": [
        "https://arxiv.org/pdf/1706.03762",
        "https://calibre-ebook.com/downloads/demos/demo.docx"
    ],
    "chunkSize": 2000
}
```

#### Google Drive link and base64 file (API)

```json
{
    "urls": ["https://drive.google.com/file/d/0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx/view"],
    "base64Documents": [{ "fileName": "hello.csv", "data": "bmFtZSxjaXR5CkFkYSxMb25kb24K" }]
}
```

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~document-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d "{\"base64Documents\":[{\"fileName\":\"report.pdf\",\"data\":\"$(base64 -w0 report.pdf)\"}]}"
```

#### Output example

Real output from a test run (Markdown shortened). The Google Drive share link was converted to a direct download:

```json
{
    "url": "https://drive.google.com/file/d/0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx/view",
    "resolvedUrl": "https://drive.google.com/uc?export=download&id=0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx&confirm=t",
    "source": "url",
    "finalUrl": "https://drive.usercontent.google.com/download?id=0B1HXnM1lBuoqMzVhZjcwNTAtZWI5OS00ZDg3LWEyMzktNzZmYWY2Y2NhNWQx&export=download",
    "fileName": "Sample.pdf",
    "contentType": "application/pdf",
    "bytes": 23567,
    "title": "Preface",
    "pageCount": 8,
    "author": "Loren",
    "createdAt": "2010-08-26T10:15:07-05:00",
    "producer": "Acrobat Distiller 9.3.3 (Windows)",
    "success": true,
    "status": "success",
    "needsOcr": false,
    "via": "direct",
    "characters": 13129,
    "wordCount": 2274,
    "markdown": "# Sample.pdf\n\n## Contents\n### Page 1\n…",
    "convertedAt": "2026-09-29T02:26:12.454Z"
}
```

A base64 CSV (`hello.csv`) comes back as `"source": "base64"`, `"url": null`, `"bytes": 21` and the table in Markdown:

```json
{ "url": null, "source": "base64", "fileName": "hello.csv", "contentType": "text/csv", "status": "success", "markdown": "# hello.csv\n\n## Sheet1\n\n| name | city   |\n|------|--------|\n| Ada  | London |\n" }
```

With **Output rows = one row per chunk**, each chunk is its own row (the attention paper, `https://arxiv.org/pdf/1706.03762`, gave 38 chunks of up to 1,500 characters):

```json
{ "url": "https://arxiv.org/pdf/1706.03762", "status": "success", "pageCount": 15, "chunkIndex": 30, "chunkCount": 38, "headingPath": ["document.pdf", "Contents", "Page 11"], "charCount": 1500, "text": "…" }
```

#### Scanned PDFs with OCR

```json
{
    "documentUrls": [{ "url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/francais.pdf" }],
    "ocr": true,
    "ocrMaxPages": 20
}
```

Real output (a scanned French page, no text layer). Charged: 1 document + 1 OCR page:

```json
{
    "url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/francais.pdf",
    "fileName": "francais.pdf",
    "pageCount": 1,
    "producer": "Adobe Photoshop for Macintosh -- Image Conversion Plug-in",
    "success": true,
    "status": "success",
    "needsOcr": false,
    "ocr": true,
    "ocrPages": 1,
    "ocrPagesCharged": 1,
    "ocrTruncated": false,
    "ocrModel": "@cf/meta/llama-4-scout-17b-16e-instruct",
    "characters": 363,
    "wordCount": 66,
    "markdown": "# Adobe Photoshop PDF\n\n## Page 1\n\nPortez ce vieux whisky au juge blond qui fume sur son île intérieure, à côté de l'alcôve ovoïde, où les bûches se consument dans l'âtre, …"
}
```

- Each transcribed page starts with `## Page N`, followed by the page's own headings, lists and tables.
- `ocrTruncated: true` means the PDF has more pages than **OCR: max pages per document** (a 7-page scan with a limit of 6 gave 6 pages and `"pageCount": 7`).
- In our tests, a skewed English scan (OCRmyPDF `skew.pdf`) came back word for word with its headings and bullet lists. A Japanese Wikipedia page came back with only a few wrong characters. A dense Traditional Chinese page had about one wrong character in twenty, so check names and figures in CJK scans.

Documents that cannot be converted are reported and **not charged**. With OCR off, a scanned PDF looks like this:

```json
{ "url": "https://github.com/ocrmypdf/OCRmyPDF/raw/main/tests/resources/skew.pdf", "fileName": "skew.pdf", "success": false, "status": "no_text", "needsOcr": true, "pageCount": 1, "error": "the document contains no extractable text (it looks scanned; turn on \"OCR for scanned PDFs\" to transcribe it). (not charged)", "errorType": "no_text", "charged": false }
```

```json
{ "url": "https://drive.google.com/file/d/0B9P1L--7Wd2vNm9zMTJWOGxobkU/view", "success": false, "status": "download_blocked", "error": "expected a document but the share link returned a web page (the file is not shared publicly, or it is too large for a direct download). Download the file yourself and send it as base64Documents (or upload it). (not charged)", "errorType": "blocked", "charged": false }
```

A site that refuses both download routes (real output, `https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf`, shortened):

```json
{ "input": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf", "inputIndex": 0, "success": false, "status": "download_blocked", "triedRoutes": ["backend", "direct"], "error": "our servers: target returned HTTP 403; Apify network: target returned HTTP 403. Download the file yourself and send it as base64Documents (or upload it). (not charged)", "errorType": "blocked", "charged": false }
```

A link that does not exist gets `"status": "not_found"` (HTTP 404 or 410). URL rows carry `input` (the line as you entered it) and `inputIndex` (its position in the URL list), so results map back to your input even when documents finish in a different order.

The key-value store's `SUMMARY` record has `status` (`SUCCESS`, `PARTIAL_RESULTS`, `FAILED`, `NO_RESULTS` or `LIMIT_REACHED`) and counts per document status (plus `ocrDocuments` and `ocrPagesCharged` when OCR is on).

### Use with AI agents (MCP)

Connect Apify's MCP server (https://mcp.apify.com?tools=tidytools/document-to-markdown) to Claude, Cursor or any MCP client, then ask e.g. "Convert https://arxiv.org/pdf/1706.03762 to Markdown and summarize section 3."

Minimal input:

```json
{ "urls": ["https://arxiv.org/pdf/1706.03762"] }
```

Failed items are not charged and carry an `errorType` and a `status`. Agents that already have the file can send it as `base64Documents`.

### Use it from code

```bash
curl -X POST "https://api.apify.com/v2/acts/tidytools~document-to-markdown/run-sync-get-dataset-items?token=YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://arxiv.org/pdf/1706.03762"]}'
```

Use the plain `urls` list from code: the older `documentUrls` field still works, but Apify rejects the whole run (HTTP 400) when it holds a blank line or a bare domain.

**Schedules and integrations:** run it daily or weekly with Apify Schedules, get a webhook when a run finishes, or send the results to Zapier, Make, n8n, Google Sheets, Slack and other apps with Apify integrations. Results can be exported as JSON, CSV, Excel or XML.

### Limitations

- **Scanned PDFs** are OCR-ed only when **OCR for scanned PDFs** is on; otherwise they come back with `status: "no_text"` and `needsOcr: true` and are not charged.
- PDF clean-up is heuristic: it only uses the document's own words as evidence, so most CamelCase names stay intact, but a rare glued word can remain, and tables are rebuilt only when each row ends with the same number of numeric cells (merged cells, multi-line headers and text tables stay as text).
- OCR applies to PDFs with **no text at all**. A PDF where only some pages are scanned is converted from its text layer, and the scanned pages stay empty.
- OCR is done by an AI vision model: it is very good on printed text, but it can misread characters (see the Chinese result above) and is not tuned for handwriting. Image files (PNG, JPG) are not supported yet.
- Large Google Drive files (over about 100 MB) show a virus-scan page instead of the file and are reported as `download_blocked`; the 15 MB limit applies anyway.
- If a site blocks our servers, **Advanced settings** → *Auto* retries the download from Apify's network; you can add an Apify proxy (billed to your Apify account).
- Files up to 15 MB; password-protected files are not supported.
- The document must be publicly downloadable (no login).

### Related tools

- **Website to Markdown Crawler**: crawl a whole website or docs portal, including linked documents.
- **AI Web Data Extractor**: turn pages into structured JSON with the fields you choose.

### FAQ

**How do I extract text from a PDF?** Paste the PDF link (or upload the file, or send base64 from your code) and run. Each document comes back as clean Markdown text with headings, lists and tables, plus `pageCount` and `wordCount`.

**Does it OCR scanned PDFs?** Yes, when you turn on **OCR for scanned PDFs**. PDFs without a text layer are transcribed page by page by an AI vision model for $5 per 1,000 pages, on top of the document price.

**Which file types can it parse?** PDF, Word (DOCX, DOC, RTF), PowerPoint (PPTX, PPT), Excel (XLSX, XLS), OpenDocument, CSV, XML and HTML pages, all at the same price.

**Can I get chunks for RAG?** Yes. Set **RAG chunk size (characters)** for a `chunks` array with `headingPath`, or use `outputMode: "chunks"` for one row per chunk, ready for a vector database or CSV.

### Support

Open an issue in the **Issues** tab with the document URL. Issues are checked regularly.

# Actor input Schema

## `urls` (type: `array`):

Main input (fill this, or `documentFile` / `base64Documents`). Public links to PDF, DOCX, DOC, RTF, PPTX, PPT, XLSX, XLS, CSV, ODT, ODS, ODP or XML files (up to 15 MB each), one per line. Google Drive, Google Docs/Sheets/Slides, Dropbox and OneDrive/SharePoint share links are turned into direct downloads automatically. HTML pages also work. Blank or invalid lines are skipped and listed in the results (not charged).

## `documentFile` (type: `string`):

Upload one PDF, DOCX, DOC, RTF, PPTX, PPT, XLSX, XLS, CSV, ODT, ODS or ODP file (up to 15 MB) instead of, or in addition to, the URLs.

## `base64Documents` (type: `array`):

For API and AI-agent calls: a list of {"fileName": "report.pdf", "data": "<base64>"}. The file name's extension tells the converter the format. Up to 15 MB per file.

## `chunkSize` (type: `integer`):

If set (e.g. 2000), each document also gets a "chunks" array split at paragraph boundaries. 0 = off.

## `chunkOverlap` (type: `integer`):

Characters repeated between consecutive chunks.

## `outputMode` (type: `string`):

"One row per chunk" splits each document (chunk size above, 2,000 if 0) and outputs every chunk as its own row with chunkIndex and headingPath. Charged per document either way. Values: document = One row per document; chunks = One row per chunk (for vector databases and CSV).

## `maxConcurrency` (type: `integer`):

How many documents are converted at the same time.

## `httpVia` (type: `string`):

Some sites block requests from data centers. Auto retries from a second network before falling back to a real browser. If both are refused, Auto also tries our second server (Oracle, different IP) before a real browser.

## `proxyConfiguration` (type: `object`):

Only used for requests sent from Apify's network. Apify proxy usage is billed to your Apify account.

## `pdfCleanup` (type: `boolean`):

For PDFs with a text layer: join words hyphenated at line ends ("repre-sentation"), put back spaces lost between text blocks ("AbstractWe") and rebuild simple number tables as Markdown tables. Each item reports what was changed in "pdfCleanup". Turn off to get the raw converter text.

## `ocr` (type: `boolean`):

Transcribe PDFs that have no text layer (scans, photos of pages) with an AI vision model, page by page. Tables become Markdown tables; most languages work, including Chinese and Japanese. Costs $0.005 per transcribed page on top of the document price, and only applies to PDFs without extractable text.

## `ocrMaxPages` (type: `integer`):

Only the first pages up to this number are transcribed (a cost cap). The item gets "ocrTruncated": true when the PDF has more pages.

## `ocrLanguage` (type: `string`):

A hint such as "Traditional Chinese", "German" or "Japanese". Leave empty to detect it.

## `documentUrls` (type: `array`):

Same as the list above, in Apify's request-list format (also accepts a link to a text file of URLs). Kept for older inputs; the plain list above is recommended, especially for API calls.

## Actor input object example

```json
{
  "urls": [
    "https://arxiv.org/pdf/1706.03762"
  ],
  "base64Documents": [
    {
      "fileName": "hello.csv",
      "data": "bmFtZSxjaXR5CkFkYSxMb25kb24K"
    }
  ],
  "chunkSize": 0,
  "chunkOverlap": 200,
  "outputMode": "document",
  "maxConcurrency": 10,
  "httpVia": "auto",
  "pdfCleanup": true,
  "ocr": false,
  "ocrMaxPages": 20
}
```

# Actor output Schema

## `documents` (type: `string`):

No description

## `full` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://arxiv.org/pdf/1706.03762"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("tidytools/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["https://arxiv.org/pdf/1706.03762"] }

# Run the Actor and wait for it to finish
run = client.actor("tidytools/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://arxiv.org/pdf/1706.03762"
  ]
}' |
apify call tidytools/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,tidytools/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3O5uaRFA9cfTdEzNn/builds/lXY7XlkcC9vYmjoFn/openapi.json
