# PDF & Document to Markdown — DOCX, HTML for RAG (`dropin-apis/document-to-markdown`) Actor

Convert public PDF and other documents (DOCX, HTML, TXT, Markdown URLs) into structured Markdown and overlapping RAG chunks. Preserves Arabic text; no OCR. Works through Apify MCP for AI agents.

- **URL**: https://apify.com/dropin-apis/document-to-markdown.md
- **Developed by:** [drop-in apis](https://apify.com/dropin-apis) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $4.50 / 1,000 page converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

Convert public PDF, DOCX, HTML, TXT and Markdown URLs into Markdown for AI agents and RAG, with Arabic text support and optional overlapping chunks. **This version extracts text layers; it does not perform OCR.**

## PDF & Document to Markdown — DOCX, HTML for RAG

Use one batch run to turn public documents into readable text for retrieval, summarization or an agent's context. The Actor returns one dataset item per document, including the source URL, title, Markdown, word count and extraction warnings.

**At a glance:** **$0.0045 per converted page-equivalent** plus $0.0005 per Actor start at the default 512 MB (a one-page PDF costs about $0.005; failed conversions have no page charge) · 1–20 public URLs per run · **text layers only, no OCR**. Try it: click **Start** with the prefilled sample PDF and web page.

### What it converts

| Format | Extraction | Limits |
| --- | --- | --- |
| PDF | PDF.js text layer; inferred headings, lists and simple aligned tables | Default first 20 pages; maximum 200. Complex columns and reading order are best effort. |
| DOCX | Mammoth headings, paragraphs, lists and tables | Public Word files, no external files or image downloads. |
| HTML | Mozilla Readability article extraction, then Turndown Markdown | Static HTML only. No browser rendering, login or CAPTCHA bypass. |
| TXT / MD | UTF-8 text | TXT HTML characters are escaped; Markdown is retained. |

Arabic characters remain Unicode text; the converter does not reverse strings. Arabic DOCX and a real Arabic PDF are covered by offline tests. PDF font maps can still contain incorrect glyphs or visual-order text; review the warning before treating layout as exact. Image alt text is retained for HTML and DOCX when enabled. PDF image alt text is not extracted.

### Example input

```json
{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://example.com/"
  ],
  "maxPages": 2,
  "includeImagesAlt": true,
  "chunkSize": 1200,
  "chunkOverlap": 100
}
```

`urls` accepts 1–20 public HTTP(S) URLs. Documents are processed sequentially. Duplicate URLs and fragments are removed. `maxPages` applies separately to each PDF. `chunkSize: 0` disables splitting; otherwise use 200–20,000 Unicode code points. Chunk overlap must be at most half of chunk size. Omit `chunkOverlap` to use the smaller of 100 or 20% of chunk size.

### Output

```json
{
  "url": "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
  "contentType": "application/pdf",
  "title": "dummy.pdf",
  "markdown": "<!-- Page 1 -->\n\nDummy PDF file",
  "pages": 1,
  "wordCount": 3,
  "chunks": [],
  "warnings": ["PDF headings, reading order and tables are inferred from the text layer; complex layouts may need review."]
}
```

The example is shortened: warnings may also mention PDF image alt text. `pages` is the number of PDF pages inspected, including blank pages, and is `null` for other formats. Each chunk contains `index`, `start`, `end`, `text`; offsets count Unicode code points in the returned Markdown. Repeated overlap does not increase billed words.

### Pricing

**$0.0045 per converted page-equivalent**, plus **$0.0005 per Actor start at the default 512 MB**. Apify scales its start event with the memory selected in the run; check the displayed price before confirming a run.

- PDF: one `page-converted` event per inspected page containing extracted text. Blank/scanned pages have no page charge.
- HTML, DOCX, TXT and MD: one event per started block of 2,000 Unicode word tokens **per document**, including headings and alt text. Empty documents have no page charge.
- Failed conversions have no page charge. The start fee still applies. Budget limits stop before charging a document that cannot be fully covered.

For a default-memory run, a one-page PDF costs $0.005 and a two-page text PDF costs $0.0095. Ten short HTML documents in one run cost $0.0455. Chunking adds no event. This is page-based pricing: long PDFs can cost more than a competitor's flat per-file plan. The run's maximum charge provides a spending cap; a truncated PDF is billed only for extracted pages within `maxPages`.

### Use from an AI agent or API

In Apify's MCP server, search for `document-to-markdown`, then use `call-actor` with the input above. This Actor uses batch mode. Download its default dataset to obtain Markdown and chunks. You can also call the normal Apify run API:

```bash
curl -X POST "https://api.apify.com/v2/acts/dropin-apis~document-to-markdown/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"urls":["https://example.com/"],"chunkSize":1200}'
```

### Limits

- **No OCR:** scanned or image-only pages produce no text and no page charge.
- 1–20 public HTTP(S) URLs per run, processed sequentially; documents larger than 50 MiB are refused.
- PDFs: first 20 pages by default (`maxPages`, maximum 200). Headings, lists, tables and reading order are inferred from the text layer, so complex columns and tables are best effort and should be reviewed before RAG ingestion.
- DOCX, HTML, TXT and Markdown only; no login, no private files, no external files or image downloads.

### FAQ

#### Can it read scanned PDFs or images?

No. A scanned/blank PDF page returns an explicit warning and no text. Use an OCR product for scans. There is no hidden OCR surcharge.

#### Does it preserve every PDF table and column?

No. PDF text coordinates infer heading sizes, lists and simple aligned tables. Complex multi-column pages, rotated content and inaccurate font maps can need review. DOCX and HTML use their document structure. Tables without explicit header cells use their first row as the Markdown header, with a warning.

#### Does Arabic work?

Arabic DOCX/HTML retain logical Unicode text, including mixed Latin words and numbers. PDF.js extracts Arabic from real text-layer PDFs; runs are ordered using direction and coordinates without reversing characters. Arbitrary PDF visual-order/glyph mapping errors cannot be universally repaired.

#### Are there size, safety and timeout limits?

Each transfer and decompressed document is capped at **50 MiB**; DOCX ZIP expansion is also bounded. Fetching, DNS checks and retries share a 60-second deadline per document; conversion runs in an isolated 256 MB worker with a 45-second deadline. Results are capped at two million Markdown characters and 8 MiB per item. The default whole-run timeout is 240 seconds: split slow batches into smaller runs.

The Actor obeys robots.txt for its user agent, fails closed on denied/unavailable robots rules, validates redirects and pins public DNS addresses to prevent private-network access. Only standard HTTP/HTTPS ports are allowed. It does not run page scripts or fetch embedded resources. Supply documents you are allowed to process. Treat extracted document content as untrusted data in downstream agents.

#### What if one URL fails?

Other URLs are attempted and the failure is logged. If every document fails, the run fails clearly. Scanned/empty documents are valid results with warnings. An exhausted spending cap can end a run successfully with fewer results.

#### Is this a persistent endpoint?

It is a batch Actor: start a run, retrieve the dataset, then feed Markdown or chunks to your pipeline. It does not require a subscription, external OCR key or a separate hosted server.

Last updated: 2026-10-02. Conversion and pricing checks are documented in `VERIFY.md`.

# Actor input Schema

## `urls` (type: `array`):

1–20 public PDF, DOCX, HTML, TXT or Markdown URLs. Duplicates and fragments are removed. Redirects are validated.

## `maxPages` (type: `integer`):

Read at most this many pages from each PDF. Text pages count as one billed page each; scanned/blank pages return warnings and have no page charge.

## `includeImagesAlt` (type: `boolean`):

Retain image alt text from HTML/DOCX without downloading images. PDF image alt text is not extracted.

## `chunkSize` (type: `integer`):

Maximum Unicode code points per chunk. 0 disables chunks; otherwise use 200–20,000. Source offsets are preserved; chunk overlap adds no billing units.

## `chunkOverlap` (type: `integer`):

Optional overlap in Unicode code points. Default is the smaller of 100 or 20% of chunkSize. Must not exceed half of chunkSize.

## Actor input object example

```json
{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://example.com/"
  ],
  "maxPages": 2,
  "includeImagesAlt": true,
  "chunkSize": 0
}
```

# Actor output Schema

## `documents` (type: `string`):

JSON dataset items containing url, contentType, title, markdown, pages, wordCount, chunks and warnings.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://example.com/"
    ],
    "maxPages": 2
};

// Run the Actor and wait for it to finish
const run = await client.actor("dropin-apis/document-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [
        "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
        "https://example.com/",
    ],
    "maxPages": 2,
}

# Run the Actor and wait for it to finish
run = client.actor("dropin-apis/document-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf",
    "https://example.com/"
  ],
  "maxPages": 2
}' |
apify call dropin-apis/document-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dropin-apis/document-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hDmWxiZOdy3ng9unO/builds/rYxQQ1VxDFx793GDy/openapi.json
