# PDF & DOCX to Markdown — Table Extraction & RAG Chunks (`ingenious_quip_bxq/pdf-docx-to-markdown`) Actor

Convert PDF and Word (DOCX) files to LLM-ready Markdown. Tables stay real Markdown tables (76% of rows exact on unseen statistical PDFs vs 31% without our repair), plus headings, page markers and optional RAG chunks with page range and section path. Batch URLs or uploads. No AI keys.

- **URL**: https://apify.com/ingenious\_quip\_bxq/pdf-docx-to-markdown.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.50 / 1,000 page converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF & DOCX to Markdown — Table Extraction & RAG Chunks

**Convert PDF to Markdown and Word (DOCX) to Markdown — LLM-ready, with tables that stay tables.**
This document parser extracts tables from PDFs, including borderless statistical tables like central-bank releases, as real Markdown tables. Each row keeps its label and every value in its own column, so your LLM, RAG pipeline or vector database reads numbers correctly instead of a run-together string. You also get the heading hierarchy, lists, page markers for citations and, optionally, RAG chunks (page range + section path) ready for embedding.

**Table accuracy, measured on documents we did not tune on:** on two unseen Federal Reserve statistical releases (H.4.1 and G.19, 14 pages, 75 table rows checked), **57 of 75 rows (76%)** came out as correct Markdown rows, versus **23 of 75 (31%)** from the same PDF engine without our table repair. Correct = row label in one cell and every value in its own cell, in order. The scoring method is described below and the source code is public (see *License & source code*).

### What you get

- 📄 **PDF to Markdown and DOCX to Markdown** in one Actor for text-based (digital) PDFs — batch URLs or file uploads
- 📊 **PDF table extraction as real Markdown tables** — words split across columns re-joined, wrapped row labels merged, captions lifted out of headers
- 🧭 **Structure kept** — `#`/`##`/`###` headings, bullet and numbered lists, two-column (academic) reading order
- 🔢 **Page markers** — `<!-- page: 7 -->` so answers can cite pages
- ✂️ **RAG chunks for your vector database** — configurable size and overlap, never split mid-table-row, each chunk tagged with `pageStart`, `pageEnd`, `headingPath`, `chunkIndex`
- 🔒 **No AI API keys, no LLM calls** — deterministic open-source parsing; documents are not sent to any third party
- 💸 **Pay per page** — $0.50 per 1,000 pages + $0.002 per document; failed files pay no document or page fee

### How accurate are the tables?

We count a table row as correct only if its label is in one cell and every value is in its own consecutive cell, in the original order. Ground truth comes from the PDF text layer, independent of the engines tested. All numbers below were measured by us, locally.

| Document | This Actor | Same engine, no table repair |
|---|---|---|
| **Unseen:** Fed H.4.1 (11 pp) + G.19 (3 pp), 75 rows | **57/75 (76%)** | 23/75 (31%)¹ |
| ResNet paper (arXiv 1512.03385), two-column academic, 13 rows | 12/13 | 12/13 |
| CA WARN report, 16 pp, 634 rows | 633/634 | 633/634 |
| Fed H.8, 22 pp, 325 rows — *used during development, so optimistic* | 325/325² | 161/325² |
| Word demo DOCX, 5 tables incl. merged cells, 29 rows | 28/29, no empty header rows, no misaligned rows | – |

¹ The unrepaired engine wraps many numbers in backticks (inline code), which this check counts as not clean. ² Row label may include the document's own line number (`34 Deposits`); with the label required alone in its cell: 255/325 vs 101/325.

**For context — plain open-source extractors on the same files (measured by us with the libraries themselves, not any particular Actor):** `pdf-parse` returns text only, so numbers run together (`8.36.7-0.24.05.3`) and 0 rows qualify; `pdfplumber`'s default table finder detects no table in the borderless Fed H.8 tables and spaces out digits in the CA WARN dates (`0 3 / 2 5 / 2 0 16`), so 0 rows qualify on either file. The ML-based `docling` scored 132/325 on Fed H.8 but needed 11–32 s per page on one CPU core, which is why this Actor doesn't use it.

Speed on one CPU core: ≈0.4–0.5 s per PDF page, ≈0.06 s per DOCX page; peak memory ≈600 MB.

**Known limits:** this Actor does not OCR (scanned pages are flagged in `stats.scannedPages`; use the Scanned OCR Actor); some tightly-kerned PDFs lose spaces between words (`LongBeach`); complex multi-row headers can still come out split.

### Use cases

- **RAG / chatbots over documents** — ingest manuals, policies, contracts, research papers and reports with page-cited chunks.
- **Financial & statistical PDFs** — extract tables from PDF statistical releases (e.g. central-bank data) so they stay tabular.
- **Academic papers** — two-column layouts (tested on one arXiv paper so far), tables and section headings preserved.
- **Word to Markdown** — convert DOCX knowledge bases, SOPs and proposals to Markdown for Notion, GitHub, wikis or LLM prompts.
- **AI agents & MCP** — call it from Claude, ChatGPT or any agent via the Apify MCP server / API to "read" a PDF with structure.
- **Automation** — n8n, Make, Zapier: drop a file URL in, get Markdown/chunks out.

### How to use

1. Add PDF/DOCX links in **Document URLs**, or upload files in **Upload files**.
2. Choose **Output**: `Markdown`, `RAG chunks`, or both.
3. Click **Start**. Results appear in the **Dataset** (views: *Documents*, *Markdown*, *RAG chunks*); each document is also saved as a `.md` file in the key-value store.

#### Input example

```json
{
  "urls": [
    { "url": "https://www.federalreserve.gov/releases/h8/current/h8.pdf" },
    { "url": "https://arxiv.org/pdf/1512.03385" },
    { "url": "https://calibre-ebook.com/downloads/demos/demo.docx" }
  ],
  "outputFormat": "markdown_and_chunks",
  "chunkSize": 800,
  "chunkOverlap": 100,
  "pageMarkers": "comment"
}
```

| Field | Default | What it does |
|---|---|---|
| `urls` | – | Direct links to PDF/DOCX files (batch) |
| `files` | – | Uploaded files (or URLs / key-value store keys `storeName/key`) |
| `outputFormat` | `markdown` | `markdown`, `chunks`, `markdown_and_chunks` |
| `pageMarkers` | `comment` | `comment` → `<!-- page: N -->`, `text` → `--- Page N ---`, `none` |
| `chunkSize` / `chunkOverlap` | 1000 / 150 | Chunk size and overlap |
| `chunkUnit` | `tokens` | `tokens` (≈4 chars each) or `characters` |
| `prependHeadingPath` | false | Prefix each chunk with `Section: A > B` |
| `pageRange` | all | e.g. `1-10`, `5`, `20-` |
| `maxPagesPerDocument` | 500 | Safety cap per PDF (0 = none) |
| `repairTables` | true | Table repair (see FAQ) |
| `removeHeadersFooters` | true | Drop running headers/footers |
| `includeImageText` | false | Keep text found inside charts |
| `pdfPassword` | – | For encrypted PDFs |

#### Output example — document item (abridged, real output)

```json
{
  "type": "document",
  "source": "https://www.federalreserve.gov/releases/h8/current/h8.pdf",
  "fileName": "h8.pdf",
  "format": "pdf",
  "status": "success",
  "title": "FEDERAL RESERVE statistical release",
  "pageCount": 22,
  "pagesConverted": 22,
  "stats": { "tables": 21, "tableWordFixes": 91, "tableRowMerges": 68, "headings": 34, "tokenEstimate": 20431 },
  "markdownKey": "h8-c1d0f36f.md",
  "markdown": "<!-- page: 1 -->\n\n# FEDERAL RESERVE statistical release\n\n..."
}
```

The Markdown for that table:

```markdown
Table 1. Selected Assets and Liabilities of Commercial Banks in the United States^1 ...
Percent change at break adjusted, seasonally adjusted, annual rate

|  | Account | 2021 | 2022 | 2023 | 2024 | 2025 | 2025 Q1 | 2025 Q2 | ... | 2026 Aug |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Bank credit | 8.3 | 6.7 | -0.2 | 4.0 | 5.3 | 2.9 | 6.9 | ... | 4.9 |
| 21 | Credit cards and other revolving plans | 6.8 | 16.7 | 9.4 | 4.7 | 3.4 | 3.0 | 3.0 | ... | 0.3 |
```

What `pdf-parse` (plain-text extraction) returned for the same row, measured by us: `1Bank credit8.36.7-0.24.05.32.96.9...`

#### Output example — RAG chunk item

```json
{
  "type": "chunk",
  "source": "https://www.federalreserve.gov/releases/h8/current/h8.pdf",
  "fileName": "h8.pdf",
  "chunkIndex": 2,
  "chunkCount": 35,
  "text": "|  | Account | 2021 | 2022 | ... |\n|---|---|...\n| 34 | Deposits | 11.8 | -0.7 | ...",
  "headingPath": ["FEDERAL RESERVE statistical release", "H.8 ASSETS AND LIABILITIES OF COMMERCIAL BANKS IN THE UNITED STATES"],
  "pageStart": 1,
  "pageEnd": 2,
  "tokenEstimate": 490,
  "containsTable": true
}
```

When a table is larger than one chunk it is split **between rows and the header row is repeated** in every piece, so each chunk is self-explanatory — useful when chunking for vector database ingestion.

### Pricing

Pay per event. Document and page fees are charged only for documents that convert successfully; the tiny Apify start fee applies to every run.

| Event | Price |
|---|---|
| Document converted | $0.002 per PDF/DOCX |
| Page converted | $0.0005 per page ($0.50 / 1,000 pages) |
| Actor start (Apify default synthetic event) | $0.00005 per GB of memory per run ($0.0001 at the default 2 GB) |
| RAG chunks, dataset items, table repair, .md files | included — no per-item charge |
| Failed downloads / unreadable files | no document or page fee |

Examples: a 10-page PDF = **$0.007**, a 20-page report = **$0.012**, 1,000 ten-page PDFs = **$7.00**. DOCX pages are counted from the page count Word stores in the file (or 3,000 characters per page if missing). Use `maxPagesPerDocument` or `pageRange` to cap spend on very long files.

### Integrations

**Python**

```python
from apify_client import ApifyClient
client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("uonrV5h4yceYKLxSK").call(run_input={
    "urls": [{"url": "https://arxiv.org/pdf/1512.03385"}],
    "outputFormat": "chunks", "chunkSize": 800, "chunkOverlap": 100,
})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    if item["type"] == "chunk":
        print(item["pageStart"], item["headingPath"], item["text"][:80])
```

**JavaScript**

```javascript
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: '<YOUR_APIFY_TOKEN>' });
const run = await client.actor('uonrV5h4yceYKLxSK').call({
  urls: [{ url: 'https://arxiv.org/pdf/1512.03385' }], outputFormat: 'markdown',
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items[0].markdown);
```

Also works with **n8n, Make, Zapier, LangChain (`ApifyDatasetLoader`), LlamaIndex** and the **Apify MCP server** for AI agents. The dataset items are plain JSON, so you can load the chunks with LangChain or LlamaIndex and embed them into Pinecone, Qdrant, pgvector or any vector database.

### FAQ

**Does it do OCR on scanned PDFs?**
Not in this Actor. Pages without a text layer are detected and listed in `stats.scannedPages` with a warning. For scanned/image PDFs, use [Scanned PDF/Image OCR to Markdown](https://apify.com/ingenious_quip_bxq/scanned-ocr-to-markdown).

**How does table repair work — can it change my data?**
It never invents text. Words are only re-joined when the joined word physically exists on that PDF page (so `B | ank credit` → `Bank credit`, but `iPhone | sales` is never glued). Wrapped row labels are merged into their data row and table captions that were glued into the header are lifted out. Turn it off with `repairTables: false`.

**Are DOCX page markers exact?**
DOCX files have no fixed pages. We place page markers where Word saved page breaks; otherwise the whole document is one section and the page count comes from the document properties.

**What about merged cells, nested tables and images in Word?**
Merged cells keep the grid aligned (value in the first cell, spanned cells empty); nested tables are flattened into their parent cell; images are dropped (alt text kept) so no base64 blobs pollute your prompts.

**Multi-column papers?**
Supported, but tested on only one two-column arXiv paper (ResNet) so far: a layout model orders the text so paragraphs aren't interleaved across columns, and 12 of its 13 table rows came out clean. Open an Issue if a multi-column file comes out in the wrong order.

**Is this a PDF text extractor too?**
Yes. Set **Output** to `Markdown` and use the text. Unlike a plain PDF text extractor or PDF parser that returns one string per page, table rows stay aligned and headings are kept, so it also works as a general document to Markdown converter.

**Is my data sent to OpenAI or any other AI API?**
No. Everything runs inside the Actor with open-source libraries. Files are only stored in your own Apify storage, under your retention settings.

**Limits?**
Up to 200 MB per file, `maxPagesPerDocument` (default 500) per PDF. Very large Markdown (> 8 MB) is only stored as a `.md` file (`markdownKey`) rather than inside the dataset item.

**Password-protected PDFs?**
Supply `pdfPassword`.

**Something converted badly?**
Open an Issue with a public link to the file — table edge cases are exactly what we want to fix.

### Related Actors / See also

- [Scanned PDF/Image OCR to Markdown](https://apify.com/ingenious_quip_bxq/scanned-ocr-to-markdown) — use when pages have **no text layer** (this Actor flags them in `stats.scannedPages`; it does not OCR).
- [Sitemap URL Extractor — PDF/DOCX Tags + robots.txt](https://apify.com/ingenious_quip_bxq/sitemap-url-discovery) — discover PDF/DOCX URLs from XML sitemaps; outputs `DOC_TO_MARKDOWN_INPUT` for this Actor.
- [Bulk URL Status Checker — 404s & Redirects](https://apify.com/ingenious_quip_bxq/url-status-checker) — optional: verify document URLs before conversion.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)** — see `LICENSE`. The full source code is public: https://github.com/xbox002000/pdf-docx-to-markdown

Open-source components: [PyMuPDF / pymupdf4llm](https://github.com/pymupdf/pymupdf4llm) (AGPL-3.0) for PDF parsing, [mammoth](https://github.com/mwilliamson/python-mammoth) (BSD-2-Clause) for DOCX, Beautiful Soup (MIT).

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/pdf-docx-to-markdown/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Direct links to PDF or DOCX files. Add as many as you like — they are processed as one batch. Tip: use the direct download link, not a viewer page.

## `files` (type: `array`):

Upload PDF or DOCX files from your computer (stored in your Apify key-value store). You can also paste URLs or key-value store references here (`recordKey` or `storeName/recordKey`).

## `outputFormat` (type: `string`):

`markdown`: one dataset item per document with the full Markdown. `chunks`: RAG-ready chunks (one dataset item per chunk) plus a document summary item. `markdown_and_chunks`: both.

## `pageMarkers` (type: `string`):

How page boundaries appear in the Markdown. `comment` inserts `<!-- page: 3 -->` (invisible when rendered, easy to parse); `text` inserts `--- Page 3 ---`; `none` omits them. DOCX files only get markers where Word stored page breaks.

## `saveMarkdownFiles` (type: `boolean`):

Save each document as a .md file in the run's key-value store (handy for downloading, and for very large documents that exceed dataset item limits).

## `chunkSize` (type: `integer`):

Maximum chunk size (in the unit below). Chunks break at headings, paragraphs and table rows — never mid-row.

## `chunkOverlap` (type: `integer`):

How much trailing context from the previous chunk is repeated at the start of the next one (same unit as chunk size, max 50% of chunk size).

## `chunkUnit` (type: `string`):

`tokens` uses a fast approximation (≈4 characters per token) so no tokenizer download is needed; `characters` is exact.

## `prependHeadingPath` (type: `boolean`):

Adds a line like `Section: 4. Experiments > 4.1 ImageNet` to each chunk's text. The heading path is always available in the `headingPath` field either way.

## `pageRange` (type: `string`):

Convert only some pages, e.g. `1-10`, `5`, or `20-` (to the end). Leave empty for all pages.

## `maxPagesPerDocument` (type: `integer`):

Safety cap on pages converted (and billed) per PDF. 0 = no cap.

## `repairTables` (type: `boolean`):

Fix table artifacts (words split across columns, wrapped row labels, table captions glued into the header, empty columns). Uses the page's own words as ground truth, so it never invents text.

## `removeHeadersFooters` (type: `boolean`):

Drop repeated page headers/footers (page numbers, running titles) so they don't pollute your chunks. Page numbers are still kept as page markers.

## `includeImageText` (type: `boolean`):

Charts often contain axis labels and tick numbers that read as noise. Off by default.

## `pdfPassword` (type: `string`):

Password for encrypted PDFs (applied to all PDFs in the run).

## Actor input object example

```json
{
  "urls": [
    {
      "url": "https://arxiv.org/pdf/1512.03385"
    }
  ],
  "outputFormat": "markdown",
  "pageMarkers": "comment",
  "saveMarkdownFiles": true,
  "chunkSize": 1000,
  "chunkOverlap": 150,
  "chunkUnit": "tokens",
  "prependHeadingPath": false,
  "maxPagesPerDocument": 500,
  "repairTables": true,
  "removeHeadersFooters": true,
  "includeImageText": false
}
```

# Actor output Schema

## `documents` (type: `string`):

Default dataset items. `type: "document"` items: source, fileName, format (pdf/docx), status (success/error), title, pageCount, pagesConverted, truncated, metadata, stats (tables, headings, words, tokenEstimate…), markdown (full Markdown when Output includes it, null if > 8 MB), markdownKey/markdownUrl (the .md record), chunkCount, processingMs, warnings, error. `type: "chunk"` items (chunk modes only): source, fileName, title, chunkIndex, text, pageStart, pageEnd, headingPath, tokenEstimate, containsTable. Views: overview (Documents), markdown, chunks (RAG chunks).

## `markdownFiles` (type: `string`):

Default key-value store keys. One `<file-name>-<hash8>.md` record (content type text/markdown) per converted document, when `saveMarkdownFiles` is true. The dataset item's `markdownKey` / `markdownUrl` point to it. The store also holds the run's INPUT record.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        {
            "url": "https://arxiv.org/pdf/1512.03385"
        }
    ],
    "outputFormat": "markdown",
    "pageMarkers": "comment",
    "repairTables": true
};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/pdf-docx-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": [{ "url": "https://arxiv.org/pdf/1512.03385" }],
    "outputFormat": "markdown",
    "pageMarkers": "comment",
    "repairTables": True,
}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/pdf-docx-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    {
      "url": "https://arxiv.org/pdf/1512.03385"
    }
  ],
  "outputFormat": "markdown",
  "pageMarkers": "comment",
  "repairTables": true
}' |
apify call ingenious_quip_bxq/pdf-docx-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/pdf-docx-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/uonrV5h4yceYKLxSK/builds/ii8Jq2MdsB1za6vqd/openapi.json
