# PDF to Markdown for RAG (`bindler/pdf-to-markdown`) Actor

Convert PDFs to clean markdown with real reading order. Handles two-column layouts, detects headings and paragraphs. Built for RAG pipelines and LLM ingestion.

- **URL**: https://apify.com/bindler/pdf-to-markdown.md
- **Developed by:** [Neil Sangwaiya](https://apify.com/bindler) (community)
- **Categories:** AI, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

## PDF to Markdown for RAG

Convert PDFs to clean markdown **in real reading order**. Built for RAG pipelines, LLM ingestion and document search.

### The problem this solves

PDFs have no concept of paragraphs, headings or reading order. They are a set of glyphs with coordinates. Most extractors dump those glyphs in drawing order, and the result is unusable for retrieval in three specific ways.

**Two-column layouts interleave.** Academic papers, reports, whitepapers and most printed documents are two columns. Read naively, line one of the left column is followed by line one of the right column, and the text becomes meaningless. This Actor detects the gutter from the horizontal distribution of glyphs and reads each column fully before moving to the next.

**Hyphenated words stay broken.** Justified text splits words across lines, so the file contains `ob-` then `jective`. That never matches a search for "objective", and every chunk containing it is quietly degraded. This Actor rejoins them, while keeping the hyphen where it belongs: `left-to-right` and `pre-train` survive intact, because stripping those is a different kind of wrong.

**Sentences run together.** PDF producers often emit `context.` and `Unlike` as separate fragments with no whitespace, giving you `context.Unlike`. This Actor restores the space.

It also detects headings from font size, and paragraph breaks from vertical spacing, so the markdown you get back has structure a chunker can actually use.

### What you get

| Field | Description |
|---|---|
| `url` | Source PDF |
| `title` | Document title from metadata, or the filename |
| `author`, `subject`, `creator`, `producer` | PDF metadata where present |
| `pageCount` | Pages in the document |
| `markdown` | **Full text as clean markdown** |
| `wordCount` | Words extracted |
| `page` | Page number, when returning one record per page |
| `scrapedAt` | ISO timestamp |

### Example input

```json
{
  "pdfUrls": [
    "https://arxiv.org/pdf/1706.03762",
    "https://example.com/annual-report.pdf"
  ],
  "perPageRecords": false
}
```

### Options

- **Max pages per PDF** — cap long documents so cost stays predictable
- **Detect two-column layouts** — on by default; turn off only if your PDFs are strictly single column
- **One record per page** — useful when you want to chunk by page rather than by document
- **Minimum line length** — raise it to strip page numbers and running headers

### Notes

- Works on any PDF with a text layer. **Scanned documents need OCR first**, which this Actor does not perform, and will return little or no text.
- No API key, no external service. Extraction happens inside the run.
- Handles Unicode, ligatures and accented characters.

# Actor input Schema

## `pdfUrls` (type: `array`):

Direct links to PDF files.

## `maxPages` (type: `integer`):

0 reads the whole document. Set a limit to cap cost on long files.

## `detectColumns` (type: `boolean`):

Read each column fully instead of interleaving lines across the page. Leave on unless your PDFs are strictly single column.

## `perPageRecords` (type: `boolean`):

Useful when you want to chunk by page. Off returns one record per document.

## `minLineChars` (type: `integer`):

Drop lines shorter than this. Raise it to strip page numbers and running headers.

## Actor input object example

```json
{
  "pdfUrls": [
    "https://arxiv.org/pdf/1706.03762"
  ],
  "maxPages": 0,
  "detectColumns": true,
  "perPageRecords": false,
  "minLineChars": 1
}
```

# Actor output Schema

## `documents` (type: `string`):

Markdown text with document metadata.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "pdfUrls": [
        "https://arxiv.org/pdf/1706.03762"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("bindler/pdf-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "pdfUrls": ["https://arxiv.org/pdf/1706.03762"] }

# Run the Actor and wait for it to finish
run = client.actor("bindler/pdf-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "pdfUrls": [
    "https://arxiv.org/pdf/1706.03762"
  ]
}' |
apify call bindler/pdf-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bindler/pdf-to-markdown"
        }
    }
}

```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/oqZpi2XfpULfINugZ/builds/a7No0rLRQ1dCnfHAB/openapi.json
