# Convert a PDF to Markdown for RAG

**Use case:** 

Headings, paragraphs in true reading order and tables become Markdown you can chunk and embed. Multi-column pages are read column by column, not straight across the page.

## Input

```json
{
  "pdfUrls": [
    "https://arxiv.org/pdf/1706.03762"
  ],
  "detectTables": true,
  "extractKeyFields": true,
  "includeFormFields": true,
  "includeLinks": true,
  "includeLines": false,
  "outputFormats": [
    "markdown",
    "json"
  ],
  "firstPage": 1,
  "maxPages": 12,
  "inlineFullResult": "auto",
  "concurrency": 3,
  "retries": 2,
  "timeoutSecs": 120,
  "maxFileSizeMb": 100,
  "headers": {},
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

## Output

```json
{
  "url": {
    "label": "PDF"
  },
  "ok": {
    "label": "OK"
  },
  "pageCount": {
    "label": "Pages"
  },
  "pagesParsed": {
    "label": "Pages parsed"
  },
  "tableCount": {
    "label": "Tables"
  },
  "isScanned": {
    "label": "Scanned"
  },
  "jsonUrl": {
    "label": "JSON",
    "format": "link"
  },
  "markdownUrl": {
    "label": "Markdown",
    "format": "link"
  },
  "tablesCsvUrl": {
    "label": "Tables CSV",
    "format": "link"
  },
  "metadata": {
    "label": "Metadata"
  },
  "fields": {
    "label": "Fields"
  },
  "pagesCharged": {
    "label": "Pages charged"
  },
  "durationMs": {
    "label": "Duration ms"
  },
  "error": {
    "label": "Error"
  }
}
```

## About this Actor

This example demonstrates how to use [PDF to JSON Extractor — Tables, Fields & Structure](https://apify.com/power_on/pdf-to-json-extractor.md) with a specific input configuration. Visit the [Actor detail page](https://apify.com/power_on/pdf-to-json-extractor.md) to learn more, explore other use cases, and run it yourself.


## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
This Task's input is already configured above — use it as-is rather than inventing a new one.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For full API examples (JavaScript, Python, CLI, MCP, OpenAPI), see this Task's Actor page: https://apify.com/power_on/pdf-to-json-extractor.md

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).
