# PDF to Markdown & RAG Chunks: Tables to CSV, OCR (`bongo-seakhoa/pdf-to-markdown-rag-chunks`) Actor

Convert PDFs into clean Markdown and RAG-ready chunks with page numbers and heading paths. Tables export to CSV and JSON, and offline OCR reads scanned pages. Failed files are never charged.

- **URL**: https://apify.com/bongo-seakhoa/pdf-to-markdown-rag-chunks.md
- **Developed by:** [Bongo Seakhoa](https://apify.com/bongo-seakhoa) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.00 / 1,000 pdf processeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Markdown & RAG Chunks: Tables to CSV, OCR

Turn PDFs into clean Markdown and retrieval-ready chunks you can cite. Every chunk carries its page span, its heading path and exact character offsets into the Markdown. It also works as a PDF to JSON converter and extracts tables from PDF files three ways: inline in the Markdown, as JSON rows and as a downloadable CSV file. Scanned pages are read with offline OCR. Nothing is sent to a language model or a third-party API.

### What the PDF to Markdown converter outputs

- **Markdown** with headings, paragraphs joined across line breaks, de-hyphenated words, bulleted and numbered lists, and tables as GitHub-flavoured Markdown. Multi-column pages are read in column order, and running headers, footers and page numbers are removed.
- **RAG chunks** sized in characters, split at headings, paragraphs and sentences, with `page_start`/`page_end`, `heading_path`, `char_start`/`char_end` and a stable `chunk_id`. The same PDF with the same settings on the same Actor build always gives the same chunk IDs, so re-indexing a vector store is idempotent.
- **Tables** as JSON rows (header plus body) and as CSV and JSON files. A table that continues on the next page is merged into one when its columns line up, and a repeated header row is dropped. A cell merged down several rows is repeated in each of them, so each CSV row stands on its own.
- **Filled-in PDF forms**: values typed into form fields appear next to their labels, and checkboxes show as `[x]` or `[ ]`.
- **OCR for scans** with Tesseract, offline, in English, German, French, Spanish, Italian, Portuguese, Dutch and Hungarian. In the default `auto` mode only pages without a usable text layer are OCR'd (including scans whose text layer is just a scanner stamp or a Bates number), and only pages whose text came from OCR are charged at the OCR price.
- **An estimate mode** that counts pages, tells you which pages would need OCR and shows the projected price before you convert anything. It charges none of this Actor's events.
- **Honest failures.** Broken, encrypted, oversized or unreachable files get a status row with an error code and are never charged.

Typical uses: a PDF chunker that feeds a vector database for retrieval-augmented generation, LangChain or LlamaIndex PDF pipelines that need page citations, extracting tables from reports and invoices into spreadsheets, and turning scanned archives into searchable text.

### How to convert PDFs to Markdown and RAG chunks

1. Paste one or more PDF links into **PDF URLs**. Links must point at the PDF file itself; redirects are followed and Dropbox `?dl=0` links are converted to downloads.
2. Or pick a **key-value store** that holds your PDFs (upload them in Apify Console under Storage), optionally naming the record keys to use.
3. Leave the defaults or adjust chunk size, OCR and limits, then start the run.
4. Open the **Output** tab. The *Documents*, *Chunks* and *Tables* views each list every row of the run, so rows of the other two types show empty cells; sort or filter on `record_type`, or use the `.md`, `.csv` and `.json` files in the run's key-value store.

Starting the Actor with no sources at all converts a small bundled synthetic sample PDF so you can see the output format first. The sample is charged none of this Actor's events; Apify's standard Actor start fee still applies, as it does to every run.

### Input: PDF URLs or files in a key-value store

```json
{
  "sources": [
    "https://example.com/reports/annual-report-2025.pdf",
    "https://www.dropbox.com/s/abc123/scan.pdf?dl=0"
  ],
  "mode": "convert",
  "outputs": ["markdown", "chunks", "tables"],
  "chunkSize": 1500,
  "chunkOverlap": 150,
  "ocrMode": "auto",
  "ocrLanguages": ["eng"],
  "pageRange": "1-50",
  "maxPages": 200,
  "maxFileSizeMb": 50,
  "perDocumentTimeoutSecs": 900,
  "includePageMarkers": false,
  "csvFormulaGuard": true
}
```

| Field | Default | What it does |
|---|---|---|
| `sources` | `[]` | PDF URLs (http or https). |
| `keyValueStoreId`, `keyValueStoreKeys` | none | Read PDFs from a key-value store. Without keys, every record whose key ends in `.pdf` is used (up to 1,000). |
| `useBundledSample` | `false` | Also convert the 3-page synthetic sample (no event charges). |
| `mode` | `convert` | `estimate` only counts pages and projects the price; it charges none of this Actor's events. |
| `outputs` | all three | Any of `markdown`, `chunks`, `tables`. |
| `chunkSize`, `chunkOverlap` | 1500, 150 | Characters (about 4 per token). Overlap is at most half the chunk size. |
| `ocrMode` | `auto` | `auto` (pages without usable text only), `force` (every page, each charged as `ocr-page`) or `off`. |
| `ocrLanguages` | `["eng"]` | Any of `eng deu fra spa ita por nld hun`. |
| `pageRange` | all pages | For example `1-5,8,10-`. |
| `maxPages` | 200 | Pages per document; later pages are not converted and not charged. |
| `maxFileSizeMb` | 50 | Larger files are rejected before download completes. Also capped at a quarter of the run's memory: 256 MB at the default 1 GB. |
| `perDocumentTimeoutSecs` | 900 | A slow document returns the pages finished so far as `partial`. |
| `includePageMarkers` | `false` | Adds a `<!-- page N -->` line before the first block that starts on each page. Blank pages, and pages whose text only continues a paragraph or table from the page before, get no marker. |
| `csvFormulaGuard` | `true` | In the table CSV files, a cell that starts with `=`, `+`, `-`, `@`, a tab or a carriage return (also after leading spaces or invisible characters, or in full-width form) gets a leading `'` so spreadsheets show it as text instead of running it as a formula. Signed numbers such as `-12.5`, `+3` or `-45%`, and cells holding only dashes or plus signs (`-`, `--`, often a nil placeholder), are left alone. JSON rows, the Markdown and the dataset always keep the raw text. Set `false` for raw CSV cells. |
| `password` | none | Password for encrypted PDFs (secret field, never logged). |

### Output: Markdown, RAG chunks and tables as JSON and CSV

The dataset holds three kinds of rows, told apart by `record_type`. Filter on it when you read the dataset through the API. The examples below are the rows of the bundled sample, trimmed where marked with `...`.

**Document row** (one per PDF):

```json
{
  "record_type": "document",
  "document_id": "f6a2ab3234b42a02",
  "document_index": 0,
  "source": "bundled-sample:sample.pdf",
  "source_type": "sample",
  "mode": "convert",
  "status": "ok",
  "error_code": null,
  "pages_total": 3,
  "pages_processed": 3,
  "text_pages": 2,
  "ocr_pages": 1,
  "empty_pages": 0,
  "table_count": 1,
  "chunk_count": 3,
  "markdown": "# Synthetic Sample Report\n\nThis is a synthetic sample PDF bundled with the PDF Markdown Actor. ...",
  "markdown_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02.md?signature=<signature>",
  "warnings": [],
  "charged_events": {}
}
```

The sample is not charged, so its `charged_events` is empty. For a PDF from your own URL it lists what was charged, for example `{"pdf-processed": 1, "page-processed": 2, "ocr-page": 1}`.

**Chunk row:**

```json
{
  "record_type": "chunk",
  "document_id": "f6a2ab3234b42a02",
  "chunk_id": "a730736c675c2a6b183d6ea1",
  "chunk_index": 2,
  "text": "## Quarterly figures\n\nThe invented quarterly figures below are for demonstration.\n\n| Quarter | Orders | Revenue (EUR) | Returns |\n| --- | --- | --- | --- |\n| Q1 | 1,204 | 18,930.00 | 31 |\n... \n\nThe Actor reads it with offline OCR.",
  "char_start": 325,
  "char_end": 750,
  "page_start": 2,
  "page_end": 3,
  "heading_path": ["Synthetic Sample Report", "Quarterly figures"],
  "table_ids": ["t1"],
  "table_header": null,
  "approx_tokens": 107
}
```

**Table row:**

```json
{
  "record_type": "table",
  "document_id": "f6a2ab3234b42a02",
  "table_id": "t1",
  "page_start": 2,
  "page_end": 2,
  "n_rows": 5,
  "n_cols": 4,
  "header": ["Quarter", "Orders", "Revenue (EUR)", "Returns"],
  "rows": [["Q1", "1,204", "18,930.00", "31"], ["Q2", "1,388", "21,115.50", "27"], ["Q3", "1,512", "23,480.25", "40"], ["Q4", "1,690", "26,002.75", "35"]],
  "csv_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02-t1.csv?signature=<signature>",
  "json_url": "https://api.apify.com/v2/key-value-stores/<store>/records/doc0001-f6a2ab3234b42a02-t1.json?signature=<signature>"
}
```

`markdown[char_start:char_end]` is exactly the chunk's `text`, so you can always point back to the source passage. A chunk that starts partway down a table also carries the table's header row and separator in `table_header` (outside `text`), so you can embed the two together and the columns keep their names. Very long Markdown is left out of the document row (`markdown` is `null`) and is always available from `markdown_url`. A very large table's rows are left out of its table row (`rows` is `null`, `rows_omitted` is `true`) and stay in its CSV and JSON files. The run's key-value store also gets a `SUMMARY` record with the counts for the whole run.

Document `status` is `ok`, `partial` (some content may be missing: the warnings say why, for example a page limit, a timeout or a page OCR could not read), `failed` (nothing extracted; `error_code` says why, for example `BLOCKED_DESTINATION` for a link to a private or local network address) or `skipped` (not started: `CHARGE_LIMIT_REACHED` when your maximum charge was reached, `RUN_TIME_LIMIT` when the run was about to time out). If Apify restarts a run (for example when it moves the run to another server) while a document is being charged, that document is not converted again: its row carries a `BILLING_UNCERTAIN_AFTER_RESTART` warning, and its `charged_events` may not list every charge made before the restart.

### Pricing per PDF and per page

This Actor is priced per event: you pay for what was converted.

| Event | Price | When it is charged |
|---|---|---|
| `pdf-processed` | $0.003 per document | Once per document that produced content (status `ok` or `partial`). |
| `page-processed` | $0.0003 per page | Each page whose text came from the PDF's text layer. |
| `ocr-page` | $0.006 per page | Each page whose text came from OCR, instead of `page-processed`. |

With `ocrMode` set to `force`, every page is charged as `ocr-page`, even pages that have a good text layer. Apify adds the Actor start fee to every run ($0.00005 per run at up to 1 GB of memory, once more for each extra GB). Chunks, tables, files and dataset rows are not charged separately.

Worked examples:

- A 10-page born-digital report: $0.003 + 10 × $0.0003 = $0.006.
- A 20-page scanned contract: $0.003 + 20 × $0.006 = $0.123.
- 1,000 single-page invoices with a text layer: 1,000 × ($0.003 + $0.0003) = $3.30.
- A 30-page document of which 4 pages are scans: $0.003 + 26 × $0.0003 + 4 × $0.006 = $0.0348.

Never charged by this Actor: documents that fail (broken, encrypted without the right password, not a PDF, too large, unreachable), blank pages and pages that produce no text, pages beyond `maxPages` or outside `pageRange`, the bundled sample, and estimate mode. Only Apify's start fee applies to those runs.

**Your maximum charge is respected.** Before each document the Actor checks what is left of your run's maximum total charge and converts only the pages it can charge for. It always keeps one text page's price ($0.0003) unspent below your limit, so a document whose exact price would just fit can still be cut short or skipped. How pages are budgeted depends on `ocrMode`:

- `off`: every page at $0.0003.
- `force`: every page at $0.006.
- `auto`: every page at $0.006 while that fits. When it does not, a quick pre-check of the document predicts which pages need OCR, and the other pages are budgeted at $0.0003. If the pre-check misses a page that needs OCR, Apify stops charging at your limit, so you never pay more than your maximum.

A document that does not fit in full is converted up to the last page that fits and marked `partial` with a `CHARGE_LIMIT_PAGES` warning. A document that cannot pay for $0.003 plus its first page, and every document after it, is listed as `skipped` with `CHARGE_LIMIT_REACHED` and is not charged; the run normally finishes as *Succeeded*. If Apify itself ends the run at the limit, documents with no row were not processed and were not charged. Rerun the skipped documents with a higher limit; for born-digital PDFs, `ocrMode` `off` also makes the most of a small limit.

### Limits: read before you buy

- **Layout is heuristic.** Reading order handles one to four columns. Sidebars next to a column, text wrapped around figures and right-to-left scripts are not reconstructed correctly. Headings are inferred from font size and bold text and go down to H3 only.
- **Tables.** Ruled tables (with lines) are detected well. Borderless tables are detected conservatively, so some are missed and stay as text. A cell merged down several rows is repeated in each row; a cell merged across columns (a section row or a note) keeps its text in its first column. A simple two-row grouped header (one spanning cell over its sub-columns) becomes one row ("Readings Min", "Readings Max"). A table continued on the next page is merged only when its columns line up with the first part. Tables on rotated pages are not detected. In the CSV files, cells that a spreadsheet would run as a formula start with `'` (see `csvFormulaGuard`).
- **Forms.** Values of fillable form fields are read on upright pages only, and not on pages that are OCR'd.
- **OCR output is plain paragraphs.** Scanned pages give text, not headings or tables. Handwriting is not supported. Accuracy depends on scan quality; low-confidence pages get an `OCR_LOW_CONFIDENCE` warning. Only the eight languages listed above are installed. With OCR off, a scan whose text layer is only a stamp gets a `SCANNED_PAGE_NOT_OCRED` warning.
- **OCR is slow at the default memory.** Apify gives a run a quarter of a CPU core at 1 GB of memory. In our tests a scanned page took 8 to 17 seconds at 1 GB, depending on how dense the scan is. At 2 GB, the maximum for this Actor, a sparse scan took about 3 seconds per page; dense scans take longer. With the defaults, a scan of 50 to 100 pages can hit the 900-second limit per document and come back `partial`. For large scans use 2 GB of memory, raise `perDocumentTimeoutSecs`, or split the work with `pageRange`. Text pages are fast (well under a second each).
- **Downloads.** Links must return the PDF itself. Pages that show a viewer, a cookie wall or a login (Google Drive previews, SharePoint pages and the like) are reported as `NOT_A_PDF`. Servers that block automated downloads return `HTTP_ERROR` with the status code. Links must point to the public internet: see "Security".
- **Sideways text** in a minority orientation on a page (a margin note or stamp) is left out, with a warning.
- **Damaged files** that lost their cross-reference table are reported as `CORRUPT`; they are not repaired.
- **Token counts** in `approx_tokens` are characters divided by four, not a tokenizer count.
- **Per run**: at most 1,000 listed sources plus up to 1,000 `.pdf` records read from a key-value store, `maxPages` up to 5,000, and files up to 500 MB but never more than a quarter of the run's memory (128 MB at 512 MB, 256 MB at 1 GB, 500 MB at 2 GB). Defaults are 200 pages and 50 MB.

### FAQ

**Is my document sent to an AI service?** No. Text extraction uses PDFium and pdfminer, and OCR uses Tesseract, all inside the Actor's container. There is no language model and no external OCR API.

**Why is a document `partial`?** Some content may be missing. The `warnings` list gives the reason codes, for example `PAGES_TRUNCATED` (page limit), `CHARGE_LIMIT_PAGES` (cut short to fit your maximum charge), `DEADLINE_REACHED` (time limit), `OCR_FAILED` or `NO_TEXT_LAYER` (a scanned page with OCR off). You pay only for pages that produced text.

**How do I check the price before converting?** Run with `"mode": "estimate"`. Each document row then shows `pages_total`, `text_pages`, `ocr_pages` and `projected_cost_usd`, and none of this Actor's events are charged (Apify's start fee still applies). The estimate reads character counts only, so a page whose text layer is unreadable may be counted as a text page, and a scan with a stamp is counted as an OCR page even if the conversion ends up billing it as a text page.

**Can I use password-protected PDFs?** Yes: set `password`. It applies to every document in the run, is stored as a secret input and is never logged.

**Are the chunk IDs stable?** Yes, within one Actor build. A chunk ID is a hash of the document content, the chunk's position and its text. The same file with the same settings on the same build always produces the same IDs. A new build that improves extraction (or updates Tesseract) can change the text or the chunk boundaries, and with them the IDs, so re-index after an update if you rely on them.

**Can I read the output from code?** Yes. Use the dataset items endpoint and filter on `record_type`, or download the `.md`, `.csv` and `.json` files from the run's key-value store.

**Something does not work.** Open an issue on the Actor's Issues tab with the run link and, if you can share it, a PDF that shows the problem. Issues are read on a best-effort basis; there is no guaranteed response time.

### Security

- **Only public internet hosts are fetched.** Before each request, and again for every redirect, the Actor looks up the link's host and refuses it if any of its addresses is private, local or reserved: `localhost` and other loopback addresses, private networks (10.x, 172.16-31.x, 192.168.x, fc00::/7), link-local addresses including the cloud metadata address 169.254.169.254, carrier-grade NAT (100.64.0.0/10), multicast and reserved ranges, and IPv6 forms that wrap such an IPv4 address. Numeric hosts such as `http://2130706433/` count as the address they stand for. Links with a user name or password in them (`https://user:pass@host/...`) are refused too. A refused link becomes a `failed` document with `error_code` `BLOCKED_DESTINATION` and is never charged. To convert a file from a private network, upload it to a key-value store and use `keyValueStoreId`.
- **CSV files are safe to open in a spreadsheet** by default (`csvFormulaGuard`, above).

### Privacy and data retention

The Actor reads only the URLs and key-value store records you give it. Your PDFs are processed inside your run and are not sent anywhere else. The log contains counts, status and error codes, never document text, URLs or passwords. The Actor keeps nothing after the run ends: the outputs live in your run's dataset and key-value store under your Apify account, and Apify's storage retention applies to them (unnamed storages are deleted automatically after a retention period).

### About the sample and the tests

The bundled sample and every test file used to develop this Actor are synthetic PDFs generated for the purpose and labelled "synthetic test fixture" in their metadata. No third-party documents were used.

### Licences

The Actor is proprietary software by Bongo Seakhoa. It is built on open-source components under permissive licences (PDFium via pypdfium2, pdfminer.six, pdfplumber, Pillow, Tesseract, httpx and the Apify SDK, among others) and uses no AGPL or GPL Python library such as PyMuPDF. certifi (MPL-2.0) is included unmodified. The Debian base image contains GPL and LGPL system tools and libraries, used unmodified as separate programs. The full list and licence texts are in the `THIRD_PARTY_NOTICES` file shipped with the Actor image.

# Actor input Schema

## `sources` (type: `array`):

Direct links to PDF files (http or https). Redirects are followed, and Dropbox '?dl=0' links are turned into downloads. Pages that only show a PDF viewer or a login form cannot be downloaded.

## `keyValueStoreId` (type: `string`):

Pick a key-value store that holds PDF files. Without record keys below, every record whose key ends in '.pdf' is converted (up to 1,000).

## `keyValueStoreKeys` (type: `array`):

Keys of PDF records in the store picked above. Leave empty to use every record whose key ends in .pdf.

## `useBundledSample` (type: `boolean`):

Adds a 3-page synthetic sample (a text page, a table page and a scanned page) to the run. None of this Actor's events are charged for it. The sample is used automatically when no sources are given.

## `mode` (type: `string`):

'convert' produces Markdown, chunks and tables. 'estimate' only counts pages, tells which pages would need OCR and shows the projected price. It charges none of this Actor's events (Apify's standard Actor start fee still applies).

## `outputs` (type: `array`):

Which outputs to produce. Markdown goes to the document row and to a .md file; chunks and tables become their own rows; tables are also saved as CSV and JSON files.

## `chunkSize` (type: `integer`):

Target maximum length of each chunk in characters (about 4 characters per token). Chunks break at headings, paragraphs and sentences, and tables stay whole when they fit.

## `chunkOverlap` (type: `integer`):

Characters repeated between consecutive chunks. At most half the chunk size.

## `includePageMarkers` (type: `boolean`):

Insert '' before the first block that starts on each page. Blank pages, and pages whose text only continues a paragraph or table from the page before, get no marker.

## `csvFormulaGuard` (type: `boolean`):

In the table CSV files, put a single quote before any cell that starts with =, +, -, @, a tab or a carriage return (also after leading spaces or invisible characters, or in full-width form), so a spreadsheet shows it as text instead of running it as a formula. Signed numbers such as -12.5, +3 or -45%, and cells holding only dashes or plus signs (a nil placeholder such as -), are left as they are. JSON rows, the Markdown and the dataset always keep the raw cell text. Turn this off only if you need the raw text in the CSV too.

## `ocrMode` (type: `string`):

'auto' runs OCR only on pages without a usable text layer (scans, including scans whose text layer is just a scanner stamp or Bates number). 'force' runs OCR on every page, and every page is then charged at the OCR price. 'off' never runs OCR; scanned pages then produce nothing and are not charged.

## `ocrLanguages` (type: `array`):

Languages for OCR (Tesseract). Choose only the languages in your documents; each extra language slows OCR down.

## `pageRange` (type: `string`):

Pages to convert, 1-based, for example '1-5,8,10-'. Empty means all pages (up to the page limit).

## `maxPages` (type: `integer`):

Pages after this limit are not converted (the document is then marked 'partial') and not charged.

## `maxFileSizeMb` (type: `integer`):

Larger files are not downloaded and are reported as failed (not charged). The limit is also capped at a quarter of the run's memory: 128 MB at 512 MB, 256 MB at the default 1 GB, 500 MB at 2 GB.

## `perDocumentTimeoutSecs` (type: `integer`):

When a document takes longer, the pages finished so far are returned as a 'partial' result; only those pages are charged.

## `password` (type: `string`):

Password for encrypted PDFs. It is used for every document in the run and is never logged or stored in the output.

## Actor input object example

```json
{
  "sources": [],
  "keyValueStoreKeys": [],
  "useBundledSample": false,
  "mode": "convert",
  "outputs": [
    "markdown",
    "chunks",
    "tables"
  ],
  "chunkSize": 1500,
  "chunkOverlap": 150,
  "includePageMarkers": false,
  "csvFormulaGuard": true,
  "ocrMode": "auto",
  "ocrLanguages": [
    "eng"
  ],
  "maxPages": 200,
  "maxFileSizeMb": 50,
  "perDocumentTimeoutSecs": 900
}
```

# Actor output Schema

## `results` (type: `string`):

One row per document (record\_type 'document'), per chunk ('chunk') and per table ('table').

## `files` (type: `string`):

Full Markdown per document and every table as CSV and JSON, in the run's key-value store.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("bongo-seakhoa/pdf-to-markdown-rag-chunks").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("bongo-seakhoa/pdf-to-markdown-rag-chunks").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call bongo-seakhoa/pdf-to-markdown-rag-chunks --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,bongo-seakhoa/pdf-to-markdown-rag-chunks"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/hY68KrMhFa0tjP6FR/builds/3p9BAFK9E3xmqWC26/openapi.json
