# PDF to Markdown & RAG Chunks: Document Parser (DOCX/PPTX/XLSX) (`lindenwerk/pdf-to-markdown-rag`) Actor

PDF to Markdown and RAG chunks: a document parser that turns PDF, Word (DOCX), PowerPoint (PPTX) and Excel (XLSX) into clean Markdown text for RAG and LLMs. PDF table extraction inline as Markdown tables, heading-aware chunks with page numbers. Scanned pages flagged and free. No start fee.

- **URL**: https://apify.com/lindenwerk/pdf-to-markdown-rag.md
- **Developed by:** [Lindenwerk Data](https://apify.com/lindenwerk) (community)
- **Categories:** AI, Developer tools, Business
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 document converteds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What is PDF to Markdown & RAG Chunks?

Convert **PDF, Word (DOCX), PowerPoint (PPTX) and Excel (XLSX)** files to clean **Markdown for RAG and LLMs**, built with pdfplumber, python-docx, python-pptx and openpyxl.
**USD 0.002 per document plus USD 0.0005 per page** (USD 0.50 per 1,000 pages). No start fee; failed, empty and scanned-only files are free.
No login, no API key of any other service, no LLM in the loop: the same file always gives the same Markdown.

Tables stay **inline, in reading order, as Markdown tables**, so a model sees "Los 1 | Wartung | 12" next to the
paragraph that explains it. Headings become `#`/`##`/`###`, lists stay lists, running headers and footers ("Seite 3
von 46") are removed, and words hyphenated across line breaks are joined again. If you want, every document is also
split into **heading-aware chunks with page numbers** (`pageStart`, `pageEnd`, `headingPath`), ready for a vector
database and for answers that cite "page 12, § 4 Eignung".

It is made for **German, French and other European documents**: umlauts, accents, ligatures and `ß` come through
cleanly, and the language from the PDF metadata is reported. A typical use is reading the **tender documents** found by
our German, French and UK tender monitors (see below).

### Who it's for

- **RAG and AI-agent builders** who need clean Markdown and citable chunks from PDFs, Word, PowerPoint and Excel in one
  place, through the Apify API or MCP.
- **Bid and proposal teams** who want to search, summarise or compare **tender documents** (Leistungsbeschreibung,
  cahier des charges (CCTP), statement of requirements) with an LLM.
- **Knowledge-base and search teams** who index policies, manuals and reports in German, French and English.
- **Developers** who want a hosted, pay-per-use converter instead of running their own PDF stack.

### What you get

- **One row per document** with the full Markdown, title, language, created and modified date, page count, pages with
  text, pages without a text layer, table count, word count, file size and a **sha256** for deduplication and caching.
- **Optional chunk rows** (`outputMode`): heading-aware, about `chunkSize` characters, with a small overlap inside a
  section (never across headings), page range and the heading path. Long tables are split by rows with the header
  repeated in every chunk.
- **Page markers** (`<!-- page: 12 -->`) in the Markdown, so you can still cite pages after your own chunking.
- **Formats**: PDF with a text layer, DOCX (headings from styles, numbered and bulleted lists, tables, page numbers
  from Word's rendered page breaks or estimated), PPTX (one section per slide, tables, speaker notes), XLSX (one
  Markdown table per sheet, values not formulas).
- **Honest flags instead of guesses**: scanned pages are reported as `pagesWithoutText` and `scannedPageCount`, not
  filled with OCR noise. They are not charged.

### How to convert PDFs to Markdown

1. Click **Start** with the prefilled example (a one-page W3C sample PDF with a table) to see real output in a few
   seconds, for about USD 0.0025.
2. Add your own files in **Document URLs** (public links), pick a key-value store in **Your files: key-value store**
   (see below), or send **Base64 documents** from your code.
3. Choose the **Output**: documents only, documents and RAG chunks, or chunks only. Set **Chunk size** and **Chunk
   overlap** if you use chunks.
4. Optional: a **Page range** such as `1-10, 15`, keep or remove **running headers and footers**, and limits under
   **Limits and performance**.
5. Read the results in the **Output** tab (views "Overview", "Markdown", "RAG chunks" and "Pages & quality") or
   download them as JSON, CSV or Excel. Very large documents keep their full Markdown in the run's key-value store.
6. Call it from the **Apify API**, from AI agents via MCP, or chain it after our tender monitors with Apify
   **integrations** (Make, Zapier, n8n or a webhook). Save the input as a **task** and add a **schedule** to re-convert a
   list of documents regularly (every run converts and charges again; `sha256` tells you what changed).

### Input

```json
{
  "documentUrls": [
    "https://dserver.bundestag.de/brd/2026/0225-26.pdf",
    "https://www.fedlex.admin.ch/filestore/fedlex.data.admin.ch/eli/cc/2020/126/20250101/fr/pdf-a/fedlex-data-admin-ch-eli-cc-2020-126-20250101-fr-pdf-a.pdf"
  ],
  "outputMode": "documents-and-chunks",
  "chunkSize": 1500,
  "chunkOverlap": 150,
  "pageRange": "1-20"
}
```

- `documentUrls`: public http(s) links. Apify key-value store record URLs are read with your run's permissions.
- `keyValueStoreId` + `keyValueStoreKeys`: files you put into one of your key-value stores (all records if no keys).
- `files`: API only, kept for compatibility; works like `documentUrls`.
- `base64Documents`: `[{"fileName": "angebot.docx", "contentBase64": "UEsDB..."}]` for files that are not public.
- `outputMode`: `documents` (default), `documents-and-chunks` or `chunks`.
- Limits: `maxDocuments` (100), `maxPagesPerDocument` (500), `maxFileSizeMb` (30), `maxRowsPerSheet` (1000).

#### Your own files (from your computer)

This Actor runs with **limited permissions**: it can only read the storages you pick in its input, nothing else in
your account. That is why the form has no separate upload button (a Console upload goes into a new store the run
isn't allowed to read). Instead:

1. In Apify Console open **Storage > Key-value stores**, open or create a store, and upload your PDF, DOCX, PPTX or
   XLSX files as records (the record key is the file name, e.g. `angebot.docx`).
2. In this Actor's input pick that store in **Your files: key-value store**. Leave **Record keys** empty to convert
   every file in it, or list the keys you want.
3. From code or an AI agent, send small files as **Base64 documents** instead, or pass Apify record URLs in
   `documentUrls` (they are read if the store is picked, readable by ID or the URL is pre-signed). If a record can't be
   read, its row says so (`status: failed`) and it is not charged.

### Output

A document row (shortened):

```json
{
  "rowType": "document",
  "source": "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf",
  "format": "pdf",
  "status": "converted",
  "title": "table",
  "language": "EN-US",
  "pageCount": 1,
  "pagesConverted": 1,
  "pagesWithoutText": [],
  "tableCount": 1,
  "wordCount": 99,
  "markdown": "<!-- page: 1 -->\n\nExample table\n\nThis is an example of a data table.\n\n| Disability Category | Participants | Ballots Completed | ... |\n|---|---|---|---|\n| Blind | 5 | 1 | ... |",
  "sha256": "a693998ff2a475d128c11644fbf02374249f08ac134d526a4d9b913d8b5834a5",
  "chargedEvents": {"document-converted": 1, "page-converted": 1}
}
```

A chunk row (from the German procurement acceleration act as passed by the Bundestag, Bundesrat-Drucksache 225/26):

```json
{
  "rowType": "chunk",
  "fileName": "0225-26.pdf",
  "chunkId": "aacf340838a3-0007",
  "pageStart": 4,
  "pageEnd": 4,
  "headingPath": ["Gesetzesbeschluss", "Gesetz zur Beschleunigung der Vergabe öffentlicher Aufträge", "unter Berücksichtigung einer Berichtigung in beigefügter Fassung angenommen.", "Änderung des Gesetzes gegen Wettbewerbsbeschränkungen"],
  "text": "...\n5. Nach § 97 wird der folgende § 97a eingefügt:\n\n„§ 97a Losgrundsatz\n\n- (1) Leistungen sind in der Menge aufgeteilt (Teillose) ...",
  "tokenEstimate": 327
}
```

Example source: Deutscher Bundestag/Bundesrat – DIP, BR-Drs. 225/26. Bundestag and Bundesrat documents are available
free of charge in DIP (https://dip.bundestag.de). The third heading in `headingPath` shows that PDF heading detection is
a heuristic (see the limits below).

`status` is one of `converted`, `no_text` (only scanned or blank pages, free), `failed`, `encrypted`,
`unsupported_format`, `too_large`, `robots_disallowed`, `refused`, `not_found`, `http_error`, `timeout` or
`unreachable`. Everything except `converted` is free. A full sample is in `sample-output.json`.

### Pricing (pay per event)

| Event | Price | When |
|---|---|---|
| `document-converted` | **USD 0.002** | A document with at least one page of text was converted |
| `page-converted` | **USD 0.0005** (USD 0.50 per 1,000 pages) | Each page (PDF, DOCX), slide (PPTX) or sheet (XLSX) with text |
| Free | 0 | Scanned or blank pages, failed, encrypted, unsupported, too large and robots.txt-blocked files, chunk rows |
| Actor start | USD 0.00005 | Apify platform minimum per run |

A 10-page PDF costs **USD 0.007**. A 46-page tender regulation costs USD 0.025. 100 documents of 20 pages cost USD 1.20.
There is **no custom start fee** and no minimum. Set **Maximum cost per run** in Apify Console: the Actor checks the
budget before each document, converts only the first pages that still fit, marks that document `truncated`, and
stops, so it never charges more than your limit. The smallest useful limit is USD 0.01.

### Use with AI agents (MCP)

```json
{
  "mcpServers": {
    "apify": {
      "url": "https://mcp.apify.com/?tools=lindenwerk/pdf-to-markdown-rag",
      "headers": {"Authorization": "Bearer <APIFY_TOKEN>"}
    }
  }
}
```

Example prompts:

- "Convert this tender PDF to Markdown and list every Eignungskriterium with its page number."
- "Turn these three DOCX offers into chunks and tell me which one has the longest warranty."
- "Read the price sheet in this XLSX and sum the lots."

For agents, `outputMode: "chunks"` with one document keeps answers short and citable. Over HTTP:
`POST https://api.apify.com/v2/acts/lindenwerk~pdf-to-markdown-rag/run-sync-get-dataset-items?token=<APIFY_TOKEN>`
with this input as the body:

```json
{
  "documentUrls": ["https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"],
  "outputMode": "chunks",
  "chunkSize": 1000,
  "maxDocuments": 1
}
```

### Tender documents: German, French and UK procurement

Our tender monitors return a `documentsUrl` for each notice (the link to the procurement documents). Where that link
is a direct, public PDF/DOCX/XLSX download, paste it into **Document URLs** or chain the runs with an integration:

- **German & EU Tenders Monitor**: Vergabeunterlagen in German (`deu`), for example Leistungsbeschreibung,
  Preisblatt (XLSX) and Eignungskriterien. Umlauts, `ß` and hyphenation such as "Ver-gabe" are handled.
- **France Tenders Monitor**: DCE documents in French (`fra`), for example the CCTP, CCAP or règlement de la
  consultation, with accents and French list styles intact.
- **UK Tenders Monitor**: specifications and framework guidance in English.
- **SAM.gov Government Contract Opportunities Monitor**: SAM.gov attachments need your own SAM.gov access, so download them and
  put them into a key-value store and pick it, or send them as **Base64 documents**.

Many tender platforms put documents behind a login or a click-through. This Actor doesn't log in or work around that:
download such files yourself and add them through a key-value store or base64. Public links download in seconds. In our test run on Apify at the
default 512 MB, the German procurement acceleration act (Bundesrat-Drucksache 225/26, 26 pages), the Swiss federal
procurement act in French (LMP, 44 pages) and the UK framework guidance (18 pages) converted in under two minutes
together, about one second per page. More memory gives the run more CPU, so large batches can finish faster.

### How it works and limits

- Text comes from the PDF's **text layer** (pdfplumber), tables from **ruling lines** in the PDF. Tables drawn only with
  whitespace (no lines) come out as text, not as a table. Multi-column page layouts are read line by line and can mix
  columns.
- **No OCR in v0.1.** Pages without a text layer are flagged (`scanned` when an image covers the page, otherwise
  `empty`) and are free. OCR is planned for a later version as a separate, clearly priced event.
- Headings are detected from font size and bold short lines in PDFs, and from styles in Word. Running headers and
  footers that repeat on at least half of the pages are removed (`removeHeadersFooters`).
- DOCX has no fixed pages: page numbers come from Word's last rendered page breaks when the file has them
  (`pageNumbering: "word-rendered"`), otherwise they are estimated (`"estimated"`), and pages are charged on that basis.
- Encrypted PDFs, legacy `.doc`/`.ppt`/`.xls`, OpenDocument, HTML pages and images are reported as free rows with a
  clear `statusNote`. Archives that unpack to huge sizes (zip bombs) are refused.
- In our test, 5 public PDFs with 239 pages took 28 s at 512 MB, peak memory about 120 MB.

### Data, privacy and responsible use

- **Honours robots.txt** for the user agent token `LindenwerkDocBot` and sends a clear user agent. If robots.txt
  can't be fetched because of a server error, the file isn't downloaded. No proxies, no login, no captcha solving.
- **Safe fetching**: only public http(s) hosts, every redirect checked, private and cloud-metadata addresses refused,
  size limits enforced while downloading, polite retries with Retry-After.
- **No personal data fields.** The Actor never extracts names, email addresses or phone numbers as fields, and it
  never outputs the document's author metadata. The Markdown contains the document's text as it is, so only convert
  documents you are allowed to process.
- Files are processed in your own Apify run and stored only in your run's storages.

### FAQ

**How is this different from asking an LLM to read the PDF?** It's deterministic, cheap and fast, and it keeps
tables and page numbers. Many teams convert first and give the Markdown or chunks to their model.

**Why is my scanned PDF free but empty?** It has no text layer. v0.1 doesn't run OCR, so it tells you which pages need
it instead of charging for guesses.

**Can it read password-protected files?** No. Encrypted files are reported as `encrypted` and are free.

**What chunk size should I use?** 1,000 to 2,000 characters (about 250 to 500 tokens) works well for most embedding
models. `tokenEstimate` is characters divided by 4.

**Do you keep my files?** No. They stay in your Apify account's run storage under your retention settings.

### Examples

- [Convert a PDF with tables to Markdown and RAG chunks](https://apify.com/lindenwerk/pdf-to-markdown-rag/examples/pdf-tables-to-markdown-rag-chunks): a published example task with a ready-made input. Open it, adjust the input and run it in your own Apify account.

### Other Lindenwerk Data Actors

- [German & EU Tenders: Public Procurement Monitor](https://apify.com/lindenwerk/german-eu-tender-matcher): Ausschreibungen from oeffentlichevergabe.de and TED, deduplicated and scored to your CPV codes and keywords.
- [France Tenders Monitor: BOAMP, TED & Marchés Publics](https://apify.com/lindenwerk/france-tender-matcher): marchés publics from BOAMP and TED in one deduplicated, scored list.
- [UK Tenders Monitor: Government Contracts & Find a Tender Alerts](https://apify.com/lindenwerk/uk-tender-matcher): Find a Tender notices (above and below threshold) as daily tender alerts, scored to your profile.
- [SAM.gov Government Contract Opportunities Monitor](https://apify.com/lindenwerk/sam-gov-contract-matcher): US federal bids and RFPs from SAM.gov, scored to your NAICS and set-asides.
- [Website Technology & Tech Stack Detector: BuiltWith Alternative](https://apify.com/lindenwerk/tech-stack-detector): CMS, shop system, analytics, consent manager and email provider of any website, in bulk.
- [SEO Audit Crawler: Broken Links, llms.txt & AI Bot Check](https://apify.com/lindenwerk/seo-audit-crawler): on-page SEO, broken links, llms.txt and AI crawler check for whole websites.

### Deutsch (Kurzfassung)

Dieser Actor wandelt PDF, Word (DOCX), PowerPoint (PPTX) und Excel (XLSX) in sauberes Markdown für RAG und LLMs um.
Tabellen bleiben als Markdown-Tabellen an ihrer Stelle im Text, Überschriften und Listen bleiben erhalten, Kopf- und
Fußzeilen werden entfernt. Optional gibt es Chunks mit Überschriftenpfad und Seitenzahlen für Zitate. Ideal für
Vergabeunterlagen auf Deutsch und Französisch. Preis: USD 0,002 pro Dokument plus USD 0,0005 pro Seite mit Text,
keine Startgebühr. Gescannte Seiten ohne Textebene, fehlerhafte und leere Dateien sind kostenlos. Der Actor beachtet
robots.txt und erhebt keine personenbezogenen Daten als Felder.

### Changelog

See the Changelog tab. Version 0.1 is the first public release.

*Made by Lindenwerk Data.*

# Changelog

This Actor's version history is a separate document: https://apify.com/lindenwerk/pdf-to-markdown-rag/changelog.md

# Actor input Schema

## `documentUrls` (type: `array`):

Direct links to PDF, DOCX, PPTX or XLSX files, one per line (for example the tender documents linked in your German, French, UK or SAM.gov tender monitor results). Only public http(s) addresses; robots.txt is honoured, private and internal addresses are refused.

## `keyValueStoreId` (type: `string`):

Files from your computer: in Apify Console open Storage > Key-value stores, open or create a store, upload the files as records, then pick that store here. The Actor gets read access to this store only (limited permissions).

## `keyValueStoreKeys` (type: `array`):

Which records of the picked store to convert (e.g. offer.pdf). Leave empty to convert every record up to Max documents.

## `files` (type: `array`):

API only, kept for compatibility: file links, same as Document URLs. Apify record URLs work if the store is picked in keyValueStoreId, readable by ID, or the URL is pre-signed.

## `base64Documents` (type: `array`):

For API and AI-agent calls: \[{"fileName": "offer.pdf", "contentBase64": "JVBERi0..."}]. Data URLs (data:application/pdf;base64,...) also work.

## `outputMode` (type: `string`):

Chunk rows are heading-aware and carry pageStart, pageEnd and headingPath for citations. Chunks are free: you pay per document and page either way.

## `chunkSize` (type: `integer`):

Target size of a chunk in characters (about 4 characters per token, so 1,500 = ~375 tokens). Chunks break at headings, paragraphs and sentence ends; long tables split by rows with the header repeated.

## `chunkOverlap` (type: `integer`):

Characters repeated from the end of the previous chunk inside the same section. Never crosses a heading. Must be below half the chunk size.

## `pageMarkers` (type: `boolean`):

Adds invisible  comments so you can map text back to PDF pages, slides or sheets.

## `pageRange` (type: `string`):

Optional, e.g. 1-10, 15, 20-. Applies to PDF pages, slides (PPTX), sheets (XLSX) and Word pages. Only pages in the range are converted and charged.

## `removeHeadersFooters` (type: `boolean`):

PDF: drops lines that repeat at the top or bottom of most pages (page numbers, document titles).

## `includeSlideNotes` (type: `boolean`):

PPTX: add the speaker notes of each slide as a quote below the slide.

## `maxRowsPerSheet` (type: `integer`):

XLSX: convert at most this many data rows per sheet (one sheet = one page).

## `maxDocuments` (type: `integer`):

Safety limit per run (hard maximum 5,000).

## `maxPagesPerDocument` (type: `integer`):

Pages after this limit are not converted and not charged.

## `maxFileSizeMb` (type: `integer`):

Larger files are skipped for free.

## `maxConcurrency` (type: `integer`):

Documents converted at the same time. 2 suits the default 512 MB of memory; raise memory with it.

## `requestTimeoutSecs` (type: `integer`):

Timeout per download request. 429/5xx answers are retried twice, politely.

## `documentTimeoutSecs` (type: `integer`):

A document that takes longer is returned with the pages converted so far (truncated).

## Actor input object example

```json
{
  "documentUrls": [
    "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"
  ],
  "outputMode": "documents-and-chunks",
  "chunkSize": 1500,
  "chunkOverlap": 150,
  "pageMarkers": true,
  "removeHeadersFooters": true,
  "includeSlideNotes": true,
  "maxRowsPerSheet": 1000,
  "maxDocuments": 100,
  "maxPagesPerDocument": 500,
  "maxFileSizeMb": 30,
  "maxConcurrency": 2,
  "requestTimeoutSecs": 60,
  "documentTimeoutSecs": 300
}
```

# Actor output Schema

## `dataset` (type: `string`):

Document rows with status and Markdown

## `chunks` (type: `string`):

Chunk rows with page citations

## `runSummary` (type: `string`):

Documents converted, pages charged, statuses, runtime

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documentUrls": [
        "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"
    ],
    "outputMode": "documents-and-chunks"
};

// Run the Actor and wait for it to finish
const run = await client.actor("lindenwerk/pdf-to-markdown-rag").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "documentUrls": ["https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"],
    "outputMode": "documents-and-chunks",
}

# Run the Actor and wait for it to finish
run = client.actor("lindenwerk/pdf-to-markdown-rag").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documentUrls": [
    "https://www.w3.org/WAI/WCAG21/working-examples/pdf-table/table.pdf"
  ],
  "outputMode": "documents-and-chunks"
}' |
apify call lindenwerk/pdf-to-markdown-rag --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,lindenwerk/pdf-to-markdown-rag"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/CwQgcGbe6Q70lZ6IM/builds/BSLqv0CU07iV4ArPj/openapi.json
