# PDF to Markdown, Tables & OCR Extractor (`cylindrical_lighthouse/document-intelligence`) Actor

Convert PDFs and scanned documents to Markdown, header-mapped tables (JSON + CSV) and schema JSON, with OCR and page-level provenance.

- **URL**: https://apify.com/cylindrical\_lighthouse/document-intelligence.md
- **Developed by:** [Lighthouse Data](https://apify.com/cylindrical_lighthouse) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.30 / 1,000 page (text layer)s

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## PDF to Markdown, Tables & OCR Extractor

Turn PDFs into **clean Markdown, header-mapped tables (JSON + CSV) and schema-guided JSON**, including **scanned documents** (OCR) and **borderless financial and statistical tables** that plain text extractors flatten into word soup. Every result carries **page numbers, SHA-256 and source URL**, so RAG pipelines, data teams and AI agents can cite exactly where a value came from.

- **Tables that survive:** ruled grids *and* borderless tables with multi-line and spanning headers, leader dots and footnote markers come out as rows keyed by column name, plus CSV.
- **Scans that read:** pages without a text layer are OCR-ed automatically (Tesseract 5); grid lines are detected and removed first, so scanned tables keep their structure.
- **Page provenance:** per-page Markdown, `pageNumber` on every table, page citations for extracted fields.
- **Chains with your crawler:** feed it the dataset of Website Content Crawler (or any Actor) or files in a key-value store.
- **Pay per page:** $0.0004 per text page, $0.008 per OCR page. No subscription.

> Try it: the default input converts the Federal Reserve's 3-page *G.19 Consumer Credit* release in about 5 seconds, for about $0.001.

### What does it do?

For each PDF it downloads (or reads from Apify storage), the Actor:

1. Reads the text layer with exact positions (pdf.js) and rebuilds the **reading order**, including multi-column layouts.
2. **Detects tables**: from drawn cell borders (ruled tables) and from column alignment (borderless tables), then maps the header rows onto column names.
3. **OCRs** pages that have no usable text layer (`ocrMode: auto`), or every page (`always`).
4. Writes **one dataset item per document**: full Markdown, per-page Markdown, tables (JSON rows + CSV), metadata (title, author, dates, page count, language) and provenance (`url`, `sha256`, `scrapedAt`).
5. Optionally sends the Markdown to **your own** Anthropic or OpenAI key to fill a **JSON Schema** you provide, with **page citations** per field.

Failed documents (404, not a PDF, password-protected, corrupt, too large) get a clear error record, and the rest of the run carries on.

### Why use it?

| | This Actor | Plain PDF-to-text tools |
|---|---|---|
| Borderless tables (financial statements, statistics) | Rows keyed by header, CSV | Numbers run together on one line |
| Scanned pages | OCR with grid-line removal | Empty text |
| Multi-column layouts | Column-aware reading order | Lines interleaved across columns |
| Provenance | Page numbers, SHA-256, source URL | Usually none |
| Structured JSON | Your JSON Schema + page citations (BYO key) | No |
| Price | $0.0004 per text page, $0.008 per OCR page | Varies |

**Measured accuracy** (open benchmark, see [Accuracy](#accuracy-benchmark)): 99–100% table cell accuracy on synthetic, Federal Reserve and Census tables, 99.6% character accuracy on a noisy 3-column scan.

### Use cases

- **RAG and AI agents:** convert PDFs found by a crawler into Markdown chunks with page numbers for citations.
- **Data extraction from reports:** statistical releases, annual reports, price lists and regulatory filings, straight into rows and CSV.
- **Scanned archives:** OCR old reports, letters and forms into searchable text.
- **Invoices and forms at scale:** schema-guided JSON (invoice number, totals, dates) with your own LLM key.
- **Monitoring:** schedule a run on a URL that publishes a new PDF each month and diff the tables.

### How to use it

1. Click **Try for free** (or open the Actor in Apify Console).
2. Paste one or more PDF URLs into **Document URLs**, or set **Dataset ID** to chain from another Actor.
3. Keep **OCR mode** on `auto`. Set **Max pages per document** to cap cost on very long files.
4. Click **Start**. Download results as JSON, CSV or Excel, or use the **Tables** and **Markdown** views.

**Example input**

```json
{
  "documentUrls": [{ "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf" }],
  "ocrMode": "auto",
  "maxPagesPerDocument": 50
}
```

#### Chain it after Website Content Crawler

Website Content Crawler can save the PDFs it finds (`saveFiles`) or list their URLs in its dataset. Pass that dataset to this Actor:

```json
{ "datasetId": "<WCC run's default dataset ID>", "datasetUrlField": "url", "maxDocuments": 500 }
```

In Apify Console you can do this automatically: on the crawler's task, add an **integration → Run Actor** with this Actor and the input above (use `{{resource.defaultDatasetId}}` as `datasetId`). Combined with a **schedule**, you get a weekly "crawl site → convert new PDFs" pipeline. Unchanged files can be skipped downstream by `sha256`.

Files already in a key-value store (for example WCC's `saveFiles` output) work too:

```json
{ "keyValueStoreId": "<store ID>", "keyValueStoreKeys": ["report-2025.pdf", "report-2026.pdf"] }
```

#### Call it from code or an AI agent

- **API:** `POST https://api.apify.com/v2/acts/<username>~document-intelligence/run-sync-get-dataset-items?token=<TOKEN>` with the input as JSON. The response is the dataset items.
- **MCP:** add the Actor to the [Apify MCP server](https://mcp.apify.com) and let the agent call it with `documentUrls`. Field descriptions are written for tool use.

### Input

| Field | Type | Description | Example |
|---|---|---|---|
| `documentUrls` | array | Public PDF URLs (text or scanned). One dataset item per URL. | `[{"url": "https://…/report.pdf"}]` |
| `ocrMode` | string | `auto` (OCR only pages without text), `always`, `never` | `auto` |
| `ocrLanguage` | string | Tesseract language codes joined with `+`. English is built in; others download on demand. | `eng`, `deu`, `eng+fra`, `chi_sim` |
| `maxPagesPerDocument` | integer | Pages processed per document, from page 1 (default 200) | `50` |
| `maxDocuments` | integer | Documents per run, including failed ones (default 100) | `100` |
| `datasetId` | string | Read URLs from an Apify dataset (chaining) | `aBcD1234EfGh5678` |
| `datasetUrlField` | string | Field with the URL; dot paths and arrays allowed (default `url`) | `url` |
| `keyValueStoreId` | string | Key-value store holding PDF files | `username~my-store` |
| `keyValueStoreKeys` | array | Record keys of the PDFs in that store | `["invoice-001.pdf"]` |
| `extractionSchema` | object | Optional JSON Schema of data to extract with your LLM key | see below |
| `extractionInstructions` | string | Extra guidance for the model | `Amounts in USD` |
| `llmProvider` | string | `anthropic` or `openai` | `anthropic` |
| `llmModel` | string | Model ID exactly as in your provider's docs (required with a schema) | |
| `llmApiKey` | secret | Your API key; stored encrypted by Apify, never logged or output | |
| `maxFileSizeMb` | integer | Skip larger files with an error record (default 100) | `100` |
| `proxyConfiguration` | object | Only if a server blocks direct downloads | `{"useApifyProxy": true}` |

**Schema extraction example**

```json
{
  "documentUrls": [{ "url": "https://example.com/invoice-0042.pdf" }],
  "extractionSchema": {
    "type": "object",
    "properties": {
      "invoiceNumber": { "type": "string" },
      "invoiceDate": { "type": "string", "description": "ISO date" },
      "total": { "type": "number" }
    },
    "required": ["invoiceNumber", "total"]
  },
  "llmProvider": "anthropic",
  "llmModel": "<model ID from your provider>",
  "llmApiKey": "<your key>"
}
```

### Output

One item per document. Real output for a one-page invoice, shortened:

```json
{
  "url": "https://example.com/invoice-0042.pdf",
  "status": "succeeded",
  "error": null,
  "fileName": "invoice-0042.pdf",
  "sha256": "5b884246dca561b0f49300e95d84dc3a1acde5552eb292ade884ec503a18b259",
  "metadata": { "title": "Invoice INV-2026-0042", "author": null, "pageCount": 1, "language": null, "creationDate": "2026-09-29T20:25:15.000Z" },
  "pageCount": 1,
  "pagesProcessed": 1,
  "pagesOcr": 0,
  "truncated": false,
  "markdown": "# INVOICE\n\nInvoice number: INV-2026-0042\n\n…\n\n| Description | Qty | Unit price | Amount |\n| --- | --- | --- | --- |\n| Widget, standard | 10 | 4.50 | 45.00 |\n…",
  "pages": [
    { "pageNumber": 1, "markdown": "# INVOICE …", "tableIds": ["p1-t1"], "ocrUsed": false, "ocrConfidence": null, "error": null }
  ],
  "tables": [
    {
      "tableId": "p1-t1",
      "pageNumber": 1,
      "detection": "aligned",
      "columns": ["Description", "Qty", "Unit price", "Amount"],
      "rows": [
        { "Description": "Widget, standard", "Qty": "10", "Unit price": "4.50", "Amount": "45.00" },
        { "Description": "Total", "Qty": null, "Unit price": null, "Amount": "76.95" }
      ],
      "csv": "Description,Qty,Unit price,Amount\r\n\"Widget, standard\",10,4.50,45.00\r\n…"
    }
  ],
  "extracted": null,
  "scrapedAt": "2026-09-29T21:33:46.868Z"
}
```

- Missing values are `null`, never empty strings. Table cells are strings exactly as printed (`"1,234.5"`), so nothing is lost. Footnote markers are kept as superscripts (`Nonrevolving³`).
- `status` is `succeeded`, `partial` (some pages failed, or the spending limit or run timeout was reached) or `failed`, with an `error.code` such as `NOT_FOUND`, `NOT_PDF`, `ENCRYPTED`, `CORRUPT_PDF` or `FILE_TOO_LARGE`.
- A run summary is saved in the key-value store as `SUMMARY`.
- Dataset views: **Overview** (one row per document), **Tables** (one row per table with CSV), **Markdown**.

### Pricing

Pay per event: you pay only for pages processed.

| Event | Price | When |
|---|---|---|
| Page (text layer) | **$0.0004** | Each page with a text layer, converted with its tables |
| Page (OCR) | **$0.008** | Each scanned or image-only page read with OCR |
| Page (schema extraction) | **$0.005** | Each page sent to your LLM for `extractionSchema`. LLM tokens are billed to your key by your provider |
| Actor start | $0.00005 per GB of memory | Apify's standard start event ($0.00005 per run at the default 1 GB) |

Failed documents and failed pages are **not charged**. Set a maximum cost per run in Console and the Actor stops cleanly when it is reached.

**Cost examples**

| Job | Pages | Cost |
|---|---|---|
| Default input (G.19, 3 text pages) | 3 text | about **$0.001** |
| 100-page annual report | 100 text | **$0.04** |
| 20-page scanned contract | 20 OCR | **$0.16** |
| 1,000 mixed pages (80% text, 20% scans) | 800 text + 200 OCR | **$2.40** |
| 50 one-page invoices with schema extraction | 50 text + 50 extraction | **$0.27** + your LLM tokens |

Platform usage (compute) is included in these prices.

New to Apify? A free Apify account includes monthly platform credit.

### Accuracy benchmark

Measured on an open corpus of public-domain and synthetic PDFs with hand-checked ground truth The same metrics run as regression tests on every change.

| Test | Metric | Result |
|---|---|---|
| Synthetic ruled grid tables (98 cells) | cell accuracy | **100%** |
| Synthetic borderless + booktabs tables (94 cells) | cell accuracy | **100%** |
| Federal Reserve G.19, rotated page, footnotes (126 cells) | cell accuracy | **99.2%** |
| Census P60-280 Table A-3, leader dots, 13 columns (156 cells) | cell accuracy | **99.4%** |
| Federal Reserve H.8 p.3, **held out** from tuning (91 cells) | cell accuracy | **92.3%** strict; every numeric cell correct, line numbers get their own column |
| Two-column article | text accuracy (reading order) | **100%** |
| Scan: 3-column IRS page, 200 dpi | character accuracy | **99.6%** |
| Scan: 1949 NACA report | word recall | **98.9%** |
| Scan: ruled table with 0.6° skew and noise | cell accuracy | **91–99%** |

In the same benchmark, a pdfplumber-based pipeline averaged 59.8% cell accuracy on the tables (it missed borderless tables entirely), and Docling averaged 80.9% while needing 2+ GB RAM and minutes per page on CPU.

### FAQ

**Which files are supported?** PDF, both text PDFs and scans or image-only PDFs. DOCX, PPTX, XLSX and HTML are planned. Other files get a `NOT_PDF` error record.

**Do I need to choose OCR?** No. `auto` OCRs only pages without a usable text layer. Use `always` for scans that carry a poor embedded text layer (common with old OCR software).

**Which OCR languages?** Any of Tesseract's ~100 languages: `eng` is built in, others (`deu`, `fra`, `spa`, `chi_sim`, `jpn`, `ara`…) are downloaded once per run. Combine them with `+`.

**Are numbers converted to numbers?** Table cells are kept as printed strings, so formats like `(1,234)`, `n.a.` and footnote markers are not lost. Convert downstream as needed.

**How are password-protected PDFs handled?** They are reported with `ENCRYPTED`. The Actor never tries to break or bypass passwords. PDFs that are only "owner-locked" (printing or copying restricted) are processed.

**What about my LLM key?** It is an encrypted secret input. It is never logged or written to the output, and it is only sent to the provider you chose. Without an `extractionSchema`, no LLM is called.

**Is it legal?** You decide which documents to process; only use documents you are allowed to access. This Actor does not log in, solve CAPTCHAs or bypass paywalls. Documents may contain personal data, which is protected by the GDPR and other regulations: process personal data only with a legitimate reason, and consult your lawyers if unsure.

**Privacy and retention:** documents are processed in your own Apify run. Nothing is stored anywhere else, and results live only in your run's storage under your account's data retention settings. When schema extraction is on, page text is sent to your chosen LLM provider under your own key.

### Limitations

- PDF only in this version.
- Very complex tables (nested tables, cells with several lines of text in borderless tables, heavily merged cells) may need clean-up. Spanning group headers are repeated into each column name they cover, which is sometimes approximate.
- Charts and images are not described; text inside images is read only when the page is OCR-ed.
- OCR is slower than text extraction: about 15–20 s per page at the default 1 GB memory (about 12 s at 2 GB). Text pages take a fraction of a second. For scans longer than about 150 pages, raise the run timeout (default 1 hour) or the memory. If the run gets close to its timeout, the Actor stops early and still delivers the finished pages (`truncatedReason: "runTimeout"`). Handwriting is not supported.
- Text in a page's text layer is used as-is: scans with a bad embedded OCR layer need `ocrMode: always`.
- Schema extraction sends at most about 400,000 characters (about 100k tokens) per document; later pages are left out and listed in `extracted.pagesUsed`.
- Dataset items are limited to 9 MB. For huge documents the Markdown (and, if needed, the tables) move to the key-value store; see `markdownStoreKey`.

### Example run

Ready-made examples are on the **Example tasks** tab, for instance "Extract tables from a PDF report to JSON and CSV" and "OCR a scanned PDF to Markdown".

### Changelog

See [CHANGELOG.md](./CHANGELOG.md).

### Support

Found a bug or need a feature, such as another document format? Open an issue on the **Issues** tab. We reply within 48 hours. Please include the run ID and, if possible, a public link to the PDF.

# Changelog

This Actor's version history is a separate document: https://apify.com/cylindrical\_lighthouse/document-intelligence/changelog.md

# Actor input Schema

## `documentUrls` (type: `array`):

Public PDF URLs to convert (text or scanned). Each URL becomes one dataset item with Markdown, pages and tables. Example: https://www.federalreserve.gov/releases/g19/current/g19.pdf. Duplicates are processed once; no login or paywalled URLs.

## `ocrMode` (type: `string`):

When to run OCR. "auto" (default): only pages without a usable text layer (scans). "always": every page, for scans with a poor embedded text layer; charged as OCR pages. "never": text layer only, cheapest. Example: auto.

## `ocrLanguage` (type: `string`):

Tesseract language code(s) of the scanned text, combined with +. Examples: eng, deu, fra, spa, eng+deu, chi\_sim. English is built in; other languages are downloaded once per run (1-10 MB).

## `maxPagesPerDocument` (type: `integer`):

Maximum pages to process per document, counted from page 1. Caps cost for long PDFs; later pages are skipped and the item is marked truncated. Example: 50.

## `maxDocuments` (type: `integer`):

Maximum documents to process in this run, including failed ones. Example: 100.

## `datasetId` (type: `string`):

Read document URLs from an Apify dataset, e.g. the output of Website Content Crawler or any crawler. Example: aBcD1234EfGh5678 or username~my-dataset.

## `datasetUrlField` (type: `string`):

Field of each dataset item that holds the document URL (dot paths allowed; arrays of URLs are expanded). Example: url, or links.pdf.

## `keyValueStoreId` (type: `string`):

Key-value store that holds PDF files, e.g. files saved by Website Content Crawler with saveFiles. Example: aBcD1234EfGh5678 or username~my-store.

## `keyValueStoreKeys` (type: `array`):

Record keys of the PDF files in the key-value store above. Example: invoice-001.pdf.

## `extractionSchema` (type: `object`):

Optional JSON Schema describing the data to extract from each document with your own LLM API key. Example: { "type": "object", "properties": { "invoiceNumber": { "type": "string" }, "total": { "type": "number" } } }. The result is in "extracted" with page citations per field. Leave empty to skip (no LLM calls, no extraction charge).

## `extractionInstructions` (type: `string`):

Optional extra guidance for the model. Example: Amounts are in USD; use ISO dates (YYYY-MM-DD).

## `llmProvider` (type: `string`):

Which provider your API key belongs to. Example: anthropic.

## `llmModel` (type: `string`):

The model ID exactly as listed in your provider's model documentation (docs.anthropic.com or platform.openai.com). Required when an extraction schema is set.

## `llmApiKey` (type: `string`):

Your Anthropic or OpenAI API key. Stored encrypted by Apify, never logged and never written to the output. Required when an extraction schema is set. Example format: sk-...

## `maxFileSizeMb` (type: `integer`):

Skip files larger than this, with a clear error record. Very large files need more memory. Example: 100.

## `proxyConfiguration` (type: `object`):

Not needed for most public documents. Enable only if a server blocks direct downloads. Example: Apify datacenter proxy.

## Actor input object example

```json
{
  "documentUrls": [
    {
      "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf"
    }
  ],
  "ocrMode": "auto",
  "ocrLanguage": "eng",
  "maxPagesPerDocument": 50,
  "maxDocuments": 100,
  "datasetUrlField": "url",
  "keyValueStoreKeys": [],
  "extractionSchema": {},
  "llmProvider": "anthropic",
  "maxFileSizeMb": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `documents` (type: `string`):

One record per document: Markdown, pages, tables, metadata, OCR usage and source provenance (or an error record for a document that could not be processed).

## `tables` (type: `string`):

Extracted tables as header-mapped JSON rows plus CSV, with page numbers.

## `markdown` (type: `string`):

Full Markdown text per document.

## `files` (type: `string`):

Key-value store holding per-document Markdown and table files, plus the SUMMARY record.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "documentUrls": [
        {
            "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf"
        }
    ],
    "ocrMode": "auto",
    "ocrLanguage": "eng",
    "maxPagesPerDocument": 50,
    "maxDocuments": 100,
    "datasetUrlField": "url",
    "keyValueStoreKeys": [],
    "extractionSchema": {},
    "llmProvider": "anthropic",
    "maxFileSizeMb": 100,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("cylindrical_lighthouse/document-intelligence").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "documentUrls": [{ "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf" }],
    "ocrMode": "auto",
    "ocrLanguage": "eng",
    "maxPagesPerDocument": 50,
    "maxDocuments": 100,
    "datasetUrlField": "url",
    "keyValueStoreKeys": [],
    "extractionSchema": {},
    "llmProvider": "anthropic",
    "maxFileSizeMb": 100,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("cylindrical_lighthouse/document-intelligence").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "documentUrls": [
    {
      "url": "https://www.federalreserve.gov/releases/g19/current/g19.pdf"
    }
  ],
  "ocrMode": "auto",
  "ocrLanguage": "eng",
  "maxPagesPerDocument": 50,
  "maxDocuments": 100,
  "datasetUrlField": "url",
  "keyValueStoreKeys": [],
  "extractionSchema": {},
  "llmProvider": "anthropic",
  "maxFileSizeMb": 100,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call cylindrical_lighthouse/document-intelligence --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,cylindrical_lighthouse/document-intelligence"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/9PUwtyVeQpZvtYxbx/builds/mk0nYBYEdXFsru53u/openapi.json
