# Scanned PDF/Image OCR to Markdown (`ingenious_quip_bxq/scanned-ocr-to-markdown`) Actor

OCR scanned PDFs and images into Markdown with RapidOCR (CJK-capable, onnxruntime). Page markers, optional RAG chunks, geometric table assembly. No external AI API keys.

- **URL**: https://apify.com/ingenious\_quip\_bxq/scanned-ocr-to-markdown.md
- **Developed by:** [新世紀書僮](https://apify.com/ingenious_quip_bxq) (community)
- **Categories:** AI, Developer tools, Open source
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.00 / 1,000 page ocr'ds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Scanned PDF/Image OCR to Markdown

**Turn scanned PDFs and images into Markdown** with RapidOCR (onnxruntime) — CJK-capable, no external AI API keys.

This Actor rasterizes each PDF page (or loads an image), runs offline OCR, and writes Markdown with optional page markers and RAG chunks. Field names for chunking mirror the digital PDF/DOCX Actor where practical.

### What you get

- OCR for **scanned PDFs** and common **images** (PNG, JPEG, WebP, TIFF, …)
- Default engine: **RapidOCR + onnxruntime** (Apache-2.0 / MIT stack; see `NOTICE`)
- Optional **page markers** and **RAG chunks** (`chunkSize` / `chunkOverlap` / `chunkUnit`)
- Best-effort **table assembly** from OCR box geometry — **not** the digital PDF word-index table repair used by the PDF/DOCX Actor
- Dataset items + optional `.md` key-value store files

### Measured cloud runs (2026-09-30, Asia/Taipei)

Own runs on this Actor, build **0.1.3**, memory **2048 MB**. Platform cost = `usageTotalUsd` re-read after finish (not PPE revenue). Public-domain LOC/WDL scans only.

| Run | Input | Pages | Wall | Peak RSS | CU | Platform USD |
|---|---|---|---|---|---|---|
| `huicge2RO7IJnpaQs` | Emancipation Proclamation (image PDF) | 5 | 109 s | 1077 MB | 0.061 | $0.0123 |
| `OdyOeoSfhaaghyr9t` | Furness sermon 1841, pages 1–8 | 8 | 42 s | 1361 MB | 0.023 | $0.0048 |
| `e5xfmZVXzIZwjQexI` | RapidOCR CJK+EN sample JPG | 1 | 9.5 s | 310 MB | 0.005 | $0.0012 |

Dense-page platform cost ≈ **$0.0025 / page** on the Emancipation scan (large page images @ 200 DPI). Sparse cover pages are cheaper. These runs confirm OCR completes and writes dataset items; they are **not** an accuracy benchmark. Old print and manuscript-style pages show typical OCR noise (misread characters). Do not expect digital-PDF table quality.

### Use cases

- Ingest **scanned** reports, pamphlets, and image-only PDFs into Markdown for RAG or note-taking
- Batch OCR of screenshots / phone photos of documents (print, not handwriting)
- CJK + Latin mixed scans where a local RapidOCR stack is enough (no cloud vision API)

### How to use

1. Paste direct URLs to scanned PDFs/images, or upload files.
2. Choose `markdown`, `chunks`, or both.
3. Optionally set `pageRange`, `dpi` (PDF), and `maxPagesPerDocument`.

#### Example input

```json
{
  "urls": [{ "url": "https://example.com/scan.pdf" }],
  "outputFormat": "markdown",
  "pageMarkers": "comment",
  "dpi": 200,
  "maxPagesPerDocument": 10
}
```

#### Example output (dataset `document` item, fields abbreviated)

```json
{
  "type": "document",
  "status": "success",
  "fileName": "scan.pdf",
  "pageCount": 5,
  "pagesConverted": 5,
  "stats": { "ocrItems": 105, "tables": 0, "words": 728, "dpi": 200 },
  "markdown": "<!-- page: 1 -->\n\n…",
  "markdownKey": "scan.md"
}
```

### Pricing

Pay-per-event:

| Event | Price |
|---|---|
| Actor start | Apify default ($0.00005 / GB memory) |
| Document OCR'd | **$0.005** per successful document |
| Page OCR'd (primary) | **$0.006** per page |

Failed downloads / unreadable files are **never** charged. RAG chunking does not add events.

Worked example: 1 document × 10 pages → $0.005 + 10 × $0.006 = **$0.065** (plus the synthetic start event).

### Known limits

- **Handwriting** is unreliable; this Actor targets print scans.
- **Japanese / Korean**: RapidOCR can return some text, but this release was **not** accuracy-tested on JP/KR corpora — treat as best-effort.
- Complex multi-column layouts may read top-to-bottom incorrectly.
- **Table assembly ≠ digital repair**: geometry from OCR boxes only; do not expect the PDF/DOCX Actor's word-index table quality.
- Very large / high-DPI PDFs need more memory and time (default run memory **2048 MB**; multipage LOC peaks observed ~1.1–1.4 GB RSS).
- Prefer the digital PDF/DOCX Actor when a text layer already exists.

### License & source code

This Actor is open source under the **GNU Affero General Public License v3.0 (AGPL-3.0)** — see `LICENSE`. The full source code is public: https://github.com/xbox002000/scanned-ocr-to-markdown

Third-party notices: `NOTICE`.

# Changelog

This Actor's version history is a separate document: https://apify.com/ingenious\_quip\_bxq/scanned-ocr-to-markdown/changelog.md

# Actor input Schema

## `urls` (type: `array`):

Direct links to scanned PDFs or images (png/jpg/webp/tiff). Prefer direct download links.

## `files` (type: `array`):

Upload scanned PDFs or images (stored in your Apify key-value store).

## `outputFormat` (type: `string`):

`markdown`: one dataset item per document. `chunks`: RAG-ready chunks. `markdown_and_chunks`: both.

## `pageMarkers` (type: `string`):

How page boundaries appear in Markdown.

## `saveMarkdownFiles` (type: `boolean`):

Save each document as a .md file in the run's key-value store.

## `chunkSize` (type: `integer`):

Maximum chunk size (unit below). Same field names as the digital PDF/DOCX Actor.

## `chunkOverlap` (type: `integer`):

How much trailing context from the previous chunk is repeated (same unit as chunk size, max 50% of chunk size).

## `chunkUnit` (type: `string`):

`tokens` uses ≈4 characters per token; `characters` is exact.

## `prependHeadingPath` (type: `boolean`):

Adds a `Section: …` line to each chunk's text when headings were detected.

## `pageRange` (type: `string`):

e.g. `1-3`, `1`, or `2-`. Empty = all pages.

## `maxPagesPerDocument` (type: `integer`):

Safety cap on pages OCR'd (and billed) per PDF. 0 = no cap.

## `dpi` (type: `integer`):

Rasterization DPI before OCR. Higher = slower/more RAM, often better on small text. Range 72–400.

## `repairTables` (type: `boolean`):

Cluster wide-gap OCR lines into Markdown tables. Best-effort geometry; not the digital-PDF word-index repair.

## `pdfPassword` (type: `string`):

Password for encrypted PDFs.

## Actor input object example

```json
{
  "outputFormat": "markdown",
  "pageMarkers": "comment",
  "saveMarkdownFiles": true,
  "chunkSize": 1000,
  "chunkOverlap": 150,
  "chunkUnit": "tokens",
  "prependHeadingPath": false,
  "maxPagesPerDocument": 100,
  "dpi": 200,
  "repairTables": true
}
```

# Actor output Schema

## `documents` (type: `string`):

Default dataset. `type: "document"`: source, fileName, format, status, title, pageCount, pagesConverted, stats (ocrItems, tables, words…), markdown, markdownKey/markdownUrl, chunkCount, processingMs, warnings, error. `type: "chunk"` when chunk modes are on. Views: overview, markdown, chunks.

## `markdownFiles` (type: `string`):

Key-value store .md records when saveMarkdownFiles is true.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {};

// Run the Actor and wait for it to finish
const run = await client.actor("ingenious_quip_bxq/scanned-ocr-to-markdown").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {}

# Run the Actor and wait for it to finish
run = client.actor("ingenious_quip_bxq/scanned-ocr-to-markdown").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{}' |
apify call ingenious_quip_bxq/scanned-ocr-to-markdown --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,ingenious_quip_bxq/scanned-ocr-to-markdown"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ql19hFjkQ9YqYbe9x/builds/7aXZ1XBoyx7sF68p3/openapi.json
