# Image & PDF OCR — Image to Text and Scanned PDF to Text (`keyman98/image-pdf-ocr`) Actor

Extract text from images and scanned PDFs with Tesseract OCR running inside the Actor: your files are never sent to a third-party OCR service. 16 languages, text per page with confidence scores, optional word boxes. $2 per 1,000 pages; blank pages and failed files are free.

- **URL**: https://apify.com/keyman98/image-pdf-ocr.md
- **Developed by:** [KeyMan98](https://apify.com/keyman98) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $2.00 / 1,000 page ocreds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Image & PDF OCR

Extract text from images and PDFs — including scanned, image-only documents with no text layer at all — using Tesseract OCR. Give it public URLs and get back the recognized text for every page, with a confidence score and word/character counts. Nothing is installed on your side, and no file is ever sent to a third-party OCR or vision API: recognition runs entirely inside this Actor's own container.

### What you get (output fields)

For each page — a plain image counts as one page — one dataset row with:

- `url` — the file URL you requested.
- `fileType` — detected type: `png`, `jpeg`, `tiff`, `webp`, `bmp`, `gif`, or `pdf`. Detected from the file's own bytes, never from the URL's extension.
- `page` — 1-based page/frame number within the document.
- `pageCount` — total pages/frames in the document, even if `maxPagesPerDocument` limited how many were actually processed.
- `text` — the OCR text recognized on this page.
- `meanConfidence` — Tesseract's mean word confidence for this page, 0-100. Null if no word was recognized on the page at all.
- `wordCount` / `charCount` — size of the recognized text.
- `languages` — language(s) Tesseract used for this page.
- `width` / `height` — pixel size of the image actually OCRed (the rendered PDF page, or the source image).
- `wordBoxes` — only when the *Include word boxes* input is on: every recognized word with its bounding box and confidence.
- `error` — null on success, or a message explaining why this page/document failed.

### Who it's for

- Anyone with scanned paperwork — old letters, forms, receipts, faxed contracts — who needs the text out of them without retyping.
- Developers who need OCR in a pipeline (search indexing, data extraction, archiving) without standing up Tesseract themselves.
- Anyone who wants OCR **without** sending private documents to Google Cloud Vision, AWS Textract, or a similar third-party API — see "Where does OCR actually happen?" below.

### Where does OCR actually happen?

Everything runs inside this Actor's own container: PDF pages are rasterized with **pypdfium2** (the same rendering engine behind Chrome's built-in PDF viewer), and the resulting page images are recognized with **Tesseract** (the open-source OCR engine originally developed at HP and now maintained by Google), called through `pytesseract`. Your files are never uploaded to Google Cloud Vision, AWS Textract, or any other third-party OCR/vision service — they exist only for the lifetime of the run, inside this Actor.

This Actor deliberately does not use PyMuPDF for PDF rendering, even though it's a common choice: PyMuPDF is licensed AGPL, which would require this Actor's own source to be AGPL too (or a commercial license purchased from its maker). pypdfium2 (Apache/BSD) and Tesseract (Apache 2.0) avoid that entirely.

### How to use

1. **List your file URLs** — public PNG, JPEG, TIFF (including multi-page), WebP, BMP, GIF, or PDF links.
2. **Pick the language(s)** the text is in (default: English). Pick more than one if a page mixes languages.
3. **Run the Actor.** Each page becomes one row, with its recognized text and confidence.
4. Optionally adjust *PDF render DPI* (sharper but slower for small print), *Max pages per document*, *Page segmentation mode* (for a single line/word instead of a full page), or turn on *Include word boxes* if you need per-word positions.

### Input example (JSON)

```json
{
  "urls": [
    "https://example.com/scanned-letter.jpg",
    "https://example.com/scanned-report.pdf"
  ],
  "languages": ["eng"],
  "maxPagesPerDocument": 20,
  "dpi": 300,
  "includeWordBoxes": false,
  "pageSegmentationMode": 3,
  "maxFileSizeMb": 50
}
```

### Output example (JSON)

```json
{
  "url": "https://example.com/scanned-report.pdf",
  "fileType": "pdf",
  "page": 1,
  "pageCount": 2,
  "text": "TOP SECRET/SENSITIVE\n\nTALKING POINTS\n\n1. AS WE DISCUSSED BEFORE MY MIDDLE EAST TRIP, I PROPOSED TO\nPRESIDENT SADAT, PRIME MINISTER BEGIN AND CROWN PRINCE FAHD...",
  "meanConfidence": 94.32,
  "wordCount": 187,
  "charCount": 1204,
  "languages": ["eng"],
  "width": 2550,
  "height": 3301,
  "wordBoxes": null,
  "error": null
}
```

### If a page or file fails

If a URL can't be downloaded, isn't one of the supported file types, or is an encrypted/corrupted document, that URL gets **one error row** (`page` and `pageCount` null, `error` explaining why) — the rest of your list still runs. If a document opens fine but one specific page inside it fails to render or to OCR, only **that page** becomes an error row; the other pages of the same document are unaffected. No row that has `error` set is ever charged.

### Pricing

Pay only for pages where text was actually found — nothing charged for blank pages or for any row with an error. Pricing model: **pay-per-event**.

| Event | When it's charged | Price |
| --- | --- | --- |
| `page-ocr` | a page/frame produced non-empty OCR text | 0.002 USD ($2 per 1,000 pages) |

### Limitations

- **Printed text works well; handwriting does not.** Tesseract is a printed-text OCR engine — cursive or handwritten pages will come back with low confidence and garbled text, not a clear error. This is a real limit of the engine, not a bug.
- **Quality is below commercial OCR services** (Google Lens, ABBYY, cloud vision APIs) on hard cases: low-resolution scans, unusual fonts, heavy skew, or dense multi-column layouts. Tesseract is free and runs locally; those trade-offs are the cost of that.
- **No automatic upright-rotation for photos without EXIF data.** A photo's EXIF orientation tag is corrected automatically, and a PDF page renders with its own declared rotation — but a scanned image with no EXIF tag that is genuinely sideways or upside-down is OCRed as-is and will return poor results. Rotate it before submitting if you know it's misoriented.
- **`maxPagesPerDocument` and `maxFileSizeMb` are hard cutoffs**, not smart sampling: a 100-page PDF with the default 20-page cap only returns its first 20 pages (`pageCount` still reports the true total).
- **Multi-page OCR is TIFF-only.** A multi-page TIFF becomes one row per page, like a PDF. An animated GIF or WebP is always read as a single still frame (its first one) — `pageCount` is 1 even if the file has more; this Actor treats a GIF/WebP as one image, not a document, so an animation is never charged once per frame.
- Very large pages are capped before OCR to keep memory and run time predictable: a PDF page is rendered at a reduced DPI if needed so it never exceeds a 40-megapixel render; an image over roughly 100 megapixels is rejected outright as an error row rather than processed.
- Mixing many unrelated languages in one request can *reduce* accuracy — Tesseract tries to match text against all requested languages' dictionaries at once.

### FAQ

#### Does this send my files to Google Cloud Vision, AWS Textract, or another OCR API?

No. OCR runs inside this Actor with the open-source Tesseract engine. See "Where does OCR actually happen?" above.

#### Can it read handwriting?

Not reliably. Tesseract is built and trained for printed/typed text; handwritten pages will get low confidence scores and often garbled output.

#### What's the difference between submitting an image vs. a PDF?

None from your side — list either kind of URL, mixed freely, in the same run. Internally, a PDF page is rendered to an image at the DPI you choose before OCR; a plain image is read at its native resolution.

#### Can I OCR a scanned PDF that has no text layer at all?

Yes — that's exactly the case this Actor is built for. Every PDF page is rasterized to an image and OCRed from scratch; this Actor never tries to reuse a PDF's existing embedded text layer (if it has one), so it works identically whether the PDF has one or not.

#### Why is a page free even though the run succeeded?

Pages with no recognizable text (a blank page, a mostly-graphical page) are not charged — you only pay for pages Tesseract actually extracted text from.

#### Does it support multi-page TIFF files?

Yes. Each page of a multi-page TIFF becomes its own dataset row, the same as a PDF's pages.

#### What about an animated GIF or WebP — does every frame get OCRed?

No, only the first frame. Multi-page OCR is offered for TIFF specifically; a GIF/WebP is treated as one image, so you're never charged once per frame of an animation.

#### Can I get the position of each word, not just the page text?

Yes — turn on *Include word boxes* in the input. Each row then also includes a `wordBoxes` array with every word's text, confidence, and pixel bounding box.

#### Can I use this through the Apify API or an MCP server?

Yes, like any Apify Actor — through the standard Apify API, or through the Apify MCP server if you use Claude, Cursor, or another MCP-enabled client.

### Export

Results can be downloaded from the Apify dataset as JSON, CSV, or Excel, or accessed via the Apify API.

# Actor input Schema

## `urls` (type: `array`):

Public URLs of the files to OCR. Supported: PNG, JPEG, TIFF (including multi-page), WebP, BMP, GIF, and PDF. The file type is detected from its content, not its URL extension.

## `languages` (type: `array`):

Language(s) Tesseract should read the text as. Pick more than one for a page mixing languages (e.g. English + Italian) - Tesseract will try both.

## `maxPagesPerDocument` (type: `integer`):

Stop after this many pages/frames per PDF or multi-page image (keeps cost and run time predictable on huge files).

## `dpi` (type: `integer`):

Resolution used to rasterize each PDF page before OCR. Higher is sharper (better for small print) but slower and more memory. Ignored for image inputs, which are OCRed at their native resolution.

## `includeWordBoxes` (type: `boolean`):

Also return each recognized word with its bounding box and confidence, not just the full page text.

## `pageSegmentationMode` (type: `string`):

Tesseract's page segmentation mode (--psm), as a number 0-13. The default works for most single-page documents; switch it for a single line, a single word, or sparse/irregular layouts.

## `maxFileSizeMb` (type: `integer`):

Skip (as an error row, not charged) any file larger than this.

## Actor input object example

```json
{
  "urls": [
    "https://upload.wikimedia.org/wikipedia/commons/d/d6/National_Security_Action_Memorandum_No._6_Plan_of_Reorganization_of_the_Foreign_Aid_Program_-_NARA_-_193406.jpg",
    "https://upload.wikimedia.org/wikipedia/commons/a/ad/Alexander_Haig_1981_memo_regarding_the_Middle_east.pdf"
  ],
  "languages": [
    "eng"
  ],
  "maxPagesPerDocument": 20,
  "dpi": 300,
  "includeWordBoxes": false,
  "pageSegmentationMode": "3",
  "maxFileSizeMb": 50
}
```

# Actor output Schema

## `results` (type: `string`):

All results in the default dataset (JSON, CSV, Excel).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://upload.wikimedia.org/wikipedia/commons/d/d6/National_Security_Action_Memorandum_No._6_Plan_of_Reorganization_of_the_Foreign_Aid_Program_-_NARA_-_193406.jpg",
        "https://upload.wikimedia.org/wikipedia/commons/a/ad/Alexander_Haig_1981_memo_regarding_the_Middle_east.pdf"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("keyman98/image-pdf-ocr").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": [
        "https://upload.wikimedia.org/wikipedia/commons/d/d6/National_Security_Action_Memorandum_No._6_Plan_of_Reorganization_of_the_Foreign_Aid_Program_-_NARA_-_193406.jpg",
        "https://upload.wikimedia.org/wikipedia/commons/a/ad/Alexander_Haig_1981_memo_regarding_the_Middle_east.pdf",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("keyman98/image-pdf-ocr").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://upload.wikimedia.org/wikipedia/commons/d/d6/National_Security_Action_Memorandum_No._6_Plan_of_Reorganization_of_the_Foreign_Aid_Program_-_NARA_-_193406.jpg",
    "https://upload.wikimedia.org/wikipedia/commons/a/ad/Alexander_Haig_1981_memo_regarding_the_Middle_east.pdf"
  ]
}' |
apify call keyman98/image-pdf-ocr --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,keyman98/image-pdf-ocr"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sEuIuTQzG2gJSeJJR/builds/wUKAVNY2awXGFubq4/openapi.json
