# Image to Text OCR (runs inside the actor) (`madrasco/image-ocr`) Actor

Extracts the text from images (receipts, scanned pages, screenshots) with an open OCR model that runs inside the actor: your images are not sent to any other service. Returns plain text plus lines and words with boxes and confidence.

- **URL**: https://apify.com/madrasco/image-ocr.md
- **Developed by:** [Jack Valmadre](https://apify.com/madrasco) (community)
- **Categories:** AI, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $3.50 / 1,000 image reads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Image to Text OCR (runs inside the actor)

Get the text out of receipts, scanned pages and screenshots with one API call. The images are read inside your own Apify run by open OCR models (PaddleOCR PP-OCRv6, run on CPU), so they are not sent to any other OCR service and you don't need a cloud vision account.

### What you get

- One result row per image: the text in reading order, plus (if you ask for them) each line and each word with its box and a confidence score.
- A `status` on every row, so a broken link doesn't stop the run: `ok`, `no_text`, `download_failed`, `not_an_image`, `too_large` or `skipped`, with the reason in `error`.
- Two modes: **accurate** (default) and **fast**, which read about four times quicker in our test runs and got 0 to 8 points fewer words right.
- Images from links (JPEG, PNG, WebP, GIF, BMP, TIFF), or from a key-value store in your Apify account where you uploaded them.

### How well it reads

We tested this build on public images we didn't use while building or tuning it, and checked the output against text that people had already transcribed or that came from the page itself.

| Image type | What we tested | Words read correctly (accurate / fast) |
|---|---|---|
| Receipt photos | 50 receipts from the CORD v2 validation set (annotated item and total lines) | 89% / 81% |
| Scanned book and document pages | 16 English pages from Wikisource, compared with their proofread text | 94% / 90% |
| Scanned French book pages | 6 pages from Wikisource | 90% / 90% (97% / 96% ignoring case and punctuation) |
| Web page screenshots | 10 screenshots of 5 public web pages, compared with the page text | 95% / 92% |

"Words read correctly" is the share of the reference words that appear, spelled exactly the same, in the output. On receipts the reference covers only the annotated lines; on book pages it is the proofread page text, which sometimes leaves out running headers or page numbers, so the true figure there may be a little higher. Going the other way, about 8% of the words the accurate mode returned on those pages did not match the reference, even ignoring case and punctuation: misreads, plus headers and page numbers the transcription leaves out.

It also reads accented Latin letters and Chinese characters, though we have tested those on small samples only: 89% of the accented words on the French pages above were found (ignoring case), and on 4 screenshots of Chinese Wikipedia pages the character error rate was about 12 to 13%.

### Limitations

- **Product packaging and labels are read poorly.** On 46 real photos of packaging (bottles, cans, boxes and shop shelves, from TextOCR / Open Images) this build read 39% of the words exactly in accurate mode and 27% in fast mode. Curved, stylised, small or distant text is the hard part. Don't rely on it for photos like these.
- **Vertical text (such as traditional Chinese book pages) comes out in the wrong order.** The characters are mostly recognised, but the columns are returned left to right instead of right to left, and punctuation is often dropped.
- **Side-by-side text can be merged.** Text that sits at the same height in two columns may be joined on one line of `text`. Use the line boxes if layout matters.
- Handwriting has not been tested.
- OCR makes mistakes. Check the output before you rely on it for anything important, and use `meanConfidence` and the per-line `confidence` to find the rows most worth checking.

### Input

```json
{
  "imageUrls": ["https://example.com/receipt.jpg", "https://example.com/page-2.png"],
  "mode": "accurate",
  "outputLevel": "lines"
}
```

- `imageUrls`: direct links to image files. `urls`, `images` and `startUrls` are accepted too, as plain strings or `{"url": ...}` objects.
- `keyValueStoreId` (optional): a key-value store in your account holding image files; every image record is read unless you list `imageKeys`.
- `mode`: `accurate` (default) or `fast`.
- `outputLevel`: `text` (plain text only), `lines` (default; text plus lines with boxes and confidence) or `words` (also each word with its box).
- `minConfidence` (default 0.5): lines the model is less sure of than this are left out.
- `maxImages` (default 1,000) and `maxFileSizeMb` (default 25): larger files are skipped with status `too_large`.

### Output

One row per image, for example (a screenshot of a GOV.UK page, `outputLevel: "lines"`, shortened):

```json
{
  "input": "https://example.com/vat-rates.png",
  "status": "ok",
  "text": "Cookies on GOV.UK\n...\nHome > Money and tax > VAT\nVAT rates\n...",
  "lineCount": 19,
  "wordCount": 94,
  "meanConfidence": 0.9936,
  "width": 1280,
  "height": 900,
  "mode": "accurate",
  "lines": [
    {"text": "Home > Money and tax > VAT", "confidence": 0.9586, "box": [155, 398, 384, 424],
     "polygon": [[155, 398], [384, 398], [384, 424], [155, 424]]}
  ]
}
```

Boxes are pixel coordinates `[left, top, right, bottom]` in the image (after any EXIF rotation is applied). A row that failed has `status` and `error` and no text. The run's `OUTPUT` record summarises how many images got each status.

### Pricing

Pay per event: US$0.0035 per image read (`image`, rows with status `ok` or `no_text`), plus a start fee of US$0.00005 per GB of run memory, charged once per run (US$0.0002 at the default 4 GB). Fast mode is charged the same price per image as accurate mode. Images that fail to download or decode are not charged. There is no extra charge for Apify platform usage. Apify shows the price before you run, and you can set a maximum cost per run: once it is reached, the remaining images get a `skipped` row and are not charged. Run it with 4 GB of memory (the default); bigger and denser images take longer.

### Using it from code or an AI agent

Call it like any Apify actor (API, Python or JavaScript client, or an MCP-connected agent) with the input above and read the default dataset. Plain `urls` work, so a tool call such as `{"urls": ["https://..."]}` is enough.

### Privacy

Your images are processed inside your own Apify run and are not sent to any other service or kept by us. Results go to your run's dataset; images are held in memory while they are read and not saved. The OCR models are built into the actor; the only downloads during a run are your images.

### Support

Open an issue on the actor's **Issues** tab and we will reply there. You can also write to support@madrasco.dev.

This actor is provided as is. It extracts text; it does not check, interpret or guarantee that text for any legal, financial, medical or compliance purpose.

OCR models: PaddleOCR PP-OCRv6 (Apache-2.0), run with RapidOCR (Apache-2.0).

Published by Madrasco. Built and supported with AI assistance; replies to issues may be AI-assisted, and a human owner can be reached on request.

# Actor input Schema

## `imageUrls` (type: `array`):

Direct links to image files (JPEG, PNG, WebP, GIF, BMP, TIFF). One result row per image.

## `keyValueStoreId` (type: `string`):

Pick a key-value store in your Apify account that holds image files. Every record whose content type is image/\* or whose key ends in an image extension is read, unless you list keys below.

## `imageKeys` (type: `array`):

Record keys to read from the key-value store above. Leave empty to read all image records.

## `mode` (type: `string`):

Accurate suits most images. Fast uses a smaller model: quicker and cheaper to run, slightly more misread words. Both read Latin-script text (including accented letters) and Chinese.

## `outputLevel` (type: `string`):

How much detail each result row carries.

## `minConfidence` (type: `string`):

Lines the model is less sure of than this (0 to 1) are left out.

## `maxImages` (type: `integer`):

Stop after this many images.

## `maxFileSizeMb` (type: `integer`):

Larger files are skipped (not charged).

## Actor input object example

```json
{
  "imageUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/f/fb/Alexander%27s_Supermarket_receipt%2C_late_twentieth_century.tif"
  ],
  "mode": "accurate",
  "outputLevel": "lines",
  "minConfidence": "0.5",
  "maxImages": 1000,
  "maxFileSizeMb": 25
}
```

# Actor output Schema

## `results` (type: `string`):

One row per image: status, text, lines (and words) with boxes and confidence.

## `summary` (type: `string`):

Images read, statuses, time per image.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "imageUrls": [
        "https://upload.wikimedia.org/wikipedia/commons/f/fb/Alexander%27s_Supermarket_receipt%2C_late_twentieth_century.tif"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("madrasco/image-ocr").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "imageUrls": ["https://upload.wikimedia.org/wikipedia/commons/f/fb/Alexander%27s_Supermarket_receipt%2C_late_twentieth_century.tif"] }

# Run the Actor and wait for it to finish
run = client.actor("madrasco/image-ocr").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "imageUrls": [
    "https://upload.wikimedia.org/wikipedia/commons/f/fb/Alexander%27s_Supermarket_receipt%2C_late_twentieth_century.tif"
  ]
}' |
apify call madrasco/image-ocr --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,madrasco/image-ocr"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WLkgRUZM1IMetQzio/builds/Kjmw26j5eNw0nTHhU/openapi.json
