# Image to Text OCR: $1.20/1K Images, 30 Languages, No Login (`conserving_celerytop/image-to-text-ocr`) Actor

Extract text from images with OCR for $1.20 per 1,000 images. Reads receipts, screenshots, scans and photos in PNG, JPEG, WebP, TIFF, BMP or GIF. Returns text, confidence, lines and word boxes. 30 languages, auto-rotate. No text found, no charge.

- **URL**: https://apify.com/conserving\_celerytop/image-to-text-ocr.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.20 / 1,000 image with text reads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Image to Text OCR

Image to Text OCR extracts text from images. Give it links to receipts, screenshots, scanned pages, photos of documents or labels, and it returns one row per image with the text it read, an average confidence score, and optional line and word positions. It reads PNG, JPEG, WebP, TIFF (including multi-page TIFF), BMP and GIF files in 30 languages. It costs **$1.20 per 1,000 images**, and images where no text is found are free.

The OCR runs inside the Actor. Your images are not sent to Google, to a cloud vision API or to any other outside service, and nothing is kept after the run apart from the results in your own Apify storage.

### What the image to text OCR returns

- **Text**: `text`, all text read from the image, with line breaks where the lines break and a blank line between blocks. Also `wordCount`, `lineCount` and `characterCount`.
- **Confidence**: `meanConfidence`, the average word confidence from 0 to 100. Clean printed text scores above 90; under about 60 usually means a blurry, tiny or handwritten image.
- **Lines** (turn on **Include lines**): each line with its `text`, `confidence` and position (`left`, `top`, `width`, `height` in pixels).
- **Word boxes** (turn on **Include word boxes**): each word with its `confidence` and bounding box, for highlighting, redaction or form work.
- **Image details**: `imageFormat`, `imageWidth`, `imageHeight`, `fileSizeBytes`, `pageCount` and `rotationApplied` (how far the page was turned to read it).
- **Status for every image**: `status` is `ok`, or says why there is no text (for example `no_text`, `not_found`, `not_image`, `unsupported_format`, `file_too_large`), with a plain-language `error`.

### How to extract text from an image, step by step

1. Open the Actor and go to the **Input** tab.
2. In **Image URLs**, paste direct links to image files, one per line. A link to a web page that shows an image is not a file link: right-click the image, copy the image address, and use that.
3. If your images are not online, paste them in **Image data (base64)** as base64 strings or data URIs. This is the easy route from Make, Zapier, n8n or your own code.
4. Pick the **Languages** printed in the images (English is the default). Choose only the ones you need.
5. Leave **Page layout** on Automatic. If a receipt or screenshot comes out in the wrong order, try **Single block of text**; for signs, labels and memes, try **Sparse text**.
6. Click **Start**. The example input reads one small image in a few seconds.
7. Open the **Output** tab. The **Overview** view shows the text per image; the **Lines** and **Word boxes** views show positions when you turned them on. Download as JSON, CSV or Excel, or read the results through the Apify API.

### How much does image OCR cost?

You pay per image read, with the pay-per-event event `image-ocr`:

| What | Price |
| --- | --- |
| Image with text found | $0.0012 ($1.20 per 1,000) |
| Each page of a multi-page TIFF file | $0.0012 |
| Image with no text, broken link, unsupported file | free |

Apify also charges a small fixed amount when a run starts.

**Worked example:** a month of expense receipts, 2,500 photos, of which 40 are blank or unreadable and 10 links are broken, costs 2,450 x $0.0012 = **$2.94**. A batch of 200 screenshots costs $0.24.

Set a spending limit on the run if you like: the Actor reads only the images that fit, lists the ones it did not start in the run statistics, and stops cleanly.

### Input example

```json
{
    "imageUrls": ["https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"],
    "languages": ["eng"],
    "pageLayout": "auto",
    "autoRotate": true,
    "includeLines": false,
    "includeWords": false,
    "maxImages": 1000
}
```

### Output example

A real row from the example input:

```json
{
    "inputIndex": 1,
    "imageUrl": "https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png",
    "source": "url",
    "status": "ok",
    "text": "This is a lot of 12 point text to test the\nocr code and see if it works on all types\nof file format.\n\nThe quick brown dog jumped over the\nlazy fox. The quick brown dog jumped\nover the lazy fox. The quick brown dog\njumped over the lazy fox. The quick\nbrown dog jumped over the lazy fox.",
    "meanConfidence": 96.1,
    "wordCount": 60,
    "lineCount": 8,
    "characterCount": 285,
    "languages": ["eng"],
    "pageLayout": "auto",
    "rotationApplied": 0,
    "imageFormat": "PNG",
    "imageWidth": 640,
    "imageHeight": 480,
    "fileSizeBytes": 23359,
    "pageCount": 1,
    "billedImages": 1,
    "charged": true,
    "error": null
}
```

With **Include lines** on, each line looks like this:

```json
{ "text": "This is a lot of 12 point text to test the", "confidence": 96.1, "left": 36, "top": 92, "width": 544, "height": 30, "page": 1 }
```

### What you can use it for

- **Receipts and invoices**: pull totals, dates and shop names into a spreadsheet or bookkeeping flow.
- **Screenshots**: make chat logs, error messages, dashboards and app screens searchable.
- **Scanned pages and faxes**: turn TIFF and JPEG scans into text, page by page.
- **Labels and packaging**: read product names, batch codes and ingredient lists from photos.
- **Content work**: read the text in memes, slides and social images for search, tagging or captions.

### Languages

English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Czech, Romanian, Hungarian, Swedish, Danish, Norwegian, Finnish, Catalan, Turkish, Greek, Russian, Ukrainian, Arabic, Hebrew, Hindi, Indonesian, Vietnamese, Thai, Japanese, Korean, and Chinese (simplified and traditional). Pick up to 4 per run.

### FAQ

**Is it legal to use?** You send your own images, or images you have the right to process. The Actor downloads only the links you give, reads each site's robots.txt first and skips files it disallows, never logs in, and keeps nothing after the run. Please do not upload identity documents or other people's private papers unless you are allowed to process them.

**How accurate is it?** Clean printed text reads best: scans, screenshots, receipts and documents. In our tests with computer-made receipts, pages, rotated pages and screenshots, every character matched; real photos score lower. Handwriting, curved labels, very small text, heavy blur and busy backgrounds read much worse than printed text. `meanConfidence` tells you which rows to check by hand.

**My text came out in the wrong order.** Set **Page layout** to **Single block of text** (receipts, screenshots) or **Sparse text** (signs, labels).

**The page was sideways.** Leave **Auto-rotate** on: pages that are sideways or upside down are turned before reading, and `rotationApplied` shows by how much. Phone photos with a camera orientation tag are turned automatically as well.

**What does it not read?** PDF files, HEIC photos and SVG images. Convert HEIC photos to JPEG first. For animated GIF and WebP files only the first frame is read.

**What are the limits?** Files up to 50 MB (20 MB by default), images up to 100 megapixels (very large images are scaled down before reading), up to 10,000 images per run and 50 pages per TIFF file.

**Why did an image fail?** The `error` field says why. `not_image` means the link opened a web page instead of the image file. `robots_disallowed` means the site does not allow automated downloads of that path. `blocked` means the server refused the download (HTTP 401 or 403); after two such refusals the Actor skips the rest of that site's images.

**How fast is it?** Each image takes about 1.3 seconds of CPU time. At the default 2 GB of memory (half a CPU core) that is roughly 1,300 images per hour. More memory gives more CPU cores and the Actor reads several images at once: at 8 GB, roughly 5,000 images per hour. The price per image is the same at any memory size.

### Related Actors

- **Audio & Podcast Transcription** by the same author turns audio and video files into text with timestamps and subtitles.

# Actor input Schema

## `imageUrls` (type: `array`):

Enter direct links to image files, one per line: PNG, JPEG, WebP, TIFF, BMP or GIF. Use your own receipts, screenshots, scans or photos.

## `imageBase64` (type: `array`):

Paste images as base64 strings or data URIs (data:image/png;base64,...), one per line. Handy from Make, Zapier, n8n or your own code. Up to about 10 MB each.

## `languages` (type: `array`):

Choose the languages printed in the images, up to 4. Pick only the ones you need: each extra language makes reading slower.

## `pageLayout` (type: `string`):

Tell the OCR how the text is laid out. Automatic suits most pages. Single block suits receipts and screenshots that come out in the wrong order. Sparse text suits signs, labels and memes.

## `autoRotate` (type: `boolean`):

Detect pages that are sideways or upside down and turn them upright before reading. Turn off for faster runs when all images are upright.

## `includeLines` (type: `boolean`):

Add each text line with its confidence and position (left, top, width, height in pixels).

## `includeWords` (type: `boolean`):

Add each word with its confidence and bounding box in pixels.

## `maxImages` (type: `integer`):

Read at most this many images in this run. Duplicate links count once.

## `maxFileSizeMb` (type: `integer`):

Download image files up to this size. Larger files get the status file\_too\_large.

## `maxTiffPages` (type: `integer`):

Read at most this many pages of each multi-page TIFF file. Each page read counts as one image.

## `upscaleSmallImages` (type: `boolean`):

Enlarge screenshots and small images 2x before reading, which helps with small text.

## Actor input object example

```json
{
  "imageUrls": [
    "https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"
  ],
  "imageBase64": [
    "data:image/png;base64,iVBORw0KGgo..."
  ],
  "languages": [
    "eng"
  ],
  "pageLayout": "auto",
  "autoRotate": true,
  "includeLines": false,
  "includeWords": false,
  "maxImages": 1000,
  "maxFileSizeMb": 20,
  "maxTiffPages": 10,
  "upscaleSmallImages": true
}
```

# Actor output Schema

## `overview` (type: `string`):

imageUrl, status, text, meanConfidence, wordCount, rotationApplied, imageFormat, billedImages and error.

## `lines` (type: `string`):

One row per line: text, confidence, left, top, width, height.

## `words` (type: `string`):

One row per word: text, confidence, left, top, width, height.

## `stats` (type: `string`):

JSON with images planned and done, statuses, images billed, CPU seconds, requests, retries and peak memory.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "imageUrls": [
        "https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"
    ],
    "languages": [
        "eng"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/image-to-text-ocr").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "imageUrls": ["https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"],
    "languages": ["eng"],
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/image-to-text-ocr").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "imageUrls": [
    "https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"
  ],
  "languages": [
    "eng"
  ]
}' |
apify call conserving_celerytop/image-to-text-ocr --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/image-to-text-ocr"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/U8fm4umlaSIdmbvA1/builds/eq2Q5crYJQpDds85r/openapi.json
