Image to Text OCR (runs inside the actor) avatar

Image to Text OCR (runs inside the actor)

Pricing

from $3.50 / 1,000 image reads

Go to Apify Store
Image to Text OCR (runs inside the actor)

Image to Text OCR (runs inside the actor)

Extracts the text from images (receipts, scanned pages, screenshots) with an open OCR model that runs inside the actor: your images are not sent to any other service. Returns plain text plus lines and words with boxes and confidence.

Pricing

from $3.50 / 1,000 image reads

Rating

0.0

(0)

Developer

Jack Valmadre

Jack Valmadre

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Share

Get the text out of receipts, scanned pages and screenshots with one API call. The images are read inside your own Apify run by open OCR models (PaddleOCR PP-OCRv6, run on CPU), so they are not sent to any other OCR service and you don't need a cloud vision account.

What you get

  • One result row per image: the text in reading order, plus (if you ask for them) each line and each word with its box and a confidence score.
  • A status on every row, so a broken link doesn't stop the run: ok, no_text, download_failed, not_an_image, too_large or skipped, with the reason in error.
  • Two modes: accurate (default) and fast, which read about four times quicker in our test runs and got 0 to 8 points fewer words right.
  • Images from links (JPEG, PNG, WebP, GIF, BMP, TIFF), or from a key-value store in your Apify account where you uploaded them.

How well it reads

We tested this build on public images we didn't use while building or tuning it, and checked the output against text that people had already transcribed or that came from the page itself.

Image typeWhat we testedWords read correctly (accurate / fast)
Receipt photos50 receipts from the CORD v2 validation set (annotated item and total lines)89% / 81%
Scanned book and document pages16 English pages from Wikisource, compared with their proofread text94% / 90%
Scanned French book pages6 pages from Wikisource90% / 90% (97% / 96% ignoring case and punctuation)
Web page screenshots10 screenshots of 5 public web pages, compared with the page text95% / 92%

"Words read correctly" is the share of the reference words that appear, spelled exactly the same, in the output. On receipts the reference covers only the annotated lines; on book pages it is the proofread page text, which sometimes leaves out running headers or page numbers, so the true figure there may be a little higher. Going the other way, about 8% of the words the accurate mode returned on those pages did not match the reference, even ignoring case and punctuation: misreads, plus headers and page numbers the transcription leaves out.

It also reads accented Latin letters and Chinese characters, though we have tested those on small samples only: 89% of the accented words on the French pages above were found (ignoring case), and on 4 screenshots of Chinese Wikipedia pages the character error rate was about 12 to 13%.

Limitations

  • Product packaging and labels are read poorly. On 46 real photos of packaging (bottles, cans, boxes and shop shelves, from TextOCR / Open Images) this build read 39% of the words exactly in accurate mode and 27% in fast mode. Curved, stylised, small or distant text is the hard part. Don't rely on it for photos like these.
  • Vertical text (such as traditional Chinese book pages) comes out in the wrong order. The characters are mostly recognised, but the columns are returned left to right instead of right to left, and punctuation is often dropped.
  • Side-by-side text can be merged. Text that sits at the same height in two columns may be joined on one line of text. Use the line boxes if layout matters.
  • Handwriting has not been tested.
  • OCR makes mistakes. Check the output before you rely on it for anything important, and use meanConfidence and the per-line confidence to find the rows most worth checking.

Input

{
"imageUrls": ["https://example.com/receipt.jpg", "https://example.com/page-2.png"],
"mode": "accurate",
"outputLevel": "lines"
}
  • imageUrls: direct links to image files. urls, images and startUrls are accepted too, as plain strings or {"url": ...} objects.
  • keyValueStoreId (optional): a key-value store in your account holding image files; every image record is read unless you list imageKeys.
  • mode: accurate (default) or fast.
  • outputLevel: text (plain text only), lines (default; text plus lines with boxes and confidence) or words (also each word with its box).
  • minConfidence (default 0.5): lines the model is less sure of than this are left out.
  • maxImages (default 1,000) and maxFileSizeMb (default 25): larger files are skipped with status too_large.

Output

One row per image, for example (a screenshot of a GOV.UK page, outputLevel: "lines", shortened):

{
"input": "https://example.com/vat-rates.png",
"status": "ok",
"text": "Cookies on GOV.UK\n...\nHome > Money and tax > VAT\nVAT rates\n...",
"lineCount": 19,
"wordCount": 94,
"meanConfidence": 0.9936,
"width": 1280,
"height": 900,
"mode": "accurate",
"lines": [
{"text": "Home > Money and tax > VAT", "confidence": 0.9586, "box": [155, 398, 384, 424],
"polygon": [[155, 398], [384, 398], [384, 424], [155, 424]]}
]
}

Boxes are pixel coordinates [left, top, right, bottom] in the image (after any EXIF rotation is applied). A row that failed has status and error and no text. The run's OUTPUT record summarises how many images got each status.

Pricing

Pay per event: US$0.0035 per image read (image, rows with status ok or no_text), plus a start fee of US$0.00005 per GB of run memory, charged once per run (US$0.0002 at the default 4 GB). Fast mode is charged the same price per image as accurate mode. Images that fail to download or decode are not charged. There is no extra charge for Apify platform usage. Apify shows the price before you run, and you can set a maximum cost per run: once it is reached, the remaining images get a skipped row and are not charged. Run it with 4 GB of memory (the default); bigger and denser images take longer.

Using it from code or an AI agent

Call it like any Apify actor (API, Python or JavaScript client, or an MCP-connected agent) with the input above and read the default dataset. Plain urls work, so a tool call such as {"urls": ["https://..."]} is enough.

Privacy

Your images are processed inside your own Apify run and are not sent to any other service or kept by us. Results go to your run's dataset; images are held in memory while they are read and not saved. The OCR models are built into the actor; the only downloads during a run are your images.

Support

Open an issue on the actor's Issues tab and we will reply there. You can also write to support@madrasco.dev.

This actor is provided as is. It extracts text; it does not check, interpret or guarantee that text for any legal, financial, medical or compliance purpose.

OCR models: PaddleOCR PP-OCRv6 (Apache-2.0), run with RapidOCR (Apache-2.0).

Published by Madrasco. Built and supported with AI assistance; replies to issues may be AI-assisted, and a human owner can be reached on request.