Image to Text OCR (runs inside the actor)
Pricing
from $3.50 / 1,000 image reads
Image to Text OCR (runs inside the actor)
Extracts the text from images (receipts, scanned pages, screenshots) with an open OCR model that runs inside the actor: your images are not sent to any other service. Returns plain text plus lines and words with boxes and confidence.
Pricing
from $3.50 / 1,000 image reads
Rating
0.0
(0)
Developer
Jack Valmadre
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Get the text out of receipts, scanned pages and screenshots with one API call. The images are read inside your own Apify run by open OCR models (PaddleOCR PP-OCRv6, run on CPU), so they are not sent to any other OCR service and you don't need a cloud vision account.
What you get
- One result row per image: the text in reading order, plus (if you ask for them) each line and each word with its box and a confidence score.
- A
statuson every row, so a broken link doesn't stop the run:ok,no_text,download_failed,not_an_image,too_largeorskipped, with the reason inerror. - Two modes: accurate (default) and fast, which read about four times quicker in our test runs and got 0 to 8 points fewer words right.
- Images from links (JPEG, PNG, WebP, GIF, BMP, TIFF), or from a key-value store in your Apify account where you uploaded them.
How well it reads
We tested this build on public images we didn't use while building or tuning it, and checked the output against text that people had already transcribed or that came from the page itself.
| Image type | What we tested | Words read correctly (accurate / fast) |
|---|---|---|
| Receipt photos | 50 receipts from the CORD v2 validation set (annotated item and total lines) | 89% / 81% |
| Scanned book and document pages | 16 English pages from Wikisource, compared with their proofread text | 94% / 90% |
| Scanned French book pages | 6 pages from Wikisource | 90% / 90% (97% / 96% ignoring case and punctuation) |
| Web page screenshots | 10 screenshots of 5 public web pages, compared with the page text | 95% / 92% |
"Words read correctly" is the share of the reference words that appear, spelled exactly the same, in the output. On receipts the reference covers only the annotated lines; on book pages it is the proofread page text, which sometimes leaves out running headers or page numbers, so the true figure there may be a little higher. Going the other way, about 8% of the words the accurate mode returned on those pages did not match the reference, even ignoring case and punctuation: misreads, plus headers and page numbers the transcription leaves out.
It also reads accented Latin letters and Chinese characters, though we have tested those on small samples only: 89% of the accented words on the French pages above were found (ignoring case), and on 4 screenshots of Chinese Wikipedia pages the character error rate was about 12 to 13%.
Limitations
- Product packaging and labels are read poorly. On 46 real photos of packaging (bottles, cans, boxes and shop shelves, from TextOCR / Open Images) this build read 39% of the words exactly in accurate mode and 27% in fast mode. Curved, stylised, small or distant text is the hard part. Don't rely on it for photos like these.
- Vertical text (such as traditional Chinese book pages) comes out in the wrong order. The characters are mostly recognised, but the columns are returned left to right instead of right to left, and punctuation is often dropped.
- Side-by-side text can be merged. Text that sits at the same height in two columns may be joined on one line of
text. Use the line boxes if layout matters. - Handwriting has not been tested.
- OCR makes mistakes. Check the output before you rely on it for anything important, and use
meanConfidenceand the per-lineconfidenceto find the rows most worth checking.
Input
{"imageUrls": ["https://example.com/receipt.jpg", "https://example.com/page-2.png"],"mode": "accurate","outputLevel": "lines"}
imageUrls: direct links to image files.urls,imagesandstartUrlsare accepted too, as plain strings or{"url": ...}objects.keyValueStoreId(optional): a key-value store in your account holding image files; every image record is read unless you listimageKeys.mode:accurate(default) orfast.outputLevel:text(plain text only),lines(default; text plus lines with boxes and confidence) orwords(also each word with its box).minConfidence(default 0.5): lines the model is less sure of than this are left out.maxImages(default 1,000) andmaxFileSizeMb(default 25): larger files are skipped with statustoo_large.
Output
One row per image, for example (a screenshot of a GOV.UK page, outputLevel: "lines", shortened):
{"input": "https://example.com/vat-rates.png","status": "ok","text": "Cookies on GOV.UK\n...\nHome > Money and tax > VAT\nVAT rates\n...","lineCount": 19,"wordCount": 94,"meanConfidence": 0.9936,"width": 1280,"height": 900,"mode": "accurate","lines": [{"text": "Home > Money and tax > VAT", "confidence": 0.9586, "box": [155, 398, 384, 424],"polygon": [[155, 398], [384, 398], [384, 424], [155, 424]]}]}
Boxes are pixel coordinates [left, top, right, bottom] in the image (after any EXIF rotation is applied). A row that failed has status and error and no text. The run's OUTPUT record summarises how many images got each status.
Pricing
Pay per event: US$0.0035 per image read (image, rows with status ok or no_text), plus a start fee of US$0.00005 per GB of run memory, charged once per run (US$0.0002 at the default 4 GB). Fast mode is charged the same price per image as accurate mode. Images that fail to download or decode are not charged. There is no extra charge for Apify platform usage. Apify shows the price before you run, and you can set a maximum cost per run: once it is reached, the remaining images get a skipped row and are not charged. Run it with 4 GB of memory (the default); bigger and denser images take longer.
Using it from code or an AI agent
Call it like any Apify actor (API, Python or JavaScript client, or an MCP-connected agent) with the input above and read the default dataset. Plain urls work, so a tool call such as {"urls": ["https://..."]} is enough.
Privacy
Your images are processed inside your own Apify run and are not sent to any other service or kept by us. Results go to your run's dataset; images are held in memory while they are read and not saved. The OCR models are built into the actor; the only downloads during a run are your images.
Support
Open an issue on the actor's Issues tab and we will reply there. You can also write to support@madrasco.dev.
This actor is provided as is. It extracts text; it does not check, interpret or guarantee that text for any legal, financial, medical or compliance purpose.
OCR models: PaddleOCR PP-OCRv6 (Apache-2.0), run with RapidOCR (Apache-2.0).
Published by Madrasco. Built and supported with AI assistance; replies to issues may be AI-assisted, and a human owner can be reached on request.