Image to Text OCR: $1.20/1K Images, 30 Languages, No Login
Pricing
from $1.20 / 1,000 image with text reads
Image to Text OCR: $1.20/1K Images, 30 Languages, No Login
Extract text from images with OCR for $1.20 per 1,000 images. Reads receipts, screenshots, scans and photos in PNG, JPEG, WebP, TIFF, BMP or GIF. Returns text, confidence, lines and word boxes. 30 languages, auto-rotate. No text found, no charge.
Pricing
from $1.20 / 1,000 image with text reads
Rating
0.0
(0)
Developer
Don Mangu
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
5 hours ago
Last modified
Categories
Share
Image to Text OCR
Image to Text OCR extracts text from images. Give it links to receipts, screenshots, scanned pages, photos of documents or labels, and it returns one row per image with the text it read, an average confidence score, and optional line and word positions. It reads PNG, JPEG, WebP, TIFF (including multi-page TIFF), BMP and GIF files in 30 languages. It costs $1.20 per 1,000 images, and images where no text is found are free.
The OCR runs inside the Actor. Your images are not sent to Google, to a cloud vision API or to any other outside service, and nothing is kept after the run apart from the results in your own Apify storage.
What the image to text OCR returns
- Text:
text, all text read from the image, with line breaks where the lines break and a blank line between blocks. AlsowordCount,lineCountandcharacterCount. - Confidence:
meanConfidence, the average word confidence from 0 to 100. Clean printed text scores above 90; under about 60 usually means a blurry, tiny or handwritten image. - Lines (turn on Include lines): each line with its
text,confidenceand position (left,top,width,heightin pixels). - Word boxes (turn on Include word boxes): each word with its
confidenceand bounding box, for highlighting, redaction or form work. - Image details:
imageFormat,imageWidth,imageHeight,fileSizeBytes,pageCountandrotationApplied(how far the page was turned to read it). - Status for every image:
statusisok, or says why there is no text (for exampleno_text,not_found,not_image,unsupported_format,file_too_large), with a plain-languageerror.
How to extract text from an image, step by step
- Open the Actor and go to the Input tab.
- In Image URLs, paste direct links to image files, one per line. A link to a web page that shows an image is not a file link: right-click the image, copy the image address, and use that.
- If your images are not online, paste them in Image data (base64) as base64 strings or data URIs. This is the easy route from Make, Zapier, n8n or your own code.
- Pick the Languages printed in the images (English is the default). Choose only the ones you need.
- Leave Page layout on Automatic. If a receipt or screenshot comes out in the wrong order, try Single block of text; for signs, labels and memes, try Sparse text.
- Click Start. The example input reads one small image in a few seconds.
- Open the Output tab. The Overview view shows the text per image; the Lines and Word boxes views show positions when you turned them on. Download as JSON, CSV or Excel, or read the results through the Apify API.
How much does image OCR cost?
You pay per image read, with the pay-per-event event image-ocr:
| What | Price |
|---|---|
| Image with text found | $0.0012 ($1.20 per 1,000) |
| Each page of a multi-page TIFF file | $0.0012 |
| Image with no text, broken link, unsupported file | free |
Apify also charges a small fixed amount when a run starts.
Worked example: a month of expense receipts, 2,500 photos, of which 40 are blank or unreadable and 10 links are broken, costs 2,450 x $0.0012 = $2.94. A batch of 200 screenshots costs $0.24.
Set a spending limit on the run if you like: the Actor reads only the images that fit, lists the ones it did not start in the run statistics, and stops cleanly.
Input example
{"imageUrls": ["https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png"],"languages": ["eng"],"pageLayout": "auto","autoRotate": true,"includeLines": false,"includeWords": false,"maxImages": 1000}
Output example
A real row from the example input:
{"inputIndex": 1,"imageUrl": "https://raw.githubusercontent.com/naptha/tesseract.js/master/tests/assets/images/testocr.png","source": "url","status": "ok","text": "This is a lot of 12 point text to test the\nocr code and see if it works on all types\nof file format.\n\nThe quick brown dog jumped over the\nlazy fox. The quick brown dog jumped\nover the lazy fox. The quick brown dog\njumped over the lazy fox. The quick\nbrown dog jumped over the lazy fox.","meanConfidence": 96.1,"wordCount": 60,"lineCount": 8,"characterCount": 285,"languages": ["eng"],"pageLayout": "auto","rotationApplied": 0,"imageFormat": "PNG","imageWidth": 640,"imageHeight": 480,"fileSizeBytes": 23359,"pageCount": 1,"billedImages": 1,"charged": true,"error": null}
With Include lines on, each line looks like this:
{ "text": "This is a lot of 12 point text to test the", "confidence": 96.1, "left": 36, "top": 92, "width": 544, "height": 30, "page": 1 }
What you can use it for
- Receipts and invoices: pull totals, dates and shop names into a spreadsheet or bookkeeping flow.
- Screenshots: make chat logs, error messages, dashboards and app screens searchable.
- Scanned pages and faxes: turn TIFF and JPEG scans into text, page by page.
- Labels and packaging: read product names, batch codes and ingredient lists from photos.
- Content work: read the text in memes, slides and social images for search, tagging or captions.
Languages
English, German, French, Spanish, Italian, Portuguese, Dutch, Polish, Czech, Romanian, Hungarian, Swedish, Danish, Norwegian, Finnish, Catalan, Turkish, Greek, Russian, Ukrainian, Arabic, Hebrew, Hindi, Indonesian, Vietnamese, Thai, Japanese, Korean, and Chinese (simplified and traditional). Pick up to 4 per run.
FAQ
Is it legal to use? You send your own images, or images you have the right to process. The Actor downloads only the links you give, reads each site's robots.txt first and skips files it disallows, never logs in, and keeps nothing after the run. Please do not upload identity documents or other people's private papers unless you are allowed to process them.
How accurate is it? Clean printed text reads best: scans, screenshots, receipts and documents. In our tests with computer-made receipts, pages, rotated pages and screenshots, every character matched; real photos score lower. Handwriting, curved labels, very small text, heavy blur and busy backgrounds read much worse than printed text. meanConfidence tells you which rows to check by hand.
My text came out in the wrong order. Set Page layout to Single block of text (receipts, screenshots) or Sparse text (signs, labels).
The page was sideways. Leave Auto-rotate on: pages that are sideways or upside down are turned before reading, and rotationApplied shows by how much. Phone photos with a camera orientation tag are turned automatically as well.
What does it not read? PDF files, HEIC photos and SVG images. Convert HEIC photos to JPEG first. For animated GIF and WebP files only the first frame is read.
What are the limits? Files up to 50 MB (20 MB by default), images up to 100 megapixels (very large images are scaled down before reading), up to 10,000 images per run and 50 pages per TIFF file.
Why did an image fail? The error field says why. not_image means the link opened a web page instead of the image file. robots_disallowed means the site does not allow automated downloads of that path. blocked means the server refused the download (HTTP 401 or 403); after two such refusals the Actor skips the rest of that site's images.
How fast is it? Each image takes about 1.3 seconds of CPU time. At the default 2 GB of memory (half a CPU core) that is roughly 1,300 images per hour. More memory gives more CPU cores and the Actor reads several images at once: at 8 GB, roughly 5,000 images per hour. The price per image is the same at any memory size.
Related Actors
- Audio & Podcast Transcription by the same author turns audio and video files into text with timestamps and subtitles.