Google Lens OCR API - Image to Text, Handwriting & Translation avatar

Google Lens OCR API - Image to Text, Handwriting & Translation

Pricing

from $6.00 / 1,000 image ocrs

Go to Apify Store
Google Lens OCR API - Image to Text, Handwriting & Translation

Google Lens OCR API - Image to Text, Handwriting & Translation

Turn image URLs into text with Google Lens OCR. Get the full text, paragraphs with lines, optional word bounding boxes and an optional translation for scans, screenshots, photos and handwriting. Built for developers, data teams and document workflows. Pay per image read.

Pricing

from $6.00 / 1,000 image ocrs

Rating

0.0

(0)

Developer

SR

SR

Maintained by Community

Actor stats

0

Bookmarked

191

Total users

71

Monthly active users

4 days ago

Last modified

Share

Google Lens OCR API — Image to Text

Google Lens OCR for image URLs: send a scan, screenshot, photo or handwritten note and get the recognized text back as JSON, with paragraph blocks, optional word bounding boxes and an optional translation. Each image that can be read becomes one structured dataset row. Recognition and optional translation use the Google Lens engine, and your original recognized text stays in the result when translation is requested.

This actor is useful when you need searchable text from document images, structured text from screenshots, or coordinates for highlighting recognized words. It accepts JPEG, PNG, WEBP, TIFF, GIF and BMP images. Recognition quality depends on the pixels in the source: clear printed text generally gives more useful results than a blurred, very small or hard-to-read handwritten image. Check the returned text against the image when exact transcription matters.

What you can do with Google Lens OCR

  • Digitize scanned documents. Turn page scans, letters and forms into plain text you can search, index or feed into a language model.
  • Read text from screenshots. Pull error messages, chat logs, receipts or product labels out of screenshots without retyping them.
  • Highlight words on the image. Use the normalized word boxes to draw overlays, build a searchable image viewer or redact specific words.
  • Translate image text. Request a target language and receive the translation next to the original text, for example to read a foreign-language sign or document.
  • Process batches. Send up to 200 image URLs in one run and receive the rows in input order, ready for export as JSON, CSV or Excel from the Apify dataset.
  • Automate pipelines. Call the actor from the Apify API, a schedule, Make, Zapier or n8n, and read the dataset when the run finishes.

Input

image_urls contains direct links to image files. A link to a web page that displays an image does not serve the image itself and may produce no row. Enter the URLs as a list; API callers can also send one URL per line in a string. Blank entries are removed. Up to 200 URLs are accepted per run, and URLs beyond that limit are reported as unprocessed. Repeated URLs are processed separately and create separate rows.

The following is the real prefill used for the baseline run:

{"image_urls":["https://raw.githubusercontent.com/tesseract-ocr/test/main/testing/phototest.tif"],"ocr_language":"en"}

ocr_language defaults to en. It is a recognition hint, not a filter: other scripts in the image can still appear in the output. The value is echoed on the result as ocr_language; that field does not claim to be the detected language.

include_word_boxes defaults to true. Set it to false to omit the words array and reduce the output size. The full text, paragraph geometry and word count remain available. This option changes the returned fields, not the number of images processed or the price.

translate_to is optional. Set it to the requested translation language, such as the nl value in the actual translation example below. Leave it absent or empty for recognition only. Translation adds separate fields and does not replace the recognized source text, paragraphs or word coordinates.

Google Lens OCR output

One row is delivered for every successfully read image, even when that image contains no recognized text. An unreadable URL creates no dataset row. The actor retains the source URL so you can connect each result to the submitted image. Rows are returned in input order, with repeated URLs retained.

full_text is a string joining the nonempty text blocks with line breaks. paragraphs is an array of objects containing text, lines and geometry. The lines field is an array of strings. Geometry is either an object or null and may include normalized center_x, center_y, width, height and angle_deg, with the angle measured in degrees.

When word boxes are enabled, each element of words contains word, separator and geometry. The integer word_count remains present when word boxes are disabled. image_width and image_height report the source dimensions in pixels. processed_in_seconds measures recognition for that image, while fetched_in_seconds measures the complete collection time.

With translation enabled, translated_text is a string or null if the engine did not provide a translation, and translate_to echoes the requested target language. Without translation, both keys are omitted. Additional common fields are source, engine, text and the UTC timestamp fetched_at; text matches full_text.

The run summary is available in both OUTPUT and summary. It includes itemCount, errorCount and statusMessage. The errors record describes failed images. Check that record when fewer rows arrive than the number of submitted URLs. Partial delivery is indicated in the summary when a run ends before every image is attempted.

With no errors, statusMessage remains N image(s) read. With errors it is N image(s) read, M failed, including when every URL fails. The former all K image URL(s) could not be downloaded wording is replaced by this count; individual causes remain in errors.

Word boxes example

This excerpt comes from the real prefill run above (phototest.tif, 640 × 480 pixels). The row contains 65 words; the first two elements of words are shown, with the other fields of the row omitted:

{
"image_url": "https://raw.githubusercontent.com/tesseract-ocr/test/main/testing/phototest.tif",
"word_count": 65,
"image_width": 640,
"image_height": 480,
"words": [
{
"word": "This",
"separator": " ",
"geometry": {
"center_x": 0.1,
"center_y": 0.22188,
"width": 0.10312,
"height": 0.06458,
"angle_deg": 0.00013
}
},
{
"word": "is",
"separator": " ",
"geometry": {
"center_x": 0.18437,
"center_y": 0.22188,
"width": 0.0375,
"height": 0.06458,
"angle_deg": 0.00013
}
}
]
}

To get pixel coordinates, multiply center_x and width by image_width, and center_y and height by image_height. Joining every word with its separator reproduces the recognized text.

Translation example

This input and the following row come from the recorded translation response. The displayed row is the Arabic image from the two-image run; word boxes were disabled.

{
"image_urls": [
"https://raw.githubusercontent.com/tesseract-ocr/test/main/testing/eurotext.tif",
"https://raw.githubusercontent.com/tesseract-ocr/test/main/testing/arabic.tif"
],
"ocr_language": "de",
"translate_to": "nl",
"include_word_boxes": false
}
{
"image_url": "https://raw.githubusercontent.com/tesseract-ocr/test/main/testing/arabic.tif",
"full_text": "العربي",
"paragraphs": [
{
"text": "العربي",
"lines": [
"العربي"
],
"geometry": {
"center_x": 0.73125,
"center_y": 0.70726,
"width": 0.20167,
"height": 0.28,
"angle_deg": -0.51728
}
}
],
"word_count": 1,
"image_width": 600,
"image_height": 200,
"ocr_language": "de",
"processed_in_seconds": 1.782,
"source": "google-lens-ocr",
"engine": "google_lens",
"text": "العربي",
"fetched_at": "2026-10-01T16:17:32.802Z",
"translated_text": "Arabisch",
"translate_to": "nl",
"fetched_in_seconds": 2.77
}

How this Google Lens OCR actor compares

  • Versus self-hosted OCR such as Tesseract. There is no model to install, tune or host, no GPU and no language packs to manage. You send URLs and receive structured rows. Self-hosting can be cheaper at very high volume when you already run the infrastructure.
  • Versus general cloud vision APIs. Those require a cloud account, credentials and your own code around the API. This actor runs on your Apify account, bills per image read and writes results straight into a dataset you can export or connect to other tools.
  • Versus Google Lens reverse image search actors. Those return visually similar images and matching web pages. This actor focuses on text: full text, paragraphs, lines, word geometry and translation, in a stable row format.
  • Versus copying text by hand from the Lens app. The app is fine for one image. For tens or hundreds of images, or for text that must land in a spreadsheet or database, a batch run with machine-readable output saves the manual work.

Pricing

The current pay-per-event price is $0.006 per successfully read image, plus the Apify actor-start event at $0.005 per GB of actor memory, with a minimum of one event. The normal actor configuration uses 1 GB. Image recognition is charged after the corresponding rows are delivered. A failed download creates no OCR event; a successful image with empty recognized text still creates one row and one OCR event. Optional translation and word coordinates do not add another image event.

Users without a paid Apify plan can process the first ten image URLs per run. The input is limited before processing, and only delivered rows are charged. When more URLs were supplied, the summary reports urlsNotProcessed. Upgrade the Apify plan or split the input across runs to process more images.

FAQ

Why did I receive fewer rows than image URLs?

Some URLs may return an error, time out, serve HTML or contain an unsupported image. Those images are described in the errors record and do not create rows. The free-plan limit and the 200-image maximum can also reduce the number processed. Compare the input, row count and summary rather than assuming every link is still usable.

Can I recognize text without keeping every word box?

Yes. Set include_word_boxes to false. You still receive the full text, paragraphs, source dimensions and word count. Only the words array is omitted, so consumers that need searchable text without individual coordinates can keep a smaller result.

Does translation overwrite the original text?

No. The original full text and paragraphs remain in the source language. The translation is returned separately as translated_text. If no translation was provided, that field is null. If you did not request translation, the translation fields are absent.

What happens to an image with no text?

A readable image without recognized characters is a successful result with an empty full text string and zero words. It is still one processed image and one chargeable row. A URL that cannot be read is different: it produces an error and no row.

Does Google Lens OCR read handwriting?

The engine reads handwritten text as well as printed text, but results depend strongly on legibility, contrast and resolution. Neat handwriting in a sharp photo usually comes back well; cursive or faint writing may be partly missed or misread. Always check handwriting results against the image.

Which languages are supported?

Recognition is not limited to the ocr_language hint: the baseline translation run recognized Arabic text with the hint set to de. Set ocr_language to the main language of your images and use translate_to when you want a translation in another language.

How should I submit large images?

Use a direct image link and keep the source below 30 megapixels. Images exceeding that pixel limit or the 128 MiB download limit should be resized before submission. The source dimensions remain available on successful rows, allowing downstream consumers to interpret the normalized coordinates.

What happens if an image takes too long to read?

An image initially has 18 seconds for downloading and recognition. The first image on a processing page can use up to about 44 seconds. If a later image exceeds its initial limit, it gets one additional attempt as the first image on the next page, with that longer budget. A timeout on the first position produces one upstream_timeout error identifying its URL; the remaining images continue. These limits are lower than the former 90-second limit. Resize slow or large images, or submit them individually. Waiting for processing capacity does not count toward the image limit: pending images are retried while the run has time left, and unfinished work is marked as partial. Undelivered images create no OCR charge. Once all pages finish, the temporary page-budget warning is removed from the run summary; actual image errors remain visible.

Can image URLs point to a private network?

Image URLs must use HTTP or HTTPS and resolve to public IP addresses. Private-network, loopback and link-local addresses are rejected, including when a public URL redirects to them. Use a publicly accessible direct image link.