Google Lens AI Mode Scraper - Ask Questions About Images avatar

Google Lens AI Mode Scraper - Ask Questions About Images

Pricing

from $5.90 / 1,000 ai mode answers

Go to Apify Store
Google Lens AI Mode Scraper - Ask Questions About Images

Google Lens AI Mode Scraper - Ask Questions About Images

Ask Google AI Mode questions about any image and get the answer with its sources, plus Google Lens OCR text and translation. Export to CSV, Excel, JSON or XML.

Pricing

from $5.90 / 1,000 ai mode answers

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

21 hours ago

Last modified

Share

ParseForge Banner

πŸ” Google Lens AI Mode Scraper

πŸš€ Ask Google AI Mode anything about an image and export the answer in seconds. Every row pairs Google's AI answer (clean Markdown, tables included) with its cited sources, the text Google Lens reads in the image and its translation into any of 133 languages. 10 answers in 81 seconds on Apify, 18 fields per row.

Give the Actor image URLs or base64 images plus the questions you want answered. Each image is uploaded to Google Lens once, then every question is asked in Google AI Mode with the image attached, exactly like tapping "AI Mode" on a Lens result. You get back what Google answered: "This photo was taken on Yuyuan Road (ζ„šε›­θ·―) in Shanghai, China", the list of web pages it cited with real URLs, and the image's own text with pixel boxes.

One image and three questions make three rows. Answers keep their structure (headings, bullet lists, bold, comparison tables) as Markdown, so they paste straight into a doc, a CMS or an LLM prompt. The OCR text, detected language and translation ride along in every row at no extra cost.

🎯 Target AudienceπŸ’‘ Primary Use Cases
E-commerce and catalog teamsIdentify products, read labels and describe item photos at scale
AI and data teamsBuild image question-answer datasets grounded in Google's answers and sources
SEO and content teamsSee which pages Google AI Mode cites for an image and a question
Travel, media and research teamsIdentify places, landmarks and signs, and translate the text in photos

πŸ“‹ What the Google Lens AI Mode Scraper does

πŸ’‘ Why it matters: Google AI Mode reads an image the way a person does and then searches the web to explain it. That answer, and the sources behind it, only exist inside a browser session. This Actor turns them into rows you can filter, compare and feed to other tools.

  • Google AI Mode answers about images. Ask anything: "What is this?", "Where was this photo taken?", "Compare this with an alternative in a table". The answer follows the language of your question.
  • Cited sources with real URLs. Every web page AI Mode cites, with title, domain, site name and snippet. Google's redirect links are resolved to the actual page.
  • Markdown answers. Headings, bullet lists, bold text and tables come through as Markdown, not a flattened blob.
  • Google Lens OCR in every row. The full text found in the image, its detected language, and each line with its pixel bounding box.
  • Lens translation into 133 languages. Pick a target language and get Google Lens's translation of the image text.
  • Image URLs or base64. Public links (JPEG, PNG, WebP, GIF, BMP, TIFF, AVIF) or base64 strings for files that are not online.
  • Self-healing runs. When Google asks for a captcha, the Actor switches to a fresh browser identity, then to a residential exit, on its own.

🎬 Full Demo (🚧 Coming soon)

πŸ“Š Output

Every row is one image and one question.

πŸ–Ό FieldDescription
πŸ–Ό imageUrlThe image you sent (base64 images are named by position, like Base64 image #1)
❓ questionThe question asked in Google AI Mode
πŸ€– answerGoogle AI Mode's answer in Markdown
πŸ‘ answerSeesImageYes when Google confirms the answer was generated with the image attached
πŸ”— sourceCount / πŸ“š sourcesHow many pages the answer cites, and each one's title, url, domain, sourceName and snippet
πŸ“ ocrTextAll text Google Lens reads in the image
🌐 ocrLanguage / πŸ—£ ocrLanguageNameDetected language of that text (code and name)
πŸ“ ocrLineCount / πŸ“ ocrLinesNumber of text lines, and each line with x, y, width, height in pixels of the original image
πŸ”€ translatedTextGoogle Lens translation of the image text
🎯 translationTargetLanguage / 🧭 translationSourceLanguageTarget language you picked and the source language Lens detected
↔ imageWidth / ↕ imageHeightSize of the original image
πŸ•’ scrapedAt / ❌ errorCapture time, and an error message on failed rows

Three real rows from a cloud run (lists shortened for reading):

A street sign, "Where was this photo taken?", translated to Spanish:

{
"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg",
"question": "Where was this photo taken?",
"answer": "This photo was taken on **Yuyuan Road (ζ„šε›­θ·―)** in **Shanghai, China**.\n\nThe distinctive blue and white sign is standard for city streets in Shanghai. This specific intersection indicator marks a location along Yuyuan Road, a historic and trendy thoroughfare that spans across the Jing'an and Changning districts.\n\nIf you are planning a trip or a city walk here, would you like recommendations for **historic buildings**, **cafΓ©s**, or **popular photo spots** along Yuyuan Road?",
"answerSeesImage": "Yes",
"sourceCount": 2,
"ocrText": "θ₯Ώ\n\n315\n\nζ„šε›­θ·―\n\n东\n\n309\n\nW\n\nYuyuan Rd.\n\nE",
"ocrLanguage": "zh-Hans",
"ocrLanguageName": "Simplified Chinese",
"ocrLineCount": 8,
"translatedText": "Oeste\nCamino Yuyuan\nEste\nEN\nCalle Yuyuan.\nY",
"translationTargetLanguage": "es",
"translationSourceLanguage": "zh",
"imageWidth": 640,
"imageHeight": 339,
"sources": [
{
"title": "Yuyuan Road | Attractions - Your Ultimate China Travel Guide",
"url": "https://chinagotrip.com/destinations/shanghai/attractions/yuyuan-road",
"domain": "chinagotrip.com",
"sourceName": "ChinaGoTrip",
"snippet": "Yuyuan Road stands as a historic thoroughfare in Shanghai, stretching 2,775 meters through Jing'an and Changning districts. Establ..."
},
{
"title": "Jing'an and Changning districts",
"url": "https://english.shanghai.gov.cn/en-HeritageZones/20231220/17ed8c2129d24d2ab691a6aa87b9a363.html",
"domain": "english.shanghai.gov.cn",
"sourceName": "english.shanghai.gov.cn",
"snippet": "N/A"
}
],
"ocrLines": [
{ "text": "θ₯Ώ", "x": 88, "y": 87, "width": 41, "height": 32 },
{ "text": "315", "x": 76, "y": 129, "width": 58, "height": 20 }
],
"scrapedAt": "2026-09-28T17:44:50.217Z",
"error": null
}

A poster, "What does this image say?":

{
"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png",
"question": "What does this image say?",
"answer": "The image provides guidelines from the **World Health Organization (WHO)** on how to **reduce your risk of coronavirus infection**.\n\nIt outlines five main preventative measures:\n\n- **Clean hands** with soap and water or alcohol-based hand rub.\n- **Cover nose and mouth** when coughing and sneezing with a tissue or flexed elbow.\n- **Avoid close contact** with anyone with cold or flu-like symptoms.\n- **Thoroughly cook** meat and eggs.\n- **No unprotected contact** with live wild or farm animals.\n\nIf you want, I can provide **more detailed instructions** on any of these steps or look up the **latest guidelines** for staying healthy. What would you like to explore next?",
"answerSeesImage": "Yes",
"sourceCount": 0,
"ocrText": "Reduce your risk of coronavirus infection:\n\nClean hands with soap and water\nor alcohol-based hand rub\n\n...",
"ocrLanguage": "en",
"ocrLanguageName": "English",
"ocrLineCount": 12,
"translatedText": "Reduzca su riesgo de infecciΓ³n por coronavirus:\nLΓ‘vese las manos con agua y jabΓ³n o con un desinfectante para manos a base de alcohol.\n...",
"translationTargetLanguage": "es",
"translationSourceLanguage": "en",
"imageWidth": 905,
"imageHeight": 480,
"sources": [],
"scrapedAt": "2026-09-28T17:45:02.317Z",
"error": null
}

The same sign, "What should I know about this?" (answer shortened, 11 sources):

{
"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg",
"question": "What should I know about this?",
"answer": "The image displays an official street sign for **Yuyuan Road (ζ„šε›­θ·―)**, a famous century-old thoroughfare located in **Shanghai, China**.\n\n...\n\n### πŸ“œ History & Cultural Significance\n\n- **Centennial Heritage:** First constructed in **1911** (the final year of the Qing Dynasty), the street was named after a historic modern garden known as \"Yuyuan\". It spans 2,775 meters across both the Jing'an and Changning districts.\n- **Protected Status:** It is designated as one of Shanghai’s **64 municipal roads that will never be widened**. ...",
"answerSeesImage": "Yes",
"sourceCount": 11,
"ocrLanguage": "zh-Hans",
"scrapedAt": "2026-09-28T17:45:06.636Z",
"error": null
}

✨ Why choose this Actor

  • The real Google AI Mode, with the image attached. Answers come from Google's own AI Mode session for your image, and every row says whether Google confirmed it saw the image.
  • Sources you can click. Cited pages come with their real URLs, not Google redirect links, plus title, site name and snippet.
  • Answers that keep their shape. Tables, lists and headings arrive as Markdown.
  • OCR and translation included in every row. No separate mode, no extra event: text, language, line boxes and a translation into 133 languages.
  • Ask many questions per image. The image is uploaded once and every question reuses it.
  • Clean rows. No null in data fields: missing values read N/A, flags read Yes or No, lists are always arrays.
  • Errors are never charged. An image that cannot be downloaded, or a question Google declines to answer, gives an error row you do not pay for.

πŸ“ˆ How it compares to alternatives

This ActorTypical Google Lens or AI Mode Actor
AI Mode answer about an imageYesText-only AI Mode, or a separate paid mode
Cited sources with resolved URLsYes, title, domain, site name, snippetOften missing or Google redirect links
Answer formattingMarkdown with tables and listsPlain text
OCR text with line boxes in the same rowYes, includedSeparate mode or separate Actor
Translation of image text133 languages, includedSeparate mode, fewer languages
Visual matches / reverse image searchNot included (see FAQ)Some Actors

πŸš€ How to use

  1. Create a free Apify account (sign up here) and get $5 in free usage credit.
  2. Open the Actor and add one or more image URLs (or paste base64 images under More image sources).
  3. Type the questions you want answered. Every image gets every question.
  4. Optionally pick a language under Translate image text to.
  5. Set Max Items (one item = one image and one question) and click Start.
  6. Download the dataset as CSV, Excel, JSON or XML, or read it from the Apify API.

A minimal input:

{
"imageUrls": [{ "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg" }],
"questions": ["Where was this photo taken?", "What does this image say?"],
"translateTo": "en",
"maxItems": 10
}

πŸ’Ό Business use cases

πŸ›οΈ Product identification and cataloging

Send product photos from suppliers or marketplaces and ask "What product is this, and what are its key specs?". Store the answer next to the OCR text of the label and the pages Google cited, and enrich your catalog without manual lookups.

🧠 AI training and evaluation data

Build question-answer pairs grounded in real images and real web sources. The Markdown answer, source URLs and OCR text make a ready-made benchmark or fine-tuning set for multimodal models.

πŸ”Ž AI search visibility (GEO)

Ask the questions your customers ask about your product images and see which pages Google AI Mode cites. Track whether your site shows up as a source over time.

🌍 Travel, compliance and field operations

Read and translate signs, menus, labels and documents in photos from the field, and ask AI Mode to explain them in the team's language.

πŸ”Œ Automating Google Lens AI Mode Scraper

Schedule runs and route the output with Apify integrations: Make, Zapier, Slack, Airbyte, GitHub and Google Drive, or plain webhooks. For example, drop new product photos into a Google Drive folder, run the Actor on them every hour and post each answer to a Slack channel, or append every row to a Google Sheet.

🌟 Beyond business use cases

  • Research: study how Google's AI describes images, which sources it trusts and how answers change over time.
  • Personal: identify plants, landmarks, artworks or objects in your own photos and keep the answers.
  • Non-profit: translate and explain signs, notices and documents for communities that need them.
  • Experimentation: compare Google AI Mode's image answers with other multimodal models on the same images.

πŸ€– Ask an AI assistant about this scraper

Paste this README into ChatGPT, Claude or any AI assistant and ask how to phrase your questions, how to join the output with your data, or how to analyze the answers. Each row is flat JSON with one image and one question, which models handle easily.

❓ Frequently Asked Questions

πŸ€– Is this really Google AI Mode? Yes. The Actor uploads your image to Google Lens and opens Google AI Mode on it, the same flow as the "AI Mode" button in Lens. The answerSeesImage field is Yes when Google's reply header confirms the image was part of the question.

πŸ—£ Which languages can I ask in? Any language AI Mode supports. It answers in the language of your question: ask in Spanish, get the answer in Spanish.

πŸ“š Are the sources real links? Yes. Google hides cited pages behind redirect links; the Actor follows each one and returns the real page URL with its title, domain, site name and snippet. Some answers cite no sources (for example when the question is only about the image's own text), and then sources is an empty list.

πŸ”€ How does translation work? Pick a target language and Google Lens translates the text it finds in the image, block by block. If the text is already in that language, translatedText says so. All 133 languages in the list were tested against Google Lens.

πŸ“ Is OCR included even if I only want answers? Yes, every row carries the OCR text, language and line boxes. Switch off Include OCR lines with boxes to keep rows small.

🚫 Why did a question get "no answer"? Sometimes Google AI Mode declines a question for a specific image. The Actor retries with a fresh Lens session and, if Google still declines, writes an error row that is never charged. Rephrasing the question usually helps.

πŸ‘ Does it return visual matches or reverse image search results? No. Google does not serve Lens visual matches to automated clients from cloud infrastructure, so this Actor focuses on what it can deliver reliably: AI Mode answers, sources, OCR and translation.

🎲 Will the same question give the same answer twice? Not word for word. AI Mode is generative, so wording changes between runs while the facts usually stay the same.

πŸ–Ό Which image formats work? JPEG, PNG, WebP, GIF, BMP, TIFF and AVIF, by URL or as base64. Images larger than 1,200 pixels are downscaled before upload, and OCR boxes are mapped back to the original size. Maximum 20 MB per image.

⚑ How fast is it? About 10 to 20 seconds per answer, several at a time. A 10-answer run finished in 81 seconds on Apify. Use Parallel tabs (1 to 5) to trade speed for gentleness.

πŸ›‘οΈ Do I need a proxy? No. The Actor runs without one and, when Google asks for a captcha, switches to a fresh browser identity and then to a residential exit by itself. Set your own proxy only if you want to force one.

πŸ“ What formats can I export? CSV, Excel, JSON and XML from the dataset view, or any format through the Apify API.

πŸ”Œ Integrate with any app

The dataset is available through the Apify API and every Apify integration, so you can pull it into a spreadsheet, database, BI tool or ML pipeline without glue code.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more AI and data Actors.

πŸ†˜ Need Help? Open our contact form

⚠️ Disclaimer: this is an independent tool and is not affiliated with, endorsed by, or connected to Google. It collects only publicly available answers and data. You are responsible for how you use the data and the images you submit, in line with Google's terms and applicable law.