Google Lens AI Mode Scraper - Ask Questions About Images
Pricing
from $5.90 / 1,000 ai mode answers
Google Lens AI Mode Scraper - Ask Questions About Images
Ask Google AI Mode questions about any image and get the answer with its sources, plus Google Lens OCR text and translation. Export to CSV, Excel, JSON or XML.
Pricing
from $5.90 / 1,000 ai mode answers
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
21 hours ago
Last modified
Categories
Share

π Google Lens AI Mode Scraper
π Ask Google AI Mode anything about an image and export the answer in seconds. Every row pairs Google's AI answer (clean Markdown, tables included) with its cited sources, the text Google Lens reads in the image and its translation into any of 133 languages. 10 answers in 81 seconds on Apify, 18 fields per row.
Give the Actor image URLs or base64 images plus the questions you want answered. Each image is uploaded to Google Lens once, then every question is asked in Google AI Mode with the image attached, exactly like tapping "AI Mode" on a Lens result. You get back what Google answered: "This photo was taken on Yuyuan Road (ζεθ·―) in Shanghai, China", the list of web pages it cited with real URLs, and the image's own text with pixel boxes.
One image and three questions make three rows. Answers keep their structure (headings, bullet lists, bold, comparison tables) as Markdown, so they paste straight into a doc, a CMS or an LLM prompt. The OCR text, detected language and translation ride along in every row at no extra cost.
| π― Target Audience | π‘ Primary Use Cases |
|---|---|
| E-commerce and catalog teams | Identify products, read labels and describe item photos at scale |
| AI and data teams | Build image question-answer datasets grounded in Google's answers and sources |
| SEO and content teams | See which pages Google AI Mode cites for an image and a question |
| Travel, media and research teams | Identify places, landmarks and signs, and translate the text in photos |
π What the Google Lens AI Mode Scraper does
π‘ Why it matters: Google AI Mode reads an image the way a person does and then searches the web to explain it. That answer, and the sources behind it, only exist inside a browser session. This Actor turns them into rows you can filter, compare and feed to other tools.
- Google AI Mode answers about images. Ask anything: "What is this?", "Where was this photo taken?", "Compare this with an alternative in a table". The answer follows the language of your question.
- Cited sources with real URLs. Every web page AI Mode cites, with title, domain, site name and snippet. Google's redirect links are resolved to the actual page.
- Markdown answers. Headings, bullet lists, bold text and tables come through as Markdown, not a flattened blob.
- Google Lens OCR in every row. The full text found in the image, its detected language, and each line with its pixel bounding box.
- Lens translation into 133 languages. Pick a target language and get Google Lens's translation of the image text.
- Image URLs or base64. Public links (JPEG, PNG, WebP, GIF, BMP, TIFF, AVIF) or base64 strings for files that are not online.
- Self-healing runs. When Google asks for a captcha, the Actor switches to a fresh browser identity, then to a residential exit, on its own.
π¬ Full Demo (π§ Coming soon)
π Output
Every row is one image and one question.
| πΌ Field | Description |
|---|---|
πΌ imageUrl | The image you sent (base64 images are named by position, like Base64 image #1) |
β question | The question asked in Google AI Mode |
π€ answer | Google AI Mode's answer in Markdown |
π answerSeesImage | Yes when Google confirms the answer was generated with the image attached |
π sourceCount / π sources | How many pages the answer cites, and each one's title, url, domain, sourceName and snippet |
π ocrText | All text Google Lens reads in the image |
π ocrLanguage / π£ ocrLanguageName | Detected language of that text (code and name) |
π ocrLineCount / π ocrLines | Number of text lines, and each line with x, y, width, height in pixels of the original image |
π€ translatedText | Google Lens translation of the image text |
π― translationTargetLanguage / π§ translationSourceLanguage | Target language you picked and the source language Lens detected |
β imageWidth / β imageHeight | Size of the original image |
π scrapedAt / β error | Capture time, and an error message on failed rows |
Three real rows from a cloud run (lists shortened for reading):
A street sign, "Where was this photo taken?", translated to Spanish:
{"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg","question": "Where was this photo taken?","answer": "This photo was taken on **Yuyuan Road (ζεθ·―)** in **Shanghai, China**.\n\nThe distinctive blue and white sign is standard for city streets in Shanghai. This specific intersection indicator marks a location along Yuyuan Road, a historic and trendy thoroughfare that spans across the Jing'an and Changning districts.\n\nIf you are planning a trip or a city walk here, would you like recommendations for **historic buildings**, **cafΓ©s**, or **popular photo spots** along Yuyuan Road?","answerSeesImage": "Yes","sourceCount": 2,"ocrText": "θ₯Ώ\n\n315\n\nζεθ·―\n\nδΈ\n\n309\n\nW\n\nYuyuan Rd.\n\nE","ocrLanguage": "zh-Hans","ocrLanguageName": "Simplified Chinese","ocrLineCount": 8,"translatedText": "Oeste\nCamino Yuyuan\nEste\nEN\nCalle Yuyuan.\nY","translationTargetLanguage": "es","translationSourceLanguage": "zh","imageWidth": 640,"imageHeight": 339,"sources": [{"title": "Yuyuan Road | Attractions - Your Ultimate China Travel Guide","url": "https://chinagotrip.com/destinations/shanghai/attractions/yuyuan-road","domain": "chinagotrip.com","sourceName": "ChinaGoTrip","snippet": "Yuyuan Road stands as a historic thoroughfare in Shanghai, stretching 2,775 meters through Jing'an and Changning districts. Establ..."},{"title": "Jing'an and Changning districts","url": "https://english.shanghai.gov.cn/en-HeritageZones/20231220/17ed8c2129d24d2ab691a6aa87b9a363.html","domain": "english.shanghai.gov.cn","sourceName": "english.shanghai.gov.cn","snippet": "N/A"}],"ocrLines": [{ "text": "θ₯Ώ", "x": 88, "y": 87, "width": 41, "height": 32 },{ "text": "315", "x": 76, "y": 129, "width": 58, "height": 20 }],"scrapedAt": "2026-09-28T17:44:50.217Z","error": null}
A poster, "What does this image say?":
{"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png","question": "What does this image say?","answer": "The image provides guidelines from the **World Health Organization (WHO)** on how to **reduce your risk of coronavirus infection**.\n\nIt outlines five main preventative measures:\n\n- **Clean hands** with soap and water or alcohol-based hand rub.\n- **Cover nose and mouth** when coughing and sneezing with a tissue or flexed elbow.\n- **Avoid close contact** with anyone with cold or flu-like symptoms.\n- **Thoroughly cook** meat and eggs.\n- **No unprotected contact** with live wild or farm animals.\n\nIf you want, I can provide **more detailed instructions** on any of these steps or look up the **latest guidelines** for staying healthy. What would you like to explore next?","answerSeesImage": "Yes","sourceCount": 0,"ocrText": "Reduce your risk of coronavirus infection:\n\nClean hands with soap and water\nor alcohol-based hand rub\n\n...","ocrLanguage": "en","ocrLanguageName": "English","ocrLineCount": 12,"translatedText": "Reduzca su riesgo de infecciΓ³n por coronavirus:\nLΓ‘vese las manos con agua y jabΓ³n o con un desinfectante para manos a base de alcohol.\n...","translationTargetLanguage": "es","translationSourceLanguage": "en","imageWidth": 905,"imageHeight": 480,"sources": [],"scrapedAt": "2026-09-28T17:45:02.317Z","error": null}
The same sign, "What should I know about this?" (answer shortened, 11 sources):
{"imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg","question": "What should I know about this?","answer": "The image displays an official street sign for **Yuyuan Road (ζεθ·―)**, a famous century-old thoroughfare located in **Shanghai, China**.\n\n...\n\n### π History & Cultural Significance\n\n- **Centennial Heritage:** First constructed in **1911** (the final year of the Qing Dynasty), the street was named after a historic modern garden known as \"Yuyuan\". It spans 2,775 meters across both the Jing'an and Changning districts.\n- **Protected Status:** It is designated as one of Shanghaiβs **64 municipal roads that will never be widened**. ...","answerSeesImage": "Yes","sourceCount": 11,"ocrLanguage": "zh-Hans","scrapedAt": "2026-09-28T17:45:06.636Z","error": null}
β¨ Why choose this Actor
- The real Google AI Mode, with the image attached. Answers come from Google's own AI Mode session for your image, and every row says whether Google confirmed it saw the image.
- Sources you can click. Cited pages come with their real URLs, not Google redirect links, plus title, site name and snippet.
- Answers that keep their shape. Tables, lists and headings arrive as Markdown.
- OCR and translation included in every row. No separate mode, no extra event: text, language, line boxes and a translation into 133 languages.
- Ask many questions per image. The image is uploaded once and every question reuses it.
- Clean rows. No
nullin data fields: missing values readN/A, flags readYesorNo, lists are always arrays. - Errors are never charged. An image that cannot be downloaded, or a question Google declines to answer, gives an error row you do not pay for.
π How it compares to alternatives
| This Actor | Typical Google Lens or AI Mode Actor | |
|---|---|---|
| AI Mode answer about an image | Yes | Text-only AI Mode, or a separate paid mode |
| Cited sources with resolved URLs | Yes, title, domain, site name, snippet | Often missing or Google redirect links |
| Answer formatting | Markdown with tables and lists | Plain text |
| OCR text with line boxes in the same row | Yes, included | Separate mode or separate Actor |
| Translation of image text | 133 languages, included | Separate mode, fewer languages |
| Visual matches / reverse image search | Not included (see FAQ) | Some Actors |
π How to use
- Create a free Apify account (sign up here) and get $5 in free usage credit.
- Open the Actor and add one or more image URLs (or paste base64 images under More image sources).
- Type the questions you want answered. Every image gets every question.
- Optionally pick a language under Translate image text to.
- Set Max Items (one item = one image and one question) and click Start.
- Download the dataset as CSV, Excel, JSON or XML, or read it from the Apify API.
A minimal input:
{"imageUrls": [{ "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg" }],"questions": ["Where was this photo taken?", "What does this image say?"],"translateTo": "en","maxItems": 10}
πΌ Business use cases
ποΈ Product identification and cataloging
Send product photos from suppliers or marketplaces and ask "What product is this, and what are its key specs?". Store the answer next to the OCR text of the label and the pages Google cited, and enrich your catalog without manual lookups.
π§ AI training and evaluation data
Build question-answer pairs grounded in real images and real web sources. The Markdown answer, source URLs and OCR text make a ready-made benchmark or fine-tuning set for multimodal models.
π AI search visibility (GEO)
Ask the questions your customers ask about your product images and see which pages Google AI Mode cites. Track whether your site shows up as a source over time.
π Travel, compliance and field operations
Read and translate signs, menus, labels and documents in photos from the field, and ask AI Mode to explain them in the team's language.
π Automating Google Lens AI Mode Scraper
Schedule runs and route the output with Apify integrations: Make, Zapier, Slack, Airbyte, GitHub and Google Drive, or plain webhooks. For example, drop new product photos into a Google Drive folder, run the Actor on them every hour and post each answer to a Slack channel, or append every row to a Google Sheet.
π Beyond business use cases
- Research: study how Google's AI describes images, which sources it trusts and how answers change over time.
- Personal: identify plants, landmarks, artworks or objects in your own photos and keep the answers.
- Non-profit: translate and explain signs, notices and documents for communities that need them.
- Experimentation: compare Google AI Mode's image answers with other multimodal models on the same images.
π€ Ask an AI assistant about this scraper
Paste this README into ChatGPT, Claude or any AI assistant and ask how to phrase your questions, how to join the output with your data, or how to analyze the answers. Each row is flat JSON with one image and one question, which models handle easily.
β Frequently Asked Questions
π€ Is this really Google AI Mode?
Yes. The Actor uploads your image to Google Lens and opens Google AI Mode on it, the same flow as the "AI Mode" button in Lens. The answerSeesImage field is Yes when Google's reply header confirms the image was part of the question.
π£ Which languages can I ask in? Any language AI Mode supports. It answers in the language of your question: ask in Spanish, get the answer in Spanish.
π Are the sources real links?
Yes. Google hides cited pages behind redirect links; the Actor follows each one and returns the real page URL with its title, domain, site name and snippet. Some answers cite no sources (for example when the question is only about the image's own text), and then sources is an empty list.
π€ How does translation work?
Pick a target language and Google Lens translates the text it finds in the image, block by block. If the text is already in that language, translatedText says so. All 133 languages in the list were tested against Google Lens.
π Is OCR included even if I only want answers? Yes, every row carries the OCR text, language and line boxes. Switch off Include OCR lines with boxes to keep rows small.
π« Why did a question get "no answer"? Sometimes Google AI Mode declines a question for a specific image. The Actor retries with a fresh Lens session and, if Google still declines, writes an error row that is never charged. Rephrasing the question usually helps.
π Does it return visual matches or reverse image search results? No. Google does not serve Lens visual matches to automated clients from cloud infrastructure, so this Actor focuses on what it can deliver reliably: AI Mode answers, sources, OCR and translation.
π² Will the same question give the same answer twice? Not word for word. AI Mode is generative, so wording changes between runs while the facts usually stay the same.
πΌ Which image formats work? JPEG, PNG, WebP, GIF, BMP, TIFF and AVIF, by URL or as base64. Images larger than 1,200 pixels are downscaled before upload, and OCR boxes are mapped back to the original size. Maximum 20 MB per image.
β‘ How fast is it? About 10 to 20 seconds per answer, several at a time. A 10-answer run finished in 81 seconds on Apify. Use Parallel tabs (1 to 5) to trade speed for gentleness.
π‘οΈ Do I need a proxy? No. The Actor runs without one and, when Google asks for a captcha, switches to a fresh browser identity and then to a residential exit by itself. Set your own proxy only if you want to force one.
π What formats can I export? CSV, Excel, JSON and XML from the dataset view, or any format through the Apify API.
π Integrate with any app
The dataset is available through the Apify API and every Apify integration, so you can pull it into a spreadsheet, database, BI tool or ML pipeline without glue code.
π Recommended Actors
- Google Lens OCR Scraper: Google Lens text recognition with word-level boxes, for OCR-only workloads.
- Manga OCR Translator: read and translate text in manga and comic pages.
- Google Search Scraper: organic and paid Google results with 23 fields.
- Google Shopping Scraper: product prices and offers from Google Shopping.
- Image Converter API: convert images between formats before sending them anywhere.
π‘ Pro Tip: browse the complete ParseForge collection for more AI and data Actors.
π Need Help? Open our contact form
β οΈ Disclaimer: this is an independent tool and is not affiliated with, endorsed by, or connected to Google. It collects only publicly available answers and data. You are responsible for how you use the data and the images you submit, in line with Google's terms and applicable law.