# Google Lens AI Mode Scraper - Ask Questions About Images (`parseforge/google-lens-ai-mode-scraper`) Actor

Ask Google AI Mode questions about any image and get the answer with its sources, plus Google Lens OCR text and translation. Export to CSV, Excel, JSON or XML.

- **URL**: https://apify.com/parseforge/google-lens-ai-mode-scraper.md
- **Developed by:** [ParseForge](https://apify.com/parseforge) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $5.90 / 1,000 ai mode answers

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

![ParseForge Banner](https://raw.githubusercontent.com/ParseForge/apify-assets/main/banner-v4.webp)

## 🔍 Google Lens AI Mode Scraper

> 🚀 **Ask Google AI Mode anything about an image and export the answer in seconds.** Every row pairs Google's AI answer (clean Markdown, tables included) with its cited sources, the text Google Lens reads in the image and its translation into any of 133 languages. 10 answers in 81 seconds on Apify, 18 fields per row.

Give the Actor image URLs or base64 images plus the questions you want answered. Each image is uploaded to Google Lens once, then every question is asked in Google AI Mode with the image attached, exactly like tapping "AI Mode" on a Lens result. You get back what Google answered: "This photo was taken on Yuyuan Road (愚园路) in Shanghai, China", the list of web pages it cited with real URLs, and the image's own text with pixel boxes.

One image and three questions make three rows. Answers keep their structure (headings, bullet lists, bold, comparison tables) as Markdown, so they paste straight into a doc, a CMS or an LLM prompt. The OCR text, detected language and translation ride along in every row at no extra cost.

| 🎯 Target Audience | 💡 Primary Use Cases |
| --- | --- |
| E-commerce and catalog teams | Identify products, read labels and describe item photos at scale |
| AI and data teams | Build image question-answer datasets grounded in Google's answers and sources |
| SEO and content teams | See which pages Google AI Mode cites for an image and a question |
| Travel, media and research teams | Identify places, landmarks and signs, and translate the text in photos |

### 📋 What the Google Lens AI Mode Scraper does

> 💡 **Why it matters:** Google AI Mode reads an image the way a person does and then searches the web to explain it. That answer, and the sources behind it, only exist inside a browser session. This Actor turns them into rows you can filter, compare and feed to other tools.

- **Google AI Mode answers about images.** Ask anything: "What is this?", "Where was this photo taken?", "Compare this with an alternative in a table". The answer follows the language of your question.
- **Cited sources with real URLs.** Every web page AI Mode cites, with title, domain, site name and snippet. Google's redirect links are resolved to the actual page.
- **Markdown answers.** Headings, bullet lists, bold text and tables come through as Markdown, not a flattened blob.
- **Google Lens OCR in every row.** The full text found in the image, its detected language, and each line with its pixel bounding box.
- **Lens translation into 133 languages.** Pick a target language and get Google Lens's translation of the image text.
- **Image URLs or base64.** Public links (JPEG, PNG, WebP, GIF, BMP, TIFF, AVIF) or base64 strings for files that are not online.
- **Self-healing runs.** When Google asks for a captcha, the Actor switches to a fresh browser identity, then to a residential exit, on its own.

### 🎬 Full Demo (🚧 Coming soon)

### 📊 Output

Every row is one image and one question.

| 🖼 Field | Description |
| --- | --- |
| 🖼 `imageUrl` | The image you sent (base64 images are named by position, like `Base64 image #1`) |
| ❓ `question` | The question asked in Google AI Mode |
| 🤖 `answer` | Google AI Mode's answer in Markdown |
| 👁 `answerSeesImage` | `Yes` when Google confirms the answer was generated with the image attached |
| 🔗 `sourceCount` / 📚 `sources` | How many pages the answer cites, and each one's `title`, `url`, `domain`, `sourceName` and `snippet` |
| 📝 `ocrText` | All text Google Lens reads in the image |
| 🌐 `ocrLanguage` / 🗣 `ocrLanguageName` | Detected language of that text (code and name) |
| 📏 `ocrLineCount` / 📐 `ocrLines` | Number of text lines, and each line with `x`, `y`, `width`, `height` in pixels of the original image |
| 🔤 `translatedText` | Google Lens translation of the image text |
| 🎯 `translationTargetLanguage` / 🧭 `translationSourceLanguage` | Target language you picked and the source language Lens detected |
| ↔ `imageWidth` / ↕ `imageHeight` | Size of the original image |
| 🕒 `scrapedAt` / ❌ `error` | Capture time, and an error message on failed rows |

Three real rows from a cloud run (lists shortened for reading):

**A street sign, "Where was this photo taken?", translated to Spanish:**

```json
{
  "imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg",
  "question": "Where was this photo taken?",
  "answer": "This photo was taken on **Yuyuan Road (愚园路)** in **Shanghai, China**.\n\nThe distinctive blue and white sign is standard for city streets in Shanghai. This specific intersection indicator marks a location along Yuyuan Road, a historic and trendy thoroughfare that spans across the Jing'an and Changning districts.\n\nIf you are planning a trip or a city walk here, would you like recommendations for **historic buildings**, **cafés**, or **popular photo spots** along Yuyuan Road?",
  "answerSeesImage": "Yes",
  "sourceCount": 2,
  "ocrText": "西\n\n315\n\n愚园路\n\n东\n\n309\n\nW\n\nYuyuan Rd.\n\nE",
  "ocrLanguage": "zh-Hans",
  "ocrLanguageName": "Simplified Chinese",
  "ocrLineCount": 8,
  "translatedText": "Oeste\nCamino Yuyuan\nEste\nEN\nCalle Yuyuan.\nY",
  "translationTargetLanguage": "es",
  "translationSourceLanguage": "zh",
  "imageWidth": 640,
  "imageHeight": 339,
  "sources": [
    {
      "title": "Yuyuan Road | Attractions - Your Ultimate China Travel Guide",
      "url": "https://chinagotrip.com/destinations/shanghai/attractions/yuyuan-road",
      "domain": "chinagotrip.com",
      "sourceName": "ChinaGoTrip",
      "snippet": "Yuyuan Road stands as a historic thoroughfare in Shanghai, stretching 2,775 meters through Jing'an and Changning districts. Establ..."
    },
    {
      "title": "Jing'an and Changning districts",
      "url": "https://english.shanghai.gov.cn/en-HeritageZones/20231220/17ed8c2129d24d2ab691a6aa87b9a363.html",
      "domain": "english.shanghai.gov.cn",
      "sourceName": "english.shanghai.gov.cn",
      "snippet": "N/A"
    }
  ],
  "ocrLines": [
    { "text": "西", "x": 88, "y": 87, "width": 41, "height": 32 },
    { "text": "315", "x": 76, "y": 129, "width": 58, "height": 20 }
  ],
  "scrapedAt": "2026-09-28T17:44:50.217Z",
  "error": null
}
```

**A poster, "What does this image say?":**

```json
{
  "imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png",
  "question": "What does this image say?",
  "answer": "The image provides guidelines from the **World Health Organization (WHO)** on how to **reduce your risk of coronavirus infection**.\n\nIt outlines five main preventative measures:\n\n- **Clean hands** with soap and water or alcohol-based hand rub.\n- **Cover nose and mouth** when coughing and sneezing with a tissue or flexed elbow.\n- **Avoid close contact** with anyone with cold or flu-like symptoms.\n- **Thoroughly cook** meat and eggs.\n- **No unprotected contact** with live wild or farm animals.\n\nIf you want, I can provide **more detailed instructions** on any of these steps or look up the **latest guidelines** for staying healthy. What would you like to explore next?",
  "answerSeesImage": "Yes",
  "sourceCount": 0,
  "ocrText": "Reduce your risk of coronavirus infection:\n\nClean hands with soap and water\nor alcohol-based hand rub\n\n...",
  "ocrLanguage": "en",
  "ocrLanguageName": "English",
  "ocrLineCount": 12,
  "translatedText": "Reduzca su riesgo de infección por coronavirus:\nLávese las manos con agua y jabón o con un desinfectante para manos a base de alcohol.\n...",
  "translationTargetLanguage": "es",
  "translationSourceLanguage": "en",
  "imageWidth": 905,
  "imageHeight": 480,
  "sources": [],
  "scrapedAt": "2026-09-28T17:45:02.317Z",
  "error": null
}
```

**The same sign, "What should I know about this?" (answer shortened, 11 sources):**

```json
{
  "imageUrl": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg",
  "question": "What should I know about this?",
  "answer": "The image displays an official street sign for **Yuyuan Road (愚园路)**, a famous century-old thoroughfare located in **Shanghai, China**.\n\n...\n\n### 📜 History & Cultural Significance\n\n- **Centennial Heritage:** First constructed in **1911** (the final year of the Qing Dynasty), the street was named after a historic modern garden known as \"Yuyuan\". It spans 2,775 meters across both the Jing'an and Changning districts.\n- **Protected Status:** It is designated as one of Shanghai’s **64 municipal roads that will never be widened**. ...",
  "answerSeesImage": "Yes",
  "sourceCount": 11,
  "ocrLanguage": "zh-Hans",
  "scrapedAt": "2026-09-28T17:45:06.636Z",
  "error": null
}
```

### ✨ Why choose this Actor

- **The real Google AI Mode, with the image attached.** Answers come from Google's own AI Mode session for your image, and every row says whether Google confirmed it saw the image.
- **Sources you can click.** Cited pages come with their real URLs, not Google redirect links, plus title, site name and snippet.
- **Answers that keep their shape.** Tables, lists and headings arrive as Markdown.
- **OCR and translation included in every row.** No separate mode, no extra event: text, language, line boxes and a translation into 133 languages.
- **Ask many questions per image.** The image is uploaded once and every question reuses it.
- **Clean rows.** No `null` in data fields: missing values read `N/A`, flags read `Yes` or `No`, lists are always arrays.
- **Errors are never charged.** An image that cannot be downloaded, or a question Google declines to answer, gives an error row you do not pay for.

### 📈 How it compares to alternatives

| | This Actor | Typical Google Lens or AI Mode Actor |
| --- | --- | --- |
| AI Mode answer about an image | Yes | Text-only AI Mode, or a separate paid mode |
| Cited sources with resolved URLs | Yes, title, domain, site name, snippet | Often missing or Google redirect links |
| Answer formatting | Markdown with tables and lists | Plain text |
| OCR text with line boxes in the same row | Yes, included | Separate mode or separate Actor |
| Translation of image text | 133 languages, included | Separate mode, fewer languages |
| Visual matches / reverse image search | Not included (see FAQ) | Some Actors |

### 🚀 How to use

1. **Create a free Apify account** ([sign up here](https://console.apify.com/sign-up?fpr=vmoqkp)) and get $5 in free usage credit.
2. Open the Actor and add one or more **image URLs** (or paste base64 images under **More image sources**).
3. Type the **questions** you want answered. Every image gets every question.
4. Optionally pick a language under **Translate image text to**.
5. Set **Max Items** (one item = one image and one question) and click **Start**.
6. Download the dataset as CSV, Excel, JSON or XML, or read it from the Apify API.

A minimal input:

```json
{
  "imageUrls": [{ "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg" }],
  "questions": ["Where was this photo taken?", "What does this image say?"],
  "translateTo": "en",
  "maxItems": 10
}
```

### 💼 Business use cases

#### 🛍️ Product identification and cataloging

Send product photos from suppliers or marketplaces and ask "What product is this, and what are its key specs?". Store the answer next to the OCR text of the label and the pages Google cited, and enrich your catalog without manual lookups.

#### 🧠 AI training and evaluation data

Build question-answer pairs grounded in real images and real web sources. The Markdown answer, source URLs and OCR text make a ready-made benchmark or fine-tuning set for multimodal models.

#### 🔎 AI search visibility (GEO)

Ask the questions your customers ask about your product images and see which pages Google AI Mode cites. Track whether your site shows up as a source over time.

#### 🌍 Travel, compliance and field operations

Read and translate signs, menus, labels and documents in photos from the field, and ask AI Mode to explain them in the team's language.

### 🔌 Automating Google Lens AI Mode Scraper

Schedule runs and route the output with Apify integrations: **Make**, **Zapier**, **Slack**, **Airbyte**, **GitHub** and **Google Drive**, or plain webhooks. For example, drop new product photos into a Google Drive folder, run the Actor on them every hour and post each answer to a Slack channel, or append every row to a Google Sheet.

### 🌟 Beyond business use cases

- **Research:** study how Google's AI describes images, which sources it trusts and how answers change over time.
- **Personal:** identify plants, landmarks, artworks or objects in your own photos and keep the answers.
- **Non-profit:** translate and explain signs, notices and documents for communities that need them.
- **Experimentation:** compare Google AI Mode's image answers with other multimodal models on the same images.

### 🤖 Ask an AI assistant about this scraper

Paste this README into ChatGPT, Claude or any AI assistant and ask how to phrase your questions, how to join the output with your data, or how to analyze the answers. Each row is flat JSON with one image and one question, which models handle easily.

### ❓ Frequently Asked Questions

**🤖 Is this really Google AI Mode?**
Yes. The Actor uploads your image to Google Lens and opens Google AI Mode on it, the same flow as the "AI Mode" button in Lens. The `answerSeesImage` field is `Yes` when Google's reply header confirms the image was part of the question.

**🗣 Which languages can I ask in?**
Any language AI Mode supports. It answers in the language of your question: ask in Spanish, get the answer in Spanish.

**📚 Are the sources real links?**
Yes. Google hides cited pages behind redirect links; the Actor follows each one and returns the real page URL with its title, domain, site name and snippet. Some answers cite no sources (for example when the question is only about the image's own text), and then `sources` is an empty list.

**🔤 How does translation work?**
Pick a target language and Google Lens translates the text it finds in the image, block by block. If the text is already in that language, `translatedText` says so. All 133 languages in the list were tested against Google Lens.

**📝 Is OCR included even if I only want answers?**
Yes, every row carries the OCR text, language and line boxes. Switch off **Include OCR lines with boxes** to keep rows small.

**🚫 Why did a question get "no answer"?**
Sometimes Google AI Mode declines a question for a specific image. The Actor retries with a fresh Lens session and, if Google still declines, writes an error row that is never charged. Rephrasing the question usually helps.

**👁 Does it return visual matches or reverse image search results?**
No. Google does not serve Lens visual matches to automated clients from cloud infrastructure, so this Actor focuses on what it can deliver reliably: AI Mode answers, sources, OCR and translation.

**🎲 Will the same question give the same answer twice?**
Not word for word. AI Mode is generative, so wording changes between runs while the facts usually stay the same.

**🖼 Which image formats work?**
JPEG, PNG, WebP, GIF, BMP, TIFF and AVIF, by URL or as base64. Images larger than 1,200 pixels are downscaled before upload, and OCR boxes are mapped back to the original size. Maximum 20 MB per image.

**⚡ How fast is it?**
About 10 to 20 seconds per answer, several at a time. A 10-answer run finished in 81 seconds on Apify. Use **Parallel tabs** (1 to 5) to trade speed for gentleness.

**🛡️ Do I need a proxy?**
No. The Actor runs without one and, when Google asks for a captcha, switches to a fresh browser identity and then to a residential exit by itself. Set your own proxy only if you want to force one.

**📁 What formats can I export?**
CSV, Excel, JSON and XML from the dataset view, or any format through the Apify API.

### 🔌 Integrate with any app

The dataset is available through the Apify API and every Apify integration, so you can pull it into a spreadsheet, database, BI tool or ML pipeline without glue code.

### 🔗 Recommended Actors

- [Google Lens OCR Scraper](https://apify.com/parseforge/google-lens-scraper): Google Lens text recognition with word-level boxes, for OCR-only workloads.
- [Manga OCR Translator](https://apify.com/parseforge/manga-ocr-translator): read and translate text in manga and comic pages.
- [Google Search Scraper](https://apify.com/parseforge/google-search-scraper): organic and paid Google results with 23 fields.
- [Google Shopping Scraper](https://apify.com/parseforge/google-shopping-prices-scraper): product prices and offers from Google Shopping.
- [Image Converter API](https://apify.com/parseforge/image-converter-api): convert images between formats before sending them anywhere.

> 💡 **Pro Tip:** browse the complete [ParseForge collection](https://apify.com/parseforge) for more AI and data Actors.

**🆘 Need Help?** [Open our contact form](https://tally.so/r/BzdKgA)

> **⚠️ Disclaimer:** this is an independent tool and is not affiliated with, endorsed by, or connected to Google. It collects only publicly available answers and data. You are responsible for how you use the data and the images you submit, in line with Google's terms and applicable law.

# Actor input Schema

## `imageUrls` (type: `array`):

Public image URLs (JPEG, PNG, WebP, GIF, BMP, TIFF, AVIF). Every image is uploaded to Google Lens once and every question below is asked about it in Google AI Mode.

## `questions` (type: `array`):

What to ask Google AI Mode about each image. One result row per image and question. AI Mode answers in the language of the question. Leave empty to ask "What is in this image?".

## `maxItems` (type: `integer`):

Free users: Limited to 10 items (preview). Paid users: Optional, max 1,000,000

## `imagesBase64` (type: `array`):

Images encoded in base64. Raw base64 or a data URI (data:image/png;base64,...). They are processed exactly like the image URLs.

## `translateTo` (type: `string`):

Adds Google Lens translation of the text found in the image. Leave on "No translation" to skip.

## `includeOcrLines` (type: `boolean`):

Adds an ocrLines array: every line of text Lens found, with its pixel bounding box in the original image.

## `concurrency` (type: `integer`):

How many AI Mode questions are asked at the same time (1-5).

## `proxyConfiguration` (type: `object`):

Optional. The Actor works without a proxy and switches to a fresh browser, then to a residential exit, by itself when Google asks for a captcha. Set a proxy here only to force your own.

## Actor input object example

```json
{
  "imageUrls": [
    {
      "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg"
    },
    {
      "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png"
    }
  ],
  "questions": [
    "What does this image say?",
    "Where was this photo taken?",
    "What should I know about this?"
  ],
  "maxItems": 10,
  "translateTo": "es",
  "includeOcrLines": true,
  "concurrency": 3,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields: image, question, AI Mode answer, sources, OCR text, translation

## `fullData` (type: `string`):

Complete dataset with all 18 fields, including the source list and OCR line boxes

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "imageUrls": [
        {
            "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg"
        },
        {
            "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png"
        }
    ],
    "questions": [
        "What does this image say?",
        "Where was this photo taken?",
        "What should I know about this?"
    ],
    "maxItems": 10,
    "translateTo": "es"
};

// Run the Actor and wait for it to finish
const run = await client.actor("parseforge/google-lens-ai-mode-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "imageUrls": [
        { "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg" },
        { "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png" },
    ],
    "questions": [
        "What does this image say?",
        "Where was this photo taken?",
        "What should I know about this?",
    ],
    "maxItems": 10,
    "translateTo": "es",
}

# Run the Actor and wait for it to finish
run = client.actor("parseforge/google-lens-ai-mode-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "imageUrls": [
    {
      "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/chinese.jpg"
    },
    {
      "url": "https://raw.githubusercontent.com/JaidedAI/EasyOCR/master/examples/english.png"
    }
  ],
  "questions": [
    "What does this image say?",
    "Where was this photo taken?",
    "What should I know about this?"
  ],
  "maxItems": 10,
  "translateTo": "es"
}' |
apify call parseforge/google-lens-ai-mode-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,parseforge/google-lens-ai-mode-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ct6l8XdhiaeC3qr7M/builds/knrWRhLzMMVSARFAv/openapi.json
