# Google AI Mode Scraper (`s-r/google-ai-mode-scraper`) Actor

Ask Google AI Mode a question and get the answer as JSON: the full answer text, the answer split into headings, paragraphs and lists with citations per block, every cited source page with domain, title and snippet, and the distinct source domains. One row per question.

- **URL**: https://apify.com/s-r/google-ai-mode-scraper.md
- **Developed by:** [SR](https://apify.com/s-r) (community)
- **Categories:** AI, Marketing
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$10.00 / 1,000 ai mode responses

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google AI Mode Scraper: the conversational answer and its citations

This google ai mode scraper returns the answer behind Google's AI tab as
structured JSON. Ask a question the way you would ask a person and you get the
full answer text, the same answer broken into headings, paragraphs and lists
with citations attached per block, and every source page the answer drew on.

AI Mode is a separate surface from the AI Overview above the search results. It
answers at length, follows the question rather than the keywords, and cites a
different, usually wider set of pages. There is no official API for it. This
actor returns one dataset row per question.

### What you get

- **The complete answer** as one plain string (`answer_text`), which is what you
  want in a spreadsheet cell or as context in a prompt.
- **The answer with its structure intact** (`answer_blocks`): headings stay
  headings, lists stay lists, and each block carries its own citations. That is
  what makes it possible to attribute one claim to one source.
- **Every cited source page** (`sources`) with site name, domain, page title, the
  snippet Google used, the destination URL, favicon and thumbnail.
- **The distinct domains** (`source_domains`) plus a count, so "who does Google
  cite for this question" is one field, not a parse.
- **Links inside the answer** (`links`) with Google's redirect already resolved
  to the real destination.
- **The thread identifier** (`thread_id`) Google assigned to the answer.
- **An honest answered flag** (`answered`). Questions that came back without an
  answer are returned marked, and are not billed.

### Why scrape Google AI Mode

AI Mode changes what a search result is. Instead of ten links it returns one
composed answer built from several sources at once, and the sources it picks
are not the ten that would have ranked. A page can sit outside the first page
of ordinary results and still be quoted in the answer, and a page that ranks
first can be absent from it entirely. Neither fact shows up anywhere in a rank
report.

That makes citation share the thing worth measuring. If Google's answer to a
question in your category quotes three competitors and not you, that is a
concrete gap with a concrete fix, and it is invisible to every tool that
measures position. Running a fixed question set on a schedule turns it into a
trend you can act on.

The answers are also useful as raw material. They are Google's own summary of
what a good answer to a question looks like, assembled from the sources it
trusts most. `answer_blocks` shows the structure it chose: what it led with,
which subheadings it used, what it decided belonged in a list. For anyone
writing the page that should have been cited, that is a better brief than a
keyword tool.

### Input

| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| `queries` | array of strings | yes | — | One question per line. |
| `country` | string | no | `us` | Two-letter country code. Sets the market the answer is written for. |
| `language` | string | no | `en` | Two-letter language code for the answer. |
| `retries` | integer | no | `3` | Extra rounds to spend when an answer does not render. Range 0 to 5. |

Phrase inputs as questions. AI Mode is built for conversational input, and a
bare keyword string produces a noticeably thinner answer than the same subject
asked as a question.

### Output

One row per question.

```json
{
  "query": "what is web scraping used for",
  "country": "us",
  "language": "en",
  "answered": true,
  "answer_text": "Web scraping is used to automatically extract large amounts of data from websites and convert it into structured formats like spreadsheets or databases...",
  "answer_blocks": [
    {
      "type": "paragraph",
      "text": "Web scraping is used to automatically extract large amounts of data from websites...",
      "citations": [
        { "uuid": "a41b09", "site_name": "Wikipedia", "domains": ["en.wikipedia.org"] }
      ]
    },
    { "type": "heading", "text": "Common uses", "level": 3, "citations": [] },
    {
      "type": "list",
      "items": [
        "Price monitoring and competitor analysis",
        "Lead generation",
        "Market research and sentiment analysis"
      ],
      "citations": []
    }
  ],
  "answer_block_count": 9,
  "sources": [
    {
      "site_name": "ScrapingBee",
      "domain": "scrapingbee.com",
      "title": "What is Web Scraping",
      "snippet": "Web scraping is the process of collecting structured data...",
      "url": "https://www.scrapingbee.com/blog/...",
      "favicon": "https://...",
      "thumbnail": "",
      "citation_id": "a41b09"
    }
  ],
  "source_domains": ["en.wikipedia.org", "reddit.com", "scrapingbee.com"],
  "source_count": 9,
  "links": [{ "text": "web scraping", "url": "https://en.wikipedia.org/wiki/Web_scraping", "kind": "site" }],
  "products": [],
  "thread_id": "rHKhasr7IsidhvcP3v7aoAU",
  "attempts": 1,
  "error": null,
  "fetched_at": "2026-09-09T15:38:11Z",
  "duration_seconds": 3.8,
  "response_bytes": 443192
}
```

### Use cases

**Measuring citation share in AI answers.** Fix a list of questions your buyers
actually ask, run it weekly, and count how often each domain appears in
`source_domains`. That count is the metric: it tells you whether Google
considers you an authority on the question, which is now a separate outcome
from ranking for it.

**Finding the competitors that rank reports miss.** AI Mode cites pages that do
not appear on page one. Run your category's questions and the domain frequency
list will surface sites you were not tracking, because the surface that
introduced them to your buyers is not the one you were watching.

**Briefing content from Google's own structure.** Before writing the page that
should be cited, pull the current answer and read `answer_blocks`. The
subheadings, the ordering and the list items are Google's own decomposition of
the question, taken from the system doing the deciding.

**Feeding a retrieval pipeline.** Each row is an answer with its sources
already attached and deduplicated. That is a usable summarisation layer for a
question set: the prose for context, `sources` for provenance, so the answer
you pass downstream carries its own citations.

### How it compares

| | This actor | Typical alternative |
|---|---|---|
| Answer structure | Headings, paragraphs and lists preserved, citations per block | Answer as one flat string |
| Sources | Full card: name, domain, title, snippet, URL, favicon, thumbnail | Domain or URL list |
| Unanswered questions | Returned and marked, not billed | Often a failed run |
| Browser required | No | Several alternatives drive a headless browser |
| Start fee | None | $0.01 per run on some listings |

Apify's own `apify/google-ai-mode-scraper` is the category anchor at roughly
4,200 runs; its per-result price is tiered by plan, so compare against your own
tier. Among independent listings, `opspilot.cc/google-ai-mode-serp` charges
$0.01 per run as a start fee; that figure is from its live pricing. This actor
has no start fee and charges only per answer returned.

### Pricing

$0.01 per AI Mode answer returned. There is no actor start fee. Questions that
come back unanswered are returned in the dataset and are not charged. All
pricing is pay-per-event, so you only pay for results you receive. There are no
per-compute-unit charges.

### Limits and gotchas

- AI Mode answers vary between runs. The same question asked twice returns the
  same substance in different words. Track `source_domains`, which is stable,
  rather than string equality on `answer_text`, which is not.
- Some answers cite nothing at all. A short conversational reply with an empty
  `sources` array is a real answer, not a parse failure. `answer_block_count`
  tells you how substantial it was.
- Set `country` and `language` together. They control the market and the
  language of the answer, and mismatching them gives you a result that suits
  neither audience.
- Questions run a few at a time inside the run. A list of 500 is fine; expect
  minutes rather than seconds.
- Free Apify plans are capped at 10 rows per run. Split larger lists across runs
  or upgrade to remove the cap.
- `products` exists in the schema for consistency with the sibling actors and is
  usually empty here. For product-grounded answers use Google AI Product Answers
  below.

### FAQ

**Can I scrape Google AI Mode without an API key?**
Yes. Run it from the Apify Store or call it through the Apify API. No Google
account, key or quota is involved.

**Is AI Mode the same as the AI Overview above the search results?**
No. They are separate surfaces and they return different answers citing
different pages. For the overview above the blue links use the Google AI
Overview Scraper below.

**Which sites does Google cite in AI Mode?**
`source_domains` gives the distinct list per question; `sources` gives the full
card per cited page including the title and the snippet used.

**Why does the same question return slightly different text each time?**
The answer is generated per request. The substance and the cited sources stay
consistent; the exact wording does not. Build tracking on the sources.

**Can I use this to track whether my domain is cited?**
Yes. Run a fixed question list on a schedule and check for your domain in
`source_domains` per run. Presence over time is the metric.

### Related Actors

- [Google AI Overview Scraper](https://apify.com/s-r/google-ai-overview-scraper)
  for the AI summary above the ordinary search results.
- [Google AI Product Answers](https://apify.com/s-r/google-ai-product-answers)
  for AI answers about a specific product, from a barcode or model number.
- [Google Search Scraper](https://apify.com/s-r/google-serp) for the ordinary
  organic results.

# Actor input Schema

## `queries` (type: `array`):

One question per line. AI Mode answers conversational questions best, so phrase them the way you would ask a person rather than as keywords.

## `country` (type: `string`):

Two-letter country code. Sets the market the answer is written for, so 'nl' returns Dutch shops and Dutch phrasing while 'us' returns the US view of the same question.

## `language` (type: `string`):

Two-letter language code for the answer, for example en, nl, de, fr, es.

## `retries` (type: `integer`):

How many extra rounds to spend when an answer does not render. Each round tries fresh sessions. Three is enough for almost every question.

## Actor input object example

```json
{
  "queries": [
    "what is the difference between an ean and a gtin"
  ],
  "country": "us",
  "language": "en",
  "retries": 3
}
```

# Actor output Schema

## `results` (type: `string`):

One row per query with the answer and its sources.

## `output` (type: `string`):

OUTPUT record with the run's counts and status flags.

## `errors` (type: `string`):

Failures with a code and a redacted message. Absent when the run had none.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "what is web scraping used for",
        "how do i choose between a vpn and a proxy"
    ],
    "country": "us",
    "language": "en",
    "retries": 3
};

// Run the Actor and wait for it to finish
const run = await client.actor("s-r/google-ai-mode-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": [
        "what is web scraping used for",
        "how do i choose between a vpn and a proxy",
    ],
    "country": "us",
    "language": "en",
    "retries": 3,
}

# Run the Actor and wait for it to finish
run = client.actor("s-r/google-ai-mode-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "what is web scraping used for",
    "how do i choose between a vpn and a proxy"
  ],
  "country": "us",
  "language": "en",
  "retries": 3
}' |
apify call s-r/google-ai-mode-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s-r/google-ai-mode-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/yCYZdJ94vilc3aCqo/builds/7skiDzqaiA897Qj3i/openapi.json
