Google AI Mode Scraper
Pricing
$10.00 / 1,000 ai mode responses
Google AI Mode Scraper
Ask Google AI Mode a question and get the answer as JSON: the full answer text, the answer split into headings, paragraphs and lists with citations per block, every cited source page with domain, title and snippet, and the distinct source domains. One row per question.
Google AI Mode Scraper: the conversational answer and its citations
This google ai mode scraper returns the answer behind Google's AI tab as structured JSON. Ask a question the way you would ask a person and you get the full answer text, the same answer broken into headings, paragraphs and lists with citations attached per block, and every source page the answer drew on.
AI Mode is a separate surface from the AI Overview above the search results. It answers at length, follows the question rather than the keywords, and cites a different, usually wider set of pages. There is no official API for it. This actor returns one dataset row per question.
What you get
- The complete answer as one plain string (
answer_text), which is what you want in a spreadsheet cell or as context in a prompt. - The answer with its structure intact (
answer_blocks): headings stay headings, lists stay lists, and each block carries its own citations. That is what makes it possible to attribute one claim to one source. - Every cited source page (
sources) with site name, domain, page title, the snippet Google used, the destination URL, favicon and thumbnail. - The distinct domains (
source_domains) plus a count, so "who does Google cite for this question" is one field, not a parse. - Links inside the answer (
links) with Google's redirect already resolved to the real destination. - The thread identifier (
thread_id) Google assigned to the answer. - An honest answered flag (
answered). Questions that came back without an answer are returned marked, and are not billed.
Why scrape Google AI Mode
AI Mode changes what a search result is. Instead of ten links it returns one composed answer built from several sources at once, and the sources it picks are not the ten that would have ranked. A page can sit outside the first page of ordinary results and still be quoted in the answer, and a page that ranks first can be absent from it entirely. Neither fact shows up anywhere in a rank report.
That makes citation share the thing worth measuring. If Google's answer to a question in your category quotes three competitors and not you, that is a concrete gap with a concrete fix, and it is invisible to every tool that measures position. Running a fixed question set on a schedule turns it into a trend you can act on.
The answers are also useful as raw material. They are Google's own summary of
what a good answer to a question looks like, assembled from the sources it
trusts most. answer_blocks shows the structure it chose: what it led with,
which subheadings it used, what it decided belonged in a list. For anyone
writing the page that should have been cited, that is a better brief than a
keyword tool.
Input
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
queries | array of strings | yes | — | One question per line. |
country | string | no | us | Two-letter country code. Sets the market the answer is written for. |
language | string | no | en | Two-letter language code for the answer. |
retries | integer | no | 3 | Extra rounds to spend when an answer does not render. Range 0 to 5. |
Phrase inputs as questions. AI Mode is built for conversational input, and a bare keyword string produces a noticeably thinner answer than the same subject asked as a question.
Output
One row per question.
{"query": "what is web scraping used for","country": "us","language": "en","answered": true,"answer_text": "Web scraping is used to automatically extract large amounts of data from websites and convert it into structured formats like spreadsheets or databases...","answer_blocks": [{"type": "paragraph","text": "Web scraping is used to automatically extract large amounts of data from websites...","citations": [{ "uuid": "a41b09", "site_name": "Wikipedia", "domains": ["en.wikipedia.org"] }]},{ "type": "heading", "text": "Common uses", "level": 3, "citations": [] },{"type": "list","items": ["Price monitoring and competitor analysis","Lead generation","Market research and sentiment analysis"],"citations": []}],"answer_block_count": 9,"sources": [{"site_name": "ScrapingBee","domain": "scrapingbee.com","title": "What is Web Scraping","snippet": "Web scraping is the process of collecting structured data...","url": "https://www.scrapingbee.com/blog/...","favicon": "https://...","thumbnail": "","citation_id": "a41b09"}],"source_domains": ["en.wikipedia.org", "reddit.com", "scrapingbee.com"],"source_count": 9,"links": [{ "text": "web scraping", "url": "https://en.wikipedia.org/wiki/Web_scraping", "kind": "site" }],"products": [],"thread_id": "rHKhasr7IsidhvcP3v7aoAU","attempts": 1,"error": null,"fetched_at": "2026-09-09T15:38:11Z","duration_seconds": 3.8,"response_bytes": 443192}
Use cases
Measuring citation share in AI answers. Fix a list of questions your buyers
actually ask, run it weekly, and count how often each domain appears in
source_domains. That count is the metric: it tells you whether Google
considers you an authority on the question, which is now a separate outcome
from ranking for it.
Finding the competitors that rank reports miss. AI Mode cites pages that do not appear on page one. Run your category's questions and the domain frequency list will surface sites you were not tracking, because the surface that introduced them to your buyers is not the one you were watching.
Briefing content from Google's own structure. Before writing the page that
should be cited, pull the current answer and read answer_blocks. The
subheadings, the ordering and the list items are Google's own decomposition of
the question, taken from the system doing the deciding.
Feeding a retrieval pipeline. Each row is an answer with its sources
already attached and deduplicated. That is a usable summarisation layer for a
question set: the prose for context, sources for provenance, so the answer
you pass downstream carries its own citations.
How it compares
| This actor | Typical alternative | |
|---|---|---|
| Answer structure | Headings, paragraphs and lists preserved, citations per block | Answer as one flat string |
| Sources | Full card: name, domain, title, snippet, URL, favicon, thumbnail | Domain or URL list |
| Unanswered questions | Returned and marked, not billed | Often a failed run |
| Browser required | No | Several alternatives drive a headless browser |
| Start fee | None | $0.01 per run on some listings |
Apify's own apify/google-ai-mode-scraper is the category anchor at roughly
4,200 runs; its per-result price is tiered by plan, so compare against your own
tier. Among independent listings, opspilot.cc/google-ai-mode-serp charges
$0.01 per run as a start fee; that figure is from its live pricing. This actor
has no start fee and charges only per answer returned.
Pricing
$0.01 per AI Mode answer returned. There is no actor start fee. Questions that come back unanswered are returned in the dataset and are not charged. All pricing is pay-per-event, so you only pay for results you receive. There are no per-compute-unit charges.
Limits and gotchas
- AI Mode answers vary between runs. The same question asked twice returns the
same substance in different words. Track
source_domains, which is stable, rather than string equality onanswer_text, which is not. - Some answers cite nothing at all. A short conversational reply with an empty
sourcesarray is a real answer, not a parse failure.answer_block_counttells you how substantial it was. - Set
countryandlanguagetogether. They control the market and the language of the answer, and mismatching them gives you a result that suits neither audience. - Questions run a few at a time inside the run. A list of 500 is fine; expect minutes rather than seconds.
- Free Apify plans are capped at 10 rows per run. Split larger lists across runs or upgrade to remove the cap.
productsexists in the schema for consistency with the sibling actors and is usually empty here. For product-grounded answers use Google AI Product Answers below.
FAQ
Can I scrape Google AI Mode without an API key? Yes. Run it from the Apify Store or call it through the Apify API. No Google account, key or quota is involved.
Is AI Mode the same as the AI Overview above the search results? No. They are separate surfaces and they return different answers citing different pages. For the overview above the blue links use the Google AI Overview Scraper below.
Which sites does Google cite in AI Mode?
source_domains gives the distinct list per question; sources gives the full
card per cited page including the title and the snippet used.
Why does the same question return slightly different text each time? The answer is generated per request. The substance and the cited sources stay consistent; the exact wording does not. Build tracking on the sources.
Can I use this to track whether my domain is cited?
Yes. Run a fixed question list on a schedule and check for your domain in
source_domains per run. Presence over time is the metric.
Related Actors
- Google AI Overview Scraper for the AI summary above the ordinary search results.
- Google AI Product Answers for AI answers about a specific product, from a barcode or model number.
- Google Search Scraper for the ordinary organic results.