# Changelog of Google AI Overview Citation Scraper (`devilscrapes/ai-overview-citations`) Actor

- **URL**: https://apify.com/devilscrapes/ai-overview-citations/changelog.md
- **Full Actor documentation**: https://apify.com/devilscrapes/ai-overview-citations.md

## Changelog

All notable changes to this Actor are documented here.

### \[0.9] — 2026-08-06

#### Fixed

- **Uncaught browser-launch crash from Camoufox's `geoip=True` self-detection
  — the actual driver of the residual ~63% public 30-day failure rate that
  survived the v0.8 fault-isolation fix.** `open_browser` passed
  `geoip=True` to Camoufox whenever a proxy was set (always, since Apify
  Proxy is mandatory). That makes Camoufox call `camoufox.ip.public_ip()`
  **before** the Firefox process is even spawned — a blind sweep of six
  third-party IP-echo services (`api.ipify.org`, `checkip.amazonaws.com`,
  `ipinfo.io`, `icanhazip.com`, `ifconfig.co`, `ipecho.net`) through the same
  proxy exit, each with only a 5s timeout. On the shared FREE-tier
  `BUYPROXIES94952` pool that sweep failing (proxy rejects the echo-service
  connection, or every one of the six times out) raises Camoufox's *own*
  exception classes (`InvalidProxy`, `InvalidIP`, `LocaleError` and
  subclasses) — none of which are `playwright.Error` subclasses, so the
  v0.8 fix's `RECOVERABLE_BROWSER_ERRORS = (PWError, PWTimeoutError)` never
  caught them. They propagated straight out of `open_browser`'s launch,
  through `_run_browser_loop`, past `open_browser`'s partially-entered
  context manager, and crashed the whole `async with Actor:` block —
  **before a single query was ever attempted** — discarding the entire run
  regardless of how solid the per-query fault isolation underneath it was.
  - `open_browser` no longer depends on Camoufox's network self-detection at
    all. It now pins Camoufox's `locale=` (e.g. `en-US`) directly from the
    caller's `country`/`language` input via the new `camoufox_locale()`
    helper — a pure local string build, no network call, so this failure
    mode is structurally eliminated rather than merely caught-and-retried.
  - `RECOVERABLE_BROWSER_ERRORS` (`src/browser.py`) is also widened to
    include Camoufox's own launch exceptions (`InvalidIP`, `InvalidProxy`,
    `LocaleError`) as a safety net for any other Camoufox-internal launch
    failure, consistent with the existing relaunch-bounded-by-
    `BROWSER_LAUNCH_ATTEMPTS` design from v0.8.
- **Geo-mismatched proxy exit vs. requested Google locale** (memory
  `feedback-pin-proxy-country`). The proxy exit country was never pinned,
  so a residential/`BUYPROXIES94952` exit could land in any country while
  the Google request still asked for `gl=<country>` and (previously) the
  Camoufox fingerprint was set to whatever country the geoip self-detection
  found — a three-way inconsistency Google can read as a bot signal, and a
  silent wrong-locale-results risk independent of any crash.
  `_resolve_proxy_or_die` now requests `Actor.create_proxy_configuration(...,
  country_code=<input country>.upper())` first, so the real exit actually
  matches the locale used everywhere else in the run; if the pool has no
  IPs for that country it gracefully retries unpinned within the same
  outer attempt rather than treating that as exhaustion.
- New tests: `tests/test_geoip_launch_crash.py` (8 cases) — widened
  exception tuple, `open_browser` no longer sets `geoip=`,
  `camoufox_locale()` tag format, proxy country pinning + unpinned
  fallback, and a Camoufox (non-Playwright) launch exception is retried by
  `_run_browser_loop` exactly like a Playwright one already was.
- **Verification**: `ruff check .` clean, `pytest -q` 84/84 passed (was 76
  before this fix's 8 new tests — 5 pre-existing proxy-resolution tests in
  `tests/test_bug_fixes.py` updated for the new country-pinned-then-unpinned
  call shape), `pyright`
  clean, `scripts/verify_cloud_entrypoint.py ai-overview-citations` → OK,
  `scripts/verify_input_prefill.py ai-overview-citations` → OK. Also
  confirmed via a real (non-mocked) local Camoufox launch that
  `locale="en-US"` boots and navigates cleanly with no network dependency,
  and via a real local `apify run` against `tests/fixtures/input.qa.json`
  that the new country-pinned-then-unpinned proxy resolution path executes
  end-to-end and degrades to a clean exit (not a crash) when Apify Proxy
  external access is unavailable (expected in this sandbox — see
  `docs/specs/ai-overview-citations/notes.md`).
- **Not done**: no live cloud reproduction of the actual geoip-lookup crash
  was attempted — this sandbox's local network cannot reach Apify Proxy
  (`x-apify-proxy-error` 403 on every CONNECT, confirmed against
  `BUYPROXIES94952` directly), so the crash can only be forced
  deterministically via Camoufox's documented exception classes in a unit
  test, exactly as v0.8's crash path was. A human should confirm via cloud
  QA (mandatory before this reaches the Store) that: (a) runs complete
  without the `InvalidProxy`/`InvalidIP` traceback that public customer
  runs were almost certainly hitting, and (b) the `Camoufox locale=` code
  path (`src/browser.py::open_browser` was verified locally in isolation,
  not against a live proxied Google fetch).

#### Published — 2026-08-06

- `apify push` → build `0.9.1` (id `6eJB9dtJ1jXQc9pWc`) SUCCEEDED.
  Cloud QA run `qVLJZRlqOPjwtRdhM` against
  `tests/fixtures/input.qa.json`: **PASS**. Confirms the crash class this
  version fixes is gone — cloud log shows
  `Launching Camoufox (headless=True, proxy=True, locale=en-US)` with no
  `geoip=True` self-detection sweep, and no `InvalidProxy` / `InvalidIP` /
  `LocaleError` traceback anywhere in the run. Both QA queries hit Google
  CAPTCHA on the shared `BUYPROXIES94952` pool (expected/known proxy-pool
  behaviour, unrelated to this fix); the run degraded gracefully — WARNING
  retry, then a Pydantic-validated `blocked_by_captcha` marker row per
  query — and exited `0` instead of crashing, which is exactly the
  fault-isolation behaviour under test. `chargedEventCounts`:
  `actor-start=1`, `result-row=2`, matching the 2 dataset rows pushed.
  Store metadata already in sync (`sync_store_metadata.py` reported
  `already in sync` — no README/SEO/category changes in this release);
  icon unchanged (`pictureUrl` already set on the Store record). No
  pricing or public/private state changes — already public and priced.

### \[0.8] — 2026-08-05

#### Fixed

- **The per-query loop had no fault isolation — a single bad query could
  still hard-FAIL the entire run.** The 0.5/0.6/0.7 fixes each swallowed a
  *teardown-time* driver crash (`browser.close()`, then `page.close()`),
  but nothing ever caught an exception raised **during** a query's fetch —
  `page.goto()` navigation timeout, a mid-navigation driver crash, or
  `open_page()` failing outright. Any of those propagated straight out of
  `_process_query`, through `_run_browser_loop`'s `for` loop, past
  `open_browser`'s context manager, and crashed the whole
  `async with Actor:` block — discarding every row already pushed for
  prior queries in the same run and marking the run FAILED. This is the
  root cause the CEO report's 36% (21/33) public 30-day success rate
  pointed at: `recent_failures` (our own runs) was empty because our QA
  fixture queries happen to navigate cleanly, but real customer queries
  (typos, rare locales, slower proxy exits) hit navigation timeouts /
  transient driver crashes often enough to fail 2/3 of runs.
  - `_process_query` now retries a fetch crash on the same budget as a
    CAPTCHA retry (`CAPTCHA_RETRY_ATTEMPTS`), and degrades to a new
    `_query_error_row` marker (distinct `query_error` field, never
    conflated with `blocked_by_captcha` or a genuine no-overview result)
    instead of raising when every attempt crashes.
  - `_run_browser_loop` / new `_drain_queries` now also survive the
    browser **process** dying mid-run: after every query it checks
    `browser.is_connected()`; if the driver disconnected, it relaunches a
    fresh Camoufox (bounded by `BROWSER_LAUNCH_ATTEMPTS = 2`: initial
    launch + one relaunch) and resumes with the queries not yet processed,
    instead of letting the crash kill the run.
  - `open_browser(proxy)` launch failures are now caught the same way
    (`RECOVERABLE_BROWSER_ERRORS`, defined once in `browser.py`) so a
    driver that dies *at launch* also gets one relaunch attempt rather
    than crashing before a single query runs.
  - `ResultRow.query_error: str | None` (new field, `.actor/
    dataset_schema.json` updated to match) and `_finalize`'s status
    message now report a three-way breakdown — clean / CAPTCHA-blocked /
    crashed — instead of only clean vs. CAPTCHA.
  - The pre-existing "no AI Overview rendered" path (`_no_overview_row`,
    `find_ai_overview_html` returning `None`) was already correct — it
    was never miscounted as a failure. No change there; verified by the
    existing `test_build_rows_no_overview` / `test_error_page_*` suite
    plus the new regression tests below.
- Reproduced locally: `_process_query` raising `PWTimeoutError` mid-fetch
  used to propagate out of `_run_browser_loop` uncaught (confirmed via a
  focused unit test before the fix, now green after). Live cloud
  reproduction against real Google traffic was not attempted for this
  fix — the code path is unconditionally exercised by mocking Playwright's
  documented exception classes, which is deterministic and doesn't burn
  proxy/compute credits chasing a timing-dependent crash.
- New tests: `tests/test_query_fault_isolation.py` — a crashing fetch on
  one query must not stop the run, must retry once, must degrade to a
  `query_error` marker row after exhausting retries, and a dead browser
  mid-run must trigger exactly one relaunch and resume the remaining
  queries.

### \[0.7] — 2026-07-30

#### Fixed

- **Per-query `page.close()` driver crash was still hard-FAILING runs.**
  The 0.6 fix swallowed the Firefox-driver "Connection closed while reading
  from the driver" crash (triggered by `FFPage._onUncaughtError` on certain
  in-page JS errors — Google-side script bugs, not ours) only at
  `browser.close()` teardown. The *identical* error was still propagating,
  uncaught, from the per-query `page.close()` call inside
  `main._process_query`'s `finally` block — discarding HTML that had
  already been fetched successfully in the same attempt and crashing the
  entire run before any dataset rows were pushed. This was the actual
  driver of the chronic 18-20% public 30-day success rate (unchanged
  across the 0.4/0.5/0.6 fixes, which addressed different crash sites).
  Reproduced via a live cloud run against `tests/fixtures/input.qa.json`
  on build 0.6.1 (`runId=mYL60RvvbhPlXxxmL`) before the fix, confirming the
  exact traceback. New `browser.close_page_quietly()` wraps `page.close()`
  the same way `_close_browser_quietly()` already wraps `browser.close()`,
  narrowed to `playwright.async_api.Error` (the actual driver-crash class)
  instead of bare `Exception`.

### \[0.6] — 2026-07-10

#### Fixed

- **Browser teardown crash after a successful run.** Cloud QA on the 0.5
  republish confirmed the `no_viewport` fix worked (navigation and
  scraping succeeded) but hit a second, distinct crash: Playwright's
  Firefox driver raised during `browser.close()`, triggered by an
  uncaught in-page JS error event it mishandles internally
  (`FFPage._onUncaughtError` / `Connection closed while reading from
  the driver`). This fired *after* every query had already succeeded
  and every row was already pushed and charged, turning an
  otherwise-successful run into a hard FAILED with zero visible rows.
  `open_browser` now catches and logs a teardown-time exception
  instead of letting it propagate — the scrape work is already done
  by the time `close()` runs.

### \[0.5] — 2026-07-10 (superseded by 0.6 — see above)

#### Fixed

- **Cloud-only crash in `open_page()`.** `browser.new_context()` was
  called with no arguments, letting Playwright's Firefox driver
  negotiate its own default viewport. The resulting
  `Browser.setDefaultViewport` CDP payload was rejected by the
  Camoufox/Firefox CDP schema baked into the published Docker image
  (`"Found property '<root>.viewport.isMobile' - false which is not
  described in this scheme"`), crashing every cloud run before a
  single query was attempted — never reproduced locally, where a
  different cached browser binary was in use. Fixed by passing
  `no_viewport=True`, which skips the redundant negotiation (Camoufox
  already fixes the window size via its injected fingerprint config).
  Caught by cloud QA on the 0.4 republish attempt.

### \[0.4] — 2026-07-10 (superseded by 0.5 — see above)

#### Fixed

- **Proxy-pool exhaustion no longer hard-fails the run.** `_resolve_proxy_or_die`
  now retries proxy resolution up to 3 times with a 3s backoff between
  attempts before giving up; on persistent exhaustion `main()` exits
  cleanly with a status message instead of raising `SystemExit(1)`. This
  was the primary driver of the public 30-day failure rate — proxy-pool
  misses on the shared FREE-tier group were previously fatal on the
  first attempt.
- **Google's own "AI Overview is not available" error page no longer
  produces false citations.** The text-based selector fallback now
  requires an exact `"AI Overview"` heading match (not `startswith`) and
  rejects any candidate container whose text carries Google's own
  error copy ("not available", "can't generate", "try again later").
  Previously the fallback treated the error block's heading as a real
  AI Overview container and harvested its ordinary SERP links as
  citations.
- **CAPTCHA marker rows are now unambiguous.** `ResultRow` gained a
  `blocked_by_captcha` field; `_marker_row` sets it `True`, every other
  row builder leaves it `False`. `_is_captcha_marker` now reads this
  field directly instead of duck-typing off `selector_used is None`,
  which could misclassify a legitimate "no AI Overview appeared" result.

#### Added

- 20 new regression tests (`tests/test_bug_fixes.py`) covering the three
  fixes above, plus `tests/fixtures/ai_overview_error_unavailable.html`.

### \[0.1.0] — 2026-05-16

#### Added

- Initial release: Google AI Overview citation tracker.
- Camoufox-rendered SERP, 8-selector priority battery for AI Overview container detection.
- Text-based selector fallback for Google markup drift (`h1/h2/h3` starts-with-"AI Overview").
- Citation extraction (URL, registrable domain, anchor title, 1-based position) with dedup and https-only filtering.
- 200-char text excerpt of the AI Overview body.
- Per-query Pydantic-validated row; one row per (query x citation) or single marker row when AI Overview did not appear.
- `selector_used` telemetry column on every row — drift detection via dataset query.
- Mandatory Apify Proxy: RESIDENTIAL → BUYPROXIES94952 fallback path.
- CAPTCHA-aware: detects `/sorry/index`, rotates session, retries once, emits marker row on persistent block.
- Per-query proxy `session_id` rotation (Apify regex-safe via `s_{uuid_hex}`).
- Pay-Per-Event: `$0.05` actor-start + `$0.005`/result-row (~$5.50 / 1,000 rows).
- Charge call regression test pins `Actor.charge(event, count=N)` shape — no `idempotency_key` (SDK 3.x rejects it).
- Fail-loud on zero rows: `SystemExit(1)` + clear status message.
- 44 unit tests; QA fixture under `tests/fixtures/input.qa.json` (2 queries).
