# Changelog of Google Ads Transparency Scraper — Advertiser Ad Library Export (`devilscrapes/google-ads-transparency`) Actor

- **URL**: https://apify.com/devilscrapes/google-ads-transparency/changelog.md
- **Full Actor documentation**: https://apify.com/devilscrapes/google-ads-transparency.md

## Changelog

All notable changes to this Actor will be documented in this file. The format
is loosely [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and we
follow [Semantic Versioning](https://semver.org/) once we hit 1.0.

### \[0.6.0] — 2026-09-20

#### Fixed

- **A 403 on the `SearchCreatives` RPC itself was not retryable — only a
  403 on the cookie warm-up GET was.** `_warm_cookies` has special-cased
  403 as a blocked-exit-IP signal (and retried it) since 0.5.0, but
  `RETRYABLE_STATUSES` — the set the actual search POST uses — never
  picked up the same status. A page-1 403 there raised immediately with
  `retryable=False`, which skipped both the page-level backoff (3
  attempts) AND the session-level fresh-proxy retry (2 attempts) that a
  429 on the identical endpoint already gets. Net effect: a customer run
  that hit a momentarily-blocked exit IP on its first request failed
  outright with zero retries, even though the exact same block on the
  warm-up GET half a request earlier would have been retried through.
  Caught by segmenting `ops/os/snapshots/` deltas by build rather than
  trusting the 30-day smear: a two-hour window on 2026-09-18
  (`16708→16716` lifetime runs, `16077→16085` still inside the 30-day
  window — i.e. zero runs aged out, so the failure count is exact, not
  modelled) measured **6 of 8 customer runs FAILED**, all on this build.
  Fix: folded 403 into the shared `RETRYABLE_STATUSES` frozenset so both
  call sites use one set; the warm-up's now-redundant `or ... in (403, 429)` special case is gone.
- Added `test_403_on_the_search_rpc_is_retried_like_the_warm_up` and
  `test_page_one_403_also_rotates_the_exit_and_retries` (`tests/
  test_page_retry.py`) — both reproduced the bug against the pre-fix code
  before the fix was applied.

### \[0.5.0] — 2026-09-13

#### Fixed

- **The start fee was charged at boot, before a single byte had been
  exchanged with Google.** Every failure downstream — a blocked exit IP,
  a 429, a wire-shape change, a rejected dataset write — therefore billed
  the customer $0.20 for a run that returned nothing. 1,771 of 16,049
  customer runs in the 30 days to 2026-09-13 finished FAILED, and that
  count is only a **floor**: since 0.4.0 a page-1 refusal on a
  single-target run is caught for fault isolation and finishes SUCCEEDED
  with zero rows, still billed, invisible on every health dashboard. Our
  own cloud run `hWoZhTGYAeMFJ6fvu` is exactly that shape —
  `chargedEventCounts {'actor-start': 1, 'ad-result': 0}`, HTTP 400 in the
  log. Billing a base fee for a run that never reached the target is what
  delisted `vrbo-vacation-rentals-scraper` at 17 runs.
  `actor-start` is now charged by `_StartFee` the first time a target's
  search **completes and its result is recorded** — rows pushed, or a
  genuine zero-match. A run where no target ever answers charges nothing
  and finishes FAILED with a status message saying so.
- **No retry policy existed anywhere in the Actor.** One non-200 on page 1
  ended the target permanently, so per-run success equalled per-*request*
  success — invisible at the ~20 runs/day baseline and catastrophic in a
  bulk burst. The 2026-09-09 burst pushed ~7,151 runs through the same
  small exit-IP pool in ~12 hours and ~21% of them failed
  (`ops/reports/BURST-NOT-A-RATE-2026-09-12.md`). Google's refusals here
  are time-windowed, not permanent, so retryable statuses (408/425/429/5xx
  and transport errors) now get exponential backoff with jitter, honouring
  `Retry-After`; a page-1 refusal additionally retries the whole session on
  a **fresh proxy exit** (`ScrapeRequest.proxy_url_factory`). A
  non-retryable status (400 wire-shape, 404) is still attempted exactly
  once — retrying it burns the customer's clock for an answer that cannot
  change.
- **A refused cookie warm-up returned an empty success.** `_warm_cookies`
  used to `return []` all the way out of `scrape`, so a blocked exit IP
  finished as a SUCCEEDED, zero-row, fully-billed run. It now raises a
  retryable `TransparencyRpcError`, which drives the exit rotation above
  and, failing that, declines to bill.
- **Fault isolation was partial.** `_scrape_target` caught only
  `TransparencyRpcError` and `ValueError`; anything else — e.g. the
  `AttributeError` from `payload.get` when the RPC returns a JSON array
  instead of an object — propagated out of the target loop and killed the
  whole multi-target run, discarding every other target's work after the
  customer had been billed. It now catches `Exception`, and a non-dict
  200 payload is classified as a page failure rather than a crash.
- **A rejected dataset write took the run down with it.** `Actor.push_data`
  was unguarded, and the cloud client validates against
  `.actor/dataset_schema.json` (see 0.2.0) — one structurally odd creative
  blob could 400 the whole push. Writes are now isolated per target,
  rows that fail to store are never charged for, and `_creative_to_item`
  coerces the five non-nullable string fields so an odd blob costs one
  imperfect row instead of the run.
- **Per-creative rows were hand-rolled dicts** (ADR-0004 forbids them, and
  this Actor is why: the cloud dataset client validates every write against
  `.actor/dataset_schema.json` while a local `apify run` does not). Rows now
  go through a new `AdCreativeRow` Pydantic model built in `parser.to_row` —
  the module that owns the wire shape, so a structurally odd blob is coerced
  where it enters rather than 400-ing the push in production. Output keys are
  unchanged and pinned by `test_row_keys_match_the_declared_dataset_schema`.
- **The terminal status message could fail a delivered run.**
  `set_status_message` ran unguarded on the last line of `main`, with an
  unbounded "Issues:" list appended — the more targets failed, the longer
  the message. It is now capped (900 chars, at most 5 notes) and
  best-effort: a rejected status message can no longer fail a run whose
  rows are already pushed and billed.

### \[0.4.0] — 2026-09-09

#### Added

- **Domain ad-screening mode** (`screeningMode: true`). Emits ONE row per
  input domain/advertiser — `domain`, `has_ads_detected`, `ad_count`,
  `first_seen_ts`, `last_seen_ts`, `sample_creative_url` — aggregated over
  the same `SearchCreatives` pull already used by per-creative mode; no new
  endpoint, no extra request. A target with zero matching creatives is a
  clean SUCCEEDED row (`has_ads_detected: false`), not a failure. Billed
  once per screening row on the existing `ad-result` event (see README
  Pricing). Per-creative mode (the default, `screeningMode: false`) is
  unchanged — existing customers and schedules see no behavior difference.
- `src/models.py` — this Actor's first Pydantic models (ADR-0004):
  `ScreeningRow` (the new output row) and a permissive `ActorInput` shape
  check that now backs `scripts/verify_input_prefill.py` for real, instead
  of the previous "pre-ADR-0004, schema-only check skipped" exemption.

#### Fixed

- **A non-200 response on the FIRST `SearchCreatives` page was silently
  treated as "no more pages."** `_paginate` used to break out of the
  pagination loop on ANY non-200 response (any page) and return whatever
  creatives it already had — so a broken first request finished as a
  SUCCEEDED run with zero rows, still charging `actor-start`. That is the
  exact silent-withholding shape that hid the `advertiserIds` wire-shape
  bug for months (see 0.3.0 below). A non-200 (or transport error) on page
  1 now raises `TransparencyRpcError` — logged loud, naming the HTTP
  status — instead of vanishing into an empty success. A failure on page
  2+ still returns the rows already collected, now with an explicit
  partial-results reason surfaced via `Actor.set_status_message`. Fault
  isolation across multiple targets in one run is unchanged: one target's
  loud failure is logged and skipped, the rest of the batch still runs.
  Regression tests with fakes for both paths:
  `tests/test_paginate_faults.py`.

### \[0.3.0] — 2026-09-09

#### Fixed

- **`advertiserIds` search has been broken since launch.** The `SearchCreatives`
  RPC's advertiser filter (field `"13"`) needs `{"1": [<id>, ...]}` — key
  `"1"` is a **list**, even for one ID. The original shape
  (`{"1": "<id>"}`) was "inferred from the SPA... not yet probed live"
  (`scripts/recon/FINDINGS.md`) and turned out wrong: Google's RPC 400s
  every advertiser-id request with it. Because the pagination loop treats
  any non-200 response as "no more pages" and returns whatever creatives
  it already collected, this shipped as a **SUCCEEDED run with zero rows**
  — the customer paid the actor-start fee ($0.20 since the 09-04 reprice)
  for a search that never actually ran. Confirmed live before the fix
  (cloud run `hWoZhTGYAeMFJ6fvu`: SUCCEEDED, `ad-result` charged 0, HTTP
  400 in the log) and after (cloud QA run `GCfxRuW2QrWyDRieq` on build
  0.3.1: SUCCEEDED, `ad-result` charged 5, 5 real Nike creative rows).
- Fixed the two `_build_search_body` unit tests that had pinned the wrong
  (broken) shape as "correct" since they were written against the
  never-verified guess, not a live probe.
- Added a live regression test
  (`test_advertiser_id_search_returns_creatives`, opt-in `-m smoke`) so a
  future RPC schema change trips CI instead of silently reintroducing
  this.

*A $5.00 / $0.02 "premium bulk" repricing was drafted here but never
submitted to Console — superseded by the smaller `ad-result` increase below,
which shipped instead. `.actor/pay_per_event.json` had drifted to reflect
the abandoned draft; reconciled to match the actually-live Console price
(2026-08-02 SEO audit caught the 8x mismatch before it could reach a push).*

#### Changed

- **Pricing raised** — `ad-result` $0.0012 → **$0.003** per creative
  (`actor-start` unchanged at $0.005). ≈$1.20/1K → ≈$3.01/1K. Notified
  2026-07-18, live 2026-08-01 (Apify's ~14-day existing-user notice period).

### \[0.2.0] — 2026-05-15

#### Fixed

- **Crash on Apify platform** when proxy was configured. The session_id
  template used a hyphen (`gat-{uuid}`), which Apify's proxy validator
  rejects (regex `^[\w._~]+$` allows only `[A-Za-z0-9_.~]`). Switched
  to underscore (`gat_{uuid}`). Local `apify run` never tripped this
  because `Actor.create_proxy_configuration()` returns `None` without
  Apify proxy credentials — the bug only surfaced in cloud runs.
- Added regression test (`test_proxy_session_id_matches_apify_regex`)
  that scans `_resolve_proxy_url` source for `session_id=...` literals
  and asserts they satisfy Apify's regex.
- **Dataset schema validation failure** on the platform. Fields
  `format_type`, `first_shown_ts`, `last_shown_ts`, `impressions`,
  `preview_image_url`, `preview_content_js_url` were declared as
  single-type (`"integer"` / `"string"`) but the parser returns
  `None` for them when the RPC omits the value. Widened to
  `["integer", "null"]` / `["string", "null"]`. Local `apify run`
  skipped this check; only the cloud dataset client enforces the
  schema strictly.

### \[0.1.0] — 2026-05-15

#### Added

- Initial release. Scrape Google Ads Transparency Center by brand domain
  or advertiser ID via direct RPC replay (`curl-cffi`, Firefox 147
  TLS+H2 fingerprint).
- Batch input: multiple `searchDomains` + `advertiserIds` in a single run.
- Apify Proxy integration with session-locked URLs for cookie continuity.
- 51 region locale labels (display-only metadata — see Limitations in README).
- Pay-per-event pricing: `$0.005` actor start + `$0.0012` per ad scraped.
- Parser tested against 4 golden fixtures (still image, rich video,
  minimal, malformed) plus opt-in live smoke tests.

#### Known limitations

- **Region is metadata, not a server-side filter.** Empirically verified —
  see `scripts/recon/FINDINGS.md`.
- Video / rich creatives expose only a `content.js` URL; rendering the
  actual frame is out of scope.
- No keyword / political-ad-specific modes yet — track via Apify Store
  listing if you want them.
