# Google Ads Transparency Center Scraper (`gravelly_caladium/google-ads-transparency`) Actor

Fetches ad creatives from Google's Ads Transparency Center by advertiser name, advertiser ID, or landing domain -- format, first/last shown dates, preview assets, and (best-effort) regional reach -- via the same public JSON-RPC endpoints the web app uses. No login, no API key.

- **URL**: https://apify.com/gravelly\_caladium/google-ads-transparency.md
- **Developed by:** [Relay Data Tools](https://apify.com/gravelly_caladium) (community)
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per usage

This Actor is paid per platform usage. The Actor is free to use, and you only pay for the Apify platform usage, which gets cheaper the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-usage

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Ads Transparency Center Scraper

Fetches ad creatives from Google's [Ads Transparency Center](https://adstransparency.google.com/)
by advertiser name, advertiser ID, or landing domain -- format, first/last shown dates, preview
assets and, best-effort, regional reach -- using the same public JSON-RPC endpoints the web app
itself calls. **No login, no API key, no cookies.**

### Why this Actor

Google publishes every advertiser's ads for anyone to browse, for free, in a browser. The
existing scrapers of this data on Apify are thin and unevenly maintained: the most-used one
(~500 users/30 days) sits at 1.9 stars, a newer one (~700 users/30 days) has a single review,
and half a dozen more sit under 50 users with no ratings at all. None of that data is personal
\-- it's public ad creatives -- so there's nothing sensitive here. The gap isn't the data, it's
reliability and field completeness: retries that actually recover from Google's rate limiting,
pagination that doesn't silently truncate, and consistent output fields whether you searched by
advertiser or by domain.

### What it does

Give it a list of advertiser names, advertiser IDs, or landing domains. For each one, it:

1. Determines whether the entry is an explicit advertiser ID (`AR...`) or free text (a name or
   domain), and calls the right search mode.
2. Pages through every creative for that advertiser/domain (up to `maxAdsPerAdvertiser`),
   de-duplicating by creative ID.
3. Normalizes each creative into a flat row: format, first/last shown dates, preview asset
   URL(s), landing domain (when searching by domain), and a direct link back to the creative's
   page on adstransparency.google.com.
4. Optionally (`includeDetails`) attempts to fetch per-region reach for each creative -- see
   **Limitations**, this is best-effort and often comes back empty.

### Input

```json
{
  "advertisers": ["nike.com", "bolt.eu", "AR18378488041124659201"],
  "region": "US",
  "dateFrom": "",
  "dateTo": "",
  "formats": [],
  "platforms": [],
  "maxAdsPerAdvertiser": 100,
  "includeDetails": false,
  "proxyConfiguration": { "useApifyProxy": true }
}
```

| Field | Type | Default | Description |
| --- | --- | --- | --- |
| `advertisers` | array of strings | *(required)* | Advertiser name, explicit advertiser ID (`AR...`), or landing domain. A domain returns ads from every advertiser account whose ads point to it -- usually what you want for competitor/brand monitoring. |
| `region` | string | *(none)* | Two-letter ISO 3166-1 country code. Unverified against every country -- see Limitations. |
| `dateFrom` / `dateTo` | string (`YYYY-MM-DD`) | *(none)* | Client-side date window on last/first shown. Google's archive starts 2018-05-31. |
| `formats` | array (`TEXT`/`IMAGE`/`VIDEO`) | `[]` (all) | Best-effort -- only `TEXT` was empirically confirmed. |
| `platforms` | array (`SEARCH`/`YOUTUBE`/`SHOPPING`/`MAPS`/`PLAY`) | `[]` (all) | Experimental -- see Limitations. |
| `maxAdsPerAdvertiser` | integer | 100 | Cap per advertisers\[] entry. |
| `includeDetails` | boolean | false | Best-effort per-creative region/impression lookup; roughly doubles request count for often-null extra fields -- see Limitations. |
| `proxyConfiguration` | object | `{"useApifyProxy": false}` | Strongly recommended for anything beyond a couple of advertisers -- see Limitations on rate limiting. |

### Output (one dataset item per creative)

```json
{
  "matchedInput": "nike.com",
  "advertiserId": "AR18378488041124659201",
  "advertiserName": "Nike Retail BV",
  "advertiserCountry": null,
  "creativeId": "CR11979533485061701633",
  "format": "TEXT",
  "formatCode": 1,
  "firstShown": "2025-10-23T00:41:40+00:00",
  "lastShown": "2026-09-28T18:17:28+00:00",
  "regions": [],
  "impressionsMin": null,
  "impressionsMax": null,
  "previewUrl": "https://tpc.googlesyndication.com/archive/simgad/12487185708595374021",
  "imageUrls": ["https://tpc.googlesyndication.com/archive/simgad/12487185708595374021"],
  "videoUrl": null,
  "text": null,
  "landingDomain": "nike.com",
  "adUrl": "https://adstransparency.google.com/advertiser/AR18378488041124659201/creative/CR11979533485061701633?region=anywhere",
  "scrapedAt": "2026-09-28T19:00:00+00:00",
  "error": null
}
```

If an `advertisers[]` entry can't be resolved or the request gets blocked/rate-limited after
retries, that entry produces exactly one row with `error` set and every other field `null` --
it never crashes the rest of the run. A ready-to-use **Overview** table view (searched-for,
advertiser, format, first/last shown, landing domain, preview link, Transparency Center link,
error) is available in the dataset UI/API.

### Step 1 feasibility findings (how this was verified)

Everything below was confirmed **live, from this machine, with plain HTTP requests** (Python's
`urllib`, then this Actor's own `httpx`-based client) -- no browser automation needed at
runtime, no auth, no proxy required to get a first successful response:

- `POST /anji/_/rpc/SearchService/SearchCreatives?authuser=` with a
  `Content-Type: application/x-www-form-urlencoded` body of `f.req=<url-encoded JSON>` lists
  creatives, either for a specific list of advertiser IDs (`{"3":{"13":{"1":[ids]}}}`) or a
  free-text/domain query (`{"3":{"12":{"1":"query"}}}`). A sibling field `{"7":{"1":1}}` is
  required -- omitting it silently returns an empty result, not an error.
- Pagination works: the response's field `"2"` is a cursor string; feeding it back as request
  field `"4"` returns the next page. Confirmed across two consecutive pages with no overlap.
- `POST /anji/_/rpc/SearchService/SearchSuggestions` resolves free text to domain suggestions
  anonymously.
- `POST /anji/_/rpc/LookupService/GetAdvertiserById` returns an advertiser's name/country
  anonymously, given an ID from a search result.
- `POST /anji/_/rpc/LookupService/GetCreativeById` **does** work anonymously (confirmed via a
  real browser session showing a 200 with rich per-region impression-range data) -- contrary to
  one public write-up's claim that it needs an XSRF token. However, this session could not
  reverse-engineer its anonymous *request* shape: every field-number guess tried (based on the
  response's own field numbers, and on the two other RPCs' conventions) returned an empty `{}`
  rather than an error, which usually means "valid shape, filtered to nothing" rather than
  "wrong shape" for this API (see next section) -- so it's implemented as a best-effort call
  behind `includeDetails` and documented as such, not asserted as solved.
- **Rate limiting is real and escalates.** A burst of roughly 15-20 requests in a couple of
  minutes went from clean 200s, to HTTP 429, to full reCAPTCHA interstitial pages served as an
  HTTP 302 redirect to `google.com/sorry/...`. Waiting ~30+ minutes and retrying did **not**
  clear the block from this machine's IP -- consistent with the one public write-up that
  explicitly warned backing off alone doesn't help and the source IP needs to change. This is
  exactly why `proxyConfiguration` is a first-class input, and why the client treats a 302 the
  same as a 429 (retry with backoff, then surface distinctly).

Because of that block, the request budget for verifying **request** shapes (as opposed to just
reading confirmed response shapes) ran out partway through development. The unconfirmed pieces
are called out explicitly in **Limitations** below, and the code is written to degrade
gracefully (never crash, never silently invent data) wherever a shape wasn't confirmed.

### Limitations

- **`GetCreativeById`'s anonymous request shape is unsolved.** The endpoint clearly exists and
  clearly returns rich per-region data (see above) -- this session just didn't find the right
  request fields before hitting the rate limit. With `includeDetails=true`, the Actor calls it
  with its best guess; in practice this currently returns `{}` for every creative, so
  `regions[]`/`impressionsMin`/`impressionsMax` will usually stay empty. This is handled as
  "no extra data," not a failure -- the base row from `SearchCreatives` is always returned
  regardless. If you get this endpoint working, `src/parsers.py::parse_creative_detail` is
  already written and unit-tested against a real captured response
  (`tests/fixtures/get_creative_by_id.json`) -- only `rpc_client.py::get_creative_by_id`'s
  request payload needs fixing.
- **Format codes are partially reverse-engineered.** Only `formatCode == 1` was empirically
  confirmed as `TEXT` (a creative's own detail page in a live browser was labelled "Format:
  Text"). `IMAGE` (2) and `VIDEO` (3) in `constants.py` are informed guesses based on public
  write-ups and on the shape of the creative payload (a static `<img>` screenshot vs. an
  interactive iframe/`content.js` URL). `formatCode` (the raw integer) is always included
  alongside `format` so you can verify/re-map it yourself.
- **Platform filtering (`platforms[]`) is unverified.** Testing was cut short by rate limiting
  before this could be confirmed against the live API. The Actor still accepts it and attempts
  a server-side filter; if Google rejects the shape, it automatically retries the same page
  once without the platform filter rather than failing the whole advertiser/domain, and logs a
  warning. Per-creative platform (e.g. "this ad ran on YouTube vs. Search") is not exposed by
  any response field this session found, so it's not in the output schema at all.
- **Region filtering (`region`) is implemented per two independent public write-ups (region
  code = 2000 + ISO 3166-1 numeric) but wasn't empirically re-verified live** -- the live test
  for it was the request that first triggered the 429 in this session. If a region filter
  returns unexpectedly few/zero results, try leaving `region` empty and compare.
- **Date filtering is client-side, not server-side.** The web UI's calendar picker clearly
  triggers a new `SearchCreatives` request when you change it, so a server-side date filter
  almost certainly exists -- but its request field wasn't captured before the rate limit hit.
  `dateFrom`/`dateTo` here fetch normally and then drop out-of-window rows, which is correct
  but not more efficient for advertisers with very long ad histories.
- **Ad text/headline is rarely extractable.** Google mostly serves creative previews as either
  a static screenshot image or an interactive iframe pointing at
  `displayads-formats.googleusercontent.com/.../content.js` (a signed, script-rendered
  preview) -- there isn't a reliable structured text field in the `SearchCreatives` response
  for either case. `text` is included in the schema for forward-compatibility but is `null`
  in practice today.
- **Advertiser-name resolution is intentionally simple.** Rather than depend on
  `SearchSuggestions`'s advertiser-suggestion branch (whose shape this session never actually
  observed live -- every suggestion response captured was domain/site suggestions only), a
  plain-text `advertisers[]` entry that isn't an explicit `AR...` ID is sent directly as a
  free-text query to `SearchCreatives`, the same call confirmed to work for domain strings.
  This is simpler and stays entirely within verified request shapes, at the cost of not
  pre-resolving to a single canonical advertiser (a name search can span multiple advertiser
  accounts, same as a domain search does).
- **Rate limiting means real runs need `proxyConfiguration`.** Without a proxy, expect roughly
  a few dozen requests (~1-2 advertisers' worth of pagination, fewer with `includeDetails`)
  before hitting a 302/429 wall that does not clear quickly from the same IP.
- **This Actor only reads what Google already serves to any visitor of the Transparency
  Center.** It does not require or use a login, and returns no personal data -- only public ad
  creatives, formats, and dates.

### Development / running locally

```
python -m venv .venv
.venv\Scripts\pip install -r requirements.txt pytest pytest-asyncio
.venv\Scripts\python -m pytest              # unit tests, fixtures only, no network
.venv\Scripts\python -m pytest -m network   # optional live smoke test against the real API
apify run                                    # full local run via storage/key_value_stores/default/INPUT.json
```

### Publishing (not run as part of this build)

```
apify login
apify push
```

Neither command was executed while building this Actor.

### FAQ

**Why is a row's `error` set instead of the row just being missing?**
So a run's dataset always accounts for every input you gave it -- you can tell "this advertiser
genuinely has zero ads matching your filters" apart from "this advertiser's request got
blocked, try again with a proxy" by reading `error`.

**Why do some rows have `landingDomain: null`?**
Google's search API only echoes back a landing domain when you searched *by* domain. Searching
by advertiser name or ID doesn't expose it in the same response -- see Limitations.

**Can I look up a single, exact advertiser account instead of everything matching a name?**
Yes -- pass its explicit ID (`AR...`, visible in any row's `advertiserId` or in the Transparency
Center's own URLs) in `advertisers[]` instead of its name.

**Why did `includeDetails=true` not add anything?**
Expected today -- see Limitations on `GetCreativeById`. The base fields from `SearchCreatives`
are unaffected either way.

**How is this priced?**
See [PRICING.md](PRICING.md) for the proposed pay-per-event plan.

# Actor input Schema

## `advertisers` (type: `array`):

Who to fetch ads for. Each entry can be an advertiser name (e.g. "Nike"), an explicit advertiser ID (e.g. "AR18378488041124659201"), or a landing domain (e.g. "nike.com"). A domain returns ads from every advertiser account whose ads point to that domain, which is usually what you want for competitor/brand monitoring. A name is resolved to the best-matching advertiser ID via Google's own search-suggestions endpoint.

## `region` (type: `string`):

Two-letter ISO 3166-1 country code to scope results to a region (e.g. "US", "DE", "LV"). Leave empty for no region filter (all regions). This filter is implemented per publicly documented reverse-engineering of the endpoint; it was not exhaustively verified against every country during development -- if a region code returns unexpectedly few results, try leaving it empty.

## `dateFrom` (type: `string`):

Only include creatives last shown on or after this date (YYYY-MM-DD). Google's archive starts 2018-05-31; leave empty for no lower bound.

## `dateTo` (type: `string`):

Only include creatives first shown on or before this date (YYYY-MM-DD). Leave empty for no upper bound (today).

## `formats` (type: `array`):

Restrict to specific creative formats. Best-effort: only the TEXT code was empirically confirmed against the live UI during development (see README Limitations); IMAGE and VIDEO are informed guesses. Leave empty for all formats.

## `platforms` (type: `array`):

Restrict to specific Google platforms (Search, YouTube, Shopping, Maps, Play). Experimental / unverified -- see README Limitations. Leave empty for all platforms.

## `maxAdsPerAdvertiser` (type: `integer`):

Stop paginating a given advertiser/domain once this many creatives have been collected.

## `includeDetails` (type: `boolean`):

For each creative, also call the creative-detail endpoint to try to fill in regions\[] and impression ranges. This roughly doubles the number of requests and is best-effort: the detail endpoint's exact request shape was not fully reverse-engineered during development, so this may return no extra data for some or all creatives (regions\[] will simply be empty; nothing fails).

## `proxyConfiguration` (type: `object`):

Optional. Recommended for larger runs -- Google rate-limits a single IP to roughly a few dozen requests before returning HTTP 429 for a while. Apify Proxy (residential or datacenter) rotates the source IP per request to avoid this.

## Actor input object example

```json
{
  "advertisers": [
    "nike.com",
    "bolt.eu",
    "AR18378488041124659201"
  ],
  "region": "",
  "dateFrom": "",
  "dateTo": "",
  "formats": [],
  "platforms": [],
  "maxAdsPerAdvertiser": 100,
  "includeDetails": false,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `results` (type: `string`):

No description

## `resultsAll` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "advertisers": [
        "nike.com",
        "bolt.eu"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("gravelly_caladium/google-ads-transparency").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "advertisers": [
        "nike.com",
        "bolt.eu",
    ] }

# Run the Actor and wait for it to finish
run = client.actor("gravelly_caladium/google-ads-transparency").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "advertisers": [
    "nike.com",
    "bolt.eu"
  ]
}' |
apify call gravelly_caladium/google-ads-transparency --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gravelly_caladium/google-ads-transparency"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/sKbSUkNedErAOmUPc/builds/mI8pfFCT5n36lwwiY/openapi.json
