# Google Ads Transparency Scraper: Ads and Creatives (`garje/google-ads-transparency-scraper`) Actor

Scrape ads from the Google Ads Transparency Center by website domain, advertiser ID or advertiser name, with the ad content itself (headline, text, images, video links, display URL), ISO dates, exact limits and a monitoring mode for new creatives.

- **URL**: https://apify.com/garje/google-ads-transparency-scraper.md
- **Developed by:** [Aniruddha Garje](https://apify.com/garje) (community)
- **Categories:** Marketing, SEO tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.50 / 1,000 ads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Google Ads Transparency Scraper: Ads and Creatives

Export ads from the Google Ads Transparency Center by website domain, advertiser ID or advertiser name. You get the ad content itself where Google provides it. I built it on the same public endpoints the Transparency Center website uses, with plain HTTP requests and no login.

### At a glance

| | |
|---|---|
| Use it when | You need the ads a website or advertiser runs on Google, from the Google Ads Transparency Center. |
| Input you need | At least one of `domains`, `advertiserIds` or `searchTerms` (advertiser names); optionally `region`, `dateFrom`, `dateTo`, `onlyNewSince` and exact limits. |
| What you get | One row per ad with advertiser, format, first and last shown dates, image and video links, and a `contentStatus`. |
| Cost | $1.50 per 1,000 ads; $3 per 1,000 downloaded images. |
| Limits | Most text ads come back as images, not separate words; the Actor does not read text out of images. |

I built this Actor and test it against what the live site shows. The check runs again every week.

### What you get

One row per ad (creative):

| Field | Example |
|---|---|
| `advertiserId`, `advertiserName` | `AR16735076323512287233`, `Nike, Inc.` |
| `creativeId`, `adUrl` | `CR00485485175946346497`, link to the ad in the Transparency Center |
| `format` | `text`, `image` or `video`, as the Transparency Center classifies it |
| `firstShown`, `lastShown`, `daysShown` | ISO 8601 UTC timestamps, for example `2026-07-16T14:20:10Z` |
| `firstShownDate`, `lastShownDate` | The calendar dates exactly as the Transparency Center displays them (US Pacific time), for example `2026-07-16` |
| `targetDomain`, `region` | `nike.com`, `US` |
| `searchTerm`, `matchScore` | `nike`, `1.0` | For ads found by an advertiser name: the name and how well the advertiser matched it |
| `platform` | `youtube` (the platform filter that found the ad, when you select specific platforms) |
| `headline`, `description`, `texts` | The words shown in the ad |
| `displayUrl` | `www.nike.com/` |
| `imageUrl`, `imageUrls` | The ad images (most text ads are published as one rendered image) |
| `videoIds`, `videoUrls` | YouTube links for video ads |
| `contentStatus` | `ok`, `imageOnly`, `externallyHosted`, `notAvailable`, `previewFailed` or `none`, so you always know what content came back |
| `assetKeys` | Keys of saved images in the run's storage, when **Download images** is on |

### What content to expect

**Most text ads come back as images.** Google publishes most text ads in the Transparency Center as one rendered image, not as separate words. For those ads the Actor returns the image URL and sets `contentStatus` to `imageOnly`. In a measured run of 4,578 ads on 7 October 2026:

| contentStatus | Share of ads |
|---|---|
| `imageOnly` | 79.6% |
| `ok` (text, images or video links read from the preview) | 14.3% |
| `notAvailable` (a preview type the Actor does not read yet) | 6.1% |
| `externallyHosted` (served by an outside ad server; preview link returned) | under 0.1% |

Of the 2,958 text ads in that run, 2,866 (96.9%) came back as images. Headline or ad text came back for about 7% of all ads (303 of 4,578), and image URLs for 83%. The Actor does not read text out of images.

### Dates by region

`firstShown`, `lastShown` and `daysShown` came back for every ad in the regions we measured (US and `anywhere`, 8 October 2026). The Transparency Center
also shows, for each region separately, a first-shown date and impression ranges, but only for EU countries; outside
the EU it shows a per-region last-shown date only (checked on 8 October 2026). This Actor does not return those
per-region figures.

### What this Actor handles for you

1. **Ad content, not only metadata.** With **Include creative content** on (the default), each ad's preview is read and the headline, text, images, video links and display URL are returned in plain fields.
2. **No silent gaps.** Content is read from the rendered preview Google itself shows. When an ad's content is a single image, the image URL is returned and `contentStatus` says `imageOnly`, so nothing is silently null.
3. **Exact limits.** `maxAdsPerAdvertiser` is enforced exactly per advertiser, and `maxAdsPerDomain` exactly per domain, including across several domains in one run.
4. **Checked inputs and retried pages.** Inputs are checked before any request, with a clear message. Rate-limit pages are detected and retried on a new IP rather than saved as empty data, and a failed page is reported in `RUN_STATS` while the run continues with the rest.
5. **Several domains, ISO dates and the site's own formats.** Give several domains in one run; each domain has its own limit. Every date is ISO 8601. `format` is the value the Transparency Center's own format filter uses.

### Input

| Field | Default | What it does |
|---|---|---|
| `domains` | `[]` | Website domains, several per run |
| `advertiserIds` | `[]` | Direct AR... IDs |
| `searchTerms` | `[]` | Advertiser names. Every advertiser Google suggests is scored and saved in the `SEARCH_MATCHES` record |
| `minMatchScore` | 1 | Score needed to scrape a suggested advertiser; 1 means the name matches exactly, ignoring legal forms such as Inc. or GmbH |
| `maxAdvertisersPerTerm` | 3 | Best-scoring advertisers scraped per name, the one with the most ads first |
| `region` | `anywhere` | Two-letter country code or `anywhere` |
| `dateFrom`, `dateTo` | last 30 days | Shown window |
| `formats` | all | `text`, `image`, `video` |
| `platforms` | all | `search`, `youtube`, `shopping`, `maps`, `play`, using the Transparency Center's own platform filter |
| `includeCreativeContent` | true | Headline, text, image and video URLs, display URL |
| `downloadAssets` | false | Saves up to 3 images per ad to storage, as a separate paid event |
| `maxAdsPerAdvertiser` | 100 | Enforced exactly; 0 means no limit |
| `maxAdsPerDomain` | 200 | Enforced exactly per domain, across all advertisers of that domain; 0 means no limit |
| `onlyNewSince` | | Monitoring mode: only ads first shown after this time |

### Pricing

Pay per result: **$1.50 per 1,000 ads** and **$3 per 1,000 downloaded images**. You pay only for rows and files actually saved, never for errors, retries or duplicates. If you set a maximum cost per run, the Actor stops cleanly when it is reached.

### Monitoring competitors

Schedule a daily run for your competitors' domains with `onlyNewSince` set to the previous run's time, so each run returns only ads first shown after it.

### Reliability

A restarted or migrated run resumes where it stopped and does not deliver or charge the same ad twice.

### Use cases

Each one is a ready task you can open, run and copy:

- [All Google ads for a company domain](https://apify.com/garje/google-ads-transparency-scraper/examples/google-ads-by-domain): every advertiser of one website, with format and first and last shown dates.
- [Google ads with image and video links](https://apify.com/garje/google-ads-transparency-scraper/examples/google-ads-image-video-links): the creatives themselves, with image URLs and YouTube links.
- [Weekly monitor of new Google ads](https://apify.com/garje/google-ads-transparency-scraper/examples/weekly-new-google-ads): only ads first shown after a date. Schedule it weekly and move the date forward.

You can also find long-running ads by sorting on `daysShown`, or look up a brand by name: every candidate advertiser is scored in `SEARCH_MATCHES`.

### Read more

I write up what the data shows, with the run ID behind every number:

- [How long Google ads run: median days shown by format](https://garje-data-notes.laude--pify.workers.dev/articles/how-long-google-ads-run.html)
- [New Google ads in one week for 10 well-known brands](https://garje-data-notes.laude--pify.workers.dev/articles/new-google-ads-in-one-week.html)
- [Most text ads in the Transparency Center come back as images](https://garje-data-notes.laude--pify.workers.dev/articles/text-ads-come-back-as-images.html)
- [Facts and limits of this Actor](https://garje-data-notes.laude--pify.workers.dev/facts/google-ads-transparency-scraper.html) and its [changelog](https://garje-data-notes.laude--pify.workers.dev/changelog/google-ads-transparency-scraper.html)

# Actor input Schema

## `domains` (type: `array`):

Look up every advertiser running ads for these domains, in bulk (for example nike.com).

## `advertiserIds` (type: `array`):

Transparency Center advertiser IDs, for example AR16735076323512287233 (shown in the advertiser's page URL).

## `searchTerms` (type: `array`):

Advertiser names to search, for example nike. Google suggests several advertisers per name; every candidate is scored (1 means the name matches exactly once legal forms such as Inc. or GmbH are ignored) and saved in the SEARCH_MATCHES record. Only candidates scoring at least Minimum match score are scraped, and each row says which term and score found it.

## `minMatchScore` (type: `number`):

Advertiser names need at least this score to be scraped: 1 means exact name matches only, 0.8 also accepts close spellings. Use advertiserIds when you know the exact advertiser.

## `maxAdvertisersPerTerm` (type: `integer`):

At most this many matching advertisers are scraped per name, best score first, then the one with the most ads.

## `region` (type: `string`):

Two-letter ISO 3166 country code where the ads were shown, for example US or DE, or "anywhere". firstShown, lastShown and daysShown come back for every region.

## `dateFrom` (type: `string`):

Only ads shown on or after this date. Defaults to 30 days before Shown until.

## `dateTo` (type: `string`):

Only ads shown on or before this date. Defaults to today.

## `formats` (type: `array`):

Ad formats to include. Each row carries the format the Transparency Center reports.

## `platforms` (type: `array`):

Where the ads were shown, using the Transparency Center's own platform filter. Leave all selected for every platform. When some are selected, each row says which platform filter found it.

## `includeCreativeContent` (type: `boolean`):

Headline, text, image and video URLs and display URL for each ad. Most text ads are published as a single rendered image, which is returned as imageUrls. <a href="https://garje-data-notes.laude--pify.workers.dev/articles/text-ads-come-back-as-images.html" target="_blank">Read an example</a>.

## `downloadAssets` (type: `boolean`):

Saves up to 3 images per ad to the run's key-value store. Each saved image is a separate paid event. Videos are YouTube links and are not downloaded.

## `maxAdsPerAdvertiser` (type: `integer`):

Enforced exactly per advertiser. Set 0 for no limit.

## `maxAdsPerDomain` (type: `integer`):

Enforced exactly per domain in Website domains, across all its advertisers. Big brands have hundreds of resellers advertising their domain, so this keeps cost predictable. Set 0 for no limit.

## `onlyNewSince` (type: `string`):

Monitoring mode: ISO date or date-time. Returns only ads first shown after it, so a scheduled run gives you just the new creatives. <a href="https://garje-data-notes.laude--pify.workers.dev/articles/new-google-ads-in-one-week.html" target="_blank">Read an example</a>.

## `canary` (type: `object`):

For scheduled monitoring runs. Example: {"baselineStore": "canary-baseline", "thresholdPoints": 5}. The first run saves the null rate of key fields (and the mix of a status field) to that named key-value store; later runs compare and fail with an alert if any value moves by more than thresholdPoints percentage points. Leave empty for normal runs.

## `proxyConfiguration` (type: `object`):

Apify Proxy. Google throttles repeated requests from one IP, so a rotating proxy is needed.

## Actor input object example

```json
{
  "domains": [
    "nike.com"
  ],
  "advertiserIds": [
    "AR16735076323512287233"
  ],
  "searchTerms": [
    "nike"
  ],
  "minMatchScore": 1,
  "maxAdvertisersPerTerm": 3,
  "region": "US",
  "dateFrom": "2026-09-01",
  "dateTo": "2026-09-30",
  "formats": [
    "text",
    "image"
  ],
  "platforms": [
    "search",
    "youtube"
  ],
  "includeCreativeContent": true,
  "downloadAssets": false,
  "maxAdsPerAdvertiser": 50,
  "maxAdsPerDomain": 200,
  "onlyNewSince": "2026-10-01T00:00:00Z",
  "canary": {
    "baselineStore": "canary-baseline",
    "thresholdPoints": 5
  },
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}
```

# Actor output Schema

## `ads` (type: `string`):

No description

## `stats` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "domains": [
        "nike.com"
    ],
    "region": "US",
    "maxAdsPerAdvertiser": 5,
    "maxAdsPerDomain": 20,
    "proxyConfiguration": {
        "useApifyProxy": true
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("garje/google-ads-transparency-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "domains": ["nike.com"],
    "region": "US",
    "maxAdsPerAdvertiser": 5,
    "maxAdsPerDomain": 20,
    "proxyConfiguration": { "useApifyProxy": True },
}

# Run the Actor and wait for it to finish
run = client.actor("garje/google-ads-transparency-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "domains": [
    "nike.com"
  ],
  "region": "US",
  "maxAdsPerAdvertiser": 5,
  "maxAdsPerDomain": 20,
  "proxyConfiguration": {
    "useApifyProxy": true
  }
}' |
apify call garje/google-ads-transparency-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,garje/google-ads-transparency-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/DKAp75JZVYAQhDl6U/builds/RdJO8k2cB2Sqf8g1u/openapi.json
