# Changelog of Google Ads Transparency Scraper — Copy, Regions, Impressions (`foxlabs/google-ads-transparency-scraper`) Actor

- **URL**: https://apify.com/foxlabs/google-ads-transparency-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/foxlabs/google-ads-transparency-scraper.md

## Changelog

### 0.1.15 — 2026-09-27

- README: measured speeds corrected. A first run with the form's starting input took 12–132 s in four runs (the 0.1.14 notes said about 12 s), and OCR took 0.6–7.5 s per picture depending on how busy the platform's server was (the README said 3–4 s). No code change.

### 0.1.14 — 2026-09-27

- **The input form starts at 10 ads per query** (`maxAdsPerQuery` prefill). API calls without the field still get 100. Apify's daily Store test runs the form's starting input and allows 5 minutes: with 100 ads and OCR at 1 GB a run took 4.5 minutes in our measurement, Apify's test runs timed out, and the Actor was marked "Under maintenance". With 10 ads the same input took 12–132 seconds in four runs.
- **Minimum memory is 512 MB** (was 256 MB). With OCR, memory use peaked at about 240–340 MB; a run at 256 MB was stopped for lack of memory and failed without a message.
- README: two more rows in the speed table and the memory note.

### 0.1.13 — 2026-09-26

- OCR: a display URL behind icon noise in the ad's header ("© \_http://www.louveinvest.com/", "a FISPAN “SPAY support.fispan.com/") is now read as the display URL. In a review of 20 OCR'd ads, 2 had missed it.
- README: OCR coverage, accuracy (20-ad review), speed and cost, measured on the platform.

### 0.1.12 — 2026-09-26

- OCR now uses Tesseract's "fast" models (tessdata\_fast, fetched at build time from a pinned commit) instead of the "best\_int" npm packages. On the 8 sample ads it was 1.5–1.7x faster with the same or better text ("Fußballschuhe" instead of "FuRballschuhe").

### 0.1.11 — 2026-09-26

- Tesseract's own notes ("Estimating resolution…", "failed to load ./ita.special-words") no longer appear in the run log.
- `SOURCE_REPORT` has an `ocr` section: images read, and the time spent downloading, waiting and recognising.

### 0.1.10 — 2026-09-26

- **Text of image-only ads (OCR).** Google keeps many ads only as a picture: all 12 hubspot.com text ads in a test run, and most image ads. A new option, `ocrImageAds`, reads the text in the picture with Tesseract. It is on by default and needs `includeAdCopy`.
  - Languages: English, German, French, Spanish, Portuguese, Italian, Dutch, Polish and Turkish. The models ship with the Actor, so nothing is downloaded at run time.
  - Output: the text lines go to `adTexts` and the display URL to `displayUrl`, with `adCopyStatus: "ocr"` and a new `ocrConfidence` field (0–100).
  - Some pictures hold no ad, only Google's "Collapsed ad on mobile / Expanded ad" placeholder. They get `adCopyStatus: "image-placeholder"`. Pictures with no readable text get `"image-no-text"`.
  - Text read this way counts as ad copy.
  - Measured on 8 archived ads: 0.1–0.6 s per image, and search ads come out close to word-for-word.
- The Docker image keeps only the language models the Actor loads.

### 0.1.9 — 2026-09-26

Fixes from an independent review of the README against measured runs, before publication.

- **Monitoring finds new ads among still-running ones.** Google orders the list by last shown, so a new ad can sit behind pages of ads the monitor has already seen. The run used to stop after two such pages and could miss new ads; it now checks up to 10,000 listed ads per query (listing only; details and ad copy are fetched for new ads) and says so in the log and `SOURCE_REPORT` when it hits that limit.
- **Monitoring remembers delivered ads after each query and when a run is aborted or migrated.** Before, the memory was written only at the end, so an aborted run's ads came back, and were charged, again.
- **Shopping ads have a `price` field**, and `merchantName` no longer holds a price or a condition label: 34 of 111 merchant values were a price ("149,99 €") or "Reconditionné". A retailer's own listings have no merchant line, so `merchantName` stays empty there. "Przez: Channable"-style labels are read as the comparison shopping service.
- **The same advertiser, creative, domain or name given twice in one run is run once** (for example an `AR…` ID and its Transparency Center URL, or `nike.com` and `https://www.nike.com/`), so its ads are not delivered and charged twice. The duplicate leaves a free row with `status` "skipped" in `SOURCE_REPORT` and an explanatory `error`.
- **Creative URLs keep the YouTube ID.** The creative record's preview draws a video player instead of the thumbnail, and its notice ("Some controls cannot be shown…") was reported as ad text; the list-style preview is used now.
- **Not ad content anymore:** Material icons and Google Maps pins in `imageUrl` (54 rows in the test runs) and invisible direction marks around text lines (44 rows; they also hid display URLs such as "www.hubspot.com/"). Localised Shopping buttons ("Acheter maintenant", "Jetzt einkaufen", "Kup teraz") count as the call to action.
- **A query that fails midway** now reports how many ads it had delivered, instead of 0.
- README rewritten against the measurements: text coverage by advertiser (search ads are often archived as pictures), impressions only for ads running about three months in the EEA or Turkey, domain queries including other advertisers' ads, agency accounts, and the exact monitoring behaviour.

### 0.1.8 — 2026-09-26

Fixes from the pre-publication checks (input fuzzing and a repeat run on another day).

- **Dates:** `dateFrom` / `dateTo` sent with "Shown during: Any time" (the default) are now applied as a custom window. Before, they counted only with "Custom dates below", and a run that set only the dates returned unfiltered ads without a word. A fixed period such as "Last 30 days" still takes precedence, and the log says the dates were ignored.
- **Invalid input stops the run at once, with the reason in the run's status message:** a query list with only blank lines, a date not in YYYY-MM-DD form, or `dateFrom` after `dateTo`. Before, blank queries ended the run as failed with only a stack trace in the log.
- **Brand names:** when no advertiser carries the exact name, the log and the `SOURCE_REPORT` record say that the closest match was used. Google's name lookup is accent-sensitive: "ülker" matches only people named Ülker. Query the brand's domain to get all of its advertisers.
- **Creative URLs:** rows now carry `format` and the advertiser name from the creative's own record; both were empty before.
- **Ad previews** refused with HTTP 429 get a second proxy attempt on a new session. In a 600-ad test run, 4 ads had kept `adCopyStatus: "http-429"`.
- **Standby:** unknown `region`, `adFormat`, `platform` or `datePreset` values, malformed dates and non-numeric `maxAds` / `maxAdvertisers` are answered with HTTP 400 and the reason. Before, they were dropped and the request returned unfiltered ads, and a non-numeric `maxAds` returned nothing. The response has `dateNote` when a fixed period overrode `dateFrom` / `dateTo`.

### 0.1 — 2026-09-23

First version.

- Search the Google Ads Transparency Center by domain, brand name, advertiser ID or Transparency Center URL.
- Filters applied by Google: region (243 countries), ad format, platform (Search, YouTube, Maps, Play, Shopping) and "shown during" dates.
- Per ad: advertiser name, legal name, billing country and (where Google publishes it) D-U-N-S number; format; first/last shown; days shown; image or preview.
- Details (one extra request per ad): first/last shown and impression range per country, platform breakdown and total for the EEA and Turkey, audience-targeting approach, topic or Shopping product category, variation count.
- Ad copy from the ad preview: text lines, call to action, display URL, landing page, YouTube video ID, Shopping product title, merchant and the comparison shopping service that placed the ad ("By Google", "Par Yteo"). Template placeholders such as "\[Price]" or "\<Rating (Reviews)>" and the renderer's "Sponsored" labels are dropped; dynamic keyword insertion such as "{Keyword: Housse}" is kept as the advertiser wrote it.
- Ad previews are read in all three shapes seen on the platform: HTML (search, Shopping, local), renderer config (app-install and YouTube layouts: headline, description, call to action, display and destination URL, plus an `app` object with ID, name, store, category and developer), and video previews without HTML (YouTube ID from the thumbnail). Previews refused with HTTP 429 are retried and then fetched through a proxy. Previews Google itself cannot draw (its `UNABLE_TO_RENDER_PREVIEW` condition) are reported as `adCopyStatus: "preview-unavailable"`.
- Brand-name queries pick among the brand's advertiser entities by the requested region, then by ad volume: "Nike" → Nike, Inc. (not a one-ad account that is merely named "Nike"), "Decathlon" in France → Decathlon France SASU.
- Monitoring mode: return only ads a named monitor has not returned before (verified on the platform over two runs). A run sees at most "Max ads per query" ads, so pair it with a "Shown during" period.
- Every query writes its outcome to the `SOURCE_REPORT` record; queries with no ads leave an explanatory, unbilled row.
- Switches to Apify residential proxy by itself when Google rate-limits direct requests, and moves to a new proxy session whenever Google refuses the current one. Eight ads are processed at a time.
- Google Maps preview buttons ("Cancel", "Add Stop") are not reported as ad copy.
- Targeting: `topicsOfInterest` and `customerLists` were swapped in private test builds 0.1.1–0.1.5; fixed in 0.1.6. All five targeting categories are now cross-checked, ad by ad, against Google's BigQuery dataset.
- Standby-ready: the same code answers `GET /ads` and `GET /advertisers` when started in Standby mode.
