# Changelog of Google Images Scraper — Image Search with Full Metadata (`foxlabs/google-images-scraper`) Actor

- **URL**: https://apify.com/foxlabs/google-images-scraper/changelog.md
- **Full Actor documentation**: https://apify.com/foxlabs/google-images-scraper.md

## Changelog

### 0.1.4 — 2026-10-01

- Documentation only, no code change: the README was checked claim by claim against the test runs and corrected (request success and timeouts, the image check covered 39 sites, related searches that change the subject, what `productBrand` holds, when `format` is empty). The sample rows now come from build 0.1.3 runs.

### 0.1.3 — 2026-10-01

- New field `imageUrlCrawlerOnly` (true or false on every image row): `true` when Google's link to the original is an Instagram or Facebook "lookaside" crawler link or a TikTok image-API link. In our checks these returned a web page or HTTP 403 instead of the file (Meta links 11 of 11, TikTok 2 of 2); they were 369 of 4,670 image rows (7.9 %) in the test runs of builds 0.1.1–0.1.2. `thumbnailUrl` and `pageUrl` still work for those rows.
- New input `skipCrawlerOnlyImages` (off by default): leaves those images out, like the other Actor-side filters (not charged), and fills the limit with other images.
- New input `relatedSearchesExclude`: words or phrases; a related search containing one of them is not used (for "red panda", Google suggested "red panda drawings", "anime red panda" and "turning red red panda"). Skipped ones are listed in `SOURCE_REPORT` (`relatedSearchesExcluded`).
- `SOURCE_REPORT` lists every Google search in `searchesUsed`, including a page read in parallel that was not needed (0.1.2 counted it in `searches` but did not list it) and a first page that was tried again. The note names a later page of the term that failed twice. Request latency now covers answered requests only; `timeouts` counts requests stopped at the 90-second limit.
- A page read alongside the page where the term's own results ran out is now used (it is already paid for) instead of dropped.
- After a related search that delivers nothing, the next batch reads only as many related searches as are left before the stop (3 in a row that deliver nothing), so an impossible filter reads at most 4 of them instead of up to 6.
- No field removed or renamed.

### 0.1.2 — 2026-10-01

- New field `licenseName`: the short name of a Creative Commons license ("CC BY-NC 2.0", "CC0 1.0", "Public Domain Mark 1.0") when the image's license link is a creativecommons.org license; `null` otherwise.
- When the Actor-side filters (aspect ratio, minimum width or height) leave nothing, the Actor stops widening after 3 related searches that deliver no image, instead of using up `maxSearchesPerQuery`. The status row still explains how many images each filter removed.
- After the first page, a term's own result pages are read two at a time when more than one page is still needed (still delivered in Google's order). Run time still depends mostly on how fast Google answers: in two test runs, 500 filtered red panda images took 326 s instead of 442 s, while 700 golden retriever images took 248 s instead of 221 s. When the last page turns out to be the end, the page read alongside it is one extra Google request.

### 0.1 — 2026-09-30

First version.

- Google Images results for any search term through Apify's Google SERP proxy: up to 100 images per Google page, the term's pages read until Google has no more, up to 1,000 images per term.
- Per image: the original image URL with width and height, aspect ratio and orientation, format (from the file name), file size (Google's label, in bytes), dominant color, Google's thumbnail, the page the image appears on with its title, site name and domain, and Google's "About this image" link.
- When Google shows them: the product on the page (title, brand, price, currency, Google Shopping link) and the image's license details (license link, creators, credit line, copyright notice).
- Related searches (`relatedSearches: auto`, default): when a term's own pages run out before `maxImagesPerQuery`, the Actor continues with the related searches Google lists that contain every word of the term. Those rows carry `sourceQuery` and `fromRelatedSearch: true`. `off` keeps only the term's own results.
- Google filters: size, color (black and white, transparent, 12 colors), type (photo, clip art, line drawing, animated), file type, time, usage rights (Creative Commons, commercial). Each value was checked against Google; Google's `il:cl` usage-rights value and its aspect-ratio value had no effect and are not used. No "face" filter.
- Actor-side filters on the image size: aspect ratio (tall, square, wide, panoramic), minimum width and height. Images left out are not charged.
- Duplicates removed by Google's image ID within a term; optionally across terms (`removeDuplicatesAcrossQueries`).
- 51 countries (the country's Google domain) and any interface language.
- Pricing: one `image` event per delivered image. Status rows (no result, everything filtered) and related-searches rows are free.
- `SOURCE_REPORT` record per run: every Google search per term, duplicates, images left out per filter, Google's related searches, request latency.
- Batch and Standby (HTTP API) modes share the same code and the same event.
- Default memory 256 MB (up to 1 GB can be chosen).
