# Bulk Image Downloader & Google Images Scraper (`s_actors/bulk-image-downloader`) Actor

Download images in bulk: search Google Images with every filter (size, color, license, date), or paste image URLs, or take the image fields of any dataset. Full-size files in a ZIP with metadata, no duplicates, resize and convert for AI datasets. Scheduled runs fetch only new images.

- **URL**: https://apify.com/s\_actors/bulk-image-downloader.md
- **Developed by:** [Superior Actors](https://apify.com/s_actors) (community)
- **Categories:** AI, Developer tools, Automation
- **Stats:** 6 total users, 5 monthly users, 100.0% runs succeeded, 1 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### Bulk Image Downloader & Google Images Scraper to ZIP

Download images in bulk: **search Google Images** with every filter (size, color, type, license, date), **paste a list of image URLs**, or **take the image fields of any dataset** (Amazon, Shopify, Instagram, real estate scrapers). You get the **full-size files in a ZIP** with a `metadata.csv`, each file also with a **direct link**, plus a table of image URL, size, rank and source page.

Built for **AI and computer-vision datasets**, product catalogs, mood boards and content work: **no duplicates**, **minimum resolution checked before download**, **convert WebP/AVIF to JPEG**, **resize to 512 or 1024 px**. Scheduled runs download **only new images**.

No browser, no API key, no proxy setup. Use it as a **Google Images API**: one call returns clean JSON.

#### Why this Actor

| | |
|---|---|
| 📦 **Files, not just links** | Other Google Images scrapers return URLs and leave the downloading to you. Here every image is downloaded, checked and saved: one ZIP (in ~200 MB parts) with `metadata.csv`, and each file in the key-value store with a link |
| 🔎 **Three sources, one run** | Google Images searches, a list of image URLs, or another Actor's dataset (image fields are found automatically) |
| 🧬 **No duplicates** | The same picture resized or recompressed on different sites is downloaded once (perceptual hash) |
| 📏 **Pay only for images you want** | Minimum width/height is checked against Google's data **before** downloading; failed downloads, duplicates and small images are free |
| 🤖 **AI-ready** | Convert to JPEG/PNG/WebP, shrink the longest side to 512/1024 px, keep ranks and source pages for labeling |
| 🛟 **Every image saved** | Downloads with a real browser fingerprint; when a site blocks the original (Instagram, Facebook), Google's thumbnail is saved instead and marked |
| 🔔 **Only new images** | Scheduled runs skip everything already downloaded, including the same picture at a new URL |

#### How it works

1. Add **Google Images searches** (e.g. `red fox`), and/or **Image URLs**, and/or **Datasets with image links**
2. Optional: Google filters (**Size**, **Color**, **Type**, **Usage rights**, **Time**, **Only from this site**), **Minimum width/height**, **Convert to**, **Resize**
3. Run. Click **ZIP archive** in the **Output** tab to download all images, or click the file link of a row to save one image

#### How to download the images

| Way | How |
|---|---|
| **All images at once** | **Output** tab → **ZIP archive**: `images.zip` downloads with every file in folders and `metadata.csv`. The run's status message has the same download link |
| **One image** | Click the file link (`fileUrl`) in the table row: the file is saved to your computer |
| **Every file separately** | **Storage** tab → **Key-value store**: each image is a record you can download |
| **From code** | `fileUrl` of each row is a direct download link; the ZIP is the record `images.zip` of the run's key-value store |

#### What you get

A cloud run in September 2026: 3 searches (`eiffel tower`, `golden retriever puppy`, `modern kitchen interior`), 200 images each.

| | |
|---|---|
| Original files saved | 559 of 600 (93%) |
| With the thumbnail fallback | 600 of 600 |
| Time | ~2 minutes for 600 files (199 MB ZIP) |
| Formats | JPEG 515, WebP 59, PNG 25, GIF 1 (originals as the site serves them) |
| Images per search | ~240 unique (Google shows ~250 per search; use more specific searches for more) |
| Image size from Google vs. the real file | equal in 331 of 331 checked |

#### Output

| Field | Description |
|---|---|
| `fileUrl` | Direct link to the saved file |
| `fileName`, `zipFile` | Path of the file inside the ZIP (`red-fox/005_a-z-animals.com.jpg`) and which ZIP part |
| `fileWidth`, `fileHeight`, `fileSizeKB`, `format` | The saved file (after conversion or resizing) |
| `imageUrl`, `pageUrl`, `domain` | The original image and the page it is on |
| `width`, `height` | Size of the original, from Google |
| `query`, `position` | The search and the rank in Google Images (or the URL / dataset the image came from) |
| `thumbnailUrl`, `imageId` | Google thumbnail and Google's image ID |
| `sourceField`, `sourceTitle` | Dataset source only: the field the link was in and the item's title or name |
| `downloadedFrom` | `original` or `thumbnail` |
| `downloaded`, `error` | Failed downloads are listed with the reason (free) |
| `hash` | Perceptual hash (dHash), for your own de-duplication |

`images.zip` contains the files in one folder per search and a `metadata.csv` with the same fields.

#### Google Images filters

| Filter | Options |
|---|---|
| Size | Large, Medium, Icon, Larger than 2/4/6/8/10/12/15/20/40/70 MP |
| Color | Full color, Black and white, Transparent, 12 colors |
| Type | Photo, Face, Clip art, Line drawing, Animated (GIF) |
| Usage rights | Creative Commons licenses, Commercial & other licenses |
| Time | Past 24 hours, week, month, year |
| Aspect ratio | Tall, Square, Wide, Panoramic |
| File type | JPG, PNG, GIF, WebP, SVG, BMP, ICO |
| Site | Only images from one website |
| Country, language, SafeSearch | Any Google country and language |

#### Images from another Actor's dataset

Put the dataset ID (or its link from Console) into **Datasets with image links**. Image links are found in fields such as `image`, `images`, `thumbnail`, `photos`, `logo`, or any link ending in `.jpg`, `.png`, `.webp`. To take specific fields only, list them in **Image fields**, dotted for nested ones: `images`, `product.mainImage`, `photos.url`.

Each row keeps the item's title and link, so product photos stay matched to their products.

#### Only new images: monitoring

1. Switch on **Only new images** and set a **Monitor name** (one per task)
2. Save as a task and schedule it (daily or weekly)
3. The first run downloads and remembers the images. Later runs download only images that are new, and skip pictures you already have even at a different URL

#### Pricing

Pay per image, no subscription:

| Event | Price |
|---|---|
| Run start | $0.002 |
| Image downloaded (file saved + row) | $0.002 |
| Image result without download (**Download image files** off) | $0.0005 |

Free: failed downloads, duplicates, images below your minimum size, already-seen images in monitoring.

Examples: 1,000 images downloaded = $2. A table of 1,000 Google Images results without files = $0.50.

#### Ready-made tasks

Open one, change the search, and run:

| Task | What you get |
|---|---|
| [Download Google Images in Bulk: Full-Size Files in a ZIP](https://apify.com/s_actors/bulk-image-downloader/examples/download-google-images-zip) | All images of a search as full-size files in one ZIP |
| [Bulk Image Downloader from a List of URLs](https://apify.com/s_actors/bulk-image-downloader/examples/bulk-image-downloader-from-url-list) | Your image links downloaded, converted to JPG |
| [AI Image Dataset Builder: Google Images to 512px JPEG](https://apify.com/s_actors/bulk-image-downloader/examples/ai-image-dataset-builder) | One folder per class, 512 px JPEG, no duplicates |

#### Input example

```json
{
    "queries": ["golden retriever", "labrador retriever", "german shepherd"],
    "maxImagesPerQuery": 200,
    "imageSize": "large",
    "minWidth": 800,
    "convertTo": "jpeg",
    "maxDimension": 512
}
```

#### Tips and limits

- **About 250 images per search** is Google's own limit. For more, split a broad search into specific ones (`red fox winter`, `red fox cub`, `red fox running`).
- **Minimum size** for Google searches is checked before downloading; for URL lists and datasets it is checked after (still free).
- **Usage rights**: Google's filter shows how an image is labeled; check the license on the source page before commercial use.
- **Big runs**: files are packed into ZIP parts of ~200 MB (`images.zip`, `images-2.zip`...). Turn off **Save each file** if you only need the ZIP.
- **Titles**: Google's lightweight results page this Actor reads has no image titles; use `domain` and `pageUrl`.

# Actor input Schema

## `queries` (type: `array`):

What to search on Google Images, one per line (e.g. "red fox", "modern kitchen interior"). Each search gives up to ~250 images (Google's own limit); use more specific searches for more.

## `maxImagesPerQuery` (type: `integer`):

How many images to take from each search, in Google's order.

## `imageUrls` (type: `array`):

Direct links to image files to download in bulk, one per line.

## `datasetIds` (type: `array`):

Datasets from runs of other Actors (Amazon, Shopify, Instagram, real estate...). Image links are found automatically in fields like image, images, thumbnail or photo.

## `imageFields` (type: `array`):

Optional: exact fields that hold image links, dotted for nested ones (e.g. "images", "product.mainImage", "photos.url"). Empty = detect automatically.

## `imageSize` (type: `string`):

Google Images size filter. "Larger than N MP" keeps only big originals (4 MP ≈ 2400×1600).

## `color` (type: `string`):

Google Images color filter; "Transparent background" finds PNG cut-outs.

## `imageType` (type: `string`):

Google Images type filter.

## `license` (type: `string`):

Google's usage rights filter. Always check the license on the source page before reuse.

## `time` (type: `string`):

When the image was found.

## `aspectRatio` (type: `string`):

Shape of the image.

## `fileType` (type: `string`):

Search only this file type (Google's filetype: operator).

## `site` (type: `string`):

Search images from one website only, e.g. "wikimedia.org" or "unsplash.com".

## `country` (type: `string`):

Two-letter country code of the Google results, e.g. us, gb, de, fr, jp.

## `language` (type: `string`):

Two-letter language code, e.g. en, de, es, ja.

## `safeSearch` (type: `boolean`):

Hide explicit results.

## `downloadImages` (type: `boolean`):

On: download every image and save the files. Off: only the table of image links (cheaper, Google Images searches only).

## `minWidth` (type: `integer`):

Skip smaller images. For Google searches this is checked before downloading, so small images cost nothing.

## `minHeight` (type: `integer`):

Skip images lower than this.

## `convertTo` (type: `string`):

Many sites serve WebP or AVIF; convert everything to one format for your tools or training pipeline.

## `maxDimension` (type: `integer`):

Shrink larger images so the longest side is at most this (e.g. 512 or 1024 for AI datasets). Empty = keep the original size.

## `removeDuplicates` (type: `boolean`):

Skip images that are the same picture as one already downloaded, even resized or recompressed (perceptual hash). Free.

## `thumbnailFallback` (type: `boolean`):

Some sites (Instagram, Facebook) block downloads: save Google's thumbnail instead. Marked in the "downloadedFrom" field.

## `createZip` (type: `boolean`):

Pack the files into images.zip (in parts of ~200 MB) with a metadata.csv inside.

## `saveFiles` (type: `boolean`):

Save every image to the key-value store with a direct link in the table ("fileUrl"), handy for pipelines and integrations.

## `maxImages` (type: `integer`):

Stop after this many images for the whole run. Empty = no limit.

## `maxFileSizeMB` (type: `integer`):

Skip files larger than this.

## `onlyNew` (type: `boolean`):

For scheduled runs. The first run downloads and remembers the images; later runs return only images that are new to these searches, URLs or datasets (and not duplicates of earlier ones). You pay only for the new ones.

## `monitorName` (type: `string`):

Separate monitors keep separate memory. Use a different name for each scheduled task.

## `maxConcurrency` (type: `integer`):

How many files download at the same time.

## Actor input object example

```json
{
  "queries": [
    "red fox"
  ],
  "maxImagesPerQuery": 20,
  "imageSize": "any",
  "color": "any",
  "imageType": "any",
  "license": "any",
  "time": "any",
  "aspectRatio": "any",
  "fileType": "any",
  "country": "us",
  "language": "en",
  "safeSearch": false,
  "downloadImages": true,
  "convertTo": "original",
  "removeDuplicates": true,
  "thumbnailFallback": true,
  "createZip": true,
  "saveFiles": true,
  "maxFileSizeMB": 25,
  "onlyNew": false,
  "monitorName": "default",
  "maxConcurrency": 16
}
```

# Actor output Schema

## `overview` (type: `string`):

One row per image: file link, size, source page.

## `results` (type: `string`):

Search results: image URL, size, rank, source page.

## `datasets` (type: `string`):

Images taken from other datasets, with the item they belong to.

## `zip` (type: `string`):

All files with metadata.csv. Large runs add images-2.zip, images-3.zip...

## `files` (type: `string`):

List of every stored file and ZIP part.

## `all` (type: `string`):

Every field of every row.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "red fox"
    ],
    "maxImagesPerQuery": 20
};

// Run the Actor and wait for it to finish
const run = await client.actor("s_actors/bulk-image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["red fox"],
    "maxImagesPerQuery": 20,
}

# Run the Actor and wait for it to finish
run = client.actor("s_actors/bulk-image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "red fox"
  ],
  "maxImagesPerQuery": 20
}' |
apify call s_actors/bulk-image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,s_actors/bulk-image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/aqqqdd5HtJcd7KmHX/builds/CczRpbhjHjy7sE0Zc/openapi.json
