# Bulk Image Downloader API: Images to ZIP, No Login (`conserving_celerytop/bulk-image-downloader`) Actor

Download images in bulk from image links or web pages into one ZIP. Filter by type, width, height and file size, skip duplicates, and get a row per image with link, format, size, dimensions and alt text. $1 per 1,000 images saved. Skipped files are free. No login.

- **URL**: https://apify.com/conserving_celerytop/bulk-image-downloader.md
- **Developed by:** [Don Mangu](https://apify.com/conserving_celerytop) (community)
- **Stats:** 1 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.60 / 1,000 image downloads

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Image Downloader: URLs and Pages to ZIP

Bulk Image Downloader saves images from a list of image links or web pages into one ZIP file. Paste direct image links, page links, or both. For each page, the Actor reads the HTML and downloads the images on it. You get a ZIP with every image, an index.csv inside it, and one dataset row per image with its link, file type, width, height, file size and alt text.

Use it to collect product photos for a catalog, archive the images of a blog or documentation site, gather reference pictures for a design project, or build an image set from pages you are allowed to copy.

### What this bulk image downloader does

- **Image links and page links in one list.** A link that returns an image is saved as it is. A link that returns a web page is read, and the images on it are downloaded: img tags, the largest size in srcset, picture sources, lazy-load attributes (data-src and similar), Open Graph and Twitter card images, links to image files and, if you turn it on, CSS background images.
- **Filters.** Keep only the types you want (JPEG, PNG, WebP, AVIF, SVG, GIF, BMP, TIFF, HEIC, ICO), and set a minimum width, minimum height, minimum file size and maximum file size. The type is read from the file itself, so a JPEG served as "application/octet-stream" is still found.
- **No duplicates.** When two links return the same bytes, the file is saved once.
- **One ZIP.** All images go into `images.zip` in the run's key-value store. Choose one folder, one folder per page, or one folder per website. Very large runs are split into `images-part-2.zip` and so on.
- **Polite by default.** robots.txt is read for every website, at most 2 downloads run against the same website at once, and a website that answers 401, 403 or 429 is not contacted again in that run.

### How to download images in bulk

1. Add your links to **Image or page URLs**. You can type them, paste them, or upload a text or CSV file of links.
2. Set **Maximum images** (start with 20 to try it).
3. Optional: choose **Image types** and set **Minimum width** to 200 or more to leave out icons and tracking pixels.
4. Click **Start**. When the run ends, open the **Output** tab and click the ZIP link, or open the dataset to see every image with its status.

### How much does it cost to download images?

You pay per image saved in the ZIP: **$1.00 per 1,000 images** ($0.001 per image). Files larger than 1 MB count once per started MB, so a 2.5 MB photo counts as 3. Images left out by your filters, duplicates, broken links and pages without images are free.

Example: 40 product pages with about 25 photos each, filtered to JPEG and WebP of at least 400 px wide, gives about 800 images under 1 MB each. That run costs about $0.80. You can set a maximum cost per run in the run options, and the Actor stops cleanly when it is reached.

### Input example

```json
{
    "startUrls": [
        { "url": "https://www.wikipedia.org/" },
        { "url": "https://example.com/photos/sunset.jpg" }
    ],
    "maxImages": 200,
    "imageTypes": ["jpeg", "png", "webp"],
    "minWidth": 300,
    "zipFolders": "byPage"
}
```

### Output example

The ZIP link is in the Output tab and in the `OUTPUT` record. Each dataset row looks like this:

```json
{
    "imageUrl": "https://example.com/photos/sunset.jpg",
    "pageUrl": "https://example.com/gallery",
    "status": "ok",
    "fileName": "example.com-gallery/0007-sunset.jpg",
    "zipFile": "images.zip",
    "zipUrl": "https://api.apify.com/v2/key-value-stores/.../records/images.zip",
    "format": "jpeg",
    "width": 1600,
    "height": 1067,
    "sizeBytes": 245112,
    "altText": "Sunset over the harbor",
    "billedUnits": 1,
    "charged": true,
    "error": null
}
```

Rows that were not saved have a `status` such as `filtered_dimensions`, `duplicate`, `too_large`, `robots_disallowed` or `not_found`, and `error` says why in plain words.

### Related tools

To resize, compress or convert the downloaded images to WebP or AVIF, pass the ZIP contents or the image links to an image converter Actor. To list every page of a website first, run a website crawler and feed its page links into this Actor.

### FAQ

**Is it legal to download images from websites?**
Downloading public images is usually allowed for personal use, research or with the owner's permission, but images are often protected by copyright. You are responsible for having the right to use the images you download. The Actor only fetches public pages without logging in, and it follows each website's robots.txt. Photos of people and their alt text can be personal data under laws such as the GDPR, so only collect them when you have a lawful reason.

**Why did a page return no images?**
Some websites build their pages with JavaScript, so the images are not in the HTML. The row for that page has the status `page_no_images`. In that case, add the image links directly.

**Why do some images have the status `host_stopped`?**
The website answered 401, 403 or 429 (access refused or too many requests), so the Actor stopped contacting it for the rest of the run. Try again later with fewer parallel downloads.

**What are the limits?**
Up to 100,000 images per run and up to 100 MB per file. The ZIP is held in memory until it is saved, so for very large runs give the run more memory or lower **Maximum ZIP size**.

**Can I get each image as its own file?**
Yes. Turn on **Also save each image as its own file** and each dataset row gets a `fileUrl` with a direct download link.

# Actor input Schema

## `startUrls` (type: `array`):

Direct image links (.jpg, .png, .webp and so on) and web page links, mixed in any order. For a web page, the images on that page are downloaded: img tags, srcset (largest size), picture sources, lazy-load attributes, Open Graph images and links to image files. You can also upload a text or CSV file of links.

## `maxImages` (type: `integer`):

Stop after this many images are saved in the ZIP. Images left out by the filters do not count.

## `imageTypes` (type: `array`):

Keep only these file types. The type is read from the file itself, not from the link or the server's label.

## `minWidth` (type: `integer`):

Keep images at least this wide. Use 200 or more to leave out icons, spacers and tracking pixels. 0 keeps all widths.

## `minHeight` (type: `integer`):

Keep images at least this tall. 0 keeps all heights.

## `minFileSizeKb` (type: `integer`):

Keep files of at least this size. 0 keeps all sizes.

## `maxFileSizeMb` (type: `integer`):

Keep files up to this size. Larger files are stopped during download and are not charged.

## `zipFolders` (type: `string`):

One flat folder, or one folder per web page (named after the page address).

## `saveIndividualFiles` (type: `boolean`):

Also store every image as a separate record in the key-value store, with its own download link in the dataset (fileUrl). Useful for feeding images to other tools one by one.

## `skipDuplicates` (type: `boolean`):

Save a file only once when several links return exactly the same bytes (compared by SHA-256). Duplicates are listed in the dataset and not charged.

## `includeLinkedImages` (type: `boolean`):

On web pages, also download links (a href) that point to image files, such as full-size versions behind thumbnails.

## `includeCssBackgrounds` (type: `boolean`):

On web pages, also download images set as CSS backgrounds in style attributes and style blocks of the page.

## `maxZipSizeMb` (type: `integer`):

When the ZIP reaches this size, it is saved and a new part is started (images-part-2.zip and so on). The Actor lowers it automatically to fit the run's memory.

## `maxConcurrency` (type: `integer`):

Downloads at the same time. At most 2 go to the same website at once.

## `requestTimeoutSecs` (type: `integer`):

Time allowed for one page or image to download completely.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Cat"
    }
  ],
  "maxImages": 5,
  "imageTypes": [
    "jpeg",
    "png",
    "gif",
    "webp",
    "avif",
    "svg",
    "bmp",
    "tiff",
    "heic"
  ],
  "minWidth": 0,
  "minHeight": 0,
  "minFileSizeKb": 0,
  "maxFileSizeMb": 20,
  "zipFolders": "flat",
  "saveIndividualFiles": false,
  "skipDuplicates": true,
  "includeLinkedImages": true,
  "includeCssBackgrounds": false,
  "maxZipSizeMb": 250,
  "maxConcurrency": 5,
  "requestTimeoutSecs": 60
}
```

# Actor output Schema

## `zip` (type: `string`):

images.zip: all saved images plus index.csv (file name, image URL, page URL, width, height, size, format, alt text). Very large runs add images-part-2.zip and so on, listed in the summary.

## `images` (type: `string`):

imageUrl, pageUrl, status, fileName, zipUrl, format, width, height, sizeBytes and altText for every image, including the ones left out by the filters.

## `summary` (type: `string`):

JSON with the ZIP files (key, link, image count, bytes), images saved, pages read and counts by status.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://en.wikipedia.org/wiki/Cat"
        }
    ],
    "maxImages": 5
};

// Run the Actor and wait for it to finish
const run = await client.actor("conserving_celerytop/bulk-image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": [{ "url": "https://en.wikipedia.org/wiki/Cat" }],
    "maxImages": 5,
}

# Run the Actor and wait for it to finish
run = client.actor("conserving_celerytop/bulk-image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Cat"
    }
  ],
  "maxImages": 5
}' |
apify call conserving_celerytop/bulk-image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,conserving_celerytop/bulk-image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/AC40ZEOQCX5Fll73P/builds/0ynPi4Zp8bdc15hnk/openapi.json
