# Bulk Image Downloader (Pages, URLs, ZIP) (`dima_kadirovich/bulk-image-downloader`) Actor

Download every image from web pages or a list of image URLs. Get public download links or ZIP files. Skips icons, tracking pixels and duplicates, keeps alt text. Low flat price, no extra platform costs.

- **URL**: https://apify.com/dima\_kadirovich/bulk-image-downloader.md
- **Developed by:** [Cronexa Data Tools](https://apify.com/dima_kadirovich) (community)
- **Categories:** Automation, E-commerce, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 images

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

### What does Bulk Image Downloader do?

Bulk Image Downloader **finds and downloads every image** from any list of web pages, or downloads a list of direct image URLs. You get:

- 🔗 **A public download link for every image**, a **ZIP archive**, or both
- 🖼️ **The largest version of every image**: it reads responsive `srcset` and `<picture>` sources, lazy-loaded images (`data-src`), CSS backgrounds, `og:image` previews, and full-size images behind thumbnails
- 🧹 **No junk**: tracking pixels, spacers, tiny icons and **duplicate images are skipped automatically**, and you don't pay for them
- 🏷️ **Useful metadata**: alt text, width × height, format, file size, source page and SHA-256 hash for every image, which is ready for AI/ML datasets
- 💸 **Simple, low price**: **$1 per 1,000 images + $2 per 1,000 pages**. Platform usage is included, so there are no extra compute or bandwidth fees.

### What can I use it for?

- **E-commerce**: download all product photos from supplier or competitor product pages
- **AI and machine learning datasets**: collect images with their alt text as captions
- **Website migration and backup**: save every image from your old site before a redesign
- **Marketing and research**: gather visuals from articles, landing pages and galleries
- **Bulk downloads from a URL list**: paste thousands of direct image links (or upload a CSV) and get them all in a ZIP

### How do I use it?

1. Add your URLs to **Start URLs**. These can be web pages, direct image links, or a mix. You can also upload a CSV/TXT file.
2. Choose the **Output**: *Files with download links*, *ZIP archive(s)*, or *Both*.
3. (Optional) Set filters such as **Minimum width** (for example `300` to keep only real photos, not logos) or **Allowed formats**.
4. Click **Start**. Open the **Images** tab for previews and links, or **Run summary** for your ZIP download links.

#### Example input

```json
{
  "startUrls": [
    { "url": "https://en.wikipedia.org/wiki/Hummingbird" },
    { "url": "https://upload.wikimedia.org/wikipedia/commons/thumb/b/b1/Eutoxeres_aquila_28748616.jpg/250px-Eutoxeres_aquila_28748616.jpg" }
  ],
  "outputMode": "filesAndZip",
  "minWidth": 300
}
```

### Output example

Each saved image is one row in the dataset:

```json
{
  "imageUrl": "https://assets.science.nasa.gov/dynamicimage/assets/science/missions/swift-observatory/misc--spacecraft-art/Swift_in_space_09.jpg?w=1024",
  "downloadUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/00003-nasa.gov-Swift_in_space_09.jpg?signature=...",
  "zipFile": "images-part1.zip",
  "fileName": "00003-nasa.gov-Swift_in_space_09.jpg",
  "pageUrl": "https://www.nasa.gov/image-of-the-day/",
  "altText": "NASA Highlights Lessons Learned From Swift Boost Mission",
  "foundIn": "img",
  "format": "jpeg",
  "contentType": "image/jpeg",
  "width": 1024,
  "height": 542,
  "fileSizeBytes": 112653,
  "sha256": "7e384316d3720ae68fae9598df78a00a590a9576f7662c0a8fc0b0dfd40de73e"
}
```

The **Run summary** (`OUTPUT`) lists your **ZIP download links**, per-page results, and how many images were skipped and why (duplicate, too small, not an image, and so on).

### How much does it cost?

| What | Price |
|---|---|
| Image saved | **$0.001** ($1 per 1,000) |
| Web page processed | **$0.002** ($2 per 1,000) |
| Platform usage (compute, bandwidth) | **Included** |

Direct image URLs are charged only as images, not as pages. **Skipped images (duplicates, pixels, filtered out) are free.** Example: a product page with 10 photos costs $0.012.

### Input options

| Field | Description |
|---|---|
| `startUrls` | Web pages and/or direct image URLs |
| `outputMode` | `files`, `zip` (50 MB parts), or `filesAndZip` |
| `minFileSizeKb` | Skip files smaller than this (default 3 KB, which removes pixels and icons) |
| `minWidth` / `minHeight` | Skip images smaller than this, in pixels |
| `allowedFormats` | Keep only these formats (JPEG, PNG, WebP, GIF, AVIF, SVG, …) |
| `deduplicate` | Skip identical images even from different URLs (default on) |
| `maxFileSizeMb` | Skip files larger than this (default 20 MB) |
| `maxImagesPerPage` / `maxImages` | Limits per page and per run |
| `includeSrcset`, `includeMetaImages`, `includeCssBackgrounds`, `includeLinkedImages` | Where to look for images |
| `proxyConfiguration` | Use a proxy if a website blocks downloads |

### Limitations

- The Actor reads the page HTML. Images that appear **only after JavaScript runs** (for example infinite-scroll galleries) may not be found. Most sites include their images in the HTML, lazy-loaded images included.
- Some websites block automated downloads. Try the **Proxy** option if you see `http403` in the run summary.
- Download links are stored in your Apify storage and follow your plan's data retention period. Download your ZIPs if you need them long-term.

### Is it legal?

The Actor downloads publicly available images from pages you choose. **Images are usually protected by copyright.** Make sure you have the right to use them for your purpose, and respect each website's terms.

### Questions or problems?

Open an issue in the **Issues** tab, and it will be answered quickly.

### Use it from AI assistants (Claude, ChatGPT, Cursor)

AI agents can run this Actor as a tool through the [Apify MCP server](https://mcp.apify.com). Add this to your MCP client (Claude Desktop, Claude Code, Cursor, VS Code…) and sign in with Apify in the browser when asked:

```json
{
  "mcpServers": {
    "bulk-image-downloader": { "url": "https://mcp.apify.com?tools=dima_kadirovich/bulk-image-downloader" }
  }
}
```

Then just ask, for example:

- *"Download all product photos from these 5 product pages as a ZIP and give me the link."*
- *"Collect every image wider than 800px from this article, with its alt text."*

The agent fills in the input, runs the Actor and reads the results. You pay the same per-result price.

### More tools from Dima Data Tools

- [Medium Articles Scraper & Monitor](https://apify.com/dima_kadirovich/medium-articles-scraper): Medium articles by tag, author, or publication as clean Markdown, with "only new" monitoring
- [Bulk Image Downloader](https://apify.com/dima_kadirovich/bulk-image-downloader): every image from any web page, as download links or ZIP, with duplicates and icons removed
- [Website SEO Audit & Broken Link Checker](https://apify.com/dima_kadirovich/website-seo-audit): crawl a site, score every page 0–100, find broken links, and get a shareable HTML report
- [Website to Markdown Crawler for AI](https://apify.com/dima_kadirovich/website-to-markdown): any website as clean main-content Markdown, with RAG chunks, llms.txt, and a cheap "only changed pages" refresh mode
- [Website Tech Stack & Domain Lookup](https://apify.com/dima_kadirovich/tech-stack-domain-lookup): technologies, email provider, SPF/DMARC, SaaS tools, SSL expiry and WHOIS for any list of domains

# Actor input Schema

## `startUrls` (type: `array`):

Web pages to collect images from, and/or direct image links. You can mix both. Upload a CSV/TXT file of URLs with the 'Bulk edit' or file upload option.

## `outputMode` (type: `string`):

Files: each image gets its own public download link. ZIP: images are packed into ZIP file(s) (up to 50 MB per part). Both: links and ZIP.

## `zipFileName` (type: `string`):

Base name for ZIP files (used only in ZIP modes).

## `minFileSizeKb` (type: `integer`):

Skip files smaller than this. The default (3 KB) removes tracking pixels, spacers and most tiny icons. Set 0 to keep everything.

## `minWidth` (type: `integer`):

Skip images narrower than this. Example: 300 keeps only real photos, not logos and buttons.

## `minHeight` (type: `integer`):

Skip images shorter than this.

## `allowedFormats` (type: `array`):

Keep only these formats. Leave empty for all formats.

## `deduplicate` (type: `boolean`):

Skip images with identical content, even if they come from different URLs. You don't pay for skipped duplicates.

## `maxFileSizeMb` (type: `integer`):

Skip files larger than this.

## `maxImagesPerPage` (type: `integer`):

Upper limit of images taken from one page.

## `maxImages` (type: `integer`):

Stop after saving this many images.

## `includeSrcset` (type: `boolean`):

When a page offers several sizes of the same image, download the largest one.

## `includeMetaImages` (type: `boolean`):

Include Open Graph / Twitter preview images and site icons.

## `includeCssBackgrounds` (type: `boolean`):

Include images used as CSS backgrounds in the page.

## `includeLinkedImages` (type: `boolean`):

Include images that the page links to (for example, a thumbnail that opens the full-size photo).

## `proxyConfiguration` (type: `object`):

Only needed if a website blocks the downloads.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Hummingbird"
    }
  ],
  "outputMode": "files",
  "zipFileName": "images",
  "minFileSizeKb": 3,
  "deduplicate": true,
  "maxFileSizeMb": 20,
  "maxImagesPerPage": 1000,
  "maxImages": 10000,
  "includeSrcset": true,
  "includeMetaImages": true,
  "includeCssBackgrounds": true,
  "includeLinkedImages": true,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `images` (type: `string`):

One row per saved image with its download link, original URL, alt text, size and dimensions.

## `files` (type: `string`):

All downloaded image files and ZIP archives.

## `summary` (type: `string`):

Images saved, ZIP download links, per-page results, and why images were skipped.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://en.wikipedia.org/wiki/Hummingbird"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("dima_kadirovich/bulk-image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://en.wikipedia.org/wiki/Hummingbird" }] }

# Run the Actor and wait for it to finish
run = client.actor("dima_kadirovich/bulk-image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://en.wikipedia.org/wiki/Hummingbird"
    }
  ]
}' |
apify call dima_kadirovich/bulk-image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dima_kadirovich/bulk-image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/3fM2DldGvg29rG2hu/builds/1c8Vb1gs8JfcAyLJZ/openapi.json
