# Bulk Image Downloader - Download All Images from URLs & Pages (`gazidev/image-downloader`) Actor

Download all images from web pages or a list of image URLs in bulk. Extracts img, srcset (largest), og:image and linked images, filters by size and format, removes SHA-1 duplicates, converts/resizes to JPG/PNG/WebP and bundles a ZIP. Public download links for every file.

- **URL**: https://apify.com/gazidev/image-downloader.md
- **Developed by:** [Cemal Atakli](https://apify.com/gazidev) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $15.00 / 1,000 image saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Image Downloader – Download All Images from URLs & Pages

**Download all images from any web page, or from a list of image URLs, in bulk.** Paste page URLs and/or direct image links. The Actor finds every image on each page, downloads the **largest available version**, skips icons and duplicates, can **convert and resize** to JPG, PNG or WebP, and gives you a **public download link for every file**, plus an optional **ZIP of all images**.

- **Finds every image on a page**: `<img>`, the largest `srcset` / `<picture>` source, lazy-loaded images (`data-src`, `data-srcset`…), `og:image` / `twitter:image`, linked full-size images (`<a href="photo.jpg">`) and, optionally, CSS background images
- **Direct image URLs work too**: mix pages and image links in one list; the type is detected automatically
- **Smart filters**: minimum width and height (no more 1×1 tracking pixels and icons), file formats, file size
- **SHA-1 deduplication**: the same picture under different URLs is saved once, and duplicates are never charged
- **Convert and resize**: WebP/AVIF → JPG for compatibility, or everything → WebP to save space, with max width/height and quality
- **ZIP bundle**: one `images.zip` with everything, ready to download
- **Custom file names**: `{index}_{name}`, `{page_host}-{sha1_8}`, `{name}-{width}x{height}`…
- **Cheap and fast**: HTTP-only (no browser), parallel downloads, **$15 per 1,000 images**, about 4.7× cheaper than the most popular image downloader in the Store

### What can I use it for?

- **Ecommerce**: download all product photos from a supplier's catalog pages or from a product-scraper dataset
- **Machine learning datasets**: collect and deduplicate images from galleries, with consistent formats and sizes
- **Content migration and backups**: move every image of a website to a new CMS or archive it
- **Marketing and SEO audits**: list every image on a page with its size, format and weight (find huge unoptimized files)
- **Research and journalism**: archive the images from news pages or public-domain collections (e.g. Wikimedia Commons)
- **Real estate, travel and classifieds**: save listing photos in bulk from URLs you already have

### How to use it

1. Paste one or more **page URLs** (e.g. a Wikipedia article, a product page, a gallery) and/or **direct image URLs**.
2. Optional: set a minimum size, choose formats, turn on **Convert to** / **Resize**, or **Create a ZIP**.
3. Click **Start**. Each image appears in the **Output** tab with a preview and a download link. The ZIP is under **Storage → Key-value store → images.zip**.

### Input

| Field | Description |
|---|---|
| `startUrls` | Web pages and/or direct image URLs (auto-detected) |
| `bulkText` | Paste many URLs at once (one per line, or comma/space separated) |
| `sourceDatasetId`, `sourceDatasetField` | **Chain with another Actor**: read image/page URLs from an Apify dataset (e.g. the `imageUrl` field from a product scraper) |
| `maxImages`, `maxImagesPerPage` | Limits (0 = no limit) |
| `minWidth`, `minHeight` | Skip small images (default 100×100 px) |
| `formats` | Keep only these formats: jpg, png, gif, webp, avif, svg, bmp, tiff, ico, heic (detected from the file content) |
| `minFileSizeKb`, `maxFileSizeMb` | File size filters. Downloads are streamed and aborted above the maximum (default 25 MB) |
| `dedupe` | Skip byte-identical files (SHA-1), on by default |
| `sameDomainImagesOnly` | Ignore third-party images (ads, trackers) |
| `preferLargestSrcset`, `includeMetaImages`, `includeLinkedImages`, `includeCssBackgrounds` | What to extract from pages |
| `convertTo`, `resizeMaxWidth`, `resizeMaxHeight`, `quality` | Convert to JPG/PNG/WebP and/or downscale (aspect ratio kept) |
| `filenameTemplate` | `{index}`, `{name}`, `{sha1}`, `{sha1_8}`, `{host}`, `{page_host}`, `{width}`, `{height}`, `{format}` |
| `createZip` | Bundle all images into `images.zip` |
| `includeSkipped` | Also list filtered images and duplicates in the dataset (free) |
| `respectRobotsTxt`, `maxConcurrency`, `requestTimeoutSecs`, `maxRetries`, `userAgent`, `proxyConfiguration` | Advanced settings |

```json
{
  "startUrls": [
    "https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States",
    "https://www.gstatic.com/webp/gallery/1.webp"
  ],
  "minWidth": 200,
  "minHeight": 200,
  "convertTo": "jpg",
  "resizeMaxWidth": 1600,
  "createZip": true,
  "filenameTemplate": "{page_host}-{index}-{name}"
}
```

### Output

One dataset row per image (export it as JSON, CSV, Excel or HTML, or read it through the API). The files themselves are in the run's key-value store, and `storedUrl` is a direct download link.

```json
{
  "pageUrl": "https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States",
  "imageUrl": "https://thumb.wikimedia.org/wikipedia/commons/thumb/0/0b/RNS_Yellowstone_13399u.jpg/960px-RNS_Yellowstone_13399u.jpg?utm_source=en.wikipedia.org&utm_campaign=index&utm_content=thumbnail",
  "storedUrl": "https://api.apify.com/v2/key-value-stores/aBcD1234EfGh5678/records/00001_960px-RNS_Yellowstone_13399u.jpg",
  "key": "00001_960px-RNS_Yellowstone_13399u.jpg",
  "fileName": "00001_960px-RNS_Yellowstone_13399u.jpg",
  "width": 960,
  "height": 1302,
  "bytes": 279214,
  "format": "jpg",
  "sha1": "…",
  "originalFormat": "jpg",
  "originalWidth": 960,
  "originalHeight": 1302,
  "converted": false,
  "alt": null,
  "foundIn": "srcset",
  "status": "saved",
  "error": null,
  "downloadedAt": "2026-09-30T10:00:00+00:00"
}
```

`status` is `saved`, `failed` (with `error`, e.g. `HTTP 404`, `File too large`, `Page disallowed by robots.txt`) or, with `includeSkipped`, `skipped-small`, `skipped-format`, `skipped-not-image` and `duplicate`. A run summary (counts and the ZIP link) is saved as the `OUTPUT` record.

### Pricing

Pay only for images that are actually saved. Failed downloads, filtered images and duplicates are free.

| Event | Price |
|---|---|
| Image saved (conversion/resize included) | **$0.015** ($15 per 1,000) |
| ZIP bundle (optional, once per run) | $0.01 |
| Actor start | $0.0005 |

| Actor | Price per 1,000 images |
|---|---|
| **This Actor** | **$15** |
| onescales/bulk-image-downloader | $70 |
| hipersoft image downloader | $56 |

*Competitor prices are from the Apify Store in September 2026 and may change.* Set **Maximum cost per run** in the run options, and the Actor stops gracefully when it is reached (the ZIP is still created).

### FAQ

**Does it download full-resolution images?** It downloads the largest version the page offers (the biggest `srcset` / `<picture>` candidate, lazy-load originals, and linked full-size files). It cannot get sizes that the site never publishes.

**Does it work on JavaScript-heavy sites?** It reads the HTML the server sends (no browser), which covers most sites, CMSs, shops, blogs and wikis. Images that exist only after client-side rendering, and sites with strong bot protection (e.g. Unsplash, Pexels), may return few images or a 403. For those, pass the image URLs directly (for example from another scraper's dataset).

**How long are files kept?** They live in the run's default key-value store and follow your Apify plan's data retention (unnamed storages are deleted after the retention period). Download the ZIP or copy the files to a named store if you need them longer.

**Can I feed it URLs from another Actor?** Yes. Set `sourceDatasetId` to the other run's dataset ID and `sourceDatasetField` to the field holding the URLs (strings or arrays of strings).

**Does it respect robots.txt?** Yes. Page URLs that robots.txt disallows are skipped (you can turn this off for your own sites). Direct image URLs are downloaded as given.

**Copyright?** You are responsible for making sure you have the right to download and use the images (your own sites, public-domain or licensed content, fair use). This tool does not grant any rights to the images it downloads. Respect the websites' terms of service.

### Use with AI agents / Apify MCP

- **Apify MCP server**: add `https://mcp.apify.com/?actors=gazidev/image-downloader` to Claude Desktop, Cursor, VS Code or any MCP client. An agent can then run requests like: "Download all images larger than 500px from this page and give me a ZIP."
- **API**: `POST https://api.apify.com/v2/acts/gazidev~image-downloader/run-sync-get-dataset-items?token=...` with `{"startUrls": ["https://example.com/gallery"], "maxImages": 20}` returns the image rows (with `storedUrl` download links) synchronously for small jobs.
- Tip for agents: filter rows on `status == "saved"` and use `storedUrl`. `width`, `height`, `bytes` and `sha1` let you choose or deduplicate images without downloading them again.

# Actor input Schema

## `startUrls` (type: `array`):

Web pages to extract all images from, and/or direct image URLs (.jpg, .png, .webp, ...). The type is detected automatically: HTML pages are scanned for images, image URLs are downloaded directly.

## `bulkText` (type: `string`):

Paste many URLs at once: one per line, or comma/space separated (e.g. a column copied from a spreadsheet).

## `sourceDatasetId` (type: `string`):

Chain with another Actor: ID of an Apify dataset (e.g. from a product or Instagram scraper). Image/page URLs are read from the fields below.

## `sourceDatasetField` (type: `string`):

Comma-separated field names holding URLs (string or array of strings).

## `maxImages` (type: `integer`):

Stop after saving this many images in total (0 = no limit). Handy for cheap test runs.

## `maxImagesPerPage` (type: `integer`):

Only take the first N image URLs found on each page (0 = all).

## `minWidth` (type: `integer`):

Skip images narrower than this, e.g. icons, spacers and tracking pixels. 0 = no limit.

## `minHeight` (type: `integer`):

Skip images shorter than this. 0 = no limit.

## `formats` (type: `array`):

Only save these formats (detected from the file content, not the extension). Empty = all.

## `minFileSizeKb` (type: `integer`):

Skip files smaller than this. 0 = no limit.

## `maxFileSizeMb` (type: `integer`):

Downloads are streamed and aborted as soon as they exceed this size.

## `dedupe` (type: `boolean`):

Skip images whose content is byte-identical to an image already saved in this run (same file under different URLs). Duplicates are not charged.

## `sameDomainImagesOnly` (type: `boolean`):

Ignore images served from other domains (ads, trackers, third-party widgets). Subdomains and CDNs on other domains are excluded too.

## `preferLargestSrcset` (type: `boolean`):

Responsive images list several sizes; download the biggest one instead of the small default `src`.

## `includeMetaImages` (type: `boolean`):

Also download the social preview image(s) declared in the page's meta tags.

## `includeLinkedImages` (type: `boolean`):

Also download images that are linked (e.g. full-size versions behind thumbnails).

## `includeCssBackgrounds` (type: `boolean`):

Also extract `background-image: url(...)` from inline styles and <style> blocks.

## `convertTo` (type: `string`):

Convert every raster image to one format (e.g. WebP/AVIF -> JPG for compatibility). SVGs are kept as-is. Transparent areas become white in JPG.

## `resizeMaxWidth` (type: `integer`):

Downscale larger images to fit this width, keeping the aspect ratio. 0 = no resizing.

## `resizeMaxHeight` (type: `integer`):

Downscale larger images to fit this height, keeping the aspect ratio. 0 = no resizing.

## `quality` (type: `integer`):

Quality used when converting or resizing (1-100).

## `filenameTemplate` (type: `string`):

Placeholders: {index} (00001...), {name} (original file name), {sha1}, {sha1\_8}, {host} (image host), {page\_host}, {width}, {height}, {format}. The extension is added automatically; unsupported characters become `_`.

## `createZip` (type: `boolean`):

Bundle all saved images into `images.zip` in the key-value store (one download link). Charged once per run.

## `includeSkipped` (type: `boolean`):

Also add rows for images that were filtered out (too small, wrong format) or duplicates. They are never charged. Failed downloads are always listed.

## `respectRobotsTxt` (type: `boolean`):

Skip pages that the site's robots.txt disallows for crawlers.

## `maxConcurrency` (type: `integer`):

How many downloads run in parallel.

## `requestTimeoutSecs` (type: `integer`):

Per-request timeout.

## `maxRetries` (type: `integer`):

Retries for timeouts, connection errors, 429 and 5xx responses (with backoff).

## `userAgent` (type: `string`):

Custom User-Agent header. Leave empty for the default, which identifies this Actor (required by some sites such as Wikimedia).

## `proxyConfiguration` (type: `object`):

Optional. Not needed for most sites.

## Actor input object example

```json
{
  "startUrls": [
    "https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States"
  ],
  "sourceDatasetField": "imageUrl,url",
  "maxImages": 10,
  "maxImagesPerPage": 0,
  "minWidth": 100,
  "minHeight": 100,
  "formats": [],
  "minFileSizeKb": 0,
  "maxFileSizeMb": 25,
  "dedupe": true,
  "sameDomainImagesOnly": false,
  "preferLargestSrcset": true,
  "includeMetaImages": true,
  "includeLinkedImages": true,
  "includeCssBackgrounds": false,
  "convertTo": "original",
  "resizeMaxWidth": 0,
  "resizeMaxHeight": 0,
  "quality": 85,
  "filenameTemplate": "{index}_{name}",
  "createZip": false,
  "includeSkipped": false,
  "respectRobotsTxt": true,
  "maxConcurrency": 10,
  "requestTimeoutSecs": 30,
  "maxRetries": 2,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `images` (type: `string`):

No description

## `results` (type: `string`):

No description

## `zip` (type: `string`):

No description

## `files` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        "https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States"
    ],
    "maxImages": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("gazidev/image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "startUrls": ["https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States"],
    "maxImages": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("gazidev/image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    "https://en.wikipedia.org/wiki/List_of_national_parks_of_the_United_States"
  ],
  "maxImages": 10
}' |
apify call gazidev/image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,gazidev/image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Ba5clHd9PzfGhjWBR/builds/kY24oD949O5zKCvXQ/openapi.json
