# Bulk Image Downloader: Image URLs and Page Images to Files (`pistachio_implementation/bulk-image-downloader`) Actor

Download images in bulk from a list of image URLs or from web pages you name. Saves each file with a download link, width, height, format, size and SHA256; skips duplicates, tiny icons and broken links. $2 per 1,000 images.

- **URL**: https://apify.com/pistachio\_implementation/bulk-image-downloader.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** E-commerce, Automation, Developer tools
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$2.00 / 1,000 image saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Bulk Image Downloader: Image URLs and Page Images to Files

Turn a list of image links, or a list of web pages, into saved image files. Each saved image gets a row with a download link, the original URL, the page it came from, its alt text, format, width, height, file size and a SHA256 fingerprint. Duplicates, icons below your size limit, broken links and files that are not images are skipped and cost nothing.

**Price: $2 per 1,000 images saved.** Nothing else is charged.

### Two ways to use it

1. **Image URLs.** You already have the links (from a product feed, a spreadsheet, another scraper's output). The actor downloads each one and saves it.
2. **Page URLs.** You name the pages. The actor reads each page's HTML and collects every `img` (taking the largest size from `srcset`, and lazy loading attributes such as `data-src`), `picture` sources, and the share image from `og:image` and `twitter:image`. Then it downloads them.

You can mix both in one run.

### Input

```json
{
  "imageUrls": ["https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"],
  "pageUrls": ["https://books.toscrape.com/"],
  "maxImagesPerPage": 50,
  "minWidth": 100,
  "allowedTypes": ["jpg", "png", "webp"],
  "keyValueStoreName": "my-product-images"
}
```

| Option | What it does |
|---|---|
| `minWidth`, `minHeight` | Skip small images such as icons, logos and tracking pixels (free) |
| `allowedTypes` | Keep only some formats: jpg, png, webp, gif, svg, avif, ico, bmp, tiff, heic |
| `maxFileSizeMb` | Skip files above this size (default 25 MB) |
| `skipDuplicates` | Save identical files once, even from different URLs (on by default) |
| `keyValueStoreName` | Save into a named store that is kept after the run's own storage expires |
| `maxImages`, `maxImagesPerPage` | Caps for the run and for each page |

### Output example

```json
{
  "imageUrl": "https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg",
  "pageUrl": "https://books.toscrape.com/",
  "alt": "A Light in the Attic",
  "foundIn": "img",
  "success": true,
  "fileName": "2cdad67c44b002e7ead0cc35693c0e8b.jpg",
  "storeKey": "4752d0ef411f-2cdad67c44b002e7ead0cc35693c0e8b.jpg",
  "fileUrl": "https://api.apify.com/v2/key-value-stores/<storeId>/records/4752d0ef411f-2cdad67c44b002e7ead0cc35693c0e8b.jpg?signature=<signature>",
  "format": "jpg",
  "contentType": "image/jpeg",
  "bytes": 9876,
  "width": 125,
  "height": 155,
  "sha256": "4752d0ef411f..."
}
```

Failed or skipped rows have `success: false` and a plain `error`, for example `The server answered HTTP 404`, `Not an image (content type text/html)` or `Same file already saved in this run`.

**Getting the files.** Each `fileUrl` is a signed link that downloads the image directly, with no API token needed, so you can pass it to a spreadsheet, a CMS or an AI agent. In Apify Console, open the run's Storage tab, Key value store, to browse or download them all. The Apify API and client libraries can list and fetch every record of the store for a bulk export.

### Pricing

Pay per event: **$0.002 per image saved ($2 per 1,000)**. No start fee and no platform usage on top. Broken links, non images, files over the size limit, images below your minimum size, formats you excluded, duplicates and robots.txt skips are all free. You can cap the spend of any run with the maximum charge setting in Apify.

### Limits

- Page mode reads the HTML the server sends. Images that a page adds later with JavaScript, CSS background images and images inside iframes are not collected.
- Some sites refuse downloads from cloud servers or require a login. Those files come back as free error rows. The actor does not try to get around blocks, logins or captchas.
- The actor identifies itself honestly as `ApifyImageDownloader` and respects robots.txt by default, including rules written for Apify crawlers. Turn this off only for sites you own or may download from.
- Requests to the same site are spaced out (4 per second for files, 1 per second for pages); many sites are fetched in parallel.
- Up to 20,000 images per run and 100 MB per file.
- Files in a run's default store follow your Apify data retention. Use `keyValueStoreName` to keep them.

### FAQ

**Do I have the right to download these images?** That is up to you and the image owner. The actor is a downloader like any browser's "save image". Use it for your own images, images you have a license for, or uses the law allows.

**Why is width or height empty for some images?** A few formats (AVIF, HEIC, some SVGs) do not state the size in a way the actor reads without decoding the whole file. The file is still saved.

**Can I get a ZIP?** Not yet; files are stored one by one so each has its own link. Fetch them through the API or the Console storage view.

**Can an AI agent use it?** Yes. Pass `imageUrls` or `pageUrls`, then read `fileUrl` from each row.

# Actor input Schema

## `imageUrls` (type: `array`):

Direct links to image files, one per line.

## `pageUrls` (type: `array`):

Web pages whose images you want. The actor reads the page HTML and collects img tags, srcset (largest size), picture sources and the Open Graph and Twitter share images.

## `maxImagesPerPage` (type: `integer`):

Stop collecting after this many images on one page.

## `maxImages` (type: `integer`):

Stop after saving this many images (up to 20,000).

## `minWidth` (type: `integer`):

Skip narrower images, such as icons and tracking pixels. Skipped images are free.

## `minHeight` (type: `integer`):

Skip images lower than this. Skipped images are free.

## `allowedTypes` (type: `array`):

Keep only these formats. Empty keeps all image formats.

## `maxFileSizeMb` (type: `integer`):

Files larger than this are not saved (free row).

## `includeMetaImages` (type: `boolean`):

In page mode, also collect the image the page uses when shared on social networks.

## `skipDuplicates` (type: `boolean`):

When two URLs return the same bytes, save the file once. Duplicates are free.

## `keyValueStoreName` (type: `string`):

Optional. Name of a key value store to save files into, so they are kept after the run's own storage expires. Empty uses the run's default store.

## `respectRobotsTxt` (type: `boolean`):

Skip pages and files that the site's robots.txt closes to automated tools (free rows). Turn off only for sites you own or may download from.

## Actor input object example

```json
{
  "imageUrls": [
    "https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"
  ],
  "pageUrls": [
    "https://books.toscrape.com/"
  ],
  "maxImagesPerPage": 50,
  "maxImages": 1000,
  "minWidth": 0,
  "minHeight": 0,
  "allowedTypes": [],
  "maxFileSizeMb": 25,
  "includeMetaImages": true,
  "skipDuplicates": true,
  "respectRobotsTxt": true
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `images` (type: `string`):

Downloaded image files saved in the run's key value store (when no named store is set).

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "imageUrls": [
        "https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"
    ],
    "pageUrls": [
        "https://books.toscrape.com/"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/bulk-image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "imageUrls": ["https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"],
    "pageUrls": ["https://books.toscrape.com/"],
}

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/bulk-image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "imageUrls": [
    "https://books.toscrape.com/media/cache/2c/da/2cdad67c44b002e7ead0cc35693c0e8b.jpg"
  ],
  "pageUrls": [
    "https://books.toscrape.com/"
  ]
}' |
apify call pistachio_implementation/bulk-image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/bulk-image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/a8LAr0CsIuExFwrAZ/builds/9mOsPoy2Zzk4WZsOO/openapi.json
