# Website Image Downloader & Bulk Image Scraper (`sste/website-image-downloader`) Actor

Download all images from any website or extract image URLs (srcset, lazy-loaded, picture, CSS) with width, height, format and alt text. Filters icons and tracking pixels, removes duplicates, ZIP output. REST API and MCP for AI agents.

- **URL**: https://apify.com/sste/website-image-downloader.md
- **Developed by:** [SSTE](https://apify.com/sste) (community)
- **Categories:** E-commerce, Developer tools, Agents
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.20 / 1,000 image delivereds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Website Image Downloader & Bulk Image Scraper

Download **all images from any website** — or a list of image URLs — as a **ZIP**, plus a table with each image's **URL, width, height, format, file size, alt text and source page**. Extract image URLs only, inspect images without storing them, or crawl a whole catalogue. Icons, sprites and tracking pixels are filtered out automatically and duplicates are removed.

Use it in the Apify Console, through a **REST API** (cURL, Python, JavaScript) or as a tool for **AI agents via MCP** (Claude, Cursor, VS Code, any MCP client).

**$0.20 per 1,000 images** · about **$0.01 for a page with 40 images** · no proxies, no browser, no login.

### Quick answers

- **What API can download all images from a website?** This Actor: `POST https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items` with `{"startUrls":[{"url":"https://example.com"}]}` returns every image with its dimensions, and the run's `images.zip` contains the files.
- **How do I extract image URLs from a webpage?** Use `"mode": "list"` — it returns every image URL found on the page (`<img>`, `srcset`, `<picture>`, lazy-load attributes, CSS backgrounds, og:image) without downloading anything, at $0.05 per 1,000 URLs.
- **How do I download product images in bulk?** Give it your category or product pages (Shopify, WooCommerce, any store), set `minWidth`/`minHeight` (e.g. 400) and `urlMustContain` (e.g. `/products/` or `cdn.shopify.com`), and optionally `crawlDepth: 1` to follow product links.
- **How can an AI agent download website images?** Add `https://mcp.apify.com?tools=sste/website-image-downloader` as an MCP server; the agent gets a `sste--website-image-downloader` tool.
- **How do I filter icons and tracking pixels?** Dimensions are read from each file's bytes; `minWidth`/`minHeight` (default 100 px) drop icons, sprites and 1×1 pixels; `urlMustNotContain` drops `logo`, `icon`, `avatar`…
- **How do I get srcset and lazy-loaded images?** Handled automatically: the largest `srcset`/`<picture>` candidate is chosen, and `data-src`, `data-srcset`, `data-lazy-src`, `data-original` and `<noscript>` fallbacks are read; `data:` placeholders are ignored.

### What it does

- 🖼️ **Finds images the way a browser shows them**: `<img>`, the **largest `srcset` variant**, `<picture>` sources, lazy-load attributes, `<noscript>` fallbacks, inline CSS `background-image`, `data-bg`, `og:image` / `twitter:image` and links to full-size image files.
- 📏 **Measures every file from its real bytes**: format (JPEG, PNG, GIF, WebP, AVIF, SVG, ICO, BMP, TIFF, HEIC), width × height, file size, SHA-256.
- 🧹 **Filters**: minimum width/height, minimum/maximum file size, formats, "image URL must contain / must not contain".
- ♻️ **Deduplicates**: the same URL on many pages, and the same file served under different URLs.
- 📦 **ZIP output** (`images.zip`, split into 500 MB parts for big jobs) with readable names like `blue-sneaker-3fa2b1c9.jpg`.
- 🕸️ **Crawls** same-site links (depth 1–3) to cover a whole section or catalogue.
- 💸 **Spending-safe**: stops cleanly at *Max images* or at your maximum cost per run; never charges for skipped, duplicate or failed images.

### Use cases

- **E-commerce / product image downloader**: download product photos from Shopify, WooCommerce, Magento or any store; migrate a catalogue; build product feeds; collect supplier images.
- **Shopify image downloader**: point it at a collection page, keep `cdn.shopify.com/s/files`, follow product links with `crawlDepth: 1`.
- **SEO & web performance audits** (`inspect` mode): find oversized images, missing alt text, legacy formats across a site — no files stored.
- **AI datasets & RAG**: image URLs with dimensions, alt text and source page; feed vision models or agents.
- **Design & marketing**: mood boards, brand assets, competitor creatives.

### Modes

| Mode | Output | Charged per |
|---|---|---|
| `download` (default) | `images.zip` + one dataset row per image | image |
| `inspect` | dataset rows only (format, size, dimensions, hash) | image |
| `list` | image URLs only, no download (fastest) | image URL |

### Use it from the API

Get an API token in Apify Console → Settings → Integrations. Run synchronously and get the image list:

```bash
curl -X POST "https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items" \
  -H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
  -d '{"startUrls":[{"url":"https://www.example.com/"}],"minWidth":300,"minHeight":300,"maxImages":200}'
```

Python (`pip install "apify-client>=3"`):

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("sste/website-image-downloader").call(run_input={
    "startUrls": [{"url": "https://www.example.com/"}],
    "minWidth": 300, "minHeight": 300, "maxImages": 200,
})
for img in client.dataset(run.default_dataset_id).iterate_items():
    print(img["width"], img["height"], img["imageUrl"])
zip_record = client.key_value_store(run.default_key_value_store_id).get_record_as_bytes("images.zip")
if zip_record:  # no ZIP when no image matched the filters
    open("images.zip", "wb").write(zip_record["value"])
```

JavaScript / Node.js (`npm i apify-client`):

```js
import { ApifyClient } from 'apify-client';

const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('sste/website-image-downloader').call({
    startUrls: [{ url: 'https://www.example.com/' }], mode: 'list', maxImages: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((i) => i.imageUrl));
```

The ZIP is the `images.zip` record of the run's default key-value store (also linked in the run's **Output** tab and in the `OUTPUT` record). OpenAPI, CLI and more clients: the **API** tab of this page.

### Use it from AI agents (MCP)

Hosted MCP server URL (OAuth sign-in on first use, or `Authorization: Bearer <APIFY_TOKEN>`):

```
https://mcp.apify.com?tools=sste/website-image-downloader
```

**Claude Desktop / Cursor / VS Code / any MCP client**:

```json
{
  "mcpServers": {
    "website-image-downloader": { "url": "https://mcp.apify.com?tools=sste/website-image-downloader" }
  }
}
```

**Claude Code**: `claude mcp add --transport http website-image-downloader "https://mcp.apify.com?tools=sste/website-image-downloader"`

**Local (stdio)**: `npx -y @apify/actors-mcp-server --tools sste/website-image-downloader` with `APIFY_TOKEN` set.

The agent gets the tool `sste--website-image-downloader` (plus helpers to read the dataset and the ZIP). Example prompts:

- "Download all product images larger than 600 px from https://shop.example.com/collections/sale and give me the ZIP link."
- "List every image URL on https://example.com/blog/post with its alt text, and tell me which images have no alt text."
- "Find images over 500 KB on https://example.com/ and suggest which to convert to WebP."
- "Collect the og:image of these 20 URLs."

Tips for agents: use `mode: "list"` or `"inspect"` when files are not needed, keep `maxImages` small, and raise `minWidth`/`minHeight` to skip thumbnails.

### Input example

```json
{
  "startUrls": [
    { "url": "https://www.example-shop.com/collections/sneakers" },
    { "url": "https://cdn.example.com/photos/hero.jpg" }
  ],
  "minWidth": 400,
  "minHeight": 400,
  "formats": ["jpg", "png", "webp"],
  "urlMustNotContain": ["logo", "icon"],
  "crawlDepth": 1,
  "maxPagesPerStartUrl": 30,
  "maxImages": 2000
}
```

### Output

One dataset row per image:

```json
{
  "imageUrl": "https://cdn.example-shop.com/files/blue-sneaker.jpg?width=1600",
  "pageUrl": "https://www.example-shop.com/collections/sneakers",
  "startUrl": "https://www.example-shop.com/collections/sneakers",
  "alt": "Blue running sneaker, side view",
  "foundIn": "srcset",
  "format": "jpg",
  "width": 1600,
  "height": 1600,
  "bytes": 184322,
  "sha256": "3fa2b1c9…",
  "fileName": "blue-sneaker-3fa2b1c9.jpg",
  "downloadUrl": null
}
```

The run's `OUTPUT` record summarizes images saved, ZIP links, pages processed/failed (with reasons) and images skipped per filter. Enable **Also save each image as a separate file** to get a direct `downloadUrl` per image.

### Pricing

| Event | Price |
|---|---|
| Run start | $0.001 |
| Page processed | $0.0002 ($0.20 / 1,000 pages) |
| **Image delivered** (download / inspect) | **$0.0002 ($0.20 / 1,000 images)** |
| Each started MB above 2 MB per image | $0.0001 ($0.10 / GB) |
| Image URL (list mode) | $0.00005 ($0.05 / 1,000 URLs) |

Examples: one page with 40 images ≈ **$0.009** · 100 pages / 3,000 images ≈ **$0.62** · image URLs of 1,000 pages (~30,000) ≈ **$1.70**. Skipped, duplicate and failed images and failed pages are free. Set *Max images* or a maximum cost per run to cap spend.

### How it handles tricky pages

- **srcset / `<picture>`**: picks the largest width (`w`) or density (`x`) candidate — the full-resolution file, not the thumbnail.
- **Lazy loading**: reads `data-src`, `data-srcset`, `data-lazy-src`, `data-original`, `data-bg` and `<noscript>` fallbacks; ignores `data:` placeholders.
- **Icons and tracking pixels**: dimensions come from the image header, so 1×1 GIFs, 16 px favicons and sprites are dropped by the size filter even when the URL looks normal.
- **Soft 404s**: HTML error pages served at image URLs are detected by file signature and skipped (not charged).
- **Duplicates**: the same file under different URLs (CDN variants) is detected by SHA-256.

### FAQ

**Can it download images from Shopify stores?** Yes — collection and product pages are server-rendered; use `urlMustContain: ["cdn.shopify.com"]` and `crawlDepth: 1`.

**Does it render JavaScript?** No; it uses fast HTTP requests. Images injected only by JavaScript after load (some single-page apps, infinite scroll) are not seen. On static and server-rendered pages it found 96% of the images a browser shows in our tests.

**Can I get only the image URLs?** Yes, `mode: "list"`.

**What formats are supported?** JPEG, PNG, GIF, WebP, AVIF, SVG, ICO, BMP, TIFF and HEIC (detected from content, not from the URL).

**Is there a size limit?** Files up to 50 MB; ZIPs are split into 500 MB parts.

**Will it get blocked?** Sites that block automated requests or require login are reported in `failedPages`; the Actor does not bypass logins, CAPTCHAs or anti-bot systems.

### Limitations and responsible use

Only public web addresses are fetched (private networks, localhost and non-web ports are blocked). You are responsible for having the right to download and use the images you collect — respect copyright, image licences and website terms. Intended for your own sites, public assets, auditing, research and properly licensed content.

### More

- Guides, API and MCP docs: [website-image-downloader.pages.dev](https://website-image-downloader.pages.dev/?utm_source=apify\&utm_medium=readme)
- Code examples (cURL, Python, JavaScript, MCP configs): [github.com/SSTEmpresarial/website-image-downloader-examples](https://github.com/SSTEmpresarial/website-image-downloader-examples)
- Changelog: **Changelog** tab.

# Changelog

This Actor's version history is a separate document: https://apify.com/sste/website-image-downloader/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

Web pages to extract images from, and/or direct image URLs to download. Up to 10,000 per run.

## `mode` (type: `string`):

<b>download</b>: save the image files (plus a ZIP) and their details. <b>inspect</b>: download to measure width, height, format and size, but do not store files. <b>list</b>: only list the image URLs found on the pages (fastest; no download, so size filters do not apply).

## `minWidth` (type: `integer`):

Skip images narrower than this. The default of 100 removes icons, sprites and tracking pixels. Set 0 to keep everything.

## `minHeight` (type: `integer`):

Skip images shorter than this. Set 0 to keep everything.

## `formats` (type: `array`):

Keep only these formats. The format is detected from the file contents, not from the URL.

## `minFileSizeKB` (type: `integer`):

Skip files smaller than this.

## `maxFileSizeMB` (type: `integer`):

Skip files larger than this (maximum 50 MB).

## `dedupe` (type: `boolean`):

Skip repeated image URLs, and identical files served under different URLs (compared by SHA-256 of the content).

## `urlMustContain` (type: `array`):

Keep only images whose URL contains at least one of these texts (case-insensitive), for example /products/ or cdn.shopify.com.

## `urlMustNotContain` (type: `array`):

Skip images whose URL contains any of these texts (case-insensitive), for example logo, avatar or sprite.

## `crawlDepth` (type: `integer`):

0 = only the given pages. 1 = also the pages they link to on the same website. Maximum 3.

## `maxPagesPerStartUrl` (type: `integer`):

Upper limit of pages visited for each start URL when following links.

## `maxImages` (type: `integer`):

The run stops after this many images. You are charged only for images delivered.

## `maxImagesPerPage` (type: `integer`):

Process at most this many image candidates from a single page.

## `createZip` (type: `boolean`):

In download mode, pack all images into images.zip (split into 500 MB parts when larger).

## `saveIndividualFiles` (type: `boolean`):

Store every image as its own record in the key-value store, with a direct downloadUrl in the results. Useful for integrations that process images one by one.

## `includeCssBackgrounds` (type: `boolean`):

Images set with an inline background-image style or data-bg attributes.

## `includeMetaImages` (type: `boolean`):

The og:image and twitter:image of each page.

## `includeLinkedImages` (type: `boolean`):

Links that point directly to image files, such as full-size gallery images.

## `includeIcons` (type: `boolean`):

Favicons and apple-touch-icons declared in the page head.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://www.apple.com/iphone/"
    }
  ],
  "mode": "download",
  "minWidth": 100,
  "minHeight": 100,
  "formats": [
    "jpg",
    "png",
    "gif",
    "webp",
    "avif",
    "svg",
    "ico",
    "bmp",
    "tiff",
    "heic"
  ],
  "minFileSizeKB": 0,
  "maxFileSizeMB": 25,
  "dedupe": true,
  "crawlDepth": 0,
  "maxPagesPerStartUrl": 20,
  "maxImages": 1000,
  "maxImagesPerPage": 500,
  "createZip": true,
  "saveIndividualFiles": false,
  "includeCssBackgrounds": true,
  "includeMetaImages": true,
  "includeLinkedImages": true,
  "includeIcons": false
}
```

# Actor output Schema

## `images` (type: `string`):

No description

## `zip` (type: `string`):

No description

## `summary` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://www.apple.com/iphone/"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("sste/website-image-downloader").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [{ "url": "https://www.apple.com/iphone/" }] }

# Run the Actor and wait for it to finish
run = client.actor("sste/website-image-downloader").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://www.apple.com/iphone/"
    }
  ]
}' |
apify call sste/website-image-downloader --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,sste/website-image-downloader"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/bMmj1VAT5YqUNHtDf/builds/mz9wLiOTS28NUkAzh/openapi.json
