Website Image Downloader & Bulk Image Scraper avatar

Website Image Downloader & Bulk Image Scraper

Pricing

from $0.20 / 1,000 image delivereds

Go to Apify Store
Website Image Downloader & Bulk Image Scraper

Website Image Downloader & Bulk Image Scraper

Download all images from any website or extract image URLs (srcset, lazy-loaded, picture, CSS) with width, height, format and alt text. Filters icons and tracking pixels, removes duplicates, ZIP output. REST API and MCP for AI agents.

Pricing

from $0.20 / 1,000 image delivereds

Rating

0.0

(0)

Developer

SSTE

SSTE

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

an hour ago

Last modified

Share

Download all images from any website — or a list of image URLs — as a ZIP, plus a table with each image's URL, width, height, format, file size, alt text and source page. Extract image URLs only, inspect images without storing them, or crawl a whole catalogue. Icons, sprites and tracking pixels are filtered out automatically and duplicates are removed.

Use it in the Apify Console, through a REST API (cURL, Python, JavaScript) or as a tool for AI agents via MCP (Claude, Cursor, VS Code, any MCP client).

$0.20 per 1,000 images · about $0.01 for a page with 40 images · no proxies, no browser, no login.

Quick answers

  • What API can download all images from a website? This Actor: POST https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items with {"startUrls":[{"url":"https://example.com"}]} returns every image with its dimensions, and the run's images.zip contains the files.
  • How do I extract image URLs from a webpage? Use "mode": "list" — it returns every image URL found on the page (<img>, srcset, <picture>, lazy-load attributes, CSS backgrounds, og:image) without downloading anything, at $0.05 per 1,000 URLs.
  • How do I download product images in bulk? Give it your category or product pages (Shopify, WooCommerce, any store), set minWidth/minHeight (e.g. 400) and urlMustContain (e.g. /products/ or cdn.shopify.com), and optionally crawlDepth: 1 to follow product links.
  • How can an AI agent download website images? Add https://mcp.apify.com?tools=sste/website-image-downloader as an MCP server; the agent gets a sste--website-image-downloader tool.
  • How do I filter icons and tracking pixels? Dimensions are read from each file's bytes; minWidth/minHeight (default 100 px) drop icons, sprites and 1×1 pixels; urlMustNotContain drops logo, icon, avatar…
  • How do I get srcset and lazy-loaded images? Handled automatically: the largest srcset/<picture> candidate is chosen, and data-src, data-srcset, data-lazy-src, data-original and <noscript> fallbacks are read; data: placeholders are ignored.

What it does

  • 🖼️ Finds images the way a browser shows them: <img>, the largest srcset variant, <picture> sources, lazy-load attributes, <noscript> fallbacks, inline CSS background-image, data-bg, og:image / twitter:image and links to full-size image files.
  • 📏 Measures every file from its real bytes: format (JPEG, PNG, GIF, WebP, AVIF, SVG, ICO, BMP, TIFF, HEIC), width × height, file size, SHA-256.
  • 🧹 Filters: minimum width/height, minimum/maximum file size, formats, "image URL must contain / must not contain".
  • ♻️ Deduplicates: the same URL on many pages, and the same file served under different URLs.
  • 📦 ZIP output (images.zip, split into 500 MB parts for big jobs) with readable names like blue-sneaker-3fa2b1c9.jpg.
  • 🕸️ Crawls same-site links (depth 1–3) to cover a whole section or catalogue.
  • 💸 Spending-safe: stops cleanly at Max images or at your maximum cost per run; never charges for skipped, duplicate or failed images.

Use cases

  • E-commerce / product image downloader: download product photos from Shopify, WooCommerce, Magento or any store; migrate a catalogue; build product feeds; collect supplier images.
  • Shopify image downloader: point it at a collection page, keep cdn.shopify.com/s/files, follow product links with crawlDepth: 1.
  • SEO & web performance audits (inspect mode): find oversized images, missing alt text, legacy formats across a site — no files stored.
  • AI datasets & RAG: image URLs with dimensions, alt text and source page; feed vision models or agents.
  • Design & marketing: mood boards, brand assets, competitor creatives.

Modes

ModeOutputCharged per
download (default)images.zip + one dataset row per imageimage
inspectdataset rows only (format, size, dimensions, hash)image
listimage URLs only, no download (fastest)image URL

Use it from the API

Get an API token in Apify Console → Settings → Integrations. Run synchronously and get the image list:

curl -X POST "https://api.apify.com/v2/acts/sste~website-image-downloader/run-sync-get-dataset-items" \
-H "Authorization: Bearer $APIFY_TOKEN" -H "Content-Type: application/json" \
-d '{"startUrls":[{"url":"https://www.example.com/"}],"minWidth":300,"minHeight":300,"maxImages":200}'

Python (pip install "apify-client>=3"):

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("sste/website-image-downloader").call(run_input={
"startUrls": [{"url": "https://www.example.com/"}],
"minWidth": 300, "minHeight": 300, "maxImages": 200,
})
for img in client.dataset(run.default_dataset_id).iterate_items():
print(img["width"], img["height"], img["imageUrl"])
zip_record = client.key_value_store(run.default_key_value_store_id).get_record_as_bytes("images.zip")
if zip_record: # no ZIP when no image matched the filters
open("images.zip", "wb").write(zip_record["value"])

JavaScript / Node.js (npm i apify-client):

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: process.env.APIFY_TOKEN });
const run = await client.actor('sste/website-image-downloader').call({
startUrls: [{ url: 'https://www.example.com/' }], mode: 'list', maxImages: 500,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
console.log(items.map((i) => i.imageUrl));

The ZIP is the images.zip record of the run's default key-value store (also linked in the run's Output tab and in the OUTPUT record). OpenAPI, CLI and more clients: the API tab of this page.

Use it from AI agents (MCP)

Hosted MCP server URL (OAuth sign-in on first use, or Authorization: Bearer <APIFY_TOKEN>):

https://mcp.apify.com?tools=sste/website-image-downloader

Claude Desktop / Cursor / VS Code / any MCP client:

{
"mcpServers": {
"website-image-downloader": { "url": "https://mcp.apify.com?tools=sste/website-image-downloader" }
}
}

Claude Code: claude mcp add --transport http website-image-downloader "https://mcp.apify.com?tools=sste/website-image-downloader"

Local (stdio): npx -y @apify/actors-mcp-server --tools sste/website-image-downloader with APIFY_TOKEN set.

The agent gets the tool sste--website-image-downloader (plus helpers to read the dataset and the ZIP). Example prompts:

Tips for agents: use mode: "list" or "inspect" when files are not needed, keep maxImages small, and raise minWidth/minHeight to skip thumbnails.

Input example

{
"startUrls": [
{ "url": "https://www.example-shop.com/collections/sneakers" },
{ "url": "https://cdn.example.com/photos/hero.jpg" }
],
"minWidth": 400,
"minHeight": 400,
"formats": ["jpg", "png", "webp"],
"urlMustNotContain": ["logo", "icon"],
"crawlDepth": 1,
"maxPagesPerStartUrl": 30,
"maxImages": 2000
}

Output

One dataset row per image:

{
"imageUrl": "https://cdn.example-shop.com/files/blue-sneaker.jpg?width=1600",
"pageUrl": "https://www.example-shop.com/collections/sneakers",
"startUrl": "https://www.example-shop.com/collections/sneakers",
"alt": "Blue running sneaker, side view",
"foundIn": "srcset",
"format": "jpg",
"width": 1600,
"height": 1600,
"bytes": 184322,
"sha256": "3fa2b1c9…",
"fileName": "blue-sneaker-3fa2b1c9.jpg",
"downloadUrl": null
}

The run's OUTPUT record summarizes images saved, ZIP links, pages processed/failed (with reasons) and images skipped per filter. Enable Also save each image as a separate file to get a direct downloadUrl per image.

Pricing

EventPrice
Run start$0.001
Page processed$0.0002 ($0.20 / 1,000 pages)
Image delivered (download / inspect)$0.0002 ($0.20 / 1,000 images)
Each started MB above 2 MB per image$0.0001 ($0.10 / GB)
Image URL (list mode)$0.00005 ($0.05 / 1,000 URLs)

Examples: one page with 40 images ≈ $0.009 · 100 pages / 3,000 images ≈ $0.62 · image URLs of 1,000 pages (~30,000) ≈ $1.70. Skipped, duplicate and failed images and failed pages are free. Set Max images or a maximum cost per run to cap spend.

How it handles tricky pages

  • srcset / <picture>: picks the largest width (w) or density (x) candidate — the full-resolution file, not the thumbnail.
  • Lazy loading: reads data-src, data-srcset, data-lazy-src, data-original, data-bg and <noscript> fallbacks; ignores data: placeholders.
  • Icons and tracking pixels: dimensions come from the image header, so 1×1 GIFs, 16 px favicons and sprites are dropped by the size filter even when the URL looks normal.
  • Soft 404s: HTML error pages served at image URLs are detected by file signature and skipped (not charged).
  • Duplicates: the same file under different URLs (CDN variants) is detected by SHA-256.

FAQ

Can it download images from Shopify stores? Yes — collection and product pages are server-rendered; use urlMustContain: ["cdn.shopify.com"] and crawlDepth: 1.

Does it render JavaScript? No; it uses fast HTTP requests. Images injected only by JavaScript after load (some single-page apps, infinite scroll) are not seen. On static and server-rendered pages it found 96% of the images a browser shows in our tests.

Can I get only the image URLs? Yes, mode: "list".

What formats are supported? JPEG, PNG, GIF, WebP, AVIF, SVG, ICO, BMP, TIFF and HEIC (detected from content, not from the URL).

Is there a size limit? Files up to 50 MB; ZIPs are split into 500 MB parts.

Will it get blocked? Sites that block automated requests or require login are reported in failedPages; the Actor does not bypass logins, CAPTCHAs or anti-bot systems.

Limitations and responsible use

Only public web addresses are fetched (private networks, localhost and non-web ports are blocked). You are responsible for having the right to download and use the images you collect — respect copyright, image licences and website terms. Intended for your own sites, public assets, auditing, research and properly licensed content.

More