Bulk Image Downloader: All Images from Web Pages to ZIP avatar

Bulk Image Downloader: All Images from Web Pages to ZIP

Pricing

from $7.00 / 1,000 page or image urls

Go to Apify Store
Bulk Image Downloader: All Images from Web Pages to ZIP

Bulk Image Downloader: All Images from Web Pages to ZIP

Download all images from web pages you list, or from direct image links, into your key-value store, optionally as a ZIP. One row per image: page, image URL, alt text, width, height, type, bytes, SHA-256. Duplicates saved once; icons and tracking pixels skipped.

Pricing

from $7.00 / 1,000 page or image urls

Rating

0.0

(0)

Developer

Michael Costa

Michael Costa

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Share

What does Bulk Image Downloader do?

Bulk Image Downloader downloads all the images from web pages, or from direct image links, into your own Apify storage: the largest size each page offers, optionally packed as ZIP files to download in one go, plus one row per image with its source page, alt text, size and hash.

For a page, it finds every image the page shows: <img> tags (the largest size in srcset), <picture> sources, lazy-loaded images, and the page's preview image (og:image). It is not a way around copyright: download only images you have the right to use.

Try it in one click: the input comes pre-filled with two direct links to public-domain photos on Wikimedia Commons (see the example below for their licence pages). That's 2 URLs, about $0.014 (2 × $0.007, plus $0.00005 for the run start), plus the run's platform usage. Then replace them with the pages or image links you actually want.

What data does Bulk Image Downloader return?

FieldExampleNotes
imageUrlhttps://www.python.org/static/img/python-logo.pngThe image's address.
sourcePagehttps://www.python.org/The page it was on; null for direct links.
foundInimgimg, srcset (the largest size), picture, og:image or direct.
altTextpython™null when the page gives none.
width, height580, 164Pixels, read from the file.
contentType, bytesimage/png, 15770The real type, read from the file itself.
sha2569c121e61...Duplicates (same bytes) are saved once.
key9c121e619bfe02ea-python-logo.pngThe file's name in the run's key-value store.
storeUrlhttps://api.apify.com/v2/key-value-stores/<store id>/records/...Downloads your saved copy.
inputhttps://www.python.org/The input line it came from.

One row per saved image; the files themselves are in the run's key-value store. The full list is under Output.

How much does it cost to download images in bulk?

You pay per URL you give it: $7.00 per 1,000 URLs (a page, however many images it has, or a direct image link), plus $0.00005 each time a run starts, plus the run's Apify platform usage (compute, storage and data transfer) at your plan's rates.

  • The example below: 2 URLs × $0.007 = $0.014, plus the start fee, plus the run's platform usage (about 180 KB stored).
  • A month, for example: 500 product pages = 500 × $0.007 = $3.50, plus platform usage, which grows with the bytes stored (large photos and the ZIP option cost more storage) and is billed at your plan's rates.
  • Caps: Max images per run in the input, and Maximum cost per run in the run options: no URL is started that it doesn't cover. The run stops cleanly at whichever comes first.

Never charged: a URL where no image was saved: pages that fail, are blocked, have no images, or whose images are all duplicates or filtered out. The images and ZIP files themselves aren't charged per file; storing them counts toward platform usage, so large photos and the ZIP option (which stores each image a second time) cost more storage.

How to download all images from a web page

  1. Open Bulk Image Downloader and click Try for free (or Start if you're signed in).
  2. Put pages or direct image links in Web pages or image URLs, one per line.
  3. Optional: set Minimum width / height and Minimum file size to skip icons, and turn on Also save a ZIP file.
  4. Click Start, then open the Output tab for the rows, and the run's key-value store (Storage) for the files and ZIPs.

Example: two public-domain images

The pre-filled input is two direct image links, both NASA photographs that are in the public domain. Their licence pages on Wikimedia Commons say so ({{PD-USGov-NASA}}: "This file is in the public domain in the United States because it was solely created by NASA"):

{"urls": ["https://upload.wikimedia.org/wikipedia/commons/thumb/9/97/The_Earth_seen_from_Apollo_17.jpg/500px-The_Earth_seen_from_Apollo_17.jpg",
"https://upload.wikimedia.org/wikipedia/commons/thumb/9/98/Aldrin_Apollo_11_original.jpg/500px-Aldrin_Apollo_11_original.jpg"],
"minWidth": 100, "minHeight": 100, "minFileSizeKb": 1, "includePreviewImage": true, "createZip": false}

It saved both images. One row (real output from a local run on 2026-09-25; storeUrl points to your run's key-value store):

{
"id": "f9e124d879181002a45939ce",
"input": "https://upload.wikimedia.org/wikipedia/commons/thumb/9/97/The_Earth_seen_from_Apollo_17.jpg/500px-The_Earth_seen_from_Apollo_17.jpg",
"sourcePage": null,
"imageUrl": "https://upload.wikimedia.org/wikipedia/commons/thumb/9/97/The_Earth_seen_from_Apollo_17.jpg/500px-The_Earth_seen_from_Apollo_17.jpg",
"foundIn": "direct",
"altText": null,
"width": 500,
"height": 500,
"contentType": "image/jpeg",
"bytes": 82862,
"sha256": "2075157c388887df11e7f4e1320717f5c0a38673d30c59340e3df6a059ae2f3e",
"key": "2075157c388887df-500px-The_Earth_seen_from_Apollo_17.jpg"
}

Give it a web page instead and it finds every image on that page (see Output below for a page's row, with its alt text and where on the page the image was found).

Input

FieldWhat it does
Web pages or image URLsOne per line: a page (all its images) or a direct image link. https:// is added if missing.
Minimum width / height (pixels)Skip images smaller than this (default 100 x 100). 0 keeps every size.
Minimum file size (KB)Skip files smaller than this (default 1 KB).
Include the page's preview imageAlso download the page's og:image / twitter:image (default on).
Max images per pageLook at only the first N images of each page (default: all, up to 1,000).
Also save a ZIP filePack the saved images into images-001.zip, images-002.zip, ... (default off).
Max images per runCap the total number of saved images.
{
"urls": ["https://www.python.org/", "https://example.com/photos/sunset.jpg"],
"minWidth": 200,
"minHeight": 200,
"createZip": true,
"maxResults": 500
}

Output

One row per saved image. Fields a page doesn't give are null.

{
"id": "20e72af68b277b1e6ea3d89e",
"input": "https://www.python.org/",
"sourcePage": "https://www.python.org/",
"imageUrl": "https://www.python.org/static/img/python-logo.png",
"finalUrl": "https://www.python.org/static/img/python-logo.png",
"foundIn": "img",
"altText": "python™",
"width": 580,
"height": 164,
"contentType": "image/png",
"bytes": 15770,
"sha256": "9c121e619bfe02eaba582d7080eea46fd53ec0b50717e6794a948fada4ae8f3c",
"key": "9c121e619bfe02ea-python-logo.png",
"storeUrl": "https://api.apify.com/v2/key-value-stores/<store id>/records/9c121e619bfe02ea-python-logo.png",
"scrapedAt": "2026-09-24T12:00:00Z"
}
  • foundIn says where the address came from: img, srcset (the largest size), picture, og:image, or direct (a link you gave). sourcePage is null for direct links.
  • key is the file's name in the run's key-value store: the start of its SHA-256, then its name from the URL. storeUrl downloads it.
  • id is stable across runs for the same image address.

Files in the key-value store

  • One file per saved image, under its key, with its real content type.
  • With Also save a ZIP file: images-001.zip, images-002.zip, ..., each about 64 MB at most (a new part starts when one fills up), holding exactly the images in the dataset.
  • RUN_STATS: per input, how many images were found, saved, and skipped for each reason (duplicates, tooSmall, notAnImage, tooLarge, blockedByRobots, optedOutOfAI, failed), with examples.

A run's default key-value store is unnamed, and Apify deletes unnamed stores after 7 days unless your plan keeps them longer, so download or copy the files you want to keep.

Run it on a schedule, or from your own code

  1. Save your input as a task (Save as a new task, top right of the actor page) and add it to a schedule (Console → Schedules → Create new): for example weekly over a list of product pages, if you need fresh copies. Every run downloads, and charges for, every URL again.
  2. Collect results: download the dataset as JSON, CSV or Excel; fetch the latest run's results from the API (GET https://api.apify.com/v2/actor-tasks/<task id>/runs/last/dataset/items?status=SUCCEEDED&format=csv, with your API token); let a webhook tell your system when a run succeeds; or connect it to Make, Zapier or n8n through Apify's integrations.

Files are saved in the run's own key-value store, which Apify deletes after your plan's retention period (7 days on the free plan): download or copy what you want to keep.

Can I use Bulk Image Downloader from an AI agent (MCP)?

Yes, through Apify's MCP server: add https://mcp.apify.com?tools=humble-echidna/bulk-image-downloader to your MCP client (or let the agent find it with the server's actor search). The agent passes pages, e.g. {"urls": ["https://example.com/product/1"], "maxResults": 20}, and gets back rows with each saved image's storeUrl.

Who it's for

Shops, marketplaces and catalogue teams moving or checking product images, researchers building image sets they have the rights to, and anyone archiving the images of pages they manage.

Why this one?

  • Only the images you want. Icons, spacers and 1x1 tracking pixels are skipped by a minimum width, height and file size you choose; an icon the page itself declares as small isn't even downloaded.
  • No duplicates. The same picture reached through two addresses, or repeated across pages, is saved once (compared by its bytes, SHA-256).
  • The real file, the real type. For responsive images it takes the largest size the page offers, not the thumbnail. The format is read from the file itself, not trusted from the server, so a web page served as "image/jpeg" is never saved as an image.
  • Useful rows, not just files. Alt text, source page, dimensions, format, bytes and hash for every image, ready for a spreadsheet, a product catalogue or a dataset.
  • Polite and safe. It identifies itself honestly (User-Agent HumbleEchidnaApify), follows each site's robots.txt (read once per site per run) and Crawl-delay, honours sites that opt out of AI crawlers in their robots.txt (GPTBot, CCBot, ClaudeBot, Google-Extended and the like), and only requests public web addresses on the standard ports (80 and 443). It never logs in or gets around blocking.
  • Reliable. One failing page or image never affects the others. The run log and the RUN_STATS record say exactly which image was skipped and why.

Limits

  • Pages as the server sends them: images that a page only adds with JavaScript aren't seen (it doesn't run a browser). CSS background images aren't collected.
  • Formats kept: JPEG, PNG, GIF, WebP, AVIF, BMP, TIFF and ICO. SVG files are skipped: they're documents that can carry scripts. HEIC/HEIF files are skipped too.
  • Up to 25 MB per image, 2 GB of downloads per run, 1,000 images per page, pages up to 5 MB.
  • Public web addresses on ports 80 and 443 only; no logins, no proxies, nothing behind bot protection.

FAQ

Can I download any image with this?

Only images you have the right to download and use. Most images on the web are protected by copyright: check the site's terms and the image's licence before you reuse it. The copies are saved only in your own Apify storage; this actor doesn't publish or share them.

Why was an image skipped?

Check RUN_STATS and the run log: each skipped image is counted by reason, and the first few problems per input are listed with the image address. A site's robots.txt or its AI opt-out can block an image even when the page itself was allowed.

How long are the downloaded images kept?

They're in the run's default key-value store, which is unnamed; Apify deletes unnamed stores after 7 days unless your plan keeps them longer. Download the files or ZIPs, or copy them, to keep them.

Something that used to work now fails or returns fewer images. Why?

The site may have changed its pages. The run log names the input and what went wrong, and every other input in the run is unaffected. Please open an issue with the input you used.

It fetches only the pages and image links you give it, as a logged-out visitor, and follows each site's robots.txt, including its opt-outs for AI crawlers. It doesn't log in, get around blocking, or collect personal data about anyone. Whether you may reuse the images is up to their owners' terms and licences, which you are responsible for checking.

ActorUse it when
Sitemap URL ExtractorYou want the list of product or article pages from a site's sitemap, with filters, to use as this actor's input.
Image to Text OCRYou want to read the text in images at public links.

Feedback and support

Found a bug, or a page whose images it misses? Open an issue on the Issues tab with the input you used.

Versions

Current version: 1.0. See the Changelog tab for what changed in each version.