Download all images from web pages you list, or from direct image links, into your key-value store, optionally as a ZIP. One row per image: page, image URL, alt text, width, height, type, bytes, SHA-256. Duplicates saved once; icons and tracking pixels skipped.
Versions follow MAJOR.MINOR.PATCH (src/version.py); Apify shows MAJOR.MINOR from .actor/actor.json.
Every run logs its version and records it in the RUN_STATS key-value record.
1.0.2 (2026-09-25)
Input field descriptions rewritten for AI agents (Apify's MCP server shows agents only the description, not the
form): each now states its default, its allowed range and how it combines with other fields. No change to field
names, types, defaults or behaviour.
1.0.1 (2026-09-25)
The default and pre-filled input are now two public-domain images, so the one-click try (and Apify's daily
test with the default input) only downloads files anyone may copy. It was the python.org home page (its logos are
the Python Software Foundation's) and a Wikimedia Commons cat photo licensed CC BY-SA 3.0
(commons.wikimedia.org/wiki/File:Cat_November_2010-1a.jpg, checked 2026-09-25). Now: 500 px versions of
commons.wikimedia.org/wiki/File:The_Earth_seen_from_Apollo_17.jpg and
commons.wikimedia.org/wiki/File:Aldrin_Apollo_11_original.jpg, both tagged {{PD-USGov-NASA}} ("public domain in
the United States because it was solely created by NASA"; licence pages read 2026-09-25). The README example cites
both licence pages; a test keeps the default, the prefill and those citations in step.
The input field's example page is now example.com. No behaviour change.
1.0.0 (2026-09-24, unreleased)
First release.
Downloads the images on web pages the user lists, or direct image links, into the run's key-value store (one
record per image, keyed by content hash + name, with its real content type) and writes one dataset row per image:
input, source page, image URL, final URL, where it was found, alt text, width, height, content type, bytes, SHA-256,
store key and download URL.
Page parsing (mms_extract): <img> src and lazy-loading attributes, the largest srcset candidate (split the way
the HTML spec does, so CDN addresses with commas survive), <picture> as one image (its largest source),
og:image / twitter:image, <base href>. data: URIs and fragment-only references skipped.
A link with an image file extension is downloaded directly; any other link is tried as a page, and a non-HTML
answer turns it into a direct image link.
Downloads go through mms_common with honour_ai_opt_outs=True (robots.txt, AI-crawler opt-outs, the
private-network guard, ports 80/443), stream to a temp file (sink), and are capped at 25 MB per image and 2 GB
per run. Format from the file's bytes (JPEG, PNG, GIF, WebP, AVIF, BMP, TIFF, ICO); SVG (can carry scripts),
HEIC and web pages are never stored.
Minimum width, height and file size (defaults 100 x 100 px, 1 KB): a page-declared icon and a server-declared tiny
file are refused before download. Duplicates by address and by SHA-256 are saved once per run.
Optional ZIP: images-001.zip, ... parts of about 64 MB (ZIP_STORED), written as images are pushed, so a part
only holds images in the dataset; at most one part on disk and in memory at a time.
Charged per saved image through Apify's apify-default-dataset-item event, $3.00 per 1,000 images plus
$0.00005 per run start. "Max images per run" and the maximum cost per run are claimed (mms_common Slots) before
each download; skipped, duplicate and failed images give their slot back and are never charged.
Failure isolation: a failing page or image only affects itself; RUN_STATS counts every skip reason per input.
Terms and demand gates (2026-09-24): no fixed source; it fetches only the pages and image links the user supplies,
as a logged-out visitor, honouring robots.txt and AI opt-outs (PLAN.md legal hold: "tools that only process URLs or
files the user supplies"). No presets for any site. The README says users must have the right to download the images
and that copies stay in their own storage. Demand (store_merged.json, 2026-09-23): ~162 users/30 days across bulk
image downloaders and extractors (onescales/bulk-image-downloader 114, getascraper 17, logiover 13, apify 7,
automation-lab 6, thirdwatch 5).
Cost measured from local runs on 2026-09-24, priced with Apify's published Free/Starter rates (key-value store
writes $0.05 per 1,000, reads $0.005 per 1,000, dataset writes $0.005 per 1,000, external transfer $0.20/GB,
internal $0.05/GB, $0.20 per compute unit):
Key-value store writes: exactly 1 per saved image, plus 1 per ZIP part (171 writes for 166 images + 5 parts):
$0.05 per 1,000 images, ten times the dataset write. Plus 1 store-metadata read per image for its signed link.
Typical web images (Wikipedia, Commons, Gutenberg, python.org, docs: 142 images, average 60-71 KB): about
$0.05 KV writes + $0.005 KV reads + $0.005 dataset + $0.018 transfer (counting the download as external) + $0.01
compute (7.6 s for 40 images locally, at 512 MB) = **$0.09 per 1,000 images**.
Full-size photos with ZIP on (the same run with NASA's image of the day, average 2 MB, 166 images, 337 MB, 179 s,
peak 247 MB RSS): about $0.05 KV writes + $0.41 external + $0.20 internal (file + ZIP) + $0.03 compute =
~$0.70 per 1,000 images.
Priced before release like the category leader: $7.00 per 1,000 input URLs (a page or a direct image link,
charged once when at least one image from it is saved; images and ZIPs aren't charged per file), with the platform
usage paid by the user. The cost gate showed a per-image price loses on large photos, because storage grows with
bytes (up to 25 MB per image, and the ZIP stores everything twice). $7 is the leader's lowest paid tier: one flat
price, because the SDK can't charge tier-priced custom events. Replaces the per-image $3.00 per 1,000 of the
unreleased build. Competitors: onescales $70 per 1,000 on the free tier, $10 / $8 / $7 on paid tiers, with
platform usage also passed to the user; getascraper $0.89, thirdwatch $1.00, automation-lab $0.79, logiover $7.50.