# Changelog of Bulk Image Downloader: All Images from Web Pages to ZIP (`humble-echidna/bulk-image-downloader`) Actor

- **URL**: https://apify.com/humble-echidna/bulk-image-downloader/changelog.md
- **Full Actor documentation**: https://apify.com/humble-echidna/bulk-image-downloader.md

## Changelog

Versions follow MAJOR.MINOR.PATCH (`src/version.py`); Apify shows MAJOR.MINOR from `.actor/actor.json`.
Every run logs its version and records it in the `RUN_STATS` key-value record.

### 1.0.2 (2026-09-25)

- Input field descriptions rewritten for AI agents (Apify's MCP server shows agents only the description, not the
  form): each now states its default, its allowed range and how it combines with other fields. No change to field
  names, types, defaults or behaviour.

### 1.0.1 (2026-09-25)

- **The default and pre-filled input are now two public-domain images**, so the one-click try (and Apify's daily
  test with the default input) only downloads files anyone may copy. It was the python.org home page (its logos are
  the Python Software Foundation's) and a Wikimedia Commons cat photo licensed CC BY-SA 3.0
  (commons.wikimedia.org/wiki/File:Cat\_November\_2010-1a.jpg, checked 2026-09-25). Now: 500 px versions of
  commons.wikimedia.org/wiki/File:The\_Earth\_seen\_from\_Apollo\_17.jpg and
  commons.wikimedia.org/wiki/File:Aldrin\_Apollo\_11\_original.jpg, both tagged `{{PD-USGov-NASA}}` ("public domain in
  the United States because it was solely created by NASA"; licence pages read 2026-09-25). The README example cites
  both licence pages; a test keeps the default, the prefill and those citations in step.
- The input field's example page is now example.com. No behaviour change.

### 1.0.0 (2026-09-24, unreleased)

First release.

- Downloads the images on web pages the user lists, or direct image links, into the run's key-value store (one
  record per image, keyed by content hash + name, with its real content type) and writes one dataset row per image:
  input, source page, image URL, final URL, where it was found, alt text, width, height, content type, bytes, SHA-256,
  store key and download URL.
- Page parsing (mms\_extract): `<img>` src and lazy-loading attributes, the largest `srcset` candidate (split the way
  the HTML spec does, so CDN addresses with commas survive), `<picture>` as one image (its largest source),
  `og:image` / `twitter:image`, `<base href>`. `data:` URIs and fragment-only references skipped.
- A link with an image file extension is downloaded directly; any other link is tried as a page, and a non-HTML
  answer turns it into a direct image link.
- Downloads go through mms\_common with `honour_ai_opt_outs=True` (robots.txt, AI-crawler opt-outs, the
  private-network guard, ports 80/443), stream to a temp file (`sink`), and are capped at 25 MB per image and 2 GB
  per run. Format from the file's bytes (JPEG, PNG, GIF, WebP, AVIF, BMP, TIFF, ICO); SVG (can carry scripts),
  HEIC and web pages are never stored.
- Minimum width, height and file size (defaults 100 x 100 px, 1 KB): a page-declared icon and a server-declared tiny
  file are refused before download. Duplicates by address and by SHA-256 are saved once per run.
- Optional ZIP: `images-001.zip`, ... parts of about 64 MB (ZIP\_STORED), written as images are pushed, so a part
  only holds images in the dataset; at most one part on disk and in memory at a time.
- Charged per saved image through Apify's `apify-default-dataset-item` event, $3.00 per 1,000 images plus
  $0.00005 per run start. "Max images per run" and the maximum cost per run are claimed (mms\_common `Slots`) before
  each download; skipped, duplicate and failed images give their slot back and are never charged.
- Failure isolation: a failing page or image only affects itself; RUN\_STATS counts every skip reason per input.

Terms and demand gates (2026-09-24): no fixed source; it fetches only the pages and image links the user supplies,
as a logged-out visitor, honouring robots.txt and AI opt-outs (PLAN.md legal hold: "tools that only process URLs or
files the user supplies"). No presets for any site. The README says users must have the right to download the images
and that copies stay in their own storage. Demand (store\_merged.json, 2026-09-23): ~162 users/30 days across bulk
image downloaders and extractors (onescales/bulk-image-downloader 114, getascraper 17, logiover 13, apify 7,
automation-lab 6, thirdwatch 5).

Cost measured from local runs on 2026-09-24, priced with Apify's published Free/Starter rates (key-value store
writes $0.05 per 1,000, reads $0.005 per 1,000, dataset writes $0.005 per 1,000, external transfer $0.20/GB,
internal $0.05/GB, $0.20 per compute unit):

- Key-value store writes: exactly 1 per saved image, plus 1 per ZIP part (171 writes for 166 images + 5 parts):
  $0.05 per 1,000 images, ten times the dataset write. Plus 1 store-metadata read per image for its signed link.
- Typical web images (Wikipedia, Commons, Gutenberg, python.org, docs: 142 images, average 60-71 KB): about
  $0.05 KV writes + $0.005 KV reads + $0.005 dataset + $0.018 transfer (counting the download as external) + ~$0.01
  compute (7.6 s for 40 images locally, at 512 MB) = **~$0.09 per 1,000 images**.
- Full-size photos with ZIP on (the same run with NASA's image of the day, average 2 MB, 166 images, 337 MB, 179 s,
  peak 247 MB RSS): about $0.05 KV writes + $0.41 external + $0.20 internal (file + ZIP) + $0.03 compute =
  **~$0.70 per 1,000 images**.
- Priced before release like the category leader: **$7.00 per 1,000 input URLs** (a page or a direct image link,
  charged once when at least one image from it is saved; images and ZIPs aren't charged per file), with the platform
  usage paid by the user. The cost gate showed a per-image price loses on large photos, because storage grows with
  bytes (up to 25 MB per image, and the ZIP stores everything twice). $7 is the leader's lowest paid tier: one flat
  price, because the SDK can't charge tier-priced custom events. Replaces the per-image $3.00 per 1,000 of the
  unreleased build. Competitors: onescales $70 per 1,000 on the free tier, $10 / $8 / $7 on paid tiers, with
  platform usage also passed to the user; getascraper $0.89, thirdwatch $1.00, automation-lab $0.79, logiover $7.50.
