Batch Image Downloader — URLs & Pages to KV avatar

Batch Image Downloader — URLs & Pages to KV

Pricing

from $2.50 / 1,000 image downloadeds

Go to Apify Store
Batch Image Downloader — URLs & Pages to KV

Batch Image Downloader — URLs & Pages to KV

Download image URL lists or extract img/og:image from pages into KV store with metadata (contentType, bytes, sha256, dims). Default under 1 GB; failed/empty free. For RAG/archives — not spam scraping.

Pricing

from $2.50 / 1,000 image downloadeds

Rating

0.0

(0)

Developer

新世紀書僮

新世紀書僮

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

12 hours ago

Last modified

Share

Download image URL lists (or extract img / og:image from pages) into Apify KV store with metadata for RAG and archives.

Give this Actor a list of image URLs, HTML page URLs, or both. It fetches each image, stores the bytes in the default key-value store, and writes a dataset row with contentType, bytes, sha256, and dimensions when readable. Failed or empty responses are recorded free by default. Positioning: batch media assets for Agent / dataset pipelines — not spam scraping.

What you get

  • Direct image URL batch download into KV (img-<sha256>.<ext> by default)
  • Optional page extraction: <img src|srcset>, <picture>, og:image, twitter:image
  • Metadata: content type, byte size, SHA-256, width/height (via Pillow when possible)
  • Optional WebP / JPEG / PNG re-encode (extra PPE event)
  • Duplicate content skip by hash (free duplicate rows)
  • Default 512 MB memory; failed / 4xx / empty / non-image free unless you opt in

Measured results

Cloud benches on build 0.1.2, 512 MB (Asia/Taipei 2026-10-02):

RunInputStoredPeak RSSPlatform cost
smoke3 URLs (2 ok / 1×404 free)287 MB~$0.00048
bench12 URLs (10 ok / 2 err free)1092 MB$0.00082 ($0.000082 / image)
pagesw3.org/Standards → 9 imgs993 MB~$0.00080

Details: docs/PRICING.md.

Use cases

  • Build an image asset pack from a sitemap-passed URL list for RAG multimodal pipelines
  • Archive product or documentation images with integrity hashes
  • Chain after sitemap discovery / URL status checker: only download live image URLs

How to use

{
"imageUrls": [
{ "url": "https://www.w3.org/Icons/w3c_home.png" },
{ "url": "https://httpbin.org/image/png" }
],
"pageUrls": [],
"maxImages": 50,
"maxConcurrency": 6,
"skipDuplicatesByHash": true
}

Example success row:

{
"sourceUrl": "https://www.w3.org/Icons/w3c_home.png",
"status": "success",
"contentType": "image/png",
"bytes": 2428,
"width": 72,
"height": 48,
"sha256": "…",
"kvKey": "img-<sha256>.png",
"durationMs": 120
}

Pricing (pay-per-event)

EventPriceWhen
Actor start$0.001 / GB (1 GB min)Once per run
Image downloaded (primary)$0.0025Each successful KV store
Image converted$0.001Extra when convertFormat succeeds
Failed / empty / non-image / duplicate$0Default (chargeFailedImages=false)

Example: 1,000 successful downloads ≈ $2.50 + start (before platform CU/storage). Failed items are never charged by default.

Known limits

  • Does not log into sites or bypass paywalls / bot challenges
  • SVG may lack width/height; dims come from Pillow for raster formats
  • Very large pages: maxImagesPerPage caps extraction
  • Copyright / robots: you are responsible for lawful use of downloaded assets
  • Not a “scrape any website’s images” spam tool — designed for owned or permitted URL lists

License

AGPL-3.0 (see LICENSE). Runtime stack: httpx (BSD), beautifulsoup4 (MIT), Pillow (MIT-CMU), Apify SDK.