Batch Image Downloader — URLs & Pages to KV
Pricing
from $2.50 / 1,000 image downloadeds
Batch Image Downloader — URLs & Pages to KV
Download image URL lists or extract img/og:image from pages into KV store with metadata (contentType, bytes, sha256, dims). Default under 1 GB; failed/empty free. For RAG/archives — not spam scraping.
Pricing
from $2.50 / 1,000 image downloadeds
Rating
0.0
(0)
Developer
新世紀書僮
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
12 hours ago
Last modified
Categories
Share
Download image URL lists (or extract img / og:image from pages) into Apify KV store with metadata for RAG and archives.
Give this Actor a list of image URLs, HTML page URLs, or both. It fetches each image, stores the bytes in the default key-value store, and writes a dataset row with contentType, bytes, sha256, and dimensions when readable. Failed or empty responses are recorded free by default. Positioning: batch media assets for Agent / dataset pipelines — not spam scraping.
What you get
- Direct image URL batch download into KV (
img-<sha256>.<ext>by default) - Optional page extraction:
<img src|srcset>,<picture>,og:image,twitter:image - Metadata: content type, byte size, SHA-256, width/height (via Pillow when possible)
- Optional WebP / JPEG / PNG re-encode (extra PPE event)
- Duplicate content skip by hash (free duplicate rows)
- Default 512 MB memory; failed / 4xx / empty / non-image free unless you opt in
Measured results
Cloud benches on build 0.1.2, 512 MB (Asia/Taipei 2026-10-02):
| Run | Input | Stored | Peak RSS | Platform cost |
|---|---|---|---|---|
| smoke | 3 URLs (2 ok / 1×404 free) | 2 | 87 MB | ~$0.00048 |
| bench | 12 URLs (10 ok / 2 err free) | 10 | 92 MB | |
| pages | w3.org/Standards → 9 imgs | 9 | 93 MB | ~$0.00080 |
Details: docs/PRICING.md.
Use cases
- Build an image asset pack from a sitemap-passed URL list for RAG multimodal pipelines
- Archive product or documentation images with integrity hashes
- Chain after sitemap discovery / URL status checker: only download live image URLs
How to use
{"imageUrls": [{ "url": "https://www.w3.org/Icons/w3c_home.png" },{ "url": "https://httpbin.org/image/png" }],"pageUrls": [],"maxImages": 50,"maxConcurrency": 6,"skipDuplicatesByHash": true}
Example success row:
{"sourceUrl": "https://www.w3.org/Icons/w3c_home.png","status": "success","contentType": "image/png","bytes": 2428,"width": 72,"height": 48,"sha256": "…","kvKey": "img-<sha256>.png","durationMs": 120}
Pricing (pay-per-event)
| Event | Price | When |
|---|---|---|
| Actor start | $0.001 / GB (1 GB min) | Once per run |
| Image downloaded (primary) | $0.0025 | Each successful KV store |
| Image converted | $0.001 | Extra when convertFormat succeeds |
| Failed / empty / non-image / duplicate | $0 | Default (chargeFailedImages=false) |
Example: 1,000 successful downloads ≈ $2.50 + start (before platform CU/storage). Failed items are never charged by default.
Known limits
- Does not log into sites or bypass paywalls / bot challenges
- SVG may lack width/height; dims come from Pillow for raster formats
- Very large pages:
maxImagesPerPagecaps extraction - Copyright / robots: you are responsible for lawful use of downloaded assets
- Not a “scrape any website’s images” spam tool — designed for owned or permitted URL lists
License
AGPL-3.0 (see LICENSE). Runtime stack: httpx (BSD), beautifulsoup4 (MIT), Pillow (MIT-CMU), Apify SDK.