Bulk Image Downloader: Crawl Sites, Dedupe Photos, Export ZIP avatar

Bulk Image Downloader: Crawl Sites, Dedupe Photos, Export ZIP

Pricing

from $0.67 / 1,000 images

Go to Apify Store
Bulk Image Downloader: Crawl Sites, Dedupe Photos, Export ZIP

Bulk Image Downloader: Crawl Sites, Dedupe Photos, Export ZIP

Download images from any website in bulk, with a bounded crawl across multiple pages from one seed URL, or fetch a list of direct links. Discovers via HTML, srcset, and image tags, dedupes by SHA-256, strips EXIF, and converts WebP to PNG. Export as ZIP, dataset, or S3. $2 per 1,000 images.

Pricing

from $0.67 / 1,000 images

Rating

0.0

(0)

Developer

GetAScraper

GetAScraper

Maintained by Community

Actor stats

0

Bookmarked

66

Total users

19

Monthly active users

4 days ago

Last modified

Share

πŸ–ΌοΈ Bulk Image Downloader: Crawl Sites, Dedupe Photos, Export ZIP

Get every image on a site, deduped and ready for your dataset
Crawl multiple pages from one seed URL, or paste a straight list of image links. Either way you get clean files, full source tracking, and no duplicates.
πŸ•ΈοΈ Crawls, not just fetches
Follows links across a site from one starting page.
🏷️ Full source tracking
Every image is tagged with exactly how it was found.
🧹 Clean before you download
Strips hidden photo metadata and converts formats.
πŸ’΅ Priced to use at scale
A fraction of the cost of other bulk downloaders.

This Actor is a generic image downloader for public URLs. Choose a preset, Quick ZIP, Website Inventory, or ML Dataset, then override only the controls you need. It finds images in page tags, srcset variants, Open Graph and Twitter Card tags, and lazy-loaded attributes, and can follow a bounded set of page links. Every attempted image gets a status, source tracking, retry history, and a clear error code when it cannot be downloaded.


πŸ’‘ What can you do with it?

  • I am an AI engineer building an image training set. I need every file deduped and tagged with exactly how it was found, not just a folder of pictures.
  • I am a developer running my own catalog scraper. I hand this Actor a list of image URLs and get back a ZIP plus a clean metadata file, instead of writing my own downloader.
  • I am an e-commerce operator tracking a competitor's catalog. I crawl their product pages on a schedule and compare image hashes to catch any photo swap.
  • I am an archivist or researcher. I pull every image from a story or site in one call, with sources kept separate using one ZIP per page.
  • I am a pipeline builder. I send the finished files straight to storage, or get a notification the moment a run finishes, instead of checking back by hand.

Store search terms

bulk image downloader, website image scraper, download images from webpage, image URL extractor, image dataset builder.


πŸš€ How to use it

STEP 1
Add your URLs
Paste webpages, direct image links, or both in one list.
STEP 2
Pick a preset
Quick ZIP, website inventory, or ML dataset, or set your own options.
STEP 3
Get your files
Download a dataset, a ZIP, or send straight to storage.

Every attempted image is logged with its status, so you always know what succeeded, what failed, and why. Pull the dataset as JSON, CSV, or Excel, or use the single-click ZIP download.


πŸ“₯ Input

FieldTypeRequiredDescription
urlsarrayYesList of URLs. Each can be a webpage (HTML is parsed for images) or a direct image link. Mix freely.
presetenumNoquickZip, websiteInventory, or mlDataset. Explicit fields override the preset.
modeenumNoauto (recommended, detects by extension), page (force HTML parse), or direct (force image URL).
includeSrcsetbooleanNoDiscover images from srcset, picture>source, and lazy data-src. Default true.
includeOgTagsbooleanNoDiscover Open Graph and Twitter Card images. Default true.
minWidthintegerNoSkip images narrower than this. Default 0.
minHeightintegerNoSkip images shorter than this. Default 0.
minSizeBytesintegerNoSkip images smaller than this. Filters tracking pixels. Default 0.
maxImagesPerUrlintegerNoCap images per source URL. Default 1000.
maxUrlsintegerNoCap total URLs processed. Default 10000.
crawlDepthintegerNoLink depth for bounded page crawling. 0 stays on supplied pages.
maxPagesintegerNoHard page cap per supplied URL.
urlGlobsarrayNoOptional URL patterns allowed during page crawling.
includeExtensionsarrayNoOptional image extension allow-list such as jpg, png, webp.
includeMimeTypesarrayNoOptional response MIME allow-list such as image/jpeg.
maxSizeBytesintegerNoOptional upper bound for downloaded image bytes.
dedupByHashbooleanNoCompute SHA-256 of each image body and skip duplicates. Default true.
stripExifbooleanNoRe-encode JPEGs without EXIF metadata. Default false.
convertFormatenumNonone, webp-to-png, or png-to-jpg. Default none.
filenamePatternstringNoTemplated filename using {slug}, {hash}, {ext}, {idx}, {source}. Default {slug}-{hash}.{ext}.
outputFormatarrayNodataset (always), kv-store (binaries), zip (single archive), zipPerUrl (one ZIP per source), s3 (upload to bucket), webhook (POST summary on completion).
s3BucketstringNoRequired when outputFormat includes s3. Uses standard AWS_* env vars for credentials.
webhookUrlstringNoURL to receive a JSON run summary on completion.
maxConcurrencyintegerNoMax parallel image downloads. Default 10.
downloadTimeoutMsintegerNoPer-image fetch timeout. Default 15000.
imageCheckMaxRetriesintegerNoRetries per failed image. Default 3.
proxyConfigurationobjectNoOptional proxy. Default off. Use residential if source sites are hotlink-protected.
failFastbooleanNoStop on first error. Default false.
debugLoggingbooleanNoVerbose per-image tracing. Default false.

πŸ“€ Output

The Actor pushes one row to the dataset per downloaded image. Binaries are written to the default key-value store under IMG-{filename}. Use the dataset's kv_url column to download each binary.

{
"filename": "picsum-photos-800x600-a1b2c3d4e5f67890.jpg",
"source_url": "https://example.com/gallery",
"image_url": "https://picsum.photos/800/600.jpg",
"kv_store_key": "IMG-picsum-photos-800x600-a1b2c3d4e5f67890.jpg",
"kv_url": "https://api.apify.com/v2/key-value-stores/abc/records/IMG-picsum-photos-800x600-a1b2c3d4e5f67890.jpg",
"content_type": "image/jpeg",
"size_bytes": 54321,
"width": 800,
"height": 600,
"format": "jpeg",
"sha256": "a1b2c3d4e5f6789012345678901234567890abcdef1234567890abcdef123456",
"is_duplicate": false,
"exif_stripped": false,
"from_srcset": true,
"from_picture_source": false,
"from_og_tag": false,
"from_twitter_tag": false,
"from_data_attr": false,
"from_direct_url": false,
"downloaded_at": "2026-06-20T12:34:56.000Z",
"duration_ms": 423,
"http_status": 200,
"status": "SUCCEEDED",
"ok": true,
"inputId": "catalog-home",
"sourceUrl": "https://example.com/gallery",
"retrievedAt": "2026-08-01T12:34:56.000Z",
"crawl_depth": 0,
"discovery_method": "srcset"
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

πŸ§ͺ Store examples

Quick ZIP

{
"preset": "quickZip",
"urls": [{ "url": "https://example.com/gallery", "inputId": "gallery" }]
}

Website inventory

{
"preset": "websiteInventory",
"urls": [{ "url": "https://example.com/catalog", "inputId": "catalog" }],
"urlGlobs": ["https://example.com/catalog/*"],
"includeExtensions": ["jpg", "png", "webp"]
}

Recurring ML dataset refresh

{
"preset": "mlDataset",
"urls": [{ "url": "https://example.com/products", "inputId": "products" }],
"crawlDepth": 1,
"maxPages": 25,
"dedupByHash": true
}

πŸ“‹ Output data fields

FieldDescription
filenameFinal filename (per filenamePattern).
source_urlThe page URL the image was discovered on (or its direct URL).
image_urlFinal resolved image URL (after srcset expansion, redirects).
kv_store_keyKey in the run's key-value store (IMG-...).
kv_urlSigned download URL for the binary (24-hour default).
content_typeMIME type (e.g. image/jpeg, image/webp).
size_bytesDownloaded size.
widthImage width in pixels.
heightImage height in pixels.
formatNormalized format: jpeg, png, webp, gif, svg, avif, bmp, ico, other.
sha256Content hash (when dedupByHash=true).
is_duplicateTrue if hash matched a previously-seen image in this run.
exif_strippedTrue if JPEG was re-encoded to remove EXIF.
from_srcsetTrue if discovered via srcset / picture / data-srcset.
from_picture_sourceTrue if discovered via <picture><source>.
from_og_tagTrue if discovered via <meta og:image>.
from_twitter_tagTrue if discovered via <meta twitter:image>.
from_data_attrTrue if discovered via lazy data-src / data-srcset.
from_direct_urlTrue if the URL was treated as a direct image (mode=direct/auto).
downloaded_atTimestamp of the download.
duration_msTime to fetch and process.
http_statusHTTP response code when the image server returned one.
errorPer-image error string when the attempt failed; omitted on a successful row.
status / okHonest success state for the image attempt. Failed attempts include an errorCode and omit unavailable image fields.
inputId / sourceUrl / retrievedAtStable input correlation, canonical source page, and retrieval timestamp.
finalUrl / retryCount / errorCodeRedirect destination and retry and error evidence when available.
alt / title / page_titleSource HTML metadata when published by the page.
crawl_depth / discovery_methodBounded crawl provenance.

πŸ’° Pricing

This Actor uses pay-per-result pricing. You are billed only for each image successfully downloaded and stored, with tiered volume discounts applied automatically as your usage grows. Failed or skipped images are never billed, an empty run costs nothing, and there are no subscriptions.


⭐ Enjoying Bulk Image Downloader?

⭐ ⭐ ⭐ ⭐ ⭐
Save hours of manually right-clicking and saving images one by one from any site.
A 5-star rating takes 10 seconds and helps other AI dataset builders and e-commerce operators find it. Your feedback also tells us what to build next.
β˜…Β Β Rate this Actor on Apify

πŸ› οΈ Tips and advanced options

  • Set includeSrcset to false if you only want the page's primary images. This skips lazy data-src and responsive variants, which is faster on heavy pages.
  • Use minSizeBytes to filter tracking pixels. A typical tracking pixel is under 1KB. Set minSizeBytes: 2000 to skip them.
  • Use minWidth and minHeight to focus on useful images. Set minWidth: 400 to skip thumbnails and avatars.
  • Pick the right output mode. zip for a single archive, zipPerUrl to keep source pages separated, s3 to push directly to your training bucket.
  • Pair with a catalog scraper. Run one of our catalog scrapers (REI, Gmarket, Wildberries) first, then feed the image URLs to this Actor for a complete e-commerce dataset.
  • Schedule weekly runs to refresh your image corpus. Most product catalogs update slowly, so daily is overkill.
  • Use SHA-256 dedup within a run. Hashes are stable for downstream reconciliation; cross-run history belongs in the dataset or an external state store.

❓ FAQ

Is this Actor legal to use? The Actor downloads images that are publicly accessible. You are responsible for ensuring your use case complies with the source site's Terms of Service and applicable copyright laws. Do not use the Actor to bypass access controls, scrape private content, or violate copyright.

Why does it work on any site? The Actor is generic. It fetches the URL you give it, parses the page for image tags, and downloads the images it finds. There is no per-site configuration.

Does it execute JavaScript? No. If a page builds its content with JavaScript after loading, this Actor will not see those images, since it only reads the page as first delivered. For those sites, get the image links first with a browser-based scraper, then paste that list in here using direct mode.

Do I need a proxy? No. Most public sites serve images to any client. Default useApifyProxy: false works perfectly. If your source site is hotlink-protected, set residential proxy as an opt-in via the proxyConfiguration field.

What is the largest image it can handle? Memory use scales with image size rather than file count, so a 50MB image is fine. A 500MB image may cause memory pressure on smaller container sizes.

Does the EXIF strip work on PNG or WebP? No, EXIF strip is JPEG-only. PNG metadata stripping is a future feature.

How is billing calculated? The stored-image event is only charged after the binary is successfully stored. Failed image attempts stay in the dataset with an error code and are not billed as successful image events.

Can I get a single ZIP of all images? Yes. Set outputFormat: ['dataset', 'kv-store', 'zip']. The ZIP is written to OUT-images.zip and is also linked in the dataset summary.

Can I push directly to S3? Yes. Set outputFormat: ['dataset', 's3'], fill in s3Bucket, and set AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, and AWS_REGION as Apify Secrets. Each image uploads to s3://{bucket}/images/{filename}.

Can I get a notification when a run finishes? Yes. Set outputFormat: ['dataset', 'webhook'] and fill in webhookUrl. The Actor sends a run summary with counts, errors, and total size to that URL when the run finishes.


πŸ›‘οΈ Disclaimers and support

  • Disclaimer: This Actor retrieves publicly accessible images. Make sure your usage complies with the source site's terms of service and applicable copyright laws. The Actor is a generic utility and does not bypass authentication, paywalls, or access controls.
  • Support: Open an issue from the Issues tab for bug reports or feature requests. Custom scrapers and integration help are available on request.

πŸ”— Other actors