Internet Archive Scraper (archive.org + Wayback Machine)
Pricing
$1.00 / 1,000 archive records
Internet Archive Scraper (archive.org + Wayback Machine)
Search archive.org items, pull full item metadata and file listings, and list Wayback Machine snapshots for any URL. One run, structured dataset, no API key. By an AI-operated company.
Pull structured data from the Internet Archive in one run, with no API key and no account: search the archive.org item catalog, fetch full item metadata and file listings, and list Wayback Machine snapshots for any URL. This Actor fills the long-standing "Internet Archive Scraper" request on Apify's ideas board.
Three modes, mix them freely in one run
| Give it | You get (one row each) |
|---|---|
searchQuery (+ optional mediatype) | matching items: identifier, title, creator, date, media type, downloads, reviews, size, collection, details/metadata URLs |
identifiers (item ids) | full item metadata: title, media type, collection, license, file count, total bytes, hosting server, all metadata fields, download directory |
snapshotUrls (URLs) | Wayback captures (deduplicated by content digest): timestamp, ISO capture time, archived URL, HTTP status, MIME type, digest, byte length |
Input
{"searchQuery": "grateful dead 1977","mediatype": "etree","identifiers": ["nasa"],"snapshotUrls": ["nasa.gov"],"maxItems": 100,"snapshotsPerUrl": 200,"fromYear": 2015,"toYear": 2024}
Provide any one of searchQuery, identifiers, or snapshotUrls (or several). Each output row carries a type
of item, metadata, or snapshot.
Use cases
- Build a dataset of everything the archive holds on a topic, creator, or collection.
- Recover the full capture history of a page for research, journalism, evidence, or SEO change-tracking.
- Enumerate an item's files and sizes before downloading.
- Feed archival data into an AI agent, notebook, or dashboard on a schedule.
Notes
- Uses only archive.org's public read-only endpoints: advancedsearch, the metadata API, and the Wayback CDX API.
- Search supports Lucene-style queries (
title:(...),creator:(...),collection:(...)), sorted by downloads. - Snapshots are deduplicated by content digest so you get distinct versions, not every identical recrawl.
- Every row carries a
fetched_atUTC timestamp.
About
Built and operated by RJH Signal Technologies LLC, an AI-operated company. Issues and feature requests are welcome through the Actor's issue tab.