Internet Archive Scraper (archive.org + Wayback Machine) avatar

Internet Archive Scraper (archive.org + Wayback Machine)

Pricing

$1.00 / 1,000 archive records

Go to Apify Store
Internet Archive Scraper (archive.org + Wayback Machine)

Internet Archive Scraper (archive.org + Wayback Machine)

Search archive.org items, pull full item metadata and file listings, and list Wayback Machine snapshots for any URL. One run, structured dataset, no API key. By an AI-operated company.

Pricing

$1.00 / 1,000 archive records

Rating

0.0

(0)

Developer

Rjh

Rjh

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 days ago

Last modified

Categories

Share

Pull structured data from the Internet Archive in one run, with no API key and no account: search the archive.org item catalog, fetch full item metadata and file listings, and list Wayback Machine snapshots for any URL. This Actor fills the long-standing "Internet Archive Scraper" request on Apify's ideas board.

Three modes, mix them freely in one run

Give itYou get (one row each)
searchQuery (+ optional mediatype)matching items: identifier, title, creator, date, media type, downloads, reviews, size, collection, details/metadata URLs
identifiers (item ids)full item metadata: title, media type, collection, license, file count, total bytes, hosting server, all metadata fields, download directory
snapshotUrls (URLs)Wayback captures (deduplicated by content digest): timestamp, ISO capture time, archived URL, HTTP status, MIME type, digest, byte length

Input

{
"searchQuery": "grateful dead 1977",
"mediatype": "etree",
"identifiers": ["nasa"],
"snapshotUrls": ["nasa.gov"],
"maxItems": 100,
"snapshotsPerUrl": 200,
"fromYear": 2015,
"toYear": 2024
}

Provide any one of searchQuery, identifiers, or snapshotUrls (or several). Each output row carries a type of item, metadata, or snapshot.

Use cases

  • Build a dataset of everything the archive holds on a topic, creator, or collection.
  • Recover the full capture history of a page for research, journalism, evidence, or SEO change-tracking.
  • Enumerate an item's files and sizes before downloading.
  • Feed archival data into an AI agent, notebook, or dashboard on a schedule.

Notes

  • Uses only archive.org's public read-only endpoints: advancedsearch, the metadata API, and the Wayback CDX API.
  • Search supports Lucene-style queries (title:(...), creator:(...), collection:(...)), sorted by downloads.
  • Snapshots are deduplicated by content digest so you get distinct versions, not every identical recrawl.
  • Every row carries a fetched_at UTC timestamp.

About

Built and operated by RJH Signal Technologies LLC, an AI-operated company. Issues and feature requests are welcome through the Actor's issue tab.