Wayback Machine Snapshots and Archived URLs avatar

Wayback Machine Snapshots and Archived URLs

Pricing

$0.50 / 1,000 archive row saveds

Go to Apify Store
Wayback Machine Snapshots and Archived URLs

Wayback Machine Snapshots and Archived URLs

Get the Wayback Machine capture history of any URL, every archived URL of a domain, or the archived copy closest to a date. Returns capture time, status, content type, size, digest and ready links to the archived and raw copy. Uses the Internet Archive's public CDX API.

Pricing

$0.50 / 1,000 archive row saveds

Rating

0.0

(0)

Developer

Hay Equipos

Hay Equipos

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 hours ago

Last modified

Share

See what a website looked like and how it changed, straight from the Internet Archive. Give the actor URLs or domains and choose one of three jobs:

  1. Snapshots: the full capture history of a page, with date, HTTP status, content type, size and a content fingerprint, optionally thinned to one capture per day, month or year, or to captures where the content actually changed.
  2. Archived URLs: every distinct URL the Wayback Machine holds under a domain or path, with its first capture. Great for recovering old pages, finding lost content and redirect planning.
  3. Closest: the one archived copy nearest to a date, for "what did this page say on 15 January 2020?"

Every row comes with two ready links: the normal Wayback page and the raw archived copy without the Wayback toolbar, which is what you want to feed an AI model or a parser.

The actor uses the Internet Archive's own public CDX and availability APIs. No page scraping, no login, no proxies. Requests are paced one at a time with backoff, as the Archive asks.

What you can use it for

  • SEO and site migrations: list every URL a domain ever had, find old pages that still earn links, and build redirect maps.
  • Expired domain research: check what a domain hosted before you buy it.
  • Competitor and pricing history: see when a competitor's page changed and open each version.
  • AI and research datasets: collect archived copies of pages from a given date range for training, evaluation or analysis.
  • Evidence and compliance: find the capture of a terms page or a claim nearest to a given date.
  • Content recovery: find lost blog posts and documents, including PDFs, with the content type filter.

Input

FieldWhat it doesDefault
URLs or domainsPages, domains or path prefixes, one per linerequired
What to getSnapshots, Archived URLs, or ClosestSnapshots
From date, To date2019, 2019-06 or 2019-06-30none
Keep one snapshot perEvery capture, hour, day, month, year, or content changeevery capture
Only successful capturesKeep HTTP 200 captures onlyoff
Content type filterFor example text/html or application/pdfnone
Include subdomainsArchived URLs mode: also list blog.example.com and similaroff
Closest to dateClosest mode: the target date (empty means the most recent)none
Maximum rows per URL or domainRows come oldest first500
Maximum rows in totalStop after this many rows5,000

Example input, one capture per year for a home page:

{
"urls": ["apify.com"],
"mode": "snapshots",
"from": "2015",
"to": "2025",
"collapse": "year",
"onlySuccessful": true
}

Example input, every archived HTML page under a docs path:

{
"urls": ["crawlee.dev/docs/"],
"mode": "urls",
"mimeType": "text/html",
"onlySuccessful": true,
"maxRowsPerInput": 2000
}

Output

{
"input": "apify.com",
"mode": "snapshots",
"url": "http://apify.com:80/",
"capturedAt": "2015-02-16T02:12:25Z",
"timestamp": "20150216021225",
"statusCode": 200,
"mimeType": "text/html",
"digest": "3Z65GIU4US7YSBJ3HD2AU4WJP3YZNWYN",
"lengthBytes": 440,
"archiveUrl": "https://web.archive.org/web/20150216021225/http://apify.com:80/",
"rawArchiveUrl": "https://web.archive.org/web/20150216021225id_/http://apify.com:80/",
"scrapedAt": "2026-09-27T07:21:05.981Z"
}
  • digest is the Archive's fingerprint of the captured content. Two captures with the same digest are identical.
  • lengthBytes is the compressed size stored by the Archive.
  • In Archived URLs mode each row also has firstCapturedAt.
  • In Closest mode each row also has requestedDate.
  • Inputs with no captures are listed in RUN_SUMMARY in the run's key value store and cost nothing.

Pricing

Pay per event. No start fee, no subscription, no platform usage charged on top.

EventPrice
Archive row saved$0.0005 (50 cents per 1,000 rows)

Example: the yearly history of 100 home pages over 10 years is about 1,000 rows, about $0.50. Inputs with no captures are free. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

Limits

  • The Internet Archive's API can be slow (10 to 30 seconds for a large domain) and sometimes answers "busy". The actor waits and retries; very large domains take minutes.
  • Rows come oldest first. For recent captures only, set the From date.
  • The actor returns capture records and links, not the archived page content itself. Open rawArchiveUrl to get the page as captured.
  • Sites that asked the Archive to exclude them return no rows.

FAQ

Snapshots or Archived URLs, which do I need? Snapshots follow one exact page through time. Archived URLs list all the different pages under a domain or folder.

How do I see only real changes to a page? Set "Keep one snapshot per" to Content change. Consecutive identical captures are dropped.

Does it include subdomains? In Archived URLs mode, turn on Include subdomains.

Can an AI agent use it? Yes. It has clear inputs, returns small rows, has no start fee, and the raw archive links can be fetched directly by the agent.

Is this allowed? The CDX API is the Internet Archive's public interface for exactly this kind of lookup. The actor paces its requests and backs off when the Archive is busy.