Wayback Machine Snapshots and Archived URLs
Pricing
$0.50 / 1,000 archive row saveds
Wayback Machine Snapshots and Archived URLs
Get the Wayback Machine capture history of any URL, every archived URL of a domain, or the archived copy closest to a date. Returns capture time, status, content type, size, digest and ready links to the archived and raw copy. Uses the Internet Archive's public CDX API.
Pricing
$0.50 / 1,000 archive row saveds
Rating
0.0
(0)
Developer
Hay Equipos
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 hours ago
Last modified
Categories
Share
See what a website looked like and how it changed, straight from the Internet Archive. Give the actor URLs or domains and choose one of three jobs:
- Snapshots: the full capture history of a page, with date, HTTP status, content type, size and a content fingerprint, optionally thinned to one capture per day, month or year, or to captures where the content actually changed.
- Archived URLs: every distinct URL the Wayback Machine holds under a domain or path, with its first capture. Great for recovering old pages, finding lost content and redirect planning.
- Closest: the one archived copy nearest to a date, for "what did this page say on 15 January 2020?"
Every row comes with two ready links: the normal Wayback page and the raw archived copy without the Wayback toolbar, which is what you want to feed an AI model or a parser.
The actor uses the Internet Archive's own public CDX and availability APIs. No page scraping, no login, no proxies. Requests are paced one at a time with backoff, as the Archive asks.
What you can use it for
- SEO and site migrations: list every URL a domain ever had, find old pages that still earn links, and build redirect maps.
- Expired domain research: check what a domain hosted before you buy it.
- Competitor and pricing history: see when a competitor's page changed and open each version.
- AI and research datasets: collect archived copies of pages from a given date range for training, evaluation or analysis.
- Evidence and compliance: find the capture of a terms page or a claim nearest to a given date.
- Content recovery: find lost blog posts and documents, including PDFs, with the content type filter.
Input
| Field | What it does | Default |
|---|---|---|
| URLs or domains | Pages, domains or path prefixes, one per line | required |
| What to get | Snapshots, Archived URLs, or Closest | Snapshots |
| From date, To date | 2019, 2019-06 or 2019-06-30 | none |
| Keep one snapshot per | Every capture, hour, day, month, year, or content change | every capture |
| Only successful captures | Keep HTTP 200 captures only | off |
| Content type filter | For example text/html or application/pdf | none |
| Include subdomains | Archived URLs mode: also list blog.example.com and similar | off |
| Closest to date | Closest mode: the target date (empty means the most recent) | none |
| Maximum rows per URL or domain | Rows come oldest first | 500 |
| Maximum rows in total | Stop after this many rows | 5,000 |
Example input, one capture per year for a home page:
{"urls": ["apify.com"],"mode": "snapshots","from": "2015","to": "2025","collapse": "year","onlySuccessful": true}
Example input, every archived HTML page under a docs path:
{"urls": ["crawlee.dev/docs/"],"mode": "urls","mimeType": "text/html","onlySuccessful": true,"maxRowsPerInput": 2000}
Output
{"input": "apify.com","mode": "snapshots","url": "http://apify.com:80/","capturedAt": "2015-02-16T02:12:25Z","timestamp": "20150216021225","statusCode": 200,"mimeType": "text/html","digest": "3Z65GIU4US7YSBJ3HD2AU4WJP3YZNWYN","lengthBytes": 440,"archiveUrl": "https://web.archive.org/web/20150216021225/http://apify.com:80/","rawArchiveUrl": "https://web.archive.org/web/20150216021225id_/http://apify.com:80/","scrapedAt": "2026-09-27T07:21:05.981Z"}
digestis the Archive's fingerprint of the captured content. Two captures with the same digest are identical.lengthBytesis the compressed size stored by the Archive.- In Archived URLs mode each row also has
firstCapturedAt. - In Closest mode each row also has
requestedDate. - Inputs with no captures are listed in
RUN_SUMMARYin the run's key value store and cost nothing.
Pricing
Pay per event. No start fee, no subscription, no platform usage charged on top.
| Event | Price |
|---|---|
| Archive row saved | $0.0005 (50 cents per 1,000 rows) |
Example: the yearly history of 100 home pages over 10 years is about 1,000 rows, about $0.50. Inputs with no captures are free. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.
Limits
- The Internet Archive's API can be slow (10 to 30 seconds for a large domain) and sometimes answers "busy". The actor waits and retries; very large domains take minutes.
- Rows come oldest first. For recent captures only, set the From date.
- The actor returns capture records and links, not the archived page content itself. Open
rawArchiveUrlto get the page as captured. - Sites that asked the Archive to exclude them return no rows.
FAQ
Snapshots or Archived URLs, which do I need? Snapshots follow one exact page through time. Archived URLs list all the different pages under a domain or folder.
How do I see only real changes to a page? Set "Keep one snapshot per" to Content change. Consecutive identical captures are dropped.
Does it include subdomains? In Archived URLs mode, turn on Include subdomains.
Can an AI agent use it? Yes. It has clear inputs, returns small rows, has no start fee, and the raw archive links can be fetched directly by the agent.
Is this allowed? The CDX API is the Internet Archive's public interface for exactly this kind of lookup. The actor paces its requests and backs off when the Archive is busy.