Wayback Machine Snapshot Checker (Internet Archive)
Pricing
from $1.00 / 1,000 results
Wayback Machine Snapshot Checker (Internet Archive)
Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
돈벼락
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
4 days ago
Last modified
Categories
Share
Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.
What does Wayback Machine Snapshot Checker (Internet Archive) do?
Check if and when a URL was archived by the Wayback Machine, with the closest snapshot link.
What can you use it for?
- Check whether a page has ever been archived before it disappears or changes.
- Find the oldest or most recent snapshot of a competitor's page for research.
- Verify that your own important pages are being preserved by the Wayback Machine.
- Get a direct archive.org link to cite an old version of a page.
Why use this Actor?
- Fast and cheap: lightweight HTTP-based Actor with no browser, so runs finish in seconds and cost very little.
- No API keys or accounts needed for the data source (see notes below).
- Clean, structured JSON that you can export as CSV, Excel, JSON or XML, or pull through the Apify API and integrations.
- Scheduling, webhooks and integrations: run it daily and send results to Google Sheets, Slack, Zapier, Make, n8n or your own API.
Input
Configure the run in the Input tab (all fields have sensible defaults).
| Field | Type | Required | Default | Description |
|---|---|---|---|---|
startUrls | array of URLs | yes | - | URLs to check (one result per URL). |
timestamp | string | no | - | YYYYMMDD (or a shorter prefix like YYYY). Leave empty to get the most recent snapshot. |
includeFullHistory | boolean | no | false | Also fetch every known snapshot timestamp for each URL via the CDX API (slower; degrades gracefully if the Internet Archive CDX service is temporarily down). |
maxSnapshotsPerUrl | integer | no | 100 | Only used when "Include full snapshot history" is on. |
Example input
{"startUrls": [{"url": "https://example.com"},{"url": "https://apify.com"}]}
Output
Results are stored in the default dataset. Download them from the Output tab or via the API.
Example result
{"url": "https://example.com/","archived": true,"closestSnapshotDate": "2026-09-23T01:54:39Z","closestSnapshotUrl": "http://web.archive.org/web/20260923015439/https://example.com/","snapshotCount": null,"snapshots": null,"historyNote": "Not requested (turn on \"Include full snapshot history\").","fetchedAt": "2026-09-23T02:25:52.393Z"}
Output fields
| Field | Type | Example |
|---|---|---|
url | string | "https://example.com/" |
archived | boolean | true |
closestSnapshotDate | string | "2026-09-23T01:54:39Z" |
closestSnapshotUrl | string | "http://web.archive.org/web/20260923015439/https://exampl... |
snapshotCount | null | null |
snapshots | null | null |
historyNote | string | "Not requested (turn on "Include full snapshot history")." |
fetchedAt | string | "2026-09-23T02:25:52.393Z" |
How much does it cost?
The Actor is billed by Apify according to the pricing shown on the Pricing tab of this page. Runs are small and fast, and the Apify Free plan includes monthly platform credits so you can try it at no cost. You can set a maximum cost per run, and the Actor stops cleanly when that limit is reached.
Notes and limitations
- Uses the free, official archive.org "available" API for the closest snapshot, and the CDX API (optional) for the full snapshot history.
- The Internet Archive occasionally has short outages; if the optional full-history lookup fails, the row still includes the closest-snapshot result with a note, instead of failing the whole run.
- Only publicly archived pages are covered - pages blocked by robots.txt at crawl time, or excluded by a takedown request, will not appear.
Tips
- Start with a small test run, then scale up the input.
- Use Schedules to automate recurring runs and Webhooks or Integrations to deliver the data where you need it.
- Call the Actor from your code with the Apify API or the official JavaScript and Python clients.
FAQ
What does archived: false mean?
The Wayback Machine has no snapshot of that exact URL. Try without a trailing slash or query string, since those are indexed as different URLs.
Can I get every snapshot, not just the closest?
Yes, turn on "Include full snapshot history" to add every known timestamp for that URL (subject to the occasional Internet Archive outage noted above).
Feedback
Found a bug or need an extra field? Open an issue from the Issues tab of this Actor and it will be looked at quickly.
More tools from the same author
- Sitemap URL Extractor: All URLs from XML Sitemaps - Extract every URL (with lastmod, priority, changefreq) from any website sitemap, index or .gz file.
- Broken Link Checker: Find 404s & Dead Links on Any Website - Crawl a site and list every broken internal and external link with the pages that contain it.
- QR Code Generator: URLs, Text & More (Bulk, PNG/SVG) - Generate QR codes in bulk as PNG or SVG - runs entirely locally, no external API, never breaks.