Wayback Machine Snapshot Checker (Internet Archive) avatar

Wayback Machine Snapshot Checker (Internet Archive)

Pricing

from $1.00 / 1,000 results

Go to Apify Store
Wayback Machine Snapshot Checker (Internet Archive)

Wayback Machine Snapshot Checker (Internet Archive)

Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.

Pricing

from $1.00 / 1,000 results

Rating

0.0

(0)

Developer

돈벼락

돈벼락

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

4 days ago

Last modified

Categories

Share

Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.

What does Wayback Machine Snapshot Checker (Internet Archive) do?

Check if and when a URL was archived by the Wayback Machine, with the closest snapshot link.

What can you use it for?

  • Check whether a page has ever been archived before it disappears or changes.
  • Find the oldest or most recent snapshot of a competitor's page for research.
  • Verify that your own important pages are being preserved by the Wayback Machine.
  • Get a direct archive.org link to cite an old version of a page.

Why use this Actor?

  • Fast and cheap: lightweight HTTP-based Actor with no browser, so runs finish in seconds and cost very little.
  • No API keys or accounts needed for the data source (see notes below).
  • Clean, structured JSON that you can export as CSV, Excel, JSON or XML, or pull through the Apify API and integrations.
  • Scheduling, webhooks and integrations: run it daily and send results to Google Sheets, Slack, Zapier, Make, n8n or your own API.

Input

Configure the run in the Input tab (all fields have sensible defaults).

FieldTypeRequiredDefaultDescription
startUrlsarray of URLsyes-URLs to check (one result per URL).
timestampstringno-YYYYMMDD (or a shorter prefix like YYYY). Leave empty to get the most recent snapshot.
includeFullHistorybooleannofalseAlso fetch every known snapshot timestamp for each URL via the CDX API (slower; degrades gracefully if the Internet Archive CDX service is temporarily down).
maxSnapshotsPerUrlintegerno100Only used when "Include full snapshot history" is on.

Example input

{
"startUrls": [
{
"url": "https://example.com"
},
{
"url": "https://apify.com"
}
]
}

Output

Results are stored in the default dataset. Download them from the Output tab or via the API.

Example result

{
"url": "https://example.com/",
"archived": true,
"closestSnapshotDate": "2026-09-23T01:54:39Z",
"closestSnapshotUrl": "http://web.archive.org/web/20260923015439/https://example.com/",
"snapshotCount": null,
"snapshots": null,
"historyNote": "Not requested (turn on \"Include full snapshot history\").",
"fetchedAt": "2026-09-23T02:25:52.393Z"
}

Output fields

FieldTypeExample
urlstring"https://example.com/"
archivedbooleantrue
closestSnapshotDatestring"2026-09-23T01:54:39Z"
closestSnapshotUrlstring"http://web.archive.org/web/20260923015439/https://exampl...
snapshotCountnullnull
snapshotsnullnull
historyNotestring"Not requested (turn on "Include full snapshot history")."
fetchedAtstring"2026-09-23T02:25:52.393Z"

How much does it cost?

The Actor is billed by Apify according to the pricing shown on the Pricing tab of this page. Runs are small and fast, and the Apify Free plan includes monthly platform credits so you can try it at no cost. You can set a maximum cost per run, and the Actor stops cleanly when that limit is reached.

Notes and limitations

  • Uses the free, official archive.org "available" API for the closest snapshot, and the CDX API (optional) for the full snapshot history.
  • The Internet Archive occasionally has short outages; if the optional full-history lookup fails, the row still includes the closest-snapshot result with a note, instead of failing the whole run.
  • Only publicly archived pages are covered - pages blocked by robots.txt at crawl time, or excluded by a takedown request, will not appear.

Tips

  • Start with a small test run, then scale up the input.
  • Use Schedules to automate recurring runs and Webhooks or Integrations to deliver the data where you need it.
  • Call the Actor from your code with the Apify API or the official JavaScript and Python clients.

FAQ

What does archived: false mean?

The Wayback Machine has no snapshot of that exact URL. Try without a trailing slash or query string, since those are indexed as different URLs.

Can I get every snapshot, not just the closest?

Yes, turn on "Include full snapshot history" to add every known timestamp for that URL (subject to the occasional Internet Archive outage noted above).

Feedback

Found a bug or need an extra field? Open an issue from the Issues tab of this Actor and it will be looked at quickly.

More tools from the same author