Wayback Machine Snapshot & Page Change Tracker avatar

Wayback Machine Snapshot & Page Change Tracker

Pricing

from $1.80 / 1,000 results

Go to Apify Store
Wayback Machine Snapshot & Page Change Tracker

Wayback Machine Snapshot & Page Change Tracker

Wayback Machine Snapshot & Page Change Tracker lists every Internet Archive capture of a URL and diffs the visible text of the oldest vs. newest snapshot in range — one scored change row per URL, or a full snapshot list.

Pricing

from $1.80 / 1,000 results

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

3

Total users

2

Monthly active users

2 days ago

Last modified

Share

What is Wayback Machine Snapshot & Page Change Tracker?

Wayback Machine Snapshot & Page Change Tracker is an Apify Actor that queries the Internet Archive's Wayback Machine for every capture of a URL and turns that history into structured data. In diff mode (default) it fetches the oldest and newest archived snapshot in your date range and returns one row per URL describing exactly what changed: title, word count, a line-based text diff with sample added/removed lines, and link churn. In snapshots mode it simply lists every capture, newest first, so you can pick specific dates yourself. No API key, no proxy, no headless browser — just the Wayback Machine's own CDX index and its id_ raw-snapshot endpoint.

Why use Wayback Machine Snapshot & Page Change Tracker?

  • Pricing and terms history — prove what a vendor's pricing page or ToS said on a given date, for legal, compliance or renewal-negotiation purposes.
  • SEO history — see what a page's title, headings and copy looked like before a ranking change or a competitor's redesign.
  • Competitor and brand monitoring — track when a competitor rewrote their homepage or landing page, without visiting the site yourself.
  • Content audits — spot-check how much a page has actually changed since it was last reviewed, with a percentage score instead of a manual read-through.

How to use Wayback Machine Snapshot & Page Change Tracker

  1. Paste the page URLs you want history for into URLs, e.g. https://apify.com/pricing.
  2. Leave Mode on diff to compare the oldest and newest capture, or switch to snapshots to list every capture instead.
  3. Optionally set From date / To date (e.g. 2023-01-01) to restrict the range considered.
  4. Click Start, then export the results as JSON, CSV, Excel or HTML from the Output tab.

Example input

{
"urls": ["https://apify.com/pricing", "https://example.com"],
"mode": "diff",
"maxConcurrency": 3
}

Example output

Diff mode (one row per URL):

{
"url": "https://apify.com/pricing",
"snapshotCount": 165,
"firstSnapshotAt": "2017-10-27T07:20:08.000Z",
"lastSnapshotAt": "2026-09-07T22:14:50.000Z",
"oldSnapshotUrl": "https://web.archive.org/web/20171027072008/https://www.apify.com/pricing",
"newSnapshotUrl": "https://web.archive.org/web/20260907221450/https://apify.com/pricing",
"oldTitle": "Pricing",
"newTitle": "Apify pricing - flexible plan + pay as you go · Apify",
"titleChanged": true,
"oldWordCount": 570,
"newWordCount": 2084,
"changedPercent": 95.3,
"addedLines": ["Apify pricing - flexible plan + pay as you go · Apify", "Skip to content", "…"],
"removedLines": ["Pricing", "Flexible pricing. Free for developers,", "…"],
"addedLinkCount": 187,
"removedLinkCount": 9,
"error": null,
"scrapedAt": "2026-09-12T17:30:00.000Z"
}

Snapshots mode (one row per capture):

{
"url": "https://example.com/",
"timestamp": "20260912162353",
"snapshotAt": "2026-09-12T16:23:53.000Z",
"snapshotUrl": "https://web.archive.org/web/20260912162353/https://example.com/",
"digest": "PKUMGV5XIIUJG5CKD4HZMCRWKUMVP5S6",
"lengthBytes": 1048,
"error": null,
"scrapedAt": "2026-09-12T17:30:00.000Z"
}

You can download the dataset in various formats such as JSON, HTML, CSV, or Excel.

Data table

FieldTypeDescription
urlstringThe input URL checked
snapshotCount, firstSnapshotAt, lastSnapshotAtnumber, dateDiff mode: how many captures exist in range and when the first/last were made
oldSnapshotUrl, newSnapshotUrllinkWayback pages for the oldest and newest capture compared
oldTitle, newTitle, titleChangedstring, boolean<title> of each capture and whether it differs
oldWordCount, newWordCountnumberVisible-text word counts
changedPercentnumber0-100 line-based change score between the two captures
addedLines, removedLinesarrayUp to 50 sample lines added / removed
addedLinkCount, removedLinkCountnumberLinks present only in the new / only in the old capture
timestamp, snapshotAt, snapshotUrl, digest, lengthBytes-Snapshots mode: one capture's Wayback timestamp, ISO date, URL, content hash and size
error, scrapedAtstring, datePer-URL failure reason (never throws the whole run) and when the row was produced

Input parameters

ParameterTypeDefaultDescription
urlsarray["https://apify.com/pricing"]Pages to look up, one result group per URL
modestringdiffdiff compares oldest vs. newest capture; snapshots lists them all
fromstring(none)Only captures on/after this date, e.g. 2023-01-01
tostring(none)Only captures on/before this date, e.g. 2024-06-30
maxSnapshotsinteger50Snapshots mode: max captures returned per URL (1-1000)
maxConcurrencyinteger3URLs processed in parallel (the run also self-throttles to ~1 req/s)

Pricing

Wayback Machine Snapshot & Page Change Tracker uses pay-per-event pricing: $0.003 per result row, i.e. $3 per 1,000 rows, plus a negligible actor-start fee. A diff-mode URL costs one row (two HTML fetches under the hood); snapshots mode charges one row per capture returned, capped by Max snapshots per URL. Set Maximum cost per run and the Actor trims the URL list to what the budget covers.

Wayback Machine Snapshot & Page Change Tracker vs. manually browsing web.archive.org

Clicking through web.archive.org's calendar UI one date at a time does not scale past a handful of pages, and comparing two captures by eye misses small copy changes. This Actor queries the same public CDX index Wayback's own UI uses, but returns a flat, scored dataset for as many URLs as you give it — with a diff percentage, sample changed lines and link churn ready for a spreadsheet, an alert, or an LLM prompt.

Using Wayback Machine Snapshot & Page Change Tracker with AI agents and MCP

This Actor is pay-per-event with limited permissions — the two requirements for an Actor to be callable through the Apify MCP server at mcp.apify.com. An agent passes urls and gets back a structured before/after diff it can summarize or act on, without ever touching a browser or writing scraping code. The same run works from n8n, Make, Zapier and LangChain through Apify's integrations.

FAQ

Why did I get error: "Only one snapshot in range"? The Wayback Machine has archived that URL only once inside your from/to window, so there is nothing to compare it against — widen the date range or drop it.

Why did I get a 429 or 503 error? The Wayback Machine rate-limits aggressively under load. This Actor already self-throttles to about one request per second and retries twice; lowering Max concurrency to 1-2 or simply re-running later usually clears it.

Does this see JavaScript-rendered content? No — it diffs the raw HTML bytes the Wayback Machine stored (via the id_ identity endpoint, no toolbar injection), the same as what the crawler that made the capture saved. Content injected client-side after page load was never archived and cannot be recovered.

Is this legal to run? Yes. The Internet Archive publishes these captures as a public historical record through an open API; no login-gated or paywalled content is accessed.

What are the limitations? Only pages the Wayback Machine actually captured are available — obscure or robots.txt-blocked URLs may have few or no captures. changedPercent is a plain line-based text diff, not a semantic one, so a full re-layout with identical copy can still score high.

Part of the webdatatools web-intelligence suite — every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data

Support and feedback

Found a page the Wayback Machine should have captured but this Actor missed, or a diff that looks wrong? Open an issue on the Issues tab.