Wayback Machine Scraper avatar

Wayback Machine Scraper

Pricing

from $2.04 / 1,000 snapshot fetched with contents

Go to Apify Store
Wayback Machine Scraper

Wayback Machine Scraper

Every archived capture of any URL, from the Internet Archive. List snapshots for one page, a path prefix or a whole domain, filter by status code, MIME type and date, and pull back the archived HTML exactly as it was originally served.

Pricing

from $2.04 / 1,000 snapshot fetched with contents

Rating

0.0

(0)

Developer

丂卩ㄖㄖҜㄚ

丂卩ㄖㄖҜㄚ

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Every archived capture of any URL from the Internet Archive, what changed between them, and the page itself exactly as it was originally served.

🔍 What does Wayback Machine Scraper do?

It queries the Internet Archive's index for a URL and returns one row per capture, with the option to fetch the archived page itself and compare it with the capture before it.

  • One page, a path, a hostname or a whole domain, and up to 50 URLs a run.
  • The original bytes. Content is fetched with the archive's id_ modifier, so you get the HTML the server actually sent, with no archive toolbar and no rewritten links. Extracted text is returned alongside it, not instead of it.
  • Change detection. Each capture is compared with the one before it using the archive's own content fingerprint, so finding out whether a page moved costs nothing extra. Turn on content retrieval and you also get what moved.
  • Only the changes. A decade of daily captures collapses to the handful of dates the page actually changed.
  • Point in time evidence. Ask for the capture closest to a date and get it with the gap reported in days.
  • Monitoring. Remembers the newest capture per URL and returns only what is new since the last run, so a schedule reports activity rather than history.
  • Presets for SEO history, compliance evidence, competitor watch and forensics, so you do not have to work the options out.

Server side filtering by status code, MIME type and date range, so the archive does the narrowing rather than you paying to download rows you then discard.

📊 What data can I extract from the Wayback Machine?

One record per capture:

FieldWhat it is
timestamp, capturedAtCapture time, raw and as ISO 8601 UTC
urlThe URL as captured
statusCode, mimeTypeWhat it was served as, or null if unrecorded
digest, lengthBytesArchive content fingerprint and size
snapshotUrl, rawUrlViewable archive link, and the original bytes link
htmlThe page as originally served, when content is retrieved
textReadable text, scripts, styles and archive toolbar removed
changed, changeTypeWhether it differs from the previous capture, and how
changeBasisdigest, length or unknown, so you can judge the claim
bytesDeltaSize difference from the previous capture
diffScore, magnitude, categories, title and price changes, sample text
distanceFromTargetDaysDays from the requested date, on closest-date runs

💡 Why use the Wayback Machine Scraper?

Point in time evidence. Prove what a page said on a particular date, with the capture time and the archive link attached.

Competitor watch. Track a competitor's pricing, claims or terms over time and see only the dates they moved.

Content recovery. Recover a page that was taken down or rewritten, and feed the archived HTML into your own parsing.

Migration repair. Rebuild a URL inventory after a migration lost the old sitemap.

🚀 How do I use Wayback Machine Scraper?

  1. Click Try for free.
  2. Put the page, path, hostname or domain you want into url, or a list into urls.
  3. Set matchType to decide how wide that goes, and narrow with from, to, statusCodes and mimeTypes.
  4. Turn on fetchContent if you want the archived pages themselves, then click Start.
  5. Download the results as JSON, CSV or Excel, or pull them from the API.

⬇️ Input

{
"url": "bbc.co.uk/news",
"from": "2024",
"statusCodes": ["200"],
"mimeTypes": ["text/html"]
}
FieldTypeDefaultWhat it does
urlstringbbc.co.uk/newsOne page, path, hostname or whole domain
urlsarrayUp to 50 URLs in one run
useCasestringPreset for SEO history, compliance evidence, competitor watch or forensics
matchTypestringexactHow to match the URL, exact, prefix, host or domain
fromstring2024Earliest capture date or year
tostringLatest capture date or year
targetDatestringReturn the capture closest to this date, with the gap in days
maxSnapshotsinteger100Cap on captures returned
statusCodesarray200Only captures with these status codes
mimeTypesarraytext/htmlOnly captures with these MIME types
detectChangesbooleanfalseCompare each capture with the one before it
onlyChangedbooleanfalseReturn only the captures where something moved
fetchContentbooleanfalseRetrieve the original page bytes and text
maxContentFetchesinteger25Cap on how many captures are retrieved in full
monitorbooleanfalseReturn only captures newer than the last run

⬆️ Output

Table view

Results arrive as a Snapshots table you can sort and filter in the Console, with the capture time, URL, status code, MIME type, retrieved content status, text length and the archive link lined up per capture.

JSON

A typical row:

{
"capturedAt": "2024-01-04T21:34:53.000Z",
"url": "https://stripe.com/pricing",
"statusCode": 200,
"mimeType": "text/html",
"contentStatus": 200,
"textChars": 23837,
"snapshotUrl": "https://web.archive.org/web/20240104213453/https://stripe.com/pricing",
"timestamp": "20240104213453"
}

Download it from the run as JSON, CSV or Excel, or read it straight from the API.

How the change detection decides

changed comes from the archive's own content digest, which is computed over the stored bytes and is the authoritative signal. Where a row has no digest, size is used instead, and changeBasis says so, because that is a weaker claim and you should be able to see which one you got. Where neither exists, nothing is claimed.

diff needs both pages, so it only appears where content was retrieved. It is scored on the share of words that moved rather than raw counts, so a paragraph edit on a short page is not ranked below a footer tweak on a long one.

A title or price change is never reported as minor. Both can move while almost no words do, and both are usually the reason someone is watching.

A capture can be changed: true with a diff of none. That is not a contradiction: the stored bytes differ but the readable text does not, so the change was in markup or scripts rather than in anything a reader sees.

Notes on accuracy

A URL with no captures is an answer, not an error. The run finishes with zero results and says what to widen. It does not fail.

Filters are checked, not assumed. Every returned row is tested against the filters that were requested, and the run warns if any row does not satisfy them rather than presenting the wrong answer with confidence.

Timestamps come from the index, never invented. A made up timestamp answers with an empty redirect, which is indistinguishable from a page that archived to nothing, so content is only fetched at a timestamp the archive returned.

The archive does not replay original response headers. If you are detecting technologies from archived pages, anything visible only in a response header, such as the CDN, web server or language, is not recoverable from an archived capture by anyone. Content in the HTML is unaffected.

Coverage is the archive's, not ours. Archive.org crawls on its own schedule, so not every change is captured.

💰 How much does it cost?

Pay per event, so a run costs what it does.

EventPrice
Snapshot indexed$0.003
Snapshot fetched with content$0.0034

Fetching replaces the indexing charge for that capture rather than adding to it. Indexing 100 captures costs $0.30. Indexing 100 and retrieving 25 of them costs $0.31.

Use maxSnapshots to cap the index side and maxContentFetches to cap the retrieval side.

🔌 Integrations

Send results straight to Google Sheets, Slack, Airtable, Zapier, Make or your own webhook using Apify integrations. You can also trigger a run whenever something happens in another tool.

🔗 Using Wayback Machine Scraper with the Apify API

curl -X POST "https://api.apify.com/v2/acts/spookyweb~wayback-machine-scraper/run-sync-get-dataset-items?token=YOUR_TOKEN" \
-H "Content-Type: application/json" \
-d '{"url": "bbc.co.uk/news", "from": "2024", "statusCodes": ["200"], "mimeTypes": ["text/html"]}'

Or with the Apify client:

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_TOKEN' });
const run = await client.actor('spookyweb/wayback-machine-scraper').call({
url: 'bbc.co.uk/news',
from: '2024',
statusCodes: ['200'],
mimeTypes: ['text/html'],
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Full detail is in the Apify API reference, and every run is also callable from the Python and JavaScript clients.

❓ FAQ

What is a capture, and how do I pick one?

A capture is one archived copy of a URL at one moment, identified by its timestamp. A busy page can have thousands. Narrow with from and to, filter to the status codes and MIME types you care about, or set targetDate to get the capture closest to a particular day, with distanceFromTargetDays telling you how far off it landed.

What does change detection do?

It compares each capture with the one before it and sets changed, changeType and changeBasis. The comparison uses the archive's own content digest, so it costs nothing extra, and onlyChanged collapses a decade of daily captures down to the handful of dates the page actually moved.

Can I get the page content as well as the index?

Yes. Turn on fetchContent and each capture carries html, the bytes the server originally sent with no archive toolbar and no rewritten links, plus text, the readable content with scripts and styles removed. maxContentFetches caps how many are retrieved so a wide index run does not turn into a wide download.

What does matchType do?

It decides how wide the URL query goes. exact returns captures of that one URL. prefix returns everything under that path. host returns everything on that exact hostname. domain returns the hostname and its subdomains. Widening it is the fastest way to turn one page into a whole site inventory.

Can I monitor a page for changes over time?

Yes. Set monitor and run it on a schedule. The Actor remembers the newest capture per URL and returns only what is new since the last run, so each run reports activity rather than repeating the history. Combine it with detectChanges and onlyChanged to be told only when something actually moved.

Why did my run return nothing?

Because the archive has no captures matching what you asked for. That is an answer rather than an error, so the run finishes successfully with zero results and says what to widen. Try a broader matchType, an earlier from, or dropping the status code and MIME filters.

This reads the Internet Archive's public Wayback Machine through its published CDX API. The archive exists to make historical web pages available to the public, and this Actor respects its rate limits.

Scraping publicly available data is legal in the UK, the EU and the US. What you do with an archived page is still governed by the copyright and terms of the original site, so treat archived content as you would the live version. If a capture contains personal data, handling it is on you under GDPR. Apify's ethical scraping guide covers the wider picture.

👍 Your feedback

Found a bug, or want a field that is not here yet? Open an issue on the Actor's Issues tab. Requests that make the data more useful get built, and problems get fixed quickly.

🔎 You might also like

ActorWhat it does
Website Contact ScraperEmails, phones, socials and addresses from company websites, one record per domain
Company Email FinderPublished company addresses, the naming pattern behind them and an MX check
Telegram Channel ScraperPosts, views and reactions from public Telegram channels, no account needed