Wayback Machine Scraper: Snapshots & Website History avatar

Wayback Machine Scraper: Snapshots & Website History

Pricing

Pay per event

Go to Apify Store
Wayback Machine Scraper: Snapshots & Website History

Wayback Machine Scraper: Snapshots & Website History

Scrape the Wayback Machine: every archived snapshot of any URL with date, HTTP status, MIME type and archive link, full archived-URL inventories per domain, and closest-snapshot checks. Dedupe by day, month or year. Export CSV, Excel, JSON, XML. No login or API key.

Pricing

Pay per event

Rating

0.0

(0)

Developer

RecordsData

RecordsData

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

16 hours ago

Last modified

Share

PunkRecordsData

๐Ÿ•ฐ๏ธ Wayback Machine Scraper: Snapshots & Website History

Wayback Machine Scraper exports archived snapshots from web.archive.org for any URL or domain: capture date, HTTP status, MIME type, size and a direct archive link per row. It also lists every URL archived under a domain and checks the closest snapshot to any date. Export CSV, Excel, JSON or XML. No login, no API key. Verified in October 2026 against live web.archive.org. Priced from $15 per 1,000 snapshots on the free tier.

Wayback Machine Scraper reads the Internet Archive's public CDX and availability APIs and returns one clean row per capture. For example, nytimes.com returns its first archived capture from 1996-11-12 and pages through 1,000 captures per request using resume keys, so histories are complete, not just page one. It is built for SEO teams, lawyers, journalists and domain buyers who need website history as a spreadsheet.

๐Ÿ”Ž What does Wayback Machine Scraper do?

  • Get website history by URL: every archived capture of a page, with ISO timestamp, status code and MIME type.
  • List all archived URLs of a domain: set the match type to domain and the dedupe mode to one per unique URL to inventory a site, subdomains included.
  • Build clean timelines: keep one capture per day, month or year.
  • Find archived PDFs, images or only HTTP 200 pages: filter by MIME type and status.
  • Check if a URL is archived: batch availability check, closest snapshot to any target date (YYYYMMDD).

๐Ÿ“‹ What data can you extract from the Wayback Machine?

FieldDescription
recordTypesnapshot or availability
targetThe URL or domain you asked about
originalUrlURL as it was archived
snapshotAtCapture time, ISO 8601 (UTC)
timestampRaw Wayback timestamp (YYYYMMDDhhmmss)
statusCodeHTTP status at capture time (when recorded)
mimeTypeContent type (when recorded)
sizeBytesArchived size in bytes
digestContent hash, equal hashes mean unchanged content
waybackUrlDirect web.archive.org link to the capture
isArchivedAvailability rows: whether any snapshot exists
closestSnapshotAtAvailability rows: nearest capture time
requestedTimestampAvailability rows: the date you asked for, or latest
scrapedAtWhen the row was collected (UTC)

Fields the source does not record for a capture are left out of that row, not filled with placeholders.

๐Ÿ“Š Sample output of the website history export

Real row from a run on nytimes.com:

{
"recordType": "snapshot",
"target": "nytimes.com",
"originalUrl": "http://www.nytimes.com:80/",
"snapshotAt": "1996-11-12T18:15:13Z",
"timestamp": "19961112181513",
"statusCode": 200,
"mimeType": "text/html",
"sizeBytes": 767,
"digest": "GY3YVZK6NIR7GKGXGGK4GPS2ZORULYDB",
"waybackUrl": "https://web.archive.org/web/19961112181513/http://www.nytimes.com:80/",
"scrapedAt": "2026-10-04T05:05:34.788Z"
}

Real availability row for apify.com with target date 20200101:

{
"recordType": "availability",
"target": "apify.com",
"isArchived": true,
"closestSnapshotAt": "2019-12-30T05:20:25Z",
"statusCode": 200,
"waybackUrl": "http://web.archive.org/web/20191230052025/https://apify.com/",
"requestedTimestamp": "20200101",
"scrapedAt": "2026-10-04T05:06:01.240Z"
}

๐Ÿ’ฐ How much does it cost to scrape the Wayback Machine?

Pay per event. You pay only for rows that were delivered.

EventFree tier priceWhen it is charged
snapshot-record$0.015 (so $15 per 1,000)One archived capture saved to the dataset
availability-record$0.010One URL that has an archived snapshot

Paid Apify plans get lower per-event prices (the live price list on the Pricing tab shows every tier; Gold is $10.09 per 1,000 snapshots). Rows for failed requests, errors, "never archived" availability answers and empty searches are not charged. If you set a maximum cost per run, the actor stops cleanly when it is reached. Free users get a 10-row preview.

๐Ÿš€ How to scrape website history in 3 steps

  1. Open the actor and click Try for free.
  2. Add URLs or domains, pick a match type, optionally set dates, a dedupe mode or filters.
  3. Click Start, then download CSV, Excel, JSON or XML.

โš™๏ธ Input

FieldMeaning
urlsPages or domains for snapshot history
matchTypeexact, prefix, host or domain (host plus subdomains)
fromDate / toDateRange as YYYY, YYYYMM or YYYYMMDD
onlySuccessfulKeep only HTTP 200 captures
mimeFilterFor example application/pdf
collapsenone, daily, monthly, yearly, unique-urls
checkAvailabilityUrlsURLs for availability checks
availabilityTimestampTarget date for availability, YYYYMMDD
maxItemsCap on rows returned and paid
{
"urls": ["nytimes.com"],
"matchType": "exact",
"collapse": "yearly",
"maxItems": 10
}

Invalid dates or an input with no URLs fail immediately with a clear message and cost nothing.

๐Ÿ“ฆ Output

One dataset row per snapshot or availability check, exportable as JSON, CSV, Excel or XML. The Overview view shows the key columns. A genuinely empty search finishes as succeeded with the status message "No snapshots matched the input" and zero charges. If Wayback Machine itself is down or blocks the request and no rows were delivered, the run fails instead of pretending success.

โš–๏ธ Wayback Machine scraper vs alternatives

Public Store prices measured in October 2026 (Free tier):

ActorEvent priceCovers
This actor$15 per 1,000 snapshotsHistory + domain inventory + availability, dedupe modes, filters
ryanclinton/wayback-machine-search$3.50 per 1,000 snapshotsSnapshot search
logiover/wayback-machine-url-extractor$5 per 1,000 itemsURL extraction
andok/wayback-machine-scraper$1 per 1,000 itemsSnapshot listing

We are the most expensive per row. If you only need a plain list of snapshots for one URL, a cheaper actor fits. What the extra price buys here is one actor for all three jobs (history, domain inventory, availability) with day/month/year/unique-URL dedupe, status and MIME filters, and no charge for error or empty rows.

๐Ÿ’ผ Use cases

  • SEO recovery: inventory every URL an expired or migrated domain ever had.
  • Legal evidence: timestamped capture lists with direct archive links.
  • Change tracking: monthly deduped timelines show when a page changed; compare digest values.
  • Domain due diligence: archive density and first-seen date before buying a domain.

๐Ÿ”Œ Run via API, schedule and integrations

import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('RecordsData/wayback-machine-scraper').call({
urls: ['example.com'], collapse: 'monthly', maxItems: 100,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();

Schedule runs from the Apify Console and send results to Zapier, Make, n8n, Google Sheets or Slack. The actor is also callable by AI agents through the Apify MCP server.

The actor uses the Internet Archive's public CDX and availability APIs at a polite rate (about one request every 1.5 seconds) with an identified user agent. It returns metadata and links, not archived page content. Respect archive.org's terms and the copyright of the archived pages when you reuse content.

โ“ Frequently asked questions

How far back do Wayback Machine snapshots go?

Each row carries its exact capture time. nytimes.com goes back to 1996-11-12 in our test.

Why did I get 0 results?

Either nothing was archived for that URL and match type, or your filters are too narrow. Try matchType: prefix or domain, remove date and status filters, and check that the URL has no typo. Empty searches cost nothing.

What is the difference between match types?

exact is one page, prefix is everything under a path, host is one hostname, domain is the host plus all subdomains.

How do I list every URL archived for a domain?

Use matchType: domain and collapse: unique-urls.

Can I find only archived PDFs?

Yes, set mimeFilter to application/pdf.

Do I pay for URLs that were never archived?

No. The availability row is returned with isArchived: false and is not charged.

Does it download the archived page content?

No. It returns metadata and the waybackUrl link to each capture.

Can I try it for free?

Yes. Free users get a 10-row preview; paid Apify plans unlock full runs.

๐Ÿ”— Want more data? Other PunkRecordsData scrapers

๐Ÿ’ฌ Support

Found a bug or a missing field? Open the Issues tab on this actor's page or write to contact.punkrecordsdata@gmail.com.

Last updated: 2026-10-03