Wayback Machine Scraper: Snapshots & Website History
Pricing
Pay per event
Wayback Machine Scraper: Snapshots & Website History
Scrape the Wayback Machine: every archived snapshot of any URL with date, HTTP status, MIME type and archive link, full archived-URL inventories per domain, and closest-snapshot checks. Dedupe by day, month or year. Export CSV, Excel, JSON, XML. No login or API key.
Pricing
Pay per event
Rating
0.0
(0)
Developer
RecordsData
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
16 hours ago
Last modified
Categories
Share
๐ฐ๏ธ Wayback Machine Scraper: Snapshots & Website History
Wayback Machine Scraper exports archived snapshots from web.archive.org for any URL or domain: capture date, HTTP status, MIME type, size and a direct archive link per row. It also lists every URL archived under a domain and checks the closest snapshot to any date. Export CSV, Excel, JSON or XML. No login, no API key. Verified in October 2026 against live web.archive.org. Priced from $15 per 1,000 snapshots on the free tier.
Wayback Machine Scraper reads the Internet Archive's public CDX and availability APIs and returns one clean row per capture. For example, nytimes.com returns its first archived capture from 1996-11-12 and pages through 1,000 captures per request using resume keys, so histories are complete, not just page one. It is built for SEO teams, lawyers, journalists and domain buyers who need website history as a spreadsheet.
๐ What does Wayback Machine Scraper do?
- Get website history by URL: every archived capture of a page, with ISO timestamp, status code and MIME type.
- List all archived URLs of a domain: set the match type to domain and the dedupe mode to one per unique URL to inventory a site, subdomains included.
- Build clean timelines: keep one capture per day, month or year.
- Find archived PDFs, images or only HTTP 200 pages: filter by MIME type and status.
- Check if a URL is archived: batch availability check, closest snapshot to any target date (YYYYMMDD).
๐ What data can you extract from the Wayback Machine?
| Field | Description |
|---|---|
recordType | snapshot or availability |
target | The URL or domain you asked about |
originalUrl | URL as it was archived |
snapshotAt | Capture time, ISO 8601 (UTC) |
timestamp | Raw Wayback timestamp (YYYYMMDDhhmmss) |
statusCode | HTTP status at capture time (when recorded) |
mimeType | Content type (when recorded) |
sizeBytes | Archived size in bytes |
digest | Content hash, equal hashes mean unchanged content |
waybackUrl | Direct web.archive.org link to the capture |
isArchived | Availability rows: whether any snapshot exists |
closestSnapshotAt | Availability rows: nearest capture time |
requestedTimestamp | Availability rows: the date you asked for, or latest |
scrapedAt | When the row was collected (UTC) |
Fields the source does not record for a capture are left out of that row, not filled with placeholders.
๐ Sample output of the website history export
Real row from a run on nytimes.com:
{"recordType": "snapshot","target": "nytimes.com","originalUrl": "http://www.nytimes.com:80/","snapshotAt": "1996-11-12T18:15:13Z","timestamp": "19961112181513","statusCode": 200,"mimeType": "text/html","sizeBytes": 767,"digest": "GY3YVZK6NIR7GKGXGGK4GPS2ZORULYDB","waybackUrl": "https://web.archive.org/web/19961112181513/http://www.nytimes.com:80/","scrapedAt": "2026-10-04T05:05:34.788Z"}
Real availability row for apify.com with target date 20200101:
{"recordType": "availability","target": "apify.com","isArchived": true,"closestSnapshotAt": "2019-12-30T05:20:25Z","statusCode": 200,"waybackUrl": "http://web.archive.org/web/20191230052025/https://apify.com/","requestedTimestamp": "20200101","scrapedAt": "2026-10-04T05:06:01.240Z"}
๐ฐ How much does it cost to scrape the Wayback Machine?
Pay per event. You pay only for rows that were delivered.
| Event | Free tier price | When it is charged |
|---|---|---|
snapshot-record | $0.015 (so $15 per 1,000) | One archived capture saved to the dataset |
availability-record | $0.010 | One URL that has an archived snapshot |
Paid Apify plans get lower per-event prices (the live price list on the Pricing tab shows every tier; Gold is $10.09 per 1,000 snapshots). Rows for failed requests, errors, "never archived" availability answers and empty searches are not charged. If you set a maximum cost per run, the actor stops cleanly when it is reached. Free users get a 10-row preview.
๐ How to scrape website history in 3 steps
- Open the actor and click Try for free.
- Add URLs or domains, pick a match type, optionally set dates, a dedupe mode or filters.
- Click Start, then download CSV, Excel, JSON or XML.
โ๏ธ Input
| Field | Meaning |
|---|---|
urls | Pages or domains for snapshot history |
matchType | exact, prefix, host or domain (host plus subdomains) |
fromDate / toDate | Range as YYYY, YYYYMM or YYYYMMDD |
onlySuccessful | Keep only HTTP 200 captures |
mimeFilter | For example application/pdf |
collapse | none, daily, monthly, yearly, unique-urls |
checkAvailabilityUrls | URLs for availability checks |
availabilityTimestamp | Target date for availability, YYYYMMDD |
maxItems | Cap on rows returned and paid |
{"urls": ["nytimes.com"],"matchType": "exact","collapse": "yearly","maxItems": 10}
Invalid dates or an input with no URLs fail immediately with a clear message and cost nothing.
๐ฆ Output
One dataset row per snapshot or availability check, exportable as JSON, CSV, Excel or XML. The Overview view shows the key columns. A genuinely empty search finishes as succeeded with the status message "No snapshots matched the input" and zero charges. If Wayback Machine itself is down or blocks the request and no rows were delivered, the run fails instead of pretending success.
โ๏ธ Wayback Machine scraper vs alternatives
Public Store prices measured in October 2026 (Free tier):
| Actor | Event price | Covers |
|---|---|---|
| This actor | $15 per 1,000 snapshots | History + domain inventory + availability, dedupe modes, filters |
| ryanclinton/wayback-machine-search | $3.50 per 1,000 snapshots | Snapshot search |
| logiover/wayback-machine-url-extractor | $5 per 1,000 items | URL extraction |
| andok/wayback-machine-scraper | $1 per 1,000 items | Snapshot listing |
We are the most expensive per row. If you only need a plain list of snapshots for one URL, a cheaper actor fits. What the extra price buys here is one actor for all three jobs (history, domain inventory, availability) with day/month/year/unique-URL dedupe, status and MIME filters, and no charge for error or empty rows.
๐ผ Use cases
- SEO recovery: inventory every URL an expired or migrated domain ever had.
- Legal evidence: timestamped capture lists with direct archive links.
- Change tracking: monthly deduped timelines show when a page changed; compare
digestvalues. - Domain due diligence: archive density and first-seen date before buying a domain.
๐ Run via API, schedule and integrations
import { ApifyClient } from 'apify-client';const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });const run = await client.actor('RecordsData/wayback-machine-scraper').call({urls: ['example.com'], collapse: 'monthly', maxItems: 100,});const { items } = await client.dataset(run.defaultDatasetId).listItems();
Schedule runs from the Apify Console and send results to Zapier, Make, n8n, Google Sheets or Slack. The actor is also callable by AI agents through the Apify MCP server.
๐ก๏ธ Is it legal to scrape the Wayback Machine?
The actor uses the Internet Archive's public CDX and availability APIs at a polite rate (about one request every 1.5 seconds) with an identified user agent. It returns metadata and links, not archived page content. Respect archive.org's terms and the copyright of the archived pages when you reuse content.
โ Frequently asked questions
How far back do Wayback Machine snapshots go?
Each row carries its exact capture time. nytimes.com goes back to 1996-11-12 in our test.
Why did I get 0 results?
Either nothing was archived for that URL and match type, or your filters are too narrow. Try matchType: prefix or domain, remove date and status filters, and check that the URL has no typo. Empty searches cost nothing.
What is the difference between match types?
exact is one page, prefix is everything under a path, host is one hostname, domain is the host plus all subdomains.
How do I list every URL archived for a domain?
Use matchType: domain and collapse: unique-urls.
Can I find only archived PDFs?
Yes, set mimeFilter to application/pdf.
Do I pay for URLs that were never archived?
No. The availability row is returned with isArchived: false and is not charged.
Does it download the archived page content?
No. It returns metadata and the waybackUrl link to each capture.
Can I try it for free?
Yes. Free users get a 10-row preview; paid Apify plans unlock full runs.
๐ Want more data? Other PunkRecordsData scrapers
- Wikipedia Scraper: Articles, Summaries & Pageviews
- Hacker News Scraper: Story & Comment Search
- USPTO Trademark Status Scraper
๐ฌ Support
Found a bug or a missing field? Open the Issues tab on this actor's page or write to contact.punkrecordsdata@gmail.com.
Last updated: 2026-10-03