# Wayback Machine Scraper: Snapshots & Website History (`recordsdata/wayback-machine-scraper`) Actor

Scrape the Wayback Machine: every archived snapshot of any URL with date, HTTP status, MIME type and archive link, full archived-URL inventories per domain, and closest-snapshot checks. Dedupe by day, month or year. Export CSV, Excel, JSON, XML. No login or API key.

- **URL**: https://apify.com/recordsdata/wayback-machine-scraper.md
- **Developed by:** [RecordsData](https://apify.com/recordsdata) (community)
- **Categories:** SEO tools, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

Pay per event

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

<p align="center">
  <img src="https://api.apify.com/v2/key-value-stores/AAm3a1h3Z9nYfrvh9/records/banner?v=3" alt="PunkRecordsData" width="100%" />
</p>

## 🕰️ Wayback Machine Scraper: Snapshots & Website History

> Wayback Machine Scraper exports archived snapshots from web.archive.org for any URL or domain: capture date, HTTP status, MIME type, size and a direct archive link per row. It also lists every URL archived under a domain and checks the closest snapshot to any date. Export CSV, Excel, JSON or XML. No login, no API key. Verified in October 2026 against live web.archive.org. Priced from $15 per 1,000 snapshots on the free tier.

Wayback Machine Scraper reads the Internet Archive's public CDX and availability APIs and returns one clean row per capture. For example, nytimes.com returns its first archived capture from 1996-11-12 and pages through 1,000 captures per request using resume keys, so histories are complete, not just page one. It is built for SEO teams, lawyers, journalists and domain buyers who need website history as a spreadsheet.

### 🔎 What does Wayback Machine Scraper do?

- **Get website history by URL**: every archived capture of a page, with ISO timestamp, status code and MIME type.
- **List all archived URLs of a domain**: set the match type to domain and the dedupe mode to one per unique URL to inventory a site, subdomains included.
- **Build clean timelines**: keep one capture per day, month or year.
- **Find archived PDFs, images or only HTTP 200 pages**: filter by MIME type and status.
- **Check if a URL is archived**: batch availability check, closest snapshot to any target date (YYYYMMDD).

### 📋 What data can you extract from the Wayback Machine?

| Field | Description |
|---|---|
| `recordType` | `snapshot` or `availability` |
| `target` | The URL or domain you asked about |
| `originalUrl` | URL as it was archived |
| `snapshotAt` | Capture time, ISO 8601 (UTC) |
| `timestamp` | Raw Wayback timestamp (YYYYMMDDhhmmss) |
| `statusCode` | HTTP status at capture time (when recorded) |
| `mimeType` | Content type (when recorded) |
| `sizeBytes` | Archived size in bytes |
| `digest` | Content hash, equal hashes mean unchanged content |
| `waybackUrl` | Direct web.archive.org link to the capture |
| `isArchived` | Availability rows: whether any snapshot exists |
| `closestSnapshotAt` | Availability rows: nearest capture time |
| `requestedTimestamp` | Availability rows: the date you asked for, or `latest` |
| `scrapedAt` | When the row was collected (UTC) |

Fields the source does not record for a capture are left out of that row, not filled with placeholders.

### 📊 Sample output of the website history export

Real row from a run on `nytimes.com`:

```json
{
  "recordType": "snapshot",
  "target": "nytimes.com",
  "originalUrl": "http://www.nytimes.com:80/",
  "snapshotAt": "1996-11-12T18:15:13Z",
  "timestamp": "19961112181513",
  "statusCode": 200,
  "mimeType": "text/html",
  "sizeBytes": 767,
  "digest": "GY3YVZK6NIR7GKGXGGK4GPS2ZORULYDB",
  "waybackUrl": "https://web.archive.org/web/19961112181513/http://www.nytimes.com:80/",
  "scrapedAt": "2026-10-04T05:05:34.788Z"
}
```

Real availability row for `apify.com` with target date 20200101:

```json
{
  "recordType": "availability",
  "target": "apify.com",
  "isArchived": true,
  "closestSnapshotAt": "2019-12-30T05:20:25Z",
  "statusCode": 200,
  "waybackUrl": "http://web.archive.org/web/20191230052025/https://apify.com/",
  "requestedTimestamp": "20200101",
  "scrapedAt": "2026-10-04T05:06:01.240Z"
}
```

### 💰 How much does it cost to scrape the Wayback Machine?

Pay per event. You pay only for rows that were delivered.

| Event | Free tier price | When it is charged |
|---|---|---|
| `snapshot-record` | $0.015 (so $15 per 1,000) | One archived capture saved to the dataset |
| `availability-record` | $0.010 | One URL that has an archived snapshot |

Paid Apify plans get lower per-event prices (the live price list on the Pricing tab shows every tier; Gold is $10.09 per 1,000 snapshots). Rows for failed requests, errors, "never archived" availability answers and empty searches are **not charged**. If you set a maximum cost per run, the actor stops cleanly when it is reached. Free users get a 10-row preview.

### 🚀 How to scrape website history in 3 steps

1. Open the actor and click **Try for free**.
2. Add URLs or domains, pick a match type, optionally set dates, a dedupe mode or filters.
3. Click **Start**, then download CSV, Excel, JSON or XML.

### ⚙️ Input

| Field | Meaning |
|---|---|
| `urls` | Pages or domains for snapshot history |
| `matchType` | `exact`, `prefix`, `host` or `domain` (host plus subdomains) |
| `fromDate` / `toDate` | Range as YYYY, YYYYMM or YYYYMMDD |
| `onlySuccessful` | Keep only HTTP 200 captures |
| `mimeFilter` | For example `application/pdf` |
| `collapse` | `none`, `daily`, `monthly`, `yearly`, `unique-urls` |
| `checkAvailabilityUrls` | URLs for availability checks |
| `availabilityTimestamp` | Target date for availability, YYYYMMDD |
| `maxItems` | Cap on rows returned and paid |

```json
{
  "urls": ["nytimes.com"],
  "matchType": "exact",
  "collapse": "yearly",
  "maxItems": 10
}
```

Invalid dates or an input with no URLs fail immediately with a clear message and cost nothing.

### 📦 Output

One dataset row per snapshot or availability check, exportable as JSON, CSV, Excel or XML. The Overview view shows the key columns. A genuinely empty search finishes as succeeded with the status message "No snapshots matched the input" and zero charges. If Wayback Machine itself is down or blocks the request and no rows were delivered, the run fails instead of pretending success.

### ⚖️ Wayback Machine scraper vs alternatives

Public Store prices measured in October 2026 (Free tier):

| Actor | Event price | Covers |
|---|---|---|
| This actor | $15 per 1,000 snapshots | History + domain inventory + availability, dedupe modes, filters |
| ryanclinton/wayback-machine-search | $3.50 per 1,000 snapshots | Snapshot search |
| logiover/wayback-machine-url-extractor | $5 per 1,000 items | URL extraction |
| andok/wayback-machine-scraper | $1 per 1,000 items | Snapshot listing |

We are the most expensive per row. If you only need a plain list of snapshots for one URL, a cheaper actor fits. What the extra price buys here is one actor for all three jobs (history, domain inventory, availability) with day/month/year/unique-URL dedupe, status and MIME filters, and no charge for error or empty rows.

### 💼 Use cases

- **SEO recovery**: inventory every URL an expired or migrated domain ever had.
- **Legal evidence**: timestamped capture lists with direct archive links.
- **Change tracking**: monthly deduped timelines show when a page changed; compare `digest` values.
- **Domain due diligence**: archive density and first-seen date before buying a domain.

### 🔌 Run via API, schedule and integrations

```js
import { ApifyClient } from 'apify-client';
const client = new ApifyClient({ token: 'YOUR_APIFY_TOKEN' });
const run = await client.actor('RecordsData/wayback-machine-scraper').call({
  urls: ['example.com'], collapse: 'monthly', maxItems: 100,
});
const { items } = await client.dataset(run.defaultDatasetId).listItems();
```

Schedule runs from the Apify Console and send results to Zapier, Make, n8n, Google Sheets or Slack. The actor is also callable by AI agents through the Apify MCP server.

### 🛡️ Is it legal to scrape the Wayback Machine?

The actor uses the Internet Archive's public CDX and availability APIs at a polite rate (about one request every 1.5 seconds) with an identified user agent. It returns metadata and links, not archived page content. Respect archive.org's terms and the copyright of the archived pages when you reuse content.

### ❓ Frequently asked questions

#### How far back do Wayback Machine snapshots go?

Each row carries its exact capture time. nytimes.com goes back to 1996-11-12 in our test.

#### Why did I get 0 results?

Either nothing was archived for that URL and match type, or your filters are too narrow. Try `matchType: prefix` or `domain`, remove date and status filters, and check that the URL has no typo. Empty searches cost nothing.

#### What is the difference between match types?

`exact` is one page, `prefix` is everything under a path, `host` is one hostname, `domain` is the host plus all subdomains.

#### How do I list every URL archived for a domain?

Use `matchType: domain` and `collapse: unique-urls`.

#### Can I find only archived PDFs?

Yes, set `mimeFilter` to `application/pdf`.

#### Do I pay for URLs that were never archived?

No. The availability row is returned with `isArchived: false` and is not charged.

#### Does it download the archived page content?

No. It returns metadata and the `waybackUrl` link to each capture.

#### Can I try it for free?

Yes. Free users get a 10-row preview; paid Apify plans unlock full runs.

### 🔗 Want more data? Other PunkRecordsData scrapers

- [Wikipedia Scraper: Articles, Summaries & Pageviews](https://apify.com/RecordsData/wikipedia-articles-scraper)
- [Hacker News Scraper: Story & Comment Search](https://apify.com/RecordsData/hackernews-search-scraper)
- [USPTO Trademark Status Scraper](https://apify.com/RecordsData/uspto-trademark-status-scraper)

### 💬 Support

Found a bug or a missing field? Open the **Issues** tab on this actor's page or write to contact.punkrecordsdata@gmail.com.

Last updated: 2026-10-03

# Actor input Schema

## `urls` (type: `array`):

Pages or domains to list archived Wayback Machine snapshots for (for example apify.com or apify.com/store). Leave empty only if you use the availability check list.

## `matchType` (type: `string`):

exact = that page only; prefix = everything under the path; host = whole host; domain = host + subdomains.

## `fromDate` (type: `string`):

Earliest snapshot, YYYY, YYYYMM or YYYYMMDD (e.g. 2015 or 20200101).

## `toDate` (type: `string`):

Latest snapshot, YYYY, YYYYMM or YYYYMMDD.

## `onlySuccessful` (type: `boolean`):

Keep only snapshots that archived successfully (status 200).

## `mimeFilter` (type: `string`):

Keep only this content type, e.g. text/html, application/pdf, image/jpeg.

## `collapse` (type: `string`):

Collapse snapshots: one per day/month/year, or one per unique URL (site inventory mode).

## `checkAvailabilityUrls` (type: `array`):

URLs to check for their closest archived Wayback Machine snapshot. One row per URL; URLs that were never archived return a free row with isArchived false.

## `availabilityTimestamp` (type: `string`):

Find the snapshot closest to this date (YYYYMMDD). Empty = most recent.

## `maxItems` (type: `integer`):

Maximum rows to return and pay for (snapshots plus availability checks). Free users are limited to 10 rows (preview). Paid users: leave empty for no limit, or set up to 1,000,000.

## Actor input object example

```json
{
  "urls": [
    "apify.com"
  ],
  "matchType": "exact",
  "fromDate": "",
  "toDate": "",
  "onlySuccessful": false,
  "mimeFilter": "",
  "collapse": "none",
  "checkAvailabilityUrls": [],
  "availabilityTimestamp": "",
  "maxItems": 10
}
```

# Actor output Schema

## `overview` (type: `string`):

Key fields per row

## `fullData` (type: `string`):

Complete dataset with all fields

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "apify.com"
    ],
    "fromDate": "",
    "toDate": "",
    "mimeFilter": "",
    "checkAvailabilityUrls": [],
    "availabilityTimestamp": "",
    "maxItems": 10
};

// Run the Actor and wait for it to finish
const run = await client.actor("recordsdata/wayback-machine-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["apify.com"],
    "fromDate": "",
    "toDate": "",
    "mimeFilter": "",
    "checkAvailabilityUrls": [],
    "availabilityTimestamp": "",
    "maxItems": 10,
}

# Run the Actor and wait for it to finish
run = client.actor("recordsdata/wayback-machine-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "apify.com"
  ],
  "fromDate": "",
  "toDate": "",
  "mimeFilter": "",
  "checkAvailabilityUrls": [],
  "availabilityTimestamp": "",
  "maxItems": 10
}' |
apify call recordsdata/wayback-machine-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,recordsdata/wayback-machine-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/8JRC0eEnnhQZcsg1n/builds/cCg4Qebcxo2QfNelN/openapi.json
