# Wayback Machine Scraper - Snapshot History of Any URL (`dami_studio/wayback-machine-scraper`) Actor

Every capture the Wayback Machine holds of the pages or sites you list, one row each: when it was taken, HTTP status, content type, size, digest and a link to the archived copy. Filter by date, status and type, or keep one per day, month or year. $0.13 per 1,000 captures.

- **URL**: https://apify.com/dami_studio/wayback-machine-scraper.md
- **Developed by:** [Dami's Studio](https://apify.com/dami_studio) (community)
- **Categories:** SEO tools, Developer tools, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.12 / 1,000 captures

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

**Give it pages or whole sites and get every capture the Wayback Machine holds of them**, one row per
capture: when it was taken, the HTTP status, the content type, the size and a link to the archived
copy. It works on one page, everything under a path, a whole host, or a domain with its subdomains,
with date, status and content type filters.

It lists captures and does not download them, so each row is a link to the copy rather than the copy
itself. And a few sites are excluded from the archive or keep their capture list private, which no
setting gets around.

| | |
|---|---|
| **Input** | Web pages or sites: one page, a path, a host or a whole domain |
| **Output** | One row per capture: capture time, the URL as captured, status, content type, digest, stored size, archive links |
| **Ceiling** | 100,000 captures per entry, 1,000,000 per run, 500 entries |
| **Speed** | 3.5 to 5 seconds per page or site looked up, with or without captures, so a full list of 500 takes about forty minutes. Give a long list a longer timeout |
| **Account needed** | None from you |
| **Price** | $0.13 per 1,000 captures, flat on every plan. The free plan's $5 a month covers about 38,000 |

### 🔍 What Wayback Machine Scraper does

For each page or site it reads the Wayback Machine's capture index, the list behind the archive's
own calendar pages, and turns each entry into a row. A busy home page has hundreds of thousands of
captures, so the run reads them a few thousand at a time, oldest first, until it has the number you
asked for or the list ends.

The archive applies your dates and filters, and every row is checked again before it is kept. A
capture outside your dates, with another status or another content type, is never charged.

To make a long history readable, keep one capture per day, month or year, keep only the captures
where the content changed, or list each archived URL once. On a path or a domain this is done page by
page, so one page's capture never hides another's.

The run keeps to a couple of dozen requests a minute, well inside what the archive asks of automated
tools. If the archive asks for a pause anyway, the run stops there and lists the entries it did not
reach.

### 📋 What data you get from each Wayback Machine capture

| What you get | Field |
|---|---|
| The page or site you gave, and how it was looked up | `inputUrl`, `matchType` |
| The URL as the archive captured it, and the archive's key for it | `originalUrl`, `urlKey` |
| When the capture was taken | `timestamp`, `capturedAt` |
| The HTTP status and the content type the archive got | `statusCode`, `mimeType` |
| A fingerprint of the content, and the stored size | `digest`, `length` |
| The capture in the archive's viewer, and as first served | `archiveUrl`, `rawArchiveUrl` |
| When the row was read | `scrapedAt` |

### ▶️ How to scrape the Wayback Machine

1. Open [Wayback Machine Scraper](https://apify.com/dami_studio/wayback-machine-scraper) and click
   **Try for free**.
2. Paste your pages or sites into **Web addresses**, one per line.
3. Pick **What each address covers**, then add dates or filters if you want them.
4. Set **Captures per address** and click **Start**.
5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

### 💰 How much does it cost to scrape the Wayback Machine?

**$0.13 per 1,000 captures.** Flat on every Apify plan, no volume tiers. On the free plan, the $5
Apify gives you each month covers about 38,000 captures.

You pay per capture delivered. The sample row, diagnostic rows, captures the checks drop and repeats
are not charged, and a page with no captures adds nothing. **Captures per address** and **Captures in
total** are the two caps on what one run can spend.

### 📥 What you give it

```json
{
  "urls": ["https://www.nasa.gov/", "apify.com/store/*"],
  "matchType": "exact",
  "from": "2015",
  "to": "2020-06",
  "statusCodes": ["200"],
  "mimeTypes": ["text/html"],
  "collapse": "month",
  "maxCapturesPerUrl": 1000,
  "maxCaptures": 100000,
  "newestFirst": false
}
```

| Field | Default | What it is |
|---|---|---|
| `urls` | none, the form starts with `https://www.nasa.gov/` | The pages or sites, one per line, up to 500. `nasa.gov` and `https://www.nasa.gov/` are the same page to the archive. End one with `/*` for everything under that path, or start it with `*.` for a domain and its subdomains. |
| `matchType` | `exact` | `exact` for that page, `prefix` for everything under the path, `host` for every page on the host, `domain` for the host and its subdomains. A `/*` or `*.` typed into an entry wins over this. |
| `from`, `to` | none | Dates in UTC, written `2015`, `2015-06` or `2015-06-01`. `to` covers the whole period, so `2019` runs to the last second of 2019. A date in any other format is refused before the run starts; an impossible one, like 2021-02-30, stops the run before it looks anything up. |
| `statusCodes` | all | HTTP codes such as `200` or `404`, or a class such as `3xx`. |
| `mimeTypes` | all | Content types in full, such as `text/html` or `application/pdf`, or a family such as `image/*`. A bare `html` matches nothing. |
| `collapse` | `none`, the form starts on `year` | `digest` drops a capture identical to the one before it. `day`, `month` and `year` keep the first capture of each period. `url` lists each archived URL once, with its first capture. |
| `maxCapturesPerUrl` | `1000`, the form starts at 100 | The most captures from one page, path, host or domain, up to 100,000. |
| `maxCaptures` | `100000` | The most captures in the whole run, up to 1,000,000. |
| `newestFirst` | `false` | Start from the latest capture and work back. Exact pages only. With thinning on, it keeps the latest capture of each period rather than the first. |

**A status filter leaves out unchanged captures.** When the archive finds the same content again it
often stores a revisit record, which has no status code and the type `warc/revisit`. Without a filter
those rows are part of the history. Ask for any status and they drop out.

### 📤 What you get back

A real row from a run on 3 October 2026: nasa.gov's first capture.

```json
{
  "recordType": "capture",
  "inputUrl": "https://www.nasa.gov/",
  "matchType": "exact",
  "originalUrl": "http://www.nasa.gov:80/",
  "urlKey": "gov,nasa)/",
  "timestamp": "19961231235847",
  "capturedAt": "1996-12-31T23:58:47Z",
  "statusCode": 200,
  "mimeType": "text/html",
  "digest": "MGIGF4GRGGF5GKV6VNCBAXOE3OR5BTZC",
  "length": 1811,
  "archiveUrl": "https://web.archive.org/web/19961231235847/http://www.nasa.gov:80/",
  "rawArchiveUrl": "https://web.archive.org/web/19961231235847id_/http://www.nasa.gov:80/",
  "scrapedAt": "2026-10-03T05:21:54.874Z"
}
```

| Field | How to read it |
|---|---|
| `inputUrl`, `matchType` | The entry as you typed it and how it was looked up, so rows group back to your list. |
| `originalUrl` | The URL as the archive captured it, with scheme, port and query string. |
| `urlKey` | The archive's own key for the page. Two spellings of one page share it. |
| `timestamp`, `capturedAt` | When the capture was taken: the archive's 14-digit stamp, and the same moment in UTC. |
| `statusCode` | The HTTP status the archive got. `null` on revisit records. |
| `mimeType` | The content type it got, or `warc/revisit`. |
| `digest` | A fingerprint of the content. Equal digests mean identical content. |
| `length` | The size of the stored record in bytes, compressed. It is not the size of the page. |
| `rawArchiveUrl` | The same capture as it was first served, without the archive's toolbar or rewritten links. |

### 🧾 Reading the output

| Row | How to spot it | Billed |
|---|---|---|
| A capture | `recordType` is `capture` | yes |
| The sample | `recordType` is `sample`, with `_sample: true`. Only when no pages were given | no |
| A diagnostic | `recordType` is `diagnostic`, with `_diagnostic: true` and an `errorCode` | no |

Keep the rows whose `recordType` is `capture` to get the captures alone.

| Code | What it means |
|---|---|
| `BAD_INPUT` | An entry that isn't a usable web page or site, or a setting the actor can't read. The `message` says which. |
| `NO_CAPTURES` | The archive has nothing for that page or site, or nothing that matches your dates and filters. |
| `EXCLUDED` | The site is excluded from the Wayback Machine, so its captures can't be listed. |
| `RESTRICTED` | The Wayback Machine keeps this site's capture list private. Seen on theguardian.com. |
| `NOT_ANSWERED` | The archive didn't answer for that page or site this time. Try it again later, and narrow a big domain with dates or a path. |
| `REFUSED` | The archive turned the lookup down. Rare, and worth telling us about. |
| `PARTIAL` | A later page of a long list couldn't be read. The captures before it are in the dataset. |
| `NONE_MATCHED` | The archive listed captures, but none of them matched what was asked. |
| `NOT_LOOKED_UP` | The run stopped before it reached that entry: a cap, the time limit, or a pause the archive asked for. |
| `MAX_CHARGE_TOO_LOW` | The maximum charge set for the run doesn't cover one capture, so nothing was looked up. |

One page can appear twice in the same second, once over http and once over https. The archive holds
both, so both are listed. The run report (`RUN_REPORT` in the key-value store) gives the outcome for
each entry, including whether it stopped at your cap.

### 💡 What people use it for

- Working out when a page changed, by listing only the captures where its content did.
- Rebuilding redirects after a migration, from a list of every page the old site ever had.
- Checking a domain's past before buying it: when it was first captured, when it went quiet, what it
  served in between.
- Finding the captures around a date you have to cite, with links that open the page as it was.

Checking a domain before you buy it, in three steps:

1. Run this actor on the domain's home page with `collapse` set to `year`, to see when it was first
   captured and when the captures stop.
2. Put the same domain into the `domains` field of
   [Domain Inspector](https://apify.com/dami_studio/domain-inspector) for today's registrar and expiry
   date, in `rdap.registrar` and `rdap.expiresAt`.
3. Open the `archiveUrl` of the last few captures before it went quiet to see what it was serving.

### 🚧 What it does not do

- **It lists captures and does not fetch them.** No HTML, text or files come back. The links do.
- **Only what the archive holds.** A page it never captured, or captured under a spelling you didn't
  give, won't be there.
- **Thinning a big path or domain can stop short.** The archive can only thin a site as one long
  list, so the actor reads the captures itself and keeps one per page per period. It stops after
  reading 25 captures for each one you asked for, 200,000 at most, and the run report says when.
- **Newest first is for exact pages only.**
- **Excluded sites stay excluded.** Some owners have asked the archive to hide their site, and those
  come back as `EXCLUDED`. A few large sites have their capture lists kept private by the archive,
  and those come back as `RESTRICTED`.
- **`length` is the stored size, compressed,** not the size of the page.
- **A whole domain with a filter that matches little can be too big to answer.** The archive gives up
  on a query after a minute, and such an entry comes back `NOT_ANSWERED`. Narrow it with dates or a
  path.

### 🧭 Which archive scraper do you need?

| If you want | Use |
|---|---|
| The captures of a page or a site in the Wayback Machine | This one |
| Books, audio, film and other items held on archive.org | [Internet Archive Scraper](https://apify.com/dami_studio/internet-archive-scraper) |
| A site's pages as they read today, as clean text | [Website Intelligence Crawler](https://apify.com/dami_studio/website-intelligence-crawler) |
| DNS, registration and certificate details for a domain | [Domain Inspector](https://apify.com/dami_studio/domain-inspector) |
| A screenshot of a page as it looks now | [Website Screenshot Generator](https://apify.com/dami_studio/website-screenshot-generator) |

### ❓ Questions people ask

#### Do I need an archive.org account or a key?

No. The capture index is public.

#### How do I see a page as it looked on a given day?

Find the row with the date you want and open its `archiveUrl`. For the original bytes without the
archive's toolbar, use `rawArchiveUrl`.

#### Why do some captures have no status code?

They are revisit records: the archive saw the same content again and stored a pointer to the earlier
copy. Their `mimeType` is `warc/revisit`.

#### What does excluded mean?

The site's owner asked the Wayback Machine not to show it. The archive answers the same way for
everyone.

#### Can I list every page a website ever had?

Yes. Give the domain, set `matchType` to `domain` and `collapse` to `url`, which lists each archived
URL once.

#### Can I call it from code or connect it to an AI assistant?

Yes. The [API tab](https://apify.com/dami_studio/wayback-machine-scraper/api/python) has ready-made
code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect
`https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper`. Either way the
run happens on your Apify account at the same price.

#### Is scraping the Wayback Machine legal?

The capture index is public and this reads it the way the archive's own calendar does. What you may
do with an archived page depends on the page. Apify's write-up on
[the legality of web scraping](https://blog.apify.com/is-web-scraping-legal/) is a good starting
point, and we are not lawyers.

### 🆘 If something breaks

Open the **Issues** tab on the actor page. Send the run ID and the pages or sites you used. The
`errorCode` on a diagnostic row usually names the problem on its own.

# Actor input Schema

## `urls` (type: `array`):

The pages or sites to look up, one per line. A full link and a bare address find the same captures: https://www.nasa.gov/ and nasa.gov are one page to the archive. End an address with /\* for everything under that path, or start it with \*. for a whole domain and its subdomains. Up to 500 addresses a run.

## `matchType` (type: `string`):

Exact is that one page. A path takes every page whose address starts with it. A host takes everything on that host, and a domain adds every subdomain. A /\* or \*. typed into an entry wins over this for that entry.

## `from` (type: `string`):

Only captures taken on or after this date, in UTC. Write 2015, 2015-06 or 2015-06-01. A date in any other format is refused before the run starts. Leave it empty to start at the first capture.

## `to` (type: `string`):

Only captures taken up to and including this date, in UTC. 2019 means through the last day of 2019. Leave it empty to run to the latest capture.

## `statusCodes` (type: `array`):

Only captures that got these answers when the archive took them, such as 200 or 404, or a whole class written 3xx. Leave it empty for all. Captures the archive stored as unchanged since the one before carry no status, so any status filter leaves them out.

## `mimeTypes` (type: `array`):

Only captures of these content types, written in full: text/html, application/pdf, or a family such as image/\*. A bare "html" matches nothing. Leave it empty for all.

## `collapse` (type: `string`):

Every capture lists them all. Only when the content changed drops a capture identical to the one before it. One per day, month or year keeps the first capture of each period, page by page. One per address lists every archived address once with its first capture, which is the quick way to list every page a site ever had.

## `maxCapturesPerUrl` (type: `integer`):

The most captures from one address, or from one path, host or domain. A busy home page has hundreds of thousands, so this is your main spending cap.

## `maxCaptures` (type: `integer`):

The most captures in the whole run, across all your entries. Each capture returned is one charge.

## `newestFirst` (type: `boolean`):

Start from the latest capture and work back. Only for exact pages. With a thinning option on, it keeps the latest capture of each period instead of the first.

## Actor input object example

```json
{
  "urls": [
    "https://www.nasa.gov/"
  ],
  "matchType": "exact",
  "collapse": "year",
  "maxCapturesPerUrl": 100,
  "maxCaptures": 100000,
  "newestFirst": false
}
```

# Actor output Schema

## `results` (type: `string`):

One row per capture, address by address, oldest first unless newest first was asked for. Free rows marked \_diagnostic say when an address had no captures, is excluded from the archive, got no answer, or was not looked up.

## `report` (type: `string`):

What happened for each address: its status, captures returned, pages read, whether a cap stopped it, and why the run stopped.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "https://www.nasa.gov/"
    ],
    "collapse": "year",
    "maxCapturesPerUrl": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("dami_studio/wayback-machine-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "urls": ["https://www.nasa.gov/"],
    "collapse": "year",
    "maxCapturesPerUrl": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("dami_studio/wayback-machine-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "https://www.nasa.gov/"
  ],
  "collapse": "year",
  "maxCapturesPerUrl": 100
}' |
apify call dami_studio/wayback-machine-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/Xypq41fIdMhioO7pY/builds/Vzr8wjVtS93i6GRak/openapi.json
