# Wayback Machine Snapshots and Archived URLs (`pistachio_implementation/wayback-machine-snapshots`) Actor

Get the Wayback Machine capture history of any URL, every archived URL of a domain, or the archived copy closest to a date. Returns capture time, status, content type, size, digest and ready links to the archived and raw copy. Uses the Internet Archive's public CDX API.

- **URL**: https://apify.com/pistachio\_implementation/wayback-machine-snapshots.md
- **Developed by:** [Hay Equipos](https://apify.com/pistachio_implementation) (community)
- **Categories:** SEO tools, Developer tools, AI
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

$0.50 / 1,000 archive row saveds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wayback Machine Snapshots and Archived URLs

See what a website looked like and how it changed, straight from the Internet Archive. Give the actor URLs or domains and choose one of three jobs:

1. **Snapshots:** the full capture history of a page, with date, HTTP status, content type, size and a content fingerprint, optionally thinned to one capture per day, month or year, or to captures where the content actually changed.
2. **Archived URLs:** every distinct URL the Wayback Machine holds under a domain or path, with its first capture. Great for recovering old pages, finding lost content and redirect planning.
3. **Closest:** the one archived copy nearest to a date, for "what did this page say on 15 January 2020?"

Every row comes with two ready links: the normal Wayback page and the raw archived copy without the Wayback toolbar, which is what you want to feed an AI model or a parser.

The actor uses the Internet Archive's own public CDX and availability APIs. No page scraping, no login, no proxies. Requests are paced one at a time with backoff, as the Archive asks.

### What you can use it for

- **SEO and site migrations:** list every URL a domain ever had, find old pages that still earn links, and build redirect maps.
- **Expired domain research:** check what a domain hosted before you buy it.
- **Competitor and pricing history:** see when a competitor's page changed and open each version.
- **AI and research datasets:** collect archived copies of pages from a given date range for training, evaluation or analysis.
- **Evidence and compliance:** find the capture of a terms page or a claim nearest to a given date.
- **Content recovery:** find lost blog posts and documents, including PDFs, with the content type filter.

### Input

| Field | What it does | Default |
|---|---|---|
| URLs or domains | Pages, domains or path prefixes, one per line | required |
| What to get | Snapshots, Archived URLs, or Closest | Snapshots |
| From date, To date | 2019, 2019-06 or 2019-06-30 | none |
| Keep one snapshot per | Every capture, hour, day, month, year, or content change | every capture |
| Only successful captures | Keep HTTP 200 captures only | off |
| Content type filter | For example text/html or application/pdf | none |
| Include subdomains | Archived URLs mode: also list blog.example.com and similar | off |
| Closest to date | Closest mode: the target date (empty means the most recent) | none |
| Maximum rows per URL or domain | Rows come oldest first | 500 |
| Maximum rows in total | Stop after this many rows | 5,000 |

Example input, one capture per year for a home page:

```json
{
  "urls": ["apify.com"],
  "mode": "snapshots",
  "from": "2015",
  "to": "2025",
  "collapse": "year",
  "onlySuccessful": true
}
```

Example input, every archived HTML page under a docs path:

```json
{
  "urls": ["crawlee.dev/docs/"],
  "mode": "urls",
  "mimeType": "text/html",
  "onlySuccessful": true,
  "maxRowsPerInput": 2000
}
```

### Output

```json
{
  "input": "apify.com",
  "mode": "snapshots",
  "url": "http://apify.com:80/",
  "capturedAt": "2015-02-16T02:12:25Z",
  "timestamp": "20150216021225",
  "statusCode": 200,
  "mimeType": "text/html",
  "digest": "3Z65GIU4US7YSBJ3HD2AU4WJP3YZNWYN",
  "lengthBytes": 440,
  "archiveUrl": "https://web.archive.org/web/20150216021225/http://apify.com:80/",
  "rawArchiveUrl": "https://web.archive.org/web/20150216021225id_/http://apify.com:80/",
  "scrapedAt": "2026-09-27T07:21:05.981Z"
}
```

- `digest` is the Archive's fingerprint of the captured content. Two captures with the same digest are identical.
- `lengthBytes` is the compressed size stored by the Archive.
- In Archived URLs mode each row also has `firstCapturedAt`.
- In Closest mode each row also has `requestedDate`.
- Inputs with no captures are listed in `RUN_SUMMARY` in the run's key value store and cost nothing.

### Pricing

Pay per event. No start fee, no subscription, no platform usage charged on top.

| Event | Price |
|---|---|
| Archive row saved | $0.0005 (50 cents per 1,000 rows) |

Example: the yearly history of 100 home pages over 10 years is about 1,000 rows, about $0.50. Inputs with no captures are free. Set a maximum charge per run in Apify and the actor stops cleanly when it is reached.

### Limits

- The Internet Archive's API can be slow (10 to 30 seconds for a large domain) and sometimes answers "busy". The actor waits and retries; very large domains take minutes.
- Rows come oldest first. For recent captures only, set the From date.
- The actor returns capture records and links, not the archived page content itself. Open `rawArchiveUrl` to get the page as captured.
- Sites that asked the Archive to exclude them return no rows.

### FAQ

**Snapshots or Archived URLs, which do I need?** Snapshots follow one exact page through time. Archived URLs list all the different pages under a domain or folder.

**How do I see only real changes to a page?** Set "Keep one snapshot per" to Content change. Consecutive identical captures are dropped.

**Does it include subdomains?** In Archived URLs mode, turn on Include subdomains.

**Can an AI agent use it?** Yes. It has clear inputs, returns small rows, has no start fee, and the raw archive links can be fetched directly by the agent.

**Is this allowed?** The CDX API is the Internet Archive's public interface for exactly this kind of lookup. The actor paces its requests and backs off when the Archive is busy.

# Actor input Schema

## `urls` (type: `array`):

One per line. A page (example.com/about), a domain (example.com) or a path prefix (example.com/blog/). The scheme and www are optional.

## `mode` (type: `string`):

Snapshots: every capture of exactly this URL. Archived URLs: every distinct URL the archive holds under this domain or path, with its first capture. Closest: the one archived copy nearest to a date.

## `from` (type: `string`):

Optional. 2019, 2019-06 or 2019-06-30. Applies to snapshots and archived URLs.

## `to` (type: `string`):

Optional. 2024, 2024-12 or 2024-12-31. A year or month covers the whole period.

## `collapse` (type: `string`):

Thin out the capture history. Content change keeps a capture only when the page content changed from the previous one. Applies to snapshots mode.

## `onlySuccessful` (type: `boolean`):

Drop redirects, errors and not found captures.

## `mimeType` (type: `string`):

Optional. For example text/html for pages only, or application/pdf for PDFs. Regular expressions such as image/.\* work too.

## `includeSubdomains` (type: `boolean`):

In archived URLs mode, also list URLs on subdomains (blog.example.com, docs.example.com).

## `closestTo` (type: `string`):

For closest mode. 2020-01-15 finds the capture nearest that day. Leave empty for the most recent capture.

## `maxRowsPerInput` (type: `integer`):

Popular sites have hundreds of thousands of captures. Rows are returned oldest first.

## `maxRows` (type: `integer`):

Stop after saving this many rows across all inputs.

## Actor input object example

```json
{
  "urls": [
    "apify.com"
  ],
  "mode": "snapshots",
  "collapse": "none",
  "onlySuccessful": false,
  "includeSubdomains": false,
  "maxRowsPerInput": 500,
  "maxRows": 5000
}
```

# Actor output Schema

## `results` (type: `string`):

All rows the run saved to the default dataset.

## `summary` (type: `string`):

The RUN\_SUMMARY record: counts and problems for the whole run.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "urls": [
        "apify.com"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("pistachio_implementation/wayback-machine-snapshots").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "urls": ["apify.com"] }

# Run the Actor and wait for it to finish
run = client.actor("pistachio_implementation/wayback-machine-snapshots").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "urls": [
    "apify.com"
  ]
}' |
apify call pistachio_implementation/wayback-machine-snapshots --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,pistachio_implementation/wayback-machine-snapshots"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/ArjhIa8Qix59buN6q/builds/dYI8PblAkUIGRNJbt/openapi.json
