# Wayback Machine Snapshot Checker (Internet Archive) (`smilemask/wayback-machine-snapshot-checker`) Actor

Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.

- **URL**: https://apify.com/smilemask/wayback-machine-snapshot-checker.md
- **Developed by:** [돈벼락](https://apify.com/smilemask) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Wayback Machine Snapshot Checker (Internet Archive)

Check whether any URL has been archived by the Internet Archive Wayback Machine: closest snapshot date, direct archived link, and (optionally) the full list of known snapshot timestamps. Uses the free, official archive.org API - no API key needed.

### What does Wayback Machine Snapshot Checker (Internet Archive) do?

Check if and when a URL was archived by the Wayback Machine, with the closest snapshot link.

### What can you use it for?

- Check whether a page has ever been archived before it disappears or changes.
- Find the oldest or most recent snapshot of a competitor's page for research.
- Verify that your own important pages are being preserved by the Wayback Machine.
- Get a direct archive.org link to cite an old version of a page.

### Why use this Actor?

- **Fast and cheap:** lightweight HTTP-based Actor with no browser, so runs finish in seconds and cost very little.
- **No API keys or accounts needed** for the data source (see notes below).
- **Clean, structured JSON** that you can export as CSV, Excel, JSON or XML, or pull through the Apify API and integrations.
- **Scheduling, webhooks and integrations:** run it daily and send results to Google Sheets, Slack, Zapier, Make, n8n or your own API.

### Input

Configure the run in the **Input** tab (all fields have sensible defaults).

| Field | Type | Required | Default | Description |
|---|---|---|---|---|
| `startUrls` | array of URLs | yes | - | URLs to check (one result per URL). |
| `timestamp` | string | no | - | YYYYMMDD (or a shorter prefix like YYYY). Leave empty to get the most recent snapshot. |
| `includeFullHistory` | boolean | no | `false` | Also fetch every known snapshot timestamp for each URL via the CDX API (slower; degrades gracefully if the Internet Archive CDX service is temporarily down). |
| `maxSnapshotsPerUrl` | integer | no | `100` | Only used when "Include full snapshot history" is on. |

#### Example input

```json
{
  "startUrls": [
    {
      "url": "https://example.com"
    },
    {
      "url": "https://apify.com"
    }
  ]
}
```

### Output

Results are stored in the default dataset. Download them from the **Output** tab or via the API.

#### Example result

```json
{
  "url": "https://example.com/",
  "archived": true,
  "closestSnapshotDate": "2026-09-23T01:54:39Z",
  "closestSnapshotUrl": "http://web.archive.org/web/20260923015439/https://example.com/",
  "snapshotCount": null,
  "snapshots": null,
  "historyNote": "Not requested (turn on \"Include full snapshot history\").",
  "fetchedAt": "2026-09-23T02:25:52.393Z"
}
```

#### Output fields

| Field | Type | Example |
|---|---|---|
| `url` | string | "https://example.com/" |
| `archived` | boolean | true |
| `closestSnapshotDate` | string | "2026-09-23T01:54:39Z" |
| `closestSnapshotUrl` | string | "http://web.archive.org/web/20260923015439/https://exampl... |
| `snapshotCount` | null | null |
| `snapshots` | null | null |
| `historyNote` | string | "Not requested (turn on "Include full snapshot history")." |
| `fetchedAt` | string | "2026-09-23T02:25:52.393Z" |

### How much does it cost?

The Actor is billed by Apify according to the pricing shown on the **Pricing** tab of this page. Runs are small and fast, and the Apify Free plan includes monthly platform credits so you can try it at no cost. You can set a maximum cost per run, and the Actor stops cleanly when that limit is reached.

### Notes and limitations

- Uses the free, official archive.org "available" API for the closest snapshot, and the CDX API (optional) for the full snapshot history.
- The Internet Archive occasionally has short outages; if the optional full-history lookup fails, the row still includes the closest-snapshot result with a note, instead of failing the whole run.
- Only publicly archived pages are covered - pages blocked by robots.txt at crawl time, or excluded by a takedown request, will not appear.

### Tips

- Start with a small test run, then scale up the input.
- Use **Schedules** to automate recurring runs and **Webhooks** or **Integrations** to deliver the data where you need it.
- Call the Actor from your code with the Apify API or the official JavaScript and Python clients.

### FAQ

**What does archived: false mean?**

The Wayback Machine has no snapshot of that exact URL. Try without a trailing slash or query string, since those are indexed as different URLs.

**Can I get every snapshot, not just the closest?**

Yes, turn on "Include full snapshot history" to add every known timestamp for that URL (subject to the occasional Internet Archive outage noted above).

### Feedback

Found a bug or need an extra field? Open an issue from the **Issues** tab of this Actor and it will be looked at quickly.

### More tools from the same author

- [Sitemap URL Extractor: All URLs from XML Sitemaps](https://apify.com/SmileMask/sitemap-url-extractor) - Extract every URL (with lastmod, priority, changefreq) from any website sitemap, index or .gz file.
- [Broken Link Checker: Find 404s & Dead Links on Any Website](https://apify.com/SmileMask/broken-link-checker) - Crawl a site and list every broken internal and external link with the pages that contain it.
- [QR Code Generator: URLs, Text & More (Bulk, PNG/SVG)](https://apify.com/SmileMask/qr-code-generator) - Generate QR codes in bulk as PNG or SVG - runs entirely locally, no external API, never breaks.

# Changelog

This Actor's version history is a separate document: https://apify.com/smilemask/wayback-machine-snapshot-checker/changelog.md

# Actor input Schema

## `startUrls` (type: `array`):

URLs to check (one result per URL).

## `timestamp` (type: `string`):

YYYYMMDD (or a shorter prefix like YYYY). Leave empty to get the most recent snapshot.

## `includeFullHistory` (type: `boolean`):

Also fetch every known snapshot timestamp for each URL via the CDX API (slower; degrades gracefully if the Internet Archive CDX service is temporarily down).

## `maxSnapshotsPerUrl` (type: `integer`):

Only used when "Include full snapshot history" is on.

## Actor input object example

```json
{
  "startUrls": [
    {
      "url": "https://example.com"
    },
    {
      "url": "https://apify.com"
    }
  ],
  "includeFullHistory": false,
  "maxSnapshotsPerUrl": 100
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "startUrls": [
        {
            "url": "https://example.com"
        },
        {
            "url": "https://apify.com"
        }
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("smilemask/wayback-machine-snapshot-checker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = { "startUrls": [
        { "url": "https://example.com" },
        { "url": "https://apify.com" },
    ] }

# Run the Actor and wait for it to finish
run = client.actor("smilemask/wayback-machine-snapshot-checker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "startUrls": [
    {
      "url": "https://example.com"
    },
    {
      "url": "https://apify.com"
    }
  ]
}' |
apify call smilemask/wayback-machine-snapshot-checker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,smilemask/wayback-machine-snapshot-checker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/SZusUrVyY8htMni8a/builds/Jje8U2DdaaFCGiz1l/openapi.json
