Wayback Machine Scraper - Snapshot History of Any URL
Pricing
from $0.12 / 1,000 captures
Wayback Machine Scraper - Snapshot History of Any URL
Every capture the Wayback Machine holds of the pages or sites you list, one row each: when it was taken, HTTP status, content type, size, digest and a link to the archived copy. Filter by date, status and type, or keep one per day, month or year. $0.13 per 1,000 captures.
Pricing
from $0.12 / 1,000 captures
Rating
0.0
(0)
Developer
Dami's Studio
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 days ago
Last modified
Categories
Share
Give it pages or whole sites and get every capture the Wayback Machine holds of them, one row per capture: when it was taken, the HTTP status, the content type, the size and a link to the archived copy. It works on one page, everything under a path, a whole host, or a domain with its subdomains, with date, status and content type filters.
It lists captures and does not download them, so each row is a link to the copy rather than the copy itself. And a few sites are excluded from the archive or keep their capture list private, which no setting gets around.
| Input | Web pages or sites: one page, a path, a host or a whole domain |
| Output | One row per capture: capture time, the URL as captured, status, content type, digest, stored size, archive links |
| Ceiling | 100,000 captures per entry, 1,000,000 per run, 500 entries |
| Speed | 3.5 to 5 seconds per page or site looked up, with or without captures, so a full list of 500 takes about forty minutes. Give a long list a longer timeout |
| Account needed | None from you |
| Price | $0.13 per 1,000 captures, flat on every plan. The free plan's $5 a month covers about 38,000 |
🔍 What Wayback Machine Scraper does
For each page or site it reads the Wayback Machine's capture index, the list behind the archive's own calendar pages, and turns each entry into a row. A busy home page has hundreds of thousands of captures, so the run reads them a few thousand at a time, oldest first, until it has the number you asked for or the list ends.
The archive applies your dates and filters, and every row is checked again before it is kept. A capture outside your dates, with another status or another content type, is never charged.
To make a long history readable, keep one capture per day, month or year, keep only the captures where the content changed, or list each archived URL once. On a path or a domain this is done page by page, so one page's capture never hides another's.
The run keeps to a couple of dozen requests a minute, well inside what the archive asks of automated tools. If the archive asks for a pause anyway, the run stops there and lists the entries it did not reach.
📋 What data you get from each Wayback Machine capture
| What you get | Field |
|---|---|
| The page or site you gave, and how it was looked up | inputUrl, matchType |
| The URL as the archive captured it, and the archive's key for it | originalUrl, urlKey |
| When the capture was taken | timestamp, capturedAt |
| The HTTP status and the content type the archive got | statusCode, mimeType |
| A fingerprint of the content, and the stored size | digest, length |
| The capture in the archive's viewer, and as first served | archiveUrl, rawArchiveUrl |
| When the row was read | scrapedAt |
▶️ How to scrape the Wayback Machine
- Open Wayback Machine Scraper and click Try for free.
- Paste your pages or sites into Web addresses, one per line.
- Pick What each address covers, then add dates or filters if you want them.
- Set Captures per address and click Start.
- Download the dataset as JSON, CSV or Excel, or read it from the Apify API.
💰 How much does it cost to scrape the Wayback Machine?
$0.13 per 1,000 captures. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 38,000 captures.
You pay per capture delivered. The sample row, diagnostic rows, captures the checks drop and repeats are not charged, and a page with no captures adds nothing. Captures per address and Captures in total are the two caps on what one run can spend.
📥 What you give it
{"urls": ["https://www.nasa.gov/", "apify.com/store/*"],"matchType": "exact","from": "2015","to": "2020-06","statusCodes": ["200"],"mimeTypes": ["text/html"],"collapse": "month","maxCapturesPerUrl": 1000,"maxCaptures": 100000,"newestFirst": false}
| Field | Default | What it is |
|---|---|---|
urls | none, the form starts with https://www.nasa.gov/ | The pages or sites, one per line, up to 500. nasa.gov and https://www.nasa.gov/ are the same page to the archive. End one with /* for everything under that path, or start it with *. for a domain and its subdomains. |
matchType | exact | exact for that page, prefix for everything under the path, host for every page on the host, domain for the host and its subdomains. A /* or *. typed into an entry wins over this. |
from, to | none | Dates in UTC, written 2015, 2015-06 or 2015-06-01. to covers the whole period, so 2019 runs to the last second of 2019. A date in any other format is refused before the run starts; an impossible one, like 2021-02-30, stops the run before it looks anything up. |
statusCodes | all | HTTP codes such as 200 or 404, or a class such as 3xx. |
mimeTypes | all | Content types in full, such as text/html or application/pdf, or a family such as image/*. A bare html matches nothing. |
collapse | none, the form starts on year | digest drops a capture identical to the one before it. day, month and year keep the first capture of each period. url lists each archived URL once, with its first capture. |
maxCapturesPerUrl | 1000, the form starts at 100 | The most captures from one page, path, host or domain, up to 100,000. |
maxCaptures | 100000 | The most captures in the whole run, up to 1,000,000. |
newestFirst | false | Start from the latest capture and work back. Exact pages only. With thinning on, it keeps the latest capture of each period rather than the first. |
A status filter leaves out unchanged captures. When the archive finds the same content again it
often stores a revisit record, which has no status code and the type warc/revisit. Without a filter
those rows are part of the history. Ask for any status and they drop out.
📤 What you get back
A real row from a run on 3 October 2026: nasa.gov's first capture.
{"recordType": "capture","inputUrl": "https://www.nasa.gov/","matchType": "exact","originalUrl": "http://www.nasa.gov:80/","urlKey": "gov,nasa)/","timestamp": "19961231235847","capturedAt": "1996-12-31T23:58:47Z","statusCode": 200,"mimeType": "text/html","digest": "MGIGF4GRGGF5GKV6VNCBAXOE3OR5BTZC","length": 1811,"archiveUrl": "https://web.archive.org/web/19961231235847/http://www.nasa.gov:80/","rawArchiveUrl": "https://web.archive.org/web/19961231235847id_/http://www.nasa.gov:80/","scrapedAt": "2026-10-03T05:21:54.874Z"}
| Field | How to read it |
|---|---|
inputUrl, matchType | The entry as you typed it and how it was looked up, so rows group back to your list. |
originalUrl | The URL as the archive captured it, with scheme, port and query string. |
urlKey | The archive's own key for the page. Two spellings of one page share it. |
timestamp, capturedAt | When the capture was taken: the archive's 14-digit stamp, and the same moment in UTC. |
statusCode | The HTTP status the archive got. null on revisit records. |
mimeType | The content type it got, or warc/revisit. |
digest | A fingerprint of the content. Equal digests mean identical content. |
length | The size of the stored record in bytes, compressed. It is not the size of the page. |
rawArchiveUrl | The same capture as it was first served, without the archive's toolbar or rewritten links. |
🧾 Reading the output
| Row | How to spot it | Billed |
|---|---|---|
| A capture | recordType is capture | yes |
| The sample | recordType is sample, with _sample: true. Only when no pages were given | no |
| A diagnostic | recordType is diagnostic, with _diagnostic: true and an errorCode | no |
Keep the rows whose recordType is capture to get the captures alone.
| Code | What it means |
|---|---|
BAD_INPUT | An entry that isn't a usable web page or site, or a setting the actor can't read. The message says which. |
NO_CAPTURES | The archive has nothing for that page or site, or nothing that matches your dates and filters. |
EXCLUDED | The site is excluded from the Wayback Machine, so its captures can't be listed. |
RESTRICTED | The Wayback Machine keeps this site's capture list private. Seen on theguardian.com. |
NOT_ANSWERED | The archive didn't answer for that page or site this time. Try it again later, and narrow a big domain with dates or a path. |
REFUSED | The archive turned the lookup down. Rare, and worth telling us about. |
PARTIAL | A later page of a long list couldn't be read. The captures before it are in the dataset. |
NONE_MATCHED | The archive listed captures, but none of them matched what was asked. |
NOT_LOOKED_UP | The run stopped before it reached that entry: a cap, the time limit, or a pause the archive asked for. |
MAX_CHARGE_TOO_LOW | The maximum charge set for the run doesn't cover one capture, so nothing was looked up. |
One page can appear twice in the same second, once over http and once over https. The archive holds
both, so both are listed. The run report (RUN_REPORT in the key-value store) gives the outcome for
each entry, including whether it stopped at your cap.
💡 What people use it for
- Working out when a page changed, by listing only the captures where its content did.
- Rebuilding redirects after a migration, from a list of every page the old site ever had.
- Checking a domain's past before buying it: when it was first captured, when it went quiet, what it served in between.
- Finding the captures around a date you have to cite, with links that open the page as it was.
Checking a domain before you buy it, in three steps:
- Run this actor on the domain's home page with
collapseset toyear, to see when it was first captured and when the captures stop. - Put the same domain into the
domainsfield of Domain Inspector for today's registrar and expiry date, inrdap.registrarandrdap.expiresAt. - Open the
archiveUrlof the last few captures before it went quiet to see what it was serving.
🚧 What it does not do
- It lists captures and does not fetch them. No HTML, text or files come back. The links do.
- Only what the archive holds. A page it never captured, or captured under a spelling you didn't give, won't be there.
- Thinning a big path or domain can stop short. The archive can only thin a site as one long list, so the actor reads the captures itself and keeps one per page per period. It stops after reading 25 captures for each one you asked for, 200,000 at most, and the run report says when.
- Newest first is for exact pages only.
- Excluded sites stay excluded. Some owners have asked the archive to hide their site, and those
come back as
EXCLUDED. A few large sites have their capture lists kept private by the archive, and those come back asRESTRICTED. lengthis the stored size, compressed, not the size of the page.- A whole domain with a filter that matches little can be too big to answer. The archive gives up
on a query after a minute, and such an entry comes back
NOT_ANSWERED. Narrow it with dates or a path.
🧭 Which archive scraper do you need?
| If you want | Use |
|---|---|
| The captures of a page or a site in the Wayback Machine | This one |
| Books, audio, film and other items held on archive.org | Internet Archive Scraper |
| A site's pages as they read today, as clean text | Website Intelligence Crawler |
| DNS, registration and certificate details for a domain | Domain Inspector |
| A screenshot of a page as it looks now | Website Screenshot Generator |
❓ Questions people ask
Do I need an archive.org account or a key?
No. The capture index is public.
How do I see a page as it looked on a given day?
Find the row with the date you want and open its archiveUrl. For the original bytes without the
archive's toolbar, use rawArchiveUrl.
Why do some captures have no status code?
They are revisit records: the archive saw the same content again and stored a pointer to the earlier
copy. Their mimeType is warc/revisit.
What does excluded mean?
The site's owner asked the Wayback Machine not to show it. The archive answers the same way for everyone.
Can I list every page a website ever had?
Yes. Give the domain, set matchType to domain and collapse to url, which lists each archived
URL once.
Can I call it from code or connect it to an AI assistant?
Yes. The API tab has ready-made
code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect
https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper. Either way the
run happens on your Apify account at the same price.
Is scraping the Wayback Machine legal?
The capture index is public and this reads it the way the archive's own calendar does. What you may do with an archived page depends on the page. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.
🆘 If something breaks
Open the Issues tab on the actor page. Send the run ID and the pages or sites you used. The
errorCode on a diagnostic row usually names the problem on its own.