Wayback Machine Scraper - Snapshot History of Any URL avatar

Wayback Machine Scraper - Snapshot History of Any URL

Pricing

from $0.12 / 1,000 captures

Go to Apify Store
Wayback Machine Scraper - Snapshot History of Any URL

Wayback Machine Scraper - Snapshot History of Any URL

Every capture the Wayback Machine holds of the pages or sites you list, one row each: when it was taken, HTTP status, content type, size, digest and a link to the archived copy. Filter by date, status and type, or keep one per day, month or year. $0.13 per 1,000 captures.

Pricing

from $0.12 / 1,000 captures

Rating

0.0

(0)

Developer

Dami's Studio

Dami's Studio

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Share

Give it pages or whole sites and get every capture the Wayback Machine holds of them, one row per capture: when it was taken, the HTTP status, the content type, the size and a link to the archived copy. It works on one page, everything under a path, a whole host, or a domain with its subdomains, with date, status and content type filters.

It lists captures and does not download them, so each row is a link to the copy rather than the copy itself. And a few sites are excluded from the archive or keep their capture list private, which no setting gets around.

InputWeb pages or sites: one page, a path, a host or a whole domain
OutputOne row per capture: capture time, the URL as captured, status, content type, digest, stored size, archive links
Ceiling100,000 captures per entry, 1,000,000 per run, 500 entries
Speed3.5 to 5 seconds per page or site looked up, with or without captures, so a full list of 500 takes about forty minutes. Give a long list a longer timeout
Account neededNone from you
Price$0.13 per 1,000 captures, flat on every plan. The free plan's $5 a month covers about 38,000

🔍 What Wayback Machine Scraper does

For each page or site it reads the Wayback Machine's capture index, the list behind the archive's own calendar pages, and turns each entry into a row. A busy home page has hundreds of thousands of captures, so the run reads them a few thousand at a time, oldest first, until it has the number you asked for or the list ends.

The archive applies your dates and filters, and every row is checked again before it is kept. A capture outside your dates, with another status or another content type, is never charged.

To make a long history readable, keep one capture per day, month or year, keep only the captures where the content changed, or list each archived URL once. On a path or a domain this is done page by page, so one page's capture never hides another's.

The run keeps to a couple of dozen requests a minute, well inside what the archive asks of automated tools. If the archive asks for a pause anyway, the run stops there and lists the entries it did not reach.

📋 What data you get from each Wayback Machine capture

What you getField
The page or site you gave, and how it was looked upinputUrl, matchType
The URL as the archive captured it, and the archive's key for itoriginalUrl, urlKey
When the capture was takentimestamp, capturedAt
The HTTP status and the content type the archive gotstatusCode, mimeType
A fingerprint of the content, and the stored sizedigest, length
The capture in the archive's viewer, and as first servedarchiveUrl, rawArchiveUrl
When the row was readscrapedAt

▶️ How to scrape the Wayback Machine

  1. Open Wayback Machine Scraper and click Try for free.
  2. Paste your pages or sites into Web addresses, one per line.
  3. Pick What each address covers, then add dates or filters if you want them.
  4. Set Captures per address and click Start.
  5. Download the dataset as JSON, CSV or Excel, or read it from the Apify API.

💰 How much does it cost to scrape the Wayback Machine?

$0.13 per 1,000 captures. Flat on every Apify plan, no volume tiers. On the free plan, the $5 Apify gives you each month covers about 38,000 captures.

You pay per capture delivered. The sample row, diagnostic rows, captures the checks drop and repeats are not charged, and a page with no captures adds nothing. Captures per address and Captures in total are the two caps on what one run can spend.

📥 What you give it

{
"urls": ["https://www.nasa.gov/", "apify.com/store/*"],
"matchType": "exact",
"from": "2015",
"to": "2020-06",
"statusCodes": ["200"],
"mimeTypes": ["text/html"],
"collapse": "month",
"maxCapturesPerUrl": 1000,
"maxCaptures": 100000,
"newestFirst": false
}
FieldDefaultWhat it is
urlsnone, the form starts with https://www.nasa.gov/The pages or sites, one per line, up to 500. nasa.gov and https://www.nasa.gov/ are the same page to the archive. End one with /* for everything under that path, or start it with *. for a domain and its subdomains.
matchTypeexactexact for that page, prefix for everything under the path, host for every page on the host, domain for the host and its subdomains. A /* or *. typed into an entry wins over this.
from, tononeDates in UTC, written 2015, 2015-06 or 2015-06-01. to covers the whole period, so 2019 runs to the last second of 2019. A date in any other format is refused before the run starts; an impossible one, like 2021-02-30, stops the run before it looks anything up.
statusCodesallHTTP codes such as 200 or 404, or a class such as 3xx.
mimeTypesallContent types in full, such as text/html or application/pdf, or a family such as image/*. A bare html matches nothing.
collapsenone, the form starts on yeardigest drops a capture identical to the one before it. day, month and year keep the first capture of each period. url lists each archived URL once, with its first capture.
maxCapturesPerUrl1000, the form starts at 100The most captures from one page, path, host or domain, up to 100,000.
maxCaptures100000The most captures in the whole run, up to 1,000,000.
newestFirstfalseStart from the latest capture and work back. Exact pages only. With thinning on, it keeps the latest capture of each period rather than the first.

A status filter leaves out unchanged captures. When the archive finds the same content again it often stores a revisit record, which has no status code and the type warc/revisit. Without a filter those rows are part of the history. Ask for any status and they drop out.

📤 What you get back

A real row from a run on 3 October 2026: nasa.gov's first capture.

{
"recordType": "capture",
"inputUrl": "https://www.nasa.gov/",
"matchType": "exact",
"originalUrl": "http://www.nasa.gov:80/",
"urlKey": "gov,nasa)/",
"timestamp": "19961231235847",
"capturedAt": "1996-12-31T23:58:47Z",
"statusCode": 200,
"mimeType": "text/html",
"digest": "MGIGF4GRGGF5GKV6VNCBAXOE3OR5BTZC",
"length": 1811,
"archiveUrl": "https://web.archive.org/web/19961231235847/http://www.nasa.gov:80/",
"rawArchiveUrl": "https://web.archive.org/web/19961231235847id_/http://www.nasa.gov:80/",
"scrapedAt": "2026-10-03T05:21:54.874Z"
}
FieldHow to read it
inputUrl, matchTypeThe entry as you typed it and how it was looked up, so rows group back to your list.
originalUrlThe URL as the archive captured it, with scheme, port and query string.
urlKeyThe archive's own key for the page. Two spellings of one page share it.
timestamp, capturedAtWhen the capture was taken: the archive's 14-digit stamp, and the same moment in UTC.
statusCodeThe HTTP status the archive got. null on revisit records.
mimeTypeThe content type it got, or warc/revisit.
digestA fingerprint of the content. Equal digests mean identical content.
lengthThe size of the stored record in bytes, compressed. It is not the size of the page.
rawArchiveUrlThe same capture as it was first served, without the archive's toolbar or rewritten links.

🧾 Reading the output

RowHow to spot itBilled
A capturerecordType is captureyes
The samplerecordType is sample, with _sample: true. Only when no pages were givenno
A diagnosticrecordType is diagnostic, with _diagnostic: true and an errorCodeno

Keep the rows whose recordType is capture to get the captures alone.

CodeWhat it means
BAD_INPUTAn entry that isn't a usable web page or site, or a setting the actor can't read. The message says which.
NO_CAPTURESThe archive has nothing for that page or site, or nothing that matches your dates and filters.
EXCLUDEDThe site is excluded from the Wayback Machine, so its captures can't be listed.
RESTRICTEDThe Wayback Machine keeps this site's capture list private. Seen on theguardian.com.
NOT_ANSWEREDThe archive didn't answer for that page or site this time. Try it again later, and narrow a big domain with dates or a path.
REFUSEDThe archive turned the lookup down. Rare, and worth telling us about.
PARTIALA later page of a long list couldn't be read. The captures before it are in the dataset.
NONE_MATCHEDThe archive listed captures, but none of them matched what was asked.
NOT_LOOKED_UPThe run stopped before it reached that entry: a cap, the time limit, or a pause the archive asked for.
MAX_CHARGE_TOO_LOWThe maximum charge set for the run doesn't cover one capture, so nothing was looked up.

One page can appear twice in the same second, once over http and once over https. The archive holds both, so both are listed. The run report (RUN_REPORT in the key-value store) gives the outcome for each entry, including whether it stopped at your cap.

💡 What people use it for

  • Working out when a page changed, by listing only the captures where its content did.
  • Rebuilding redirects after a migration, from a list of every page the old site ever had.
  • Checking a domain's past before buying it: when it was first captured, when it went quiet, what it served in between.
  • Finding the captures around a date you have to cite, with links that open the page as it was.

Checking a domain before you buy it, in three steps:

  1. Run this actor on the domain's home page with collapse set to year, to see when it was first captured and when the captures stop.
  2. Put the same domain into the domains field of Domain Inspector for today's registrar and expiry date, in rdap.registrar and rdap.expiresAt.
  3. Open the archiveUrl of the last few captures before it went quiet to see what it was serving.

🚧 What it does not do

  • It lists captures and does not fetch them. No HTML, text or files come back. The links do.
  • Only what the archive holds. A page it never captured, or captured under a spelling you didn't give, won't be there.
  • Thinning a big path or domain can stop short. The archive can only thin a site as one long list, so the actor reads the captures itself and keeps one per page per period. It stops after reading 25 captures for each one you asked for, 200,000 at most, and the run report says when.
  • Newest first is for exact pages only.
  • Excluded sites stay excluded. Some owners have asked the archive to hide their site, and those come back as EXCLUDED. A few large sites have their capture lists kept private by the archive, and those come back as RESTRICTED.
  • length is the stored size, compressed, not the size of the page.
  • A whole domain with a filter that matches little can be too big to answer. The archive gives up on a query after a minute, and such an entry comes back NOT_ANSWERED. Narrow it with dates or a path.

🧭 Which archive scraper do you need?

If you wantUse
The captures of a page or a site in the Wayback MachineThis one
Books, audio, film and other items held on archive.orgInternet Archive Scraper
A site's pages as they read today, as clean textWebsite Intelligence Crawler
DNS, registration and certificate details for a domainDomain Inspector
A screenshot of a page as it looks nowWebsite Screenshot Generator

❓ Questions people ask

Do I need an archive.org account or a key?

No. The capture index is public.

How do I see a page as it looked on a given day?

Find the row with the date you want and open its archiveUrl. For the original bytes without the archive's toolbar, use rawArchiveUrl.

Why do some captures have no status code?

They are revisit records: the archive saw the same content again and stored a pointer to the earlier copy. Their mimeType is warc/revisit.

What does excluded mean?

The site's owner asked the Wayback Machine not to show it. The archive answers the same way for everyone.

Can I list every page a website ever had?

Yes. Give the domain, set matchType to domain and collapse to url, which lists each archived URL once.

Can I call it from code or connect it to an AI assistant?

Yes. The API tab has ready-made code for Python, JavaScript and the command line. For Claude, ChatGPT or another MCP client, connect https://mcp.apify.com/?tools=fetch-actor-details,dami_studio/wayback-machine-scraper. Either way the run happens on your Apify account at the same price.

The capture index is public and this reads it the way the archive's own calendar does. What you may do with an archived page depends on the page. Apify's write-up on the legality of web scraping is a good starting point, and we are not lawyers.

🆘 If something breaks

Open the Issues tab on the actor page. Send the run ID and the pages or sites you used. The errorCode on a diagnostic row usually names the problem on its own.