Websites Archiver (Wayback Machine)
Pricing
from $0.80 / 1,000 archived urls
Websites Archiver (Wayback Machine)
Effortlessly archive any website with our Automated Website Archiving Tool. It leverages the power of the Wayback Machine at web.archive.org to ensure your sites are preserved for future reference.
Pricing
from $0.80 / 1,000 archived urls
Rating
5.0
(1)
Developer
Web Harvester
Maintained by CommunityActor stats
3
Bookmarked
97
Total users
1
Monthly active users
11 hours ago
Last modified
Categories
Share
Save any list of web pages to the Internet Archive's Wayback Machine and get a permanent snapshot link for each one. It works like clicking Save Page Now on web.archive.org, but for a whole list of URLs in one run, and returns the results as a dataset you can download or connect to other tools.
What does Websites Archiver do?
For every URL you give it, the Actor:
- Asks the Wayback Machine to capture the page (Save Page Now).
- Waits until the capture is finished.
- Returns the snapshot link (
https://web.archive.org/web/<timestamp>/<url>), the capture time, the HTTP status of the page and how many resources were saved.
If the Wayback Machine refuses a URL (unknown host, excluded site, crawler trap), the result says why.
Why archive pages in the Wayback Machine?
- Evidence and compliance: keep a timestamped, third-party copy of terms of service, pricing pages, ads or claims.
- Link rot protection: archive the pages your articles, docs or research cite, so the links keep working.
- Monitoring: schedule the Actor to snapshot competitor or product pages every day or week.
- Before a migration or redesign: preserve the old version of your site.
How to archive websites
- Click Try for free.
- Add the pages in URLs to archive. You can type them, paste a list, or link a text file that contains the URLs.
- Click Start.
- When the run finishes, open the Output tab and download the results as JSON, CSV, Excel or HTML.
Each URL is archived as one page. Links on the page are not followed. To archive a whole site, add all its page URLs (for example from its sitemap).
Input
| Field | Default | Description |
|---|---|---|
URLs to archive (startUrls) | required | Pages to save. Add { "url": "..." } items, or { "requestsFromUrl": "..." } pointing to a text file that contains URLs. Only http(s) URLs are read from the file; other text is skipped. Duplicates are archived once. |
Fast mode (fastArchiveMode) | false | Only submit the URLs and return their job IDs, without waiting for the result. Results then have no snapshot link and archiving is not confirmed. |
Save error pages (archiveErrorPages) | true | Also save pages that answer with HTTP 4xx/5xx, such as 404. When off, those URLs are reported as failed. |
Save list of archived resources (storeArchivedResources) | false | Add the URLs of all saved resources (HTML, scripts, styles, images) to each result. Ignored in fast mode. |
Proxy configuration (proxyConfiguration) | Apify Proxy (datacenter) | Keep the default. The Wayback Machine allows about 3 captures at a time per IP address, so a proxy makes runs faster and more reliable. See Good to know. |
Example input:
{"startUrls": [{ "url": "https://crawlee.dev" },{ "url": "https://docs.apify.com/platform" },{ "requestsFromUrl": "https://example.com/urls-to-archive.txt" }],"archiveErrorPages": true,"storeArchivedResources": false,"proxyConfiguration": { "useApifyProxy": true }}
Output
One result per URL. Example of an archived page:
{"url": "https://docs.apify.com/platform/storage/dataset","status": "archived","archived": true,"archivedUrl": "https://web.archive.org/web/20261003151218/https://docs.apify.com/storage/dataset","archivedAt": "2026-10-03T15:12:18.000Z","capturedUrl": "https://docs.apify.com/storage/dataset","httpStatus": 200,"jobId": "spn2-8c01349fdb95d9bc065af32743f2d4ebc9d3343d","archivedResourcesCount": 127,"outlinksCount": 103,"note": null,"errorCode": null}
Example of a URL the Wayback Machine refused:
{"url": "https://nonexistent-domain-qq55443.com/","status": "failed","archived": false,"archivedUrl": null,"archivedAt": null,"capturedUrl": null,"httpStatus": null,"jobId": null,"archivedResourcesCount": null,"outlinksCount": null,"note": "Cannot resolve host nonexistent-domain-qq55443.com. If the site is online, it may be blocking access from our service.","errorCode": "rejected"}
| Field | Description |
|---|---|
url | URL from your input. |
status | archived: snapshot confirmed. submitted: capture started but not confirmed (fast mode, or still running after 5 minutes). failed: not archived. |
archived | true only when the snapshot is confirmed. |
archivedUrl | Permanent Wayback Machine link. |
archivedAt | Capture time from the Wayback Machine (UTC). |
capturedUrl | URL that was actually captured, after redirects. |
httpStatus | HTTP status the page returned to the Wayback Machine. |
jobId | Save Page Now job ID. |
archivedResourcesCount | Number of saved resources (HTML, scripts, styles, images). |
outlinksCount | Number of links found on the page. |
note | Message from the Wayback Machine, such as why archiving failed or that a recent snapshot was reused. |
errorCode | Short reason code when the URL was not archived or not confirmed. See the table below. |
archivedResources | List of saved resource URLs, only with Save list of archived resources. |
The Overview view shows the snapshot links. The Failed view lists the URLs that were not archived and why.
Error codes
errorCode | status | Meaning |
|---|---|---|
rejected | failed | The Wayback Machine refused the URL: unknown host, site excluded from the archive, or a crawler trap. note has its exact message. |
rate-limited | failed | The Wayback Machine kept refusing new captures after all retries. Run the URL again later or use a proxy. |
capture-timeout | submitted | The capture was still running after 5 minutes. It may still finish later. Check the page's capture list (see below). |
not-found, service-unavailable, blocked-url, … | failed | Codes from the Wayback Machine. Usually the error the page itself returned when Save error pages is off. note has the details. |
job-failed, submit-http-503, unexpected-status-response, … | failed | A temporary Wayback Machine problem that persisted through all 8 retries. Run the URL again later. |
url-list-failed | failed | A linked URL file could not be downloaded or contains no URLs. The other URLs are still archived. url is the file's address. |
invalid-url | failed | An input entry is not an http(s) URL (for example ftp://…). |
request-failed | failed | An unexpected network error after all retries. |
Good to know
- Recent snapshots are reused. If the page was captured shortly before, the Wayback Machine returns that snapshot instead of making a new one, with a note such as "The same snapshot had been made 2 minutes ago. You can make new capture of this URL after 1 hour." The result then has that note, and
archivedAtis the time of the earlier capture. - Speed. A capture usually takes 15–30 seconds. With the default proxy the Actor archives up to 10 URLs in parallel. In our tests, runs of 10–12 URLs took 2–3.5 minutes.
- Proxy. The Wayback Machine allows only about 3 captures at a time per IP address. With a proxy, a rate-limited URL is retried right away from a new IP. Without a proxy, the Actor archives 3 URLs at a time and waits a minute after each rate limit, so runs are slower. Residential proxy also works but costs about 3× more and is not needed.
- Retries. Rate limits and temporary Wayback Machine failures are retried automatically, up to 8 times per URL. Failures caused by the page itself (unknown host, excluded site) are reported right away.
- Checking a page's captures later. Open
https://web.archive.org/web/*/<your URL>to see every snapshot of a page, for example after fast mode or acapture-timeout. If the page redirects, look up thecapturedUrlinstead. - Some sites cannot be archived. Site owners can exclude their site from the Wayback Machine, and some sites block its crawler.
How much does it cost?
The Actor sends plain HTTP requests and runs without a browser on 256 MB of memory, so platform usage is very low. Most of the run time is spent waiting for the Wayback Machine to finish each capture. In our tests, archiving 10–12 URLs used about $0.003–0.004 of platform usage with the default datacenter proxy, and about $0.011 with residential proxy.
Integrations and API
- Schedule it: use Apify Schedules to snapshot the same pages every day, week or month.
- Use the results elsewhere: send them to Google Sheets, Slack, Zapier, Make or webhooks with Apify integrations.
- Run it from code: start the Actor and read its dataset with the Apify API or the Apify JavaScript and Python clients.
Is it legal?
The Actor only uses the public Save Page Now service of the Internet Archive, the same as clicking the button on web.archive.org. Use it responsibly: the Wayback Machine is a free, non-profit service. Don't submit the same pages over and over, and respect the Internet Archive's Terms of Use.
Feedback
Found a bug or missing a feature? Open an issue in the Issues tab.