Go to example tasks

Python standard library docs pages in the archive

Distinct URLs archived under docs.python.org/3/library/, one row per page, with when the archive first and last captured it and the status codes it returned over time.

Try for free
Wayback Machine Scraper - Snapshots & Every Archived URL
Wayback Machine Scraper - Snapshots & Every Archived URLneverempty/wayback-machine-scraper
Input
Url
Captured at
Status code
+22 fields
Text
Number
Boolean
List
Object

Input

URLs or domains:docs.python.org/3/library/*
What to return:urls
URL list scope:domain
Thin out snapshots:month
Order:newest
HTTP status:any
Maximum rows per URL:300

Output fields

Input
Url
Captured at
Status code
Capture type
Mime type
Same content as previous
Length bytes
Archive url
First captured at
Last captured at
Capture count
Status codes seen
Complete
Row type
Status
Note
Status class
Timestamp
Digest
Raw archive url
First archive url
Url key
Query
Source
Checked at

How it works

Sign up on Apify01

Create your Apify account to access the Wayback Machine Scraper - Snapshots & Every Archived URL.

Start the run02

The Actor will start running based on the input automatically.

Receive the output03

Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.

Integrate into your workflow04

The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.

Image

Integrate Actor directly into your workflow

Choose from one of 100+ integration options we provide or integrate via API

Webhook

Webhook

n8n

n8n

Make

Make

Zapier

Zapier

Airbyte

Airbyte

Keboola

Keboola

IFTTT

IFTTT

Hubspot

Hubspot

GDrive

GDrive

Gmail

Gmail

Apify MCP

Apify MCP

GitHub

GitHub

Slack

Slack

LangChain

LangChain

LlamaIndex

LlamaIndex

Flowise

Flowise

Pinecone

Pinecone

OpenAI

OpenAI

Mastra

Mastra

Clay

Clay