Go to example tasks
Every archived URL on python.org
Distinct URLs the Internet Archive holds for python.org and its subdomains, each with when it was first and last captured, how many captures exist and every HTTP status code it has returned.
Wayback Machine Scraper - Snapshots & Every Archived URLneverempty/wayback-machine-scraper
Input
Url
Captured at
Status code
+22 fieldsTextNumberBooleanListObject
Input
URLs or domains:python.org
What to return:urls
URL list scope:domain
Thin out snapshots:month
Order:newest
HTTP status:any
Maximum rows per URL:300
Output fields
Input
Url
Captured at
Status code
Capture type
Mime type
Same content as previous
Length bytes
Archive url
First captured at
Last captured at
Capture count
Status codes seen
Complete
Row type
Status
Note
Status class
Timestamp
Digest
Raw archive url
First archive url
Url key
Query
Source
Checked at
Sign up on Apify01
Create your Apify account to access the Wayback Machine Scraper - Snapshots & Every Archived URL.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
