Memento TimeMap Harvester
Pricing
from $0.35 / 1,000 memento parseds
Memento TimeMap Harvester
Point at ANY Memento (RFC 7089) TimeMap and extract every archived snapshot as flat rows — memento URI, datetime (ISO 8601), rel, original URI, collection, raw link line. Works across the Internet Archive, arquivo.pt and every Memento-compliant web archive. Pay per memento.
Pricing
from $0.35 / 1,000 memento parseds
Rating
0.0
(0)
Developer
Datamule
Maintained by CommunityActor stats
0
Bookmarked
1
Total users
0
Monthly active users
4 days ago
Last modified
Categories
Share
Point at any Memento (RFC 7089) TimeMap and get every archived snapshot of a URL as clean, flat rows — across the Internet Archive, arquivo.pt, and every Memento-compliant web archive worldwide.
Memento is the time-travel-for-the-web standard. A TimeMap is a machine-readable application/link-format list of every archived snapshot (a memento) of a URL. Unlike a Wayback-only tool, this actor speaks the cross-archive Memento standard, so one actor works against archives all over the world — you just supply the TimeMap URL.
What you get
One dataset row per memento, with these columns:
| Field | Description |
|---|---|
memento_uri | URL of the archived snapshot |
datetime | Capture time, parsed to ISO 8601 (2014-10-09T10:19:54+00:00) |
datetime_raw | The original RFC 1123 datetime string (lossless) |
rel | The link relation verbatim (memento, first memento, last memento) |
type | Content type param, when the archive supplies one |
collection | Archive collection tag, when present (e.g. arquivo.pt's $root) |
original_uri | The original (live-web) URL, from the TimeMap's rel="original" link |
timemap_self | The canonical TimeMap URL, from the rel="self" link |
timemap_from / timemap_until | The archival span of the URL (ISO 8601), when the archive reports it |
source_timemap | The TimeMap URL you supplied |
_raw | The raw link-format line for this memento (lossless) |
Input
{"timemapUrls": ["https://web.archive.org/web/timemap/link/https://www.example.com","https://arquivo.pt/wayback/timemap/link/https://www.example.com/"],"maxMementosPerSource": 50}
timemapUrls(required) — one or more full TimeMap endpoint URLs. The target URL is embedded in the path. Examples of TimeMap endpoints:- Internet Archive —
https://web.archive.org/web/timemap/link/<URL> - arquivo.pt (Portugal) —
https://arquivo.pt/wayback/timemap/link/<URL> - Icelandic web archive —
https://vefsafn.is/timemap/link/<URL> - …and any other Memento-compliant archive.
- Internet Archive —
maxMementosPerSource(optional) — cap the number of mementos emitted per TimeMap. A busy URL on the Internet Archive can have hundreds of thousands of snapshots; this keeps a run small and cheap. Leave empty to emit them all.userAgent(optional) — override the request User-Agent.
Why this is not a naive comma-splitter
The application/link-format grammar (RFC 6690 / RFC 8288) is deceptively tricky: every memento's datetime is an RFC 1123 date that contains commas ("Thu, 09 Oct 2014 10:19:54 GMT"), and rels can be multi-token (rel="first memento"). This actor uses a proper quote- and angle-bracket-aware tokenizer, so links are split only on true top-level commas, multi-token rels are handled, and archive-specific params (collection, from, until) are all captured.
How it behaves
- A URL an archive has never captured → 0 rows, clean success (never fabricated).
- A source that returns HTML / a 404 / anything that is not a TimeMap → skipped with a warning; the rest of the batch continues. A run where every source fails that way fails loudly.
- Public archives serve TimeMaps anonymously — no credentials needed.
Pricing
Pay-per-event: you are charged once per memento row returned. A run that returns nothing costs nothing.
The Actor respects your maximum cost per run: it checks the remaining budget before every TimeMap request and before writing each row, so it stops cleanly when the budget runs out and never delivers unbilled rows. If billing fails, the run fails instead of finishing silently.
Common uses
- Web archiving / digital preservation research and QA.
- Building a capture timeline for a URL across multiple archives.
- Link-rot and reference-persistence studies (scholarly citation, journalism).
- Competitive / brand monitoring over historical snapshots.