Wayback CDX → Markdown for RAG
Pricing
from $3.00 / 1,000 snapshot markdowns
Wayback CDX → Markdown for RAG
Query the Wayback Machine CDX index, fetch selected snapshot HTML, and extract clean Markdown with trafilatura for RAG pipelines. Optional heading-aware chunks. Competitors often stop at CDX URL lists — this Actor returns page body Markdown.
Pricing
from $3.00 / 1,000 snapshot markdowns
Rating
0.0
(0)
Developer
新世紀書僮
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
13 hours ago
Last modified
Categories
Share
Query the Wayback Machine CDX index, fetch selected snapshot HTML, and extract clean Markdown for RAG — not just a list of archive URLs.
Many tools stop after listing CDX captures. This Actor uses open-source waybackpack (MIT) for CDX + snapshot fetch and trafilatura (Apache-2.0) to turn archived HTML into Markdown dataset rows, with optional heading-aware RAG chunks. Broken captures are reported, not fatal. Default memory: 256 MB. No browser, no AI keys.
What you get
- 📚 CDX search — page URL(s) and/or domain prefix (
example.com/*) via waybackpack - 🗓️ Filters — date range, collapse (day/month/digest), status codes, MIME substrings
- 🎯 Selection — latest, oldest, or even sample; per-target and global caps
- 📝 Markdown body — trafilatura extraction (tables optional; comments off by default)
- 🧩 Optional RAG chunks — heading-aware chunks with token estimate
- 🧯 Per-target report —
CDX_REPORTin the key-value store; failed fetches are free rows - 💾 HTTP only — 256 MB default
Measured results
Local + private cloud (2026-09-30 Asia/Taipei). Settled usageTotalUsd on own runs; see docs/PRICING.md.
| Test | Result |
|---|---|
example.com/ CDX 2024, collapse month, 2 snapshots + chunks (local) | 2 Markdown items + 2 chunks in ~47 s; CDX 12→6→2 |
example.com/ + info.cern.ch/ (2020, 1 snap each, local) | 2 Markdown items in ~40 s; 0 fetch errors |
Cloud 1 snapshot bGbenU8jxfkK90zMh (256 MB) | 1 item; ~66 s; memMax ~97 MB; settled $0.001083 |
Cloud 5 URLs × sample, max 10 N3u4tjN3UeWN4J6hH (build 0.1.2) | 10 items; ~202 s; memMax |
Cloud 2 URLs + domain, items+chunks tnMfOC9kJbluFSncc | 7 items + 10 free chunks; ~96 s; memMax ~107 MB; settled $0.001599 |
Use cases
- Rebuild historical page text for RAG / knowledge bases from Wayback captures
- Sample a site’s archived paths under a domain and extract readable Markdown
- Compare snapshots over time (collapse by day/month, then extract)
How to use
- Add page URLs and/or Domains.
- Optional: set From/To date, Collapse, Max snapshots per URL/domain, Output.
- Start the Actor. Rows appear in the Dataset;
CDX_REPORTandOUTPUTin the Key-value store.
Input example
{"urls": [{ "url": "https://example.com/" }],"fromDate": "202401","toDate": "202412","maxSnapshotsPerTarget": 2,"maxItems": 5,"collapse": "timestamp:6","selection": "latest","outputFormat": "items","delaySecs": 0.5}
Output example (one dataset item per snapshot)
{"kind": "snapshot","status": "ok","originalUrl": "https://example.com/","archiveUrl": "https://web.archive.org/web/20241201001928id_/https://example.com/","timestamp": "20241201001928","title": "Example Domain","contentMarkdown": "This domain is for use in documentation examples…","contentHash": "sha256…","wordCount": 26,"mimeType": "text/html","statusCode": "200"}
Key-value store records
| Key | Content |
|---|---|
CDX_REPORT | Per-target CDX hits / selected / fetch ok-fail |
OUTPUT | Run summary (counts, duration, charged events) |
Pricing
Pay per event (private Actor; Store publish pending coordinator):
| Event | Price |
|---|---|
snapshot-item (primary) | $0.003 per successfully extracted snapshot Markdown |
apify-actor-start | Apify default ($0.00005 / GB) |
Failed CDX queries and failed fetches are never charged. RAG chunk rows do not add extra events. Example: 1,000 snapshots ≈ $3.00 list (+ actor-start). Details: docs/PRICING.md.
Known limits
- Wayback CDX and snapshot endpoints rate-limit; use Delay between snapshot fetches and modest caps.
- Domain prefix queries (
example.com/*) can return very large CDX sets — keep Max snapshots per URL/domain low and prefer date bounds. - Some captures are redirects, soft-404s, or binary; trafilatura may return empty text (reported as error, not charged).
- waybackpack uses synchronous
requestsunder the hood (run in a worker thread); not a headless browser.
FAQ
Does this publish captures to a public GitHub repo? No. This project stays private unless you deliberately publish later.
Do you copy closed Apify Actors? No. Stack is waybackpack (MIT) + trafilatura (Apache-2.0) + the in-house Apify Python template patterns.
License & source
Actor source: AGPL-3.0 (see LICENSE). Third-party notices: NOTICE.