Internet Archive Search Scraper
Pricing
from $16.00 / 1,000 result items
Internet Archive Search Scraper
Search the Internet Archive's 50M+ item catalog of texts, audio, movies, software, web pages, and images. Filter by collection, media type, creator, and date. Pull identifiers, titles, descriptions, downloads, and rich metadata.
Pricing
from $16.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 hours ago
Last modified
Categories
Share

π Internet Archive Search Scraper
π Export the world's largest open library in seconds. Search 50M+ items across texts, audio, movies, software, web, images, and data. No login, no manual paging, no Lucene crash courses required.
The Internet Archive Search Scraper exports the open library catalog and returns 21 fields per record, including identifier, title, full description, creator, language, subject tags, collection memberships, publish date, lifetime and weekly download counts, file inventories, total byte size, license URL, and direct links to the item details page and metadata feed. The underlying source is the world's largest publicly accessible digital library, maintained since 1996.
The catalog covers 50 million+ items across eight media types (texts, audio, movies, software, web captures, images, datasets, and collections). This Actor lets you slice the corpus with Lucene-style queries plus structured filters for collection, media type, creator, and date range, then download the result as CSV, Excel, JSON, or XML in under five minutes.
| π― Target Audience | π‘ Primary Use Cases |
|---|---|
| Librarians, digital archivists, journalists, OSINT researchers, academic historians, ML dataset curators, documentary filmmakers | Citation discovery, training-corpus assembly, historical media research, source verification, public-domain media sourcing, archival preservation audits |
π What the Archive Search Scraper does
A single configurable workflow with four filter layers:
- π Lucene query. Free-text or fielded queries like
subject:photography AND mediatype:image. - π¦ Collection filter. Restrict to one Internet Archive collection like
nasaorlibrivoxaudio. - π¬ Media-type filter. Texts, audio, movies, software, web, image, data, or collection.
- π
Date range. Filter by item publish date with
dateFromanddateTo(YYYY-MM-DD). - π Per-item metadata. Optional deep fetch returns the full file list, rich subject tags, and license URL.
Each record bundles identifiers (Archive ID, details URL, metadata URL), descriptive metadata (title, creator, language, description, subject tags), classification (media type, collection memberships), engagement (lifetime, weekly, and monthly download counts), file inventory (count and total byte size), and licensing.
π‘ Why it matters: the Archive is the largest public corpus of cultural and reference material on Earth, but its native search interface assumes you already know Lucene. This Actor exposes that same query layer with structured filters and clean records, ready for analysis, ingestion, or archival back-up.
π Data fields
Each record includes: collection, creator, date, description, detailsUrl, downloads, filesCount, identifier, language, licenseUrl, mediaType, metadataUrl, month, publishDate, scrapedAt, subject, thumbnailUrl, title, totalSizeBytes, week. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
π How to use
- π Sign up. Create a free account with $5 credit (takes 2 minutes).
- π Open the Actor. Go to the Internet Archive Search Scraper page on the Apify Store.
- π― Set input. Type a search query, optionally restrict to a collection or media type, and set
maxItems. - π Run it. Click Start and let the Actor collect your dataset.
- π₯ Download. Grab results in the Dataset tab as CSV, Excel, JSON, or XML.
β±οΈ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
π Recommended Actors
- π arXiv Scraper - Preprint papers across physics, math, and CS
- π RFC Editor Index Scraper - IETF Internet standards catalog
- ποΈ Met Museum Scraper - Metropolitan Museum of Art open-access objects
- π¬ ClinicalTrials.gov Scraper - Registered medical trials with outcomes
- π REST Countries Info Scraper - 250+ countries with population, currencies, languages
π‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
β οΈ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Internet Archive or any of its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open archival data is collected.
π Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.