Internet Archive Search Scraper avatar

Internet Archive Search Scraper

Pricing

from $16.00 / 1,000 result items

Go to Apify Store
Internet Archive Search Scraper

Internet Archive Search Scraper

Search the Internet Archive's 50M+ item catalog of texts, audio, movies, software, web pages, and images. Filter by collection, media type, creator, and date. Pull identifiers, titles, descriptions, downloads, and rich metadata.

Pricing

from $16.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

8 hours ago

Last modified

Share

ParseForge Banner

πŸ“š Internet Archive Search Scraper

πŸš€ Export the world's largest open library in seconds. Search 50M+ items across texts, audio, movies, software, web, images, and data. No login, no manual paging, no Lucene crash courses required.

The Internet Archive Search Scraper exports the open library catalog and returns 21 fields per record, including identifier, title, full description, creator, language, subject tags, collection memberships, publish date, lifetime and weekly download counts, file inventories, total byte size, license URL, and direct links to the item details page and metadata feed. The underlying source is the world's largest publicly accessible digital library, maintained since 1996.

The catalog covers 50 million+ items across eight media types (texts, audio, movies, software, web captures, images, datasets, and collections). This Actor lets you slice the corpus with Lucene-style queries plus structured filters for collection, media type, creator, and date range, then download the result as CSV, Excel, JSON, or XML in under five minutes.

🎯 Target AudienceπŸ’‘ Primary Use Cases
Librarians, digital archivists, journalists, OSINT researchers, academic historians, ML dataset curators, documentary filmmakersCitation discovery, training-corpus assembly, historical media research, source verification, public-domain media sourcing, archival preservation audits

πŸ“‹ What the Archive Search Scraper does

A single configurable workflow with four filter layers:

  • πŸ”Ž Lucene query. Free-text or fielded queries like subject:photography AND mediatype:image.
  • πŸ“¦ Collection filter. Restrict to one Internet Archive collection like nasa or librivoxaudio.
  • 🎬 Media-type filter. Texts, audio, movies, software, web, image, data, or collection.
  • πŸ“… Date range. Filter by item publish date with dateFrom and dateTo (YYYY-MM-DD).
  • πŸ“š Per-item metadata. Optional deep fetch returns the full file list, rich subject tags, and license URL.

Each record bundles identifiers (Archive ID, details URL, metadata URL), descriptive metadata (title, creator, language, description, subject tags), classification (media type, collection memberships), engagement (lifetime, weekly, and monthly download counts), file inventory (count and total byte size), and licensing.

πŸ’‘ Why it matters: the Archive is the largest public corpus of cultural and reference material on Earth, but its native search interface assumes you already know Lucene. This Actor exposes that same query layer with structured filters and clean records, ready for analysis, ingestion, or archival back-up.

πŸ“Š Data fields

Each record includes: collection, creator, date, description, detailsUrl, downloads, filesCount, identifier, language, licenseUrl, mediaType, metadataUrl, month, publishDate, scrapedAt, subject, thumbnailUrl, title, totalSizeBytes, week. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

πŸš€ How to use

  1. πŸ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. 🌐 Open the Actor. Go to the Internet Archive Search Scraper page on the Apify Store.
  3. 🎯 Set input. Type a search query, optionally restrict to a collection or media type, and set maxItems.
  4. πŸš€ Run it. Click Start and let the Actor collect your dataset.
  5. πŸ“₯ Download. Grab results in the Dataset tab as CSV, Excel, JSON, or XML.

⏱️ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

πŸ’‘ Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

⚠️ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the Internet Archive or any of its contributors. All trademarks mentioned are the property of their respective owners. Only publicly available open archival data is collected.

πŸ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.