Internet Archive Scraper & API - archive.org Books & Media
Pricing
from $1.00 / 1,000 results
Internet Archive Scraper & API - archive.org Books & Media
Use this to search archive.org (Internet Archive): books, video, audio, software. Input: queries or item IDs, optional media type, collection, year filters. One result = one item: title, creator, year, media type, collections, license, downloads, URL, optional file links. $1.00 per 1,000 results.
Pricing
from $1.00 / 1,000 results
Rating
0.0
(0)
Developer
Giovanni Rich
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
8 hours ago
Last modified
Categories
Share
Internet Archive Scraper & API - archive.org Books, Movies, Audio and Software
Internet Archive Scraper is an archive.org scraper and unofficial Internet Archive search API: search books, movies, audio, software and images on archive.org and export every item's title, creator, date, subjects, collections, license, download counts, ratings and file formats as JSON, CSV or Excel, with optional direct download links for every PDF, EPUB, MP3 or MP4.
It uses archive.org's own public search and metadata APIs over plain HTTP, with no headless browser and no login, so a 100-item search finishes in a few seconds and costs very little to run.
How to use
- Enter one or more searches in Search queries (e.g.
apollo 11,moby dick, or advanced syntax likecreator:"Mark Twain"). - Optionally pick Media types (texts, movies, audio, software, ...), a Collection (e.g.
gutenberg,prelinger,nasa), a year range, a language and a sort order. - Turn on Include file list and download links if you want every file with its format, size, MD5 and download URL.
- Click Start and download the results as JSON, CSV, Excel or HTML from the Output tab, or pull them through the Apify API.
You can also skip searching and paste item identifiers or URLs (e.g. https://archive.org/details/Apollo11Audio) to get their full metadata and file lists.
What you get
- Item metadata: identifier, title, media type, creator(s), publisher, description, date, year, language, subjects, collections, ISBN, page count (texts) and runtime (video/audio).
- Rights: license URL (Creative Commons, public domain mark) and rights statement when the uploader set one.
- Popularity: all-time downloads, downloads in the last week and month, favorites, review count and average rating. Great for ranking what people actually use.
- Files (optional): every file in the item with format, source (original/derivative), size, length in seconds, MD5 and a direct
https://archive.org/download/...URL. Filter to justpdf,epub,mp3,h.264, ... - Links: item page, thumbnail image and download directory for each result.
- Search control: search in title/subjects/creator (precise, the default) or in archive.org's full text, sort by relevance, downloads, weekly downloads, upload date, publication date, title or rating.
Use cases
- Build reading lists, datasets or catalogs of public-domain books, films and recordings
- Find openly licensed media (Creative Commons, public domain) for videos, podcasts and AI training sets
- Research: track what's available on a topic, author or collection, and how popular it is
- Bulk-download pipelines: get the exact PDF/EPUB/MP3 URLs for a collection and feed them to your downloader
- Monitor new uploads to a collection on a schedule (sort by
Newest uploads)
Input example
| Field | Default | Description |
|---|---|---|
queries | - | Searches, one per line. Plain words must all match. Advanced syntax works: title:(moby dick), creator:"Mark Twain", subject:jazz |
searchIn | metadata | metadata = title, subjects and creator (precise); title; everything = archive.org full-text search (broad, noisier) |
mediaTypes | all | texts, movies, audio, software, image, data, web, collection, etree |
sortBy | relevance | downloads, downloadsThisWeek, newestUploads, oldestUploads, dateNewest, dateOldest, titleAZ, rating |
maxResults | 100 | Items per query, up to 10,000 |
collection | - | Collection ID(s), e.g. gutenberg, prelinger, nasa |
yearFrom / yearTo | - | Publication year range |
language | - | e.g. eng, fre, English |
includeFiles | false | Adds the file list with download links (one extra request per item) |
fileFormats | all | With file lists on: keep only these formats/extensions, e.g. pdf, epub, mp3 |
identifiers | - | Item IDs or archive.org URLs to look up directly |
Example:
{"queries": ["apollo 11", "moby dick", "creator:\"Mark Twain\""],"mediaTypes": ["movies", "texts", "audio"],"sortBy": "downloads","maxResults": 100,"includeFiles": true,"fileFormats": ["pdf", "epub", "mp3", "h.264"]}
Output example
One dataset item per archive.org item. This one is from a real run (moby dick, sorted by downloads, with fileFormats: ["pdf","epub"]; lists shortened):
{"query": "moby dick","position": 2,"identifier": "mobydickorwhale01melvuoft","title": "Moby-Dick; or, The whale","mediaType": "texts","creator": "Melville, Herman, 1819-1891","date": "1922-01-01","year": 1922,"publisher": "London Constable","language": ["eng"],"collections": ["robarts", "toronto", "university_of_toronto"],"licenseUrl": null,"downloads": 135030,"downloadsLastWeek": 678,"downloadsLastMonth": 3738,"favorites": 208,"itemSizeBytes": 648030033,"fileFormats": ["DjVuTXT", "EPUB", "MARC", "Text PDF", "hOCR"],"pageCount": 398,"publicDate": "2006-12-06T16:57:58Z","url": "https://archive.org/details/mobydickorwhale01melvuoft","thumbnailUrl": "https://archive.org/services/img/mobydickorwhale01melvuoft","downloadUrl": "https://archive.org/download/mobydickorwhale01melvuoft","filesCount": 2,"files": [{"name": "mobydickorwhale01melvuoft.epub","format": "EPUB","source": "derivative","sizeBytes": 1433807,"md5": "c78305e972609a88a73bf61242926249","url": "https://archive.org/download/mobydickorwhale01melvuoft/mobydickorwhale01melvuoft.epub"}],"scrapedAt": "2026-09-25T04:55:01.353Z"}
The dataset has two ready-made views: Overview (thumbnail, title, type, creator, year, downloads, rating, license, link) and Files (formats, file count and file list).
Field fill rates (real test run: 300 items for "apollo 11", "moby dick" and creator:"Mark Twain", movies/texts/audio)
| Field | Fill rate |
|---|---|
| identifier, title, media type, collections, downloads, links, dates added | 100% |
| subjects | 97% |
| description | 99% |
| creator | 89% |
| date / year | 78% |
| language | 68% |
| license URL | 44% (only items whose uploader chose a license) |
| runtime (video/audio) | 37% |
| publisher | 26% |
| file list with download URLs | 100% of items when includeFiles is on (82% had files matching the format filter) |
Pricing
Pay per result: $1.00 per 1,000 results (one result = one archive.org item saved to the dataset, with or without its file list).
- 100 items = $0.10
- 1,000 items = $1.00
- 10,000 items = $10
You're never charged for failed requests, items that don't exist, or duplicates. If you set a maximum cost per run, the scraper stops cleanly when it reaches it. Apify's free plan includes $5 of monthly platform credit, enough for thousands of items.
Integrations
- Make, Zapier and n8n: start runs and send new items to your other apps.
- Google Sheets: export the dataset straight into a spreadsheet, or refresh it on a schedule.
- Apify API: run the actor and fetch results over REST, or with the official JavaScript and Python clients.
- Webhooks: get notified when a run finishes and start your download pipeline.
- Schedules: run a
Newest uploadssearch daily to watch a collection. - MCP for AI agents: through the Apify MCP server (https://mcp.apify.com), Claude, ChatGPT, Cursor and other agents can search the Internet Archive with this actor and read the results.
Limits (read before large runs)
- 10,000 results per query. archive.org's search API won't page deeper. Split big searches by year (
yearFrom/yearTo), media type or collection. - File lists cost one extra request per item. Items like full audiobooks have hundreds of files, so file-list runs take about 1-2 seconds per item at the default concurrency.
- Metadata quality varies. Everything comes from what uploaders entered: some items have no date, creator or license. The actor returns
nullrather than guessing. - Full-text search is noisy.
searchIn: "everything"matches words anywhere, including reviews and OCR text, so popular but unrelated items can rank high when sorting by downloads. The default (title, subjects, creator) avoids this. - The actor doesn't download files. It gives you the URLs; download them with your own tool, and respect each item's license.
- Wayback Machine snapshots (archived websites) are not covered; this actor is for archive.org items.
FAQ
Is it legal to scrape the Internet Archive? The actor uses archive.org's public search and metadata APIs, which archive.org provides for exactly this kind of access, and it returns item metadata, not personal data. It even drops users' personal "favorites" lists from the collections field. What you do with the files is up to you: many items are public domain or Creative Commons, but others are in copyright, so check each item's license and rights before reusing it. This is not legal advice.
Will I get blocked?
Unlikely. archive.org's APIs are open and the actor keeps a polite default concurrency. Failed requests (archive.org sometimes returns 5xx under load) are retried with exponential backoff, up to 5 times. If you ever see HTTP 429, lower maxConcurrency or turn on Apify Proxy.
Do I need an archive.org account or API key? No. Everything is public and the actor needs no login.
Why did a query return fewer results than maxResults?
Fewer items matched. The log and the RUN_SUMMARY record show archive.org's total match count and the stop reason for each query.
Can I list a whole collection?
Yes. Leave queries empty and set collection (e.g. prelinger), optionally with a media type and year range. Up to 10,000 items per run; split by year for bigger collections.
How do I get only PDFs or MP3s?
Turn on includeFiles and set fileFormats to ["pdf"] or ["mp3"]. Each result then lists only the matching files with direct download URLs.
How it works (for developers)
Searches go to archive.org/advancedsearch.php (the Solr-backed search API) with the fields, sort and paging archive.org documents. Plain-word queries are turned into title:(...) OR subject:(...) OR creator:(...) so results stay on topic; advanced Lucene queries are passed through unchanged. File lists come from archive.org/metadata/{identifier}. Requests run through Crawlee's HttpCrawler with a session pool and exponential-backoff retries, and progress is saved so a migrated run resumes without duplicates.
Run it locally:
npm installnpm test # parser testsAPIFY_LOCAL_STORAGE_DIR=./storage node src/main.js # input in storage/key_value_stores/default/INPUT.json