Internet Archive Scraper — Search archive.org, Download Links
Pricing
from $0.35 / 1,000 items
Internet Archive Scraper — Search archive.org, Download Links
Search the Internet Archive by text, collection, media type, creator, subject, language and year, and export one row per item with downloads, size, rating, licence and links — plus a row per file with a direct download URL. No API key.
Pricing
from $0.35 / 1,000 items
Rating
0.0
(0)
Developer
Chorelet
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
3 hours ago
Last modified
Categories
Share
Search the Internet Archive and get the download links. Query 50 million public items — books, audiobooks, concerts, films, images, software and web collections — by text, collection, media type, creator, subject, language and year, and export one row per item: title, creator, description, media type, dates, collections, subjects, language, licence, downloads, size, rating and the link to the item page.
Turn on one row per file and every item brings its files with it: name, format, size in bytes, duration for audio and video, MD5, and a direct download URL you can hand to curl, a downloader or a script.
No API key, no account, no login — these are archive.org's own search and metadata endpoints.
Why this Actor
- Download links, not just search results. Each item's files come back with a direct URL, format, size, duration and MD5 — the part that turns a catalogue search into something you can actually fetch.
- Derivatives filtered out by default. Thumbnails, torrents and checksum files are dropped, so a 79-file item returns the handful you wanted rather than the noise.
- Every way archive.org indexes an item: collection, media type, creator, subject, language and the year of the work, combined into one query for you.
- One of the most requested Actors on the Apify ideas board, built on the archive's own endpoints — no key, no account, no scraping of HTML.
Sample output
One item of the dataset (long values shortened):
{"identifier": "art_of_war_librivox","title": "The Art of War","creator": "Sun Tzu","mediaType": "audio","year": 2006,"downloads": 24460268,"rating": 4.57,"detailsUrl": "https://archive.org/details/art_of_war_librivox"}
What you get
- The whole catalogue, searchable: plain words or Lucene syntax, narrowed by collection (
librivoxaudio,prelinger,nasa), media type, creator, subject, language and a year range on the work itself - Direct download links for every file, with the derivatives — thumbnails, torrents, checksums — filtered out unless you ask for them, and a format filter for
pdf,mp3,epubormp4 - Sorted the way you need it: most downloaded, newest upload, top rated, title or the date of the work
- Popularity and licence on every row, so a run tells you both what exists and what is safe to reuse
- Items and files share one schema with a
typecolumn, so a single CSV holds both
archive.org serves at most 10,000 results for one search; the run says so in the log and splitting by year or collection goes deeper.
Input example
{"queries": ["apollo 11"],"sortBy": "most downloaded","includeFiles": false,"maxFilesPerItem": 10,"originalFilesOnly": true,"maxResultsPerQuery": 100}
How much does it cost?
Pay per item — no subscription, no minimum, no charge for platform usage.
| Volume | Price |
|---|---|
| 1,000 items | $0.50 (+ $0.20 with file) |
| 10,000 items | $5.00 (+ $2.00 with file) |
| 100,000 items | $50.00 (+ $20.00 with file) |
The Apify free plan includes $5 of usage every month — about 10,000 items with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.
Use it from code, n8n, Make, Zapier or an AI agent
Run the Actor and download the dataset in one call (JSON by default; add &format=csv or xlsx):
curl -X POST "https://api.apify.com/v2/acts/chorelet~internet-archive-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \-H "Content-Type: application/json" \-d '{"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100}'
Python:
from apify_client import ApifyClientclient = ApifyClient("YOUR_APIFY_TOKEN")run = client.actor("chorelet/internet-archive-scraper").call(run_input={"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100})for item in client.dataset(run["defaultDatasetId"]).iterate_items():print(item)
- n8n, Make, Zapier — use the Apify node/module: run the Actor, then "get dataset items".
- Google Sheets, Slack, webhooks — add an integration on the run's Integrations tab.
- AI agents — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
- Schedules — run it hourly, daily or weekly from the Schedules tab.
FAQ
Do I need an API key?
No. archive.org publishes its search and metadata endpoints openly; the Actor uses those.
How many results can one search return?
Up to 10,000. The run logs it when a search is wider than that — narrow by year, collection or media type to reach the rest.
How do I make a search more precise?
A plain multi-word search is run as a phrase, so apollo 11 already excludes items that only mention one of the words. For titles only, use the registry's own syntax: title:("apollo 11") — it goes through untouched.
Can I download the files too?
The Actor returns a direct download URL for each file; downloading them is a separate step with curl, a browser or your own script. It deliberately does not pull gigabytes into a dataset.
What do the file rows cost?
A file row is charged at a fraction of an item row, so expanding a hundred items into their files stays cheap.
Is everything on archive.org free to reuse?
No. The licenseUrl column carries the licence the uploader declared, and many items are public domain while others are not. Check it before republishing.
Can I search inside books?
Not in this Actor. It searches item metadata — title, creator, subject, description — not the full text of scanned pages.
Support
Questions, missing fields or a source that changed? Open an issue on the Issues tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.