Internet Archive Scraper — Search archive.org, Download Links avatar

Internet Archive Scraper — Search archive.org, Download Links

Pricing

from $0.35 / 1,000 items

Go to Apify Store
Internet Archive Scraper — Search archive.org, Download Links

Internet Archive Scraper — Search archive.org, Download Links

Search the Internet Archive by text, collection, media type, creator, subject, language and year, and export one row per item with downloads, size, rating, licence and links — plus a row per file with a direct download URL. No API key.

Pricing

from $0.35 / 1,000 items

Rating

0.0

(0)

Developer

Chorelet

Chorelet

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

Search the Internet Archive and get the download links. Query 50 million public items — books, audiobooks, concerts, films, images, software and web collections — by text, collection, media type, creator, subject, language and year, and export one row per item: title, creator, description, media type, dates, collections, subjects, language, licence, downloads, size, rating and the link to the item page.

Turn on one row per file and every item brings its files with it: name, format, size in bytes, duration for audio and video, MD5, and a direct download URL you can hand to curl, a downloader or a script.

No API key, no account, no login — these are archive.org's own search and metadata endpoints.

Why this Actor

  • Download links, not just search results. Each item's files come back with a direct URL, format, size, duration and MD5 — the part that turns a catalogue search into something you can actually fetch.
  • Derivatives filtered out by default. Thumbnails, torrents and checksum files are dropped, so a 79-file item returns the handful you wanted rather than the noise.
  • Every way archive.org indexes an item: collection, media type, creator, subject, language and the year of the work, combined into one query for you.
  • One of the most requested Actors on the Apify ideas board, built on the archive's own endpoints — no key, no account, no scraping of HTML.

Sample output

One item of the dataset (long values shortened):

{
"identifier": "art_of_war_librivox",
"title": "The Art of War",
"creator": "Sun Tzu",
"mediaType": "audio",
"year": 2006,
"downloads": 24460268,
"rating": 4.57,
"detailsUrl": "https://archive.org/details/art_of_war_librivox"
}

What you get

  • The whole catalogue, searchable: plain words or Lucene syntax, narrowed by collection (librivoxaudio, prelinger, nasa), media type, creator, subject, language and a year range on the work itself
  • Direct download links for every file, with the derivatives — thumbnails, torrents, checksums — filtered out unless you ask for them, and a format filter for pdf, mp3, epub or mp4
  • Sorted the way you need it: most downloaded, newest upload, top rated, title or the date of the work
  • Popularity and licence on every row, so a run tells you both what exists and what is safe to reuse
  • Items and files share one schema with a type column, so a single CSV holds both

archive.org serves at most 10,000 results for one search; the run says so in the log and splitting by year or collection goes deeper.

Input example

{
"queries": [
"apollo 11"
],
"sortBy": "most downloaded",
"includeFiles": false,
"maxFilesPerItem": 10,
"originalFilesOnly": true,
"maxResultsPerQuery": 100
}

How much does it cost?

Pay per item — no subscription, no minimum, no charge for platform usage.

VolumePrice
1,000 items$0.50 (+ $0.20 with file)
10,000 items$5.00 (+ $2.00 with file)
100,000 items$50.00 (+ $20.00 with file)

The Apify free plan includes $5 of usage every month — about 10,000 items with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add &format=csv or xlsx):

curl -X POST "https://api.apify.com/v2/acts/chorelet~internet-archive-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
-H "Content-Type: application/json" \
-d '{"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100}'

Python:

from apify_client import ApifyClient
client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/internet-archive-scraper").call(run_input={"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
print(item)
  • n8n, Make, Zapier — use the Apify node/module: run the Actor, then "get dataset items".
  • Google Sheets, Slack, webhooks — add an integration on the run's Integrations tab.
  • AI agents — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
  • Schedules — run it hourly, daily or weekly from the Schedules tab.

FAQ

Do I need an API key?

No. archive.org publishes its search and metadata endpoints openly; the Actor uses those.

How many results can one search return?

Up to 10,000. The run logs it when a search is wider than that — narrow by year, collection or media type to reach the rest.

How do I make a search more precise?

A plain multi-word search is run as a phrase, so apollo 11 already excludes items that only mention one of the words. For titles only, use the registry's own syntax: title:("apollo 11") — it goes through untouched.

Can I download the files too?

The Actor returns a direct download URL for each file; downloading them is a separate step with curl, a browser or your own script. It deliberately does not pull gigabytes into a dataset.

What do the file rows cost?

A file row is charged at a fraction of an item row, so expanding a hundred items into their files stays cheap.

Is everything on archive.org free to reuse?

No. The licenseUrl column carries the licence the uploader declared, and many items are public domain while others are not. Check it before republishing.

Can I search inside books?

Not in this Actor. It searches item metadata — title, creator, subject, description — not the full text of scanned pages.

Support

Questions, missing fields or a source that changed? Open an issue on the Issues tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.