# Internet Archive Scraper — Search archive.org, Download Links (`chorelet/internet-archive-scraper`) Actor

Search the Internet Archive by text, collection, media type, creator, subject, language and year, and export one row per item with downloads, size, rating, licence and links — plus a row per file with a direct download URL. No API key.

- **URL**: https://apify.com/chorelet/internet-archive-scraper.md
- **Developed by:** [Chorelet](https://apify.com/chorelet) (community)
- **Categories:** Education, Developer tools, Automation
- **Stats:** 2 total users, 1 monthly users, 0.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.35 / 1,000 items

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Internet Archive Scraper — Search archive.org, Download Links

**Search the Internet Archive and get the download links.** Query 50 million public items — books, audiobooks, concerts, films, images, software and web collections — by text, collection, media type, creator, subject, language and year, and export one row per item: title, creator, description, media type, dates, collections, subjects, language, licence, downloads, size, rating and the link to the item page.

Turn on **one row per file** and every item brings its files with it: name, format, size in bytes, duration for audio and video, MD5, and a **direct download URL** you can hand to curl, a downloader or a script.

No API key, no account, no login — these are archive.org's own search and metadata endpoints.

### Why this Actor

- **Download links, not just search results.** Each item's files come back with a direct URL, format, size, duration and MD5 — the part that turns a catalogue search into something you can actually fetch.
- **Derivatives filtered out by default.** Thumbnails, torrents and checksum files are dropped, so a 79-file item returns the handful you wanted rather than the noise.
- **Every way archive.org indexes an item**: collection, media type, creator, subject, language and the year of the work, combined into one query for you.
- **One of the most requested Actors on the Apify ideas board**, built on the archive's own endpoints — no key, no account, no scraping of HTML.

### Sample output

One item of the dataset (long values shortened):

```json
{
  "identifier": "art_of_war_librivox",
  "title": "The Art of War",
  "creator": "Sun Tzu",
  "mediaType": "audio",
  "year": 2006,
  "downloads": 24460268,
  "rating": 4.57,
  "detailsUrl": "https://archive.org/details/art_of_war_librivox"
}
```

### What you get

- **The whole catalogue, searchable**: plain words or Lucene syntax, narrowed by collection (`librivoxaudio`, `prelinger`, `nasa`), media type, creator, subject, language and a year range on the work itself
- **Direct download links** for every file, with the derivatives — thumbnails, torrents, checksums — filtered out unless you ask for them, and a format filter for `pdf`, `mp3`, `epub` or `mp4`
- **Sorted the way you need it**: most downloaded, newest upload, top rated, title or the date of the work
- **Popularity and licence on every row**, so a run tells you both what exists and what is safe to reuse
- Items and files share one schema with a `type` column, so a single CSV holds both

archive.org serves at most 10,000 results for one search; the run says so in the log and splitting by year or collection goes deeper.

### Input example

```json
{
  "queries": [
    "apollo 11"
  ],
  "sortBy": "most downloaded",
  "includeFiles": false,
  "maxFilesPerItem": 10,
  "originalFilesOnly": true,
  "maxResultsPerQuery": 100
}
```

### How much does it cost?

Pay per item — no subscription, no minimum, no charge for platform usage.

| Volume | Price |
|---|---|
| 1,000 items | $0.50 (+ $0.20 with `file`) |
| 10,000 items | $5.00 (+ $2.00 with `file`) |
| 100,000 items | $50.00 (+ $20.00 with `file`) |

The Apify **free plan includes $5 of usage every month** — about 10,000 items with this Actor, no card needed. Nothing else is charged: platform usage is included in the price, and Apify Bronze, Silver and Gold subscribers get 10%, 20% and 30% off these prices.

### Use it from code, n8n, Make, Zapier or an AI agent

Run the Actor and download the dataset in one call (JSON by default; add `&format=csv` or `xlsx`):

```bash
curl -X POST "https://api.apify.com/v2/acts/chorelet~internet-archive-scraper/run-sync-get-dataset-items?token=$APIFY_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100}'
```

Python:

```python
from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")
run = client.actor("chorelet/internet-archive-scraper").call(run_input={"queries": ["apollo 11"], "sortBy": "most downloaded", "includeFiles": false, "maxFilesPerItem": 10, "originalFilesOnly": true, "maxResultsPerQuery": 100})
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)
```

- **n8n, Make, Zapier** — use the Apify node/module: run the Actor, then "get dataset items".
- **Google Sheets, Slack, webhooks** — add an integration on the run's *Integrations* tab.
- **AI agents** — the Actor is available as a tool through the Apify MCP server; the dataset schema describes every field for the model.
- **Schedules** — run it hourly, daily or weekly from the *Schedules* tab.

### FAQ

**Do I need an API key?**

No. archive.org publishes its search and metadata endpoints openly; the Actor uses those.

**How many results can one search return?**

Up to 10,000. The run logs it when a search is wider than that — narrow by year, collection or media type to reach the rest.

**How do I make a search more precise?**

A plain multi-word search is run as a phrase, so `apollo 11` already excludes items that only mention one of the words. For titles only, use the registry's own syntax: `title:("apollo 11")` — it goes through untouched.

**Can I download the files too?**

The Actor returns a direct download URL for each file; downloading them is a separate step with curl, a browser or your own script. It deliberately does not pull gigabytes into a dataset.

**What do the file rows cost?**

A file row is charged at a fraction of an item row, so expanding a hundred items into their files stays cheap.

**Is everything on archive.org free to reuse?**

No. The `licenseUrl` column carries the licence the uploader declared, and many items are public domain while others are not. Check it before republishing.

**Can I search inside books?**

Not in this Actor. It searches item metadata — title, creator, subject, description — not the full text of scanned pages.

### Support

Questions, missing fields or a source that changed? Open an issue on the *Issues* tab or write to support@chorelet.app — problems are usually fixed within a day, and the Actor is checked every morning by an automated test run. If the Actor saved you time, a short review on its Store page helps other people find it.

# Actor input Schema

## `queries` (type: `array`):

Plain words or Lucene syntax, one search per line: `apollo 11`, `title:("public domain")`. Leave empty to search by the filters below alone.

## `collections` (type: `array`):

archive.org collection ids: `librivoxaudio`, `prelinger`, `nasa`, `opensource_movies`, `gutenberg`.

## `mediaTypes` (type: `array`):

`texts`, `audio`, `movies`, `image`, `software`, `data`, `web`, `etree`, `collection`.

## `creators` (type: `array`):

Author, band, studio or agency as the item credits it.

## `subjects` (type: `array`):

Subject tags as archive.org stores them.

## `languages` (type: `array`):

`English`, `French`, `Spanish`…

## `yearFrom` (type: `string`):

Year of the work itself, not of the upload.

## `yearTo` (type: `string`):

Leave either end empty for an open range.

## `sortBy` (type: `string`):

What comes first.

## `includeFiles` (type: `boolean`):

Opens each item's file list and returns a direct download URL, format, size and duration for every file. One extra request per item.

## `maxFilesPerItem` (type: `integer`):

0 = every file. Items can hold hundreds of derivatives.

## `fileFormats` (type: `array`):

Partial matches against the format: `pdf`, `mp3`, `epub`, `mp4`. Empty keeps all of them.

## `originalFilesOnly` (type: `boolean`):

Drops thumbnails, torrents, checksums and other generated files.

## `maxResultsPerQuery` (type: `integer`):

archive.org serves at most 10,000 results for one search; narrow by year or collection to go deeper.

## Actor input object example

```json
{
  "queries": [
    "apollo 11"
  ],
  "collections": [],
  "mediaTypes": [],
  "creators": [],
  "subjects": [],
  "languages": [],
  "sortBy": "most downloaded",
  "includeFiles": false,
  "maxFilesPerItem": 10,
  "fileFormats": [],
  "originalFilesOnly": true,
  "maxResultsPerQuery": 100
}
```

# Actor output Schema

## `items` (type: `string`):

Everything scraped — items of the default dataset. Use ?format=csv or xlsx on this URL for spreadsheets.

## `summary` (type: `string`):

Items per search, how many matched and any errors.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "apollo 11"
    ],
    "collections": [],
    "mediaTypes": [],
    "creators": [],
    "subjects": [],
    "languages": [],
    "yearFrom": "",
    "yearTo": "",
    "sortBy": "most downloaded",
    "includeFiles": false,
    "maxFilesPerItem": 10,
    "fileFormats": [],
    "originalFilesOnly": true,
    "maxResultsPerQuery": 100
};

// Run the Actor and wait for it to finish
const run = await client.actor("chorelet/internet-archive-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["apollo 11"],
    "collections": [],
    "mediaTypes": [],
    "creators": [],
    "subjects": [],
    "languages": [],
    "yearFrom": "",
    "yearTo": "",
    "sortBy": "most downloaded",
    "includeFiles": False,
    "maxFilesPerItem": 10,
    "fileFormats": [],
    "originalFilesOnly": True,
    "maxResultsPerQuery": 100,
}

# Run the Actor and wait for it to finish
run = client.actor("chorelet/internet-archive-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "apollo 11"
  ],
  "collections": [],
  "mediaTypes": [],
  "creators": [],
  "subjects": [],
  "languages": [],
  "yearFrom": "",
  "yearTo": "",
  "sortBy": "most downloaded",
  "includeFiles": false,
  "maxFilesPerItem": 10,
  "fileFormats": [],
  "originalFilesOnly": true,
  "maxResultsPerQuery": 100
}' |
apify call chorelet/internet-archive-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,chorelet/internet-archive-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/WGMjqZHBaE9ijS6eg/builds/Bt54tqCzOFTzlP3SL/openapi.json
