# Internet Archive Scraper & API - archive.org Books & Media (`rel8ble/internet-archive-scraper`) Actor

Use this to search archive.org (Internet Archive): books, video, audio, software. Input: queries or item IDs, optional media type, collection, year filters. One result = one item: title, creator, year, media type, collections, license, downloads, URL, optional file links. $1.00 per 1,000 results.

- **URL**: https://apify.com/rel8ble/internet-archive-scraper.md
- **Developed by:** [Giovanni Rich](https://apify.com/rel8ble) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $1.00 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Internet Archive Scraper & API - archive.org Books, Movies, Audio and Software

**Internet Archive Scraper** is an archive.org scraper and unofficial Internet Archive search API: search books, movies, audio, software and images on archive.org and export every item's title, creator, date, subjects, collections, license, download counts, ratings and file formats as JSON, CSV or Excel, with optional direct download links for every PDF, EPUB, MP3 or MP4.

It uses archive.org's own public search and metadata APIs over plain HTTP, with **no headless browser and no login**, so a 100-item search finishes in a few seconds and costs very little to run.

### How to use

1. Enter one or more searches in **Search queries** (e.g. `apollo 11`, `moby dick`, or advanced syntax like `creator:"Mark Twain"`).
2. Optionally pick **Media types** (texts, movies, audio, software, ...), a **Collection** (e.g. `gutenberg`, `prelinger`, `nasa`), a year range, a language and a sort order.
3. Turn on **Include file list and download links** if you want every file with its format, size, MD5 and download URL.
4. Click **Start** and download the results as JSON, CSV, Excel or HTML from the **Output** tab, or pull them through the Apify API.

You can also skip searching and paste **item identifiers or URLs** (e.g. `https://archive.org/details/Apollo11Audio`) to get their full metadata and file lists.

### What you get

- **Item metadata**: identifier, title, media type, creator(s), publisher, description, date, year, language, subjects, collections, ISBN, page count (texts) and runtime (video/audio).
- **Rights**: license URL (Creative Commons, public domain mark) and rights statement when the uploader set one.
- **Popularity**: all-time downloads, downloads in the last week and month, favorites, review count and average rating. Great for ranking what people actually use.
- **Files** (optional): every file in the item with format, source (original/derivative), size, length in seconds, MD5 and a direct `https://archive.org/download/...` URL. Filter to just `pdf`, `epub`, `mp3`, `h.264`, ...
- **Links**: item page, thumbnail image and download directory for each result.
- **Search control**: search in title/subjects/creator (precise, the default) or in archive.org's full text, sort by relevance, downloads, weekly downloads, upload date, publication date, title or rating.

### Use cases

- Build reading lists, datasets or catalogs of public-domain books, films and recordings
- Find openly licensed media (Creative Commons, public domain) for videos, podcasts and AI training sets
- Research: track what's available on a topic, author or collection, and how popular it is
- Bulk-download pipelines: get the exact PDF/EPUB/MP3 URLs for a collection and feed them to your downloader
- Monitor new uploads to a collection on a schedule (sort by `Newest uploads`)

### Input example

| Field | Default | Description |
|---|---|---|
| `queries` | - | Searches, one per line. Plain words must all match. Advanced syntax works: `title:(moby dick)`, `creator:"Mark Twain"`, `subject:jazz` |
| `searchIn` | metadata | `metadata` = title, subjects and creator (precise); `title`; `everything` = archive.org full-text search (broad, noisier) |
| `mediaTypes` | all | `texts`, `movies`, `audio`, `software`, `image`, `data`, `web`, `collection`, `etree` |
| `sortBy` | relevance | `downloads`, `downloadsThisWeek`, `newestUploads`, `oldestUploads`, `dateNewest`, `dateOldest`, `titleAZ`, `rating` |
| `maxResults` | 100 | Items per query, up to 10,000 |
| `collection` | - | Collection ID(s), e.g. `gutenberg`, `prelinger`, `nasa` |
| `yearFrom` / `yearTo` | - | Publication year range |
| `language` | - | e.g. `eng`, `fre`, `English` |
| `includeFiles` | false | Adds the file list with download links (one extra request per item) |
| `fileFormats` | all | With file lists on: keep only these formats/extensions, e.g. `pdf`, `epub`, `mp3` |
| `identifiers` | - | Item IDs or archive.org URLs to look up directly |

Example:

```json
{
    "queries": ["apollo 11", "moby dick", "creator:\"Mark Twain\""],
    "mediaTypes": ["movies", "texts", "audio"],
    "sortBy": "downloads",
    "maxResults": 100,
    "includeFiles": true,
    "fileFormats": ["pdf", "epub", "mp3", "h.264"]
}
```

### Output example

One dataset item per archive.org item. This one is from a real run (`moby dick`, sorted by downloads, with `fileFormats: ["pdf","epub"]`; lists shortened):

```json
{
    "query": "moby dick",
    "position": 2,
    "identifier": "mobydickorwhale01melvuoft",
    "title": "Moby-Dick; or, The whale",
    "mediaType": "texts",
    "creator": "Melville, Herman, 1819-1891",
    "date": "1922-01-01",
    "year": 1922,
    "publisher": "London Constable",
    "language": ["eng"],
    "collections": ["robarts", "toronto", "university_of_toronto"],
    "licenseUrl": null,
    "downloads": 135030,
    "downloadsLastWeek": 678,
    "downloadsLastMonth": 3738,
    "favorites": 208,
    "itemSizeBytes": 648030033,
    "fileFormats": ["DjVuTXT", "EPUB", "MARC", "Text PDF", "hOCR"],
    "pageCount": 398,
    "publicDate": "2006-12-06T16:57:58Z",
    "url": "https://archive.org/details/mobydickorwhale01melvuoft",
    "thumbnailUrl": "https://archive.org/services/img/mobydickorwhale01melvuoft",
    "downloadUrl": "https://archive.org/download/mobydickorwhale01melvuoft",
    "filesCount": 2,
    "files": [
        {
            "name": "mobydickorwhale01melvuoft.epub",
            "format": "EPUB",
            "source": "derivative",
            "sizeBytes": 1433807,
            "md5": "c78305e972609a88a73bf61242926249",
            "url": "https://archive.org/download/mobydickorwhale01melvuoft/mobydickorwhale01melvuoft.epub"
        }
    ],
    "scrapedAt": "2026-09-25T04:55:01.353Z"
}
```

The dataset has two ready-made views: **Overview** (thumbnail, title, type, creator, year, downloads, rating, license, link) and **Files** (formats, file count and file list).

#### Field fill rates (real test run: 300 items for "apollo 11", "moby dick" and creator:"Mark Twain", movies/texts/audio)

| Field | Fill rate |
|---|---|
| identifier, title, media type, collections, downloads, links, dates added | 100% |
| subjects | 97% |
| description | 99% |
| creator | 89% |
| date / year | 78% |
| language | 68% |
| license URL | 44% (only items whose uploader chose a license) |
| runtime (video/audio) | 37% |
| publisher | 26% |
| file list with download URLs | 100% of items when `includeFiles` is on (82% had files matching the format filter) |

### Pricing

Pay per result: **$1.00 per 1,000 results** (one result = one archive.org item saved to the dataset, with or without its file list).

- 100 items = $0.10
- 1,000 items = $1.00
- 10,000 items = $10

You're never charged for failed requests, items that don't exist, or duplicates. If you set a maximum cost per run, the scraper stops cleanly when it reaches it. Apify's free plan includes $5 of monthly platform credit, enough for thousands of items.

### Integrations

- **Make, Zapier and n8n**: start runs and send new items to your other apps.
- **Google Sheets**: export the dataset straight into a spreadsheet, or refresh it on a schedule.
- **Apify API**: run the actor and fetch results over REST, or with the official JavaScript and Python clients.
- **Webhooks**: get notified when a run finishes and start your download pipeline.
- **Schedules**: run a `Newest uploads` search daily to watch a collection.
- **MCP for AI agents**: through the Apify MCP server (https://mcp.apify.com), Claude, ChatGPT, Cursor and other agents can search the Internet Archive with this actor and read the results.

### Limits (read before large runs)

- **10,000 results per query.** archive.org's search API won't page deeper. Split big searches by year (`yearFrom`/`yearTo`), media type or collection.
- **File lists cost one extra request per item.** Items like full audiobooks have hundreds of files, so file-list runs take about 1-2 seconds per item at the default concurrency.
- **Metadata quality varies.** Everything comes from what uploaders entered: some items have no date, creator or license. The actor returns `null` rather than guessing.
- **Full-text search is noisy.** `searchIn: "everything"` matches words anywhere, including reviews and OCR text, so popular but unrelated items can rank high when sorting by downloads. The default (title, subjects, creator) avoids this.
- **The actor doesn't download files.** It gives you the URLs; download them with your own tool, and respect each item's license.
- **Wayback Machine snapshots** (archived websites) are not covered; this actor is for archive.org items.

### FAQ

**Is it legal to scrape the Internet Archive?**
The actor uses archive.org's public search and metadata APIs, which archive.org provides for exactly this kind of access, and it returns item metadata, not personal data. It even drops users' personal "favorites" lists from the collections field. What you do with the files is up to you: many items are public domain or Creative Commons, but others are in copyright, so check each item's license and rights before reusing it. This is not legal advice.

**Will I get blocked?**
Unlikely. archive.org's APIs are open and the actor keeps a polite default concurrency. Failed requests (archive.org sometimes returns 5xx under load) are retried with exponential backoff, up to 5 times. If you ever see HTTP 429, lower `maxConcurrency` or turn on Apify Proxy.

**Do I need an archive.org account or API key?**
No. Everything is public and the actor needs no login.

**Why did a query return fewer results than `maxResults`?**
Fewer items matched. The log and the `RUN_SUMMARY` record show archive.org's total match count and the stop reason for each query.

**Can I list a whole collection?**
Yes. Leave `queries` empty and set `collection` (e.g. `prelinger`), optionally with a media type and year range. Up to 10,000 items per run; split by year for bigger collections.

**How do I get only PDFs or MP3s?**
Turn on `includeFiles` and set `fileFormats` to `["pdf"]` or `["mp3"]`. Each result then lists only the matching files with direct download URLs.

### How it works (for developers)

Searches go to `archive.org/advancedsearch.php` (the Solr-backed search API) with the fields, sort and paging archive.org documents. Plain-word queries are turned into `title:(...) OR subject:(...) OR creator:(...)` so results stay on topic; advanced Lucene queries are passed through unchanged. File lists come from `archive.org/metadata/{identifier}`. Requests run through Crawlee's `HttpCrawler` with a session pool and exponential-backoff retries, and progress is saved so a migrated run resumes without duplicates.

Run it locally:

```bash
npm install
npm test                                  # parser tests
APIFY_LOCAL_STORAGE_DIR=./storage node src/main.js   # input in storage/key_value_stores/default/INPUT.json
```

# Actor input Schema

## `queries` (type: `array`):

Optional. List of archive.org searches. Plain words must all match ("apollo 11" = apollo AND 11). Advanced syntax works: title:(moby dick), creator:"Mark Twain", subject:jazz, collection:gutenberg. Omit and set collection or mediaTypes to list a whole collection. Give queries, collection or identifiers.

## `searchIn` (type: `string`):

Optional. Where plain words must match: "metadata" (default: title, subjects, creator; precise), "title" (title only) or "everything" (full text incl. descriptions and OCR; broad and noisy). field:value queries are used as written.

## `mediaTypes` (type: `array`):

Optional. Only these item types, list of: "texts", "movies", "audio", "software", "image", "data", "web", "collection", "etree" (live concerts). Omit for all types.

## `sortBy` (type: `string`):

Optional. Result order: "relevance" (default), "downloads" (all time), "downloadsThisWeek", "newestUploads", "oldestUploads", "dateNewest", "dateOldest" (publication date), "titleAZ" or "rating".

## `maxResults` (type: `integer`):

Optional. Maximum items per query, integer 0-10000. Default 100; 0 = up to 10,000 (the API limit). You pay per item, so use 10-20 for a quick answer.

## `collection` (type: `string`):

Optional. Only items from this archive.org collection ID (from the collection URL), e.g. "gutenberg", "prelinger", "nasa". Several: comma-separated.

## `yearFrom` (type: `integer`):

Optional. Only items published in or after this year, integer 0-2100, e.g. 1950.

## `yearTo` (type: `integer`):

Optional. Only items published in or before this year, integer 0-2100, e.g. 1999.

## `language` (type: `string`):

Optional. Only items in this language as archive.org stores it: usually a 3-letter code ("eng", "fre", "ger", "spa") or a name ("English"). Several: comma-separated.

## `includeFiles` (type: `boolean`):

Optional boolean. true = add every file in each item (PDF, EPUB, MP3, MP4...) with format, size, MD5 and direct download URL, plus the full description; one extra request per item. Default false.

## `fileFormats` (type: `array`):

Optional, used with includeFiles. Keep only files whose format or extension matches, e.g. \["pdf", "epub", "mp3", "h.264"]. Omit for all files.

## `identifiers` (type: `array`):

Optional. Look up specific items instead of searching: list of archive.org identifiers ("Apollo11Audio") or item URLs ("https://archive.org/details/Apollo11Audio"). Always includes the file list.

## `maxConcurrency` (type: `integer`):

Optional, advanced. Number of parallel requests, integer 1-20. Default 10. Leave unset unless the run is very large.

## `maxRequestRetries` (type: `integer`):

Optional, advanced. Retries per failed or blocked request, integer 0-20; each retry uses a new proxy session. Default 5. Leave unset.

## `proxyConfiguration` (type: `object`):

Optional, advanced. Apify Proxy settings object. Default {"useApifyProxy": false}: no proxy is needed for this public API. Leave unset.

## Actor input object example

```json
{
  "queries": [
    "apollo 11"
  ],
  "searchIn": "metadata",
  "mediaTypes": [
    "movies"
  ],
  "sortBy": "downloads",
  "maxResults": 20,
  "includeFiles": false,
  "maxConcurrency": 10,
  "maxRequestRetries": 5,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}
```

# Actor output Schema

## `items` (type: `string`):

All archive.org items found, with metadata, download stats and (optionally) file download links.

## `summary` (type: `string`):

Per-query counts, archive.org totals and stop reasons.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "apollo 11"
    ],
    "mediaTypes": [
        "movies"
    ],
    "sortBy": "downloads",
    "maxResults": 20,
    "proxyConfiguration": {
        "useApifyProxy": false
    }
};

// Run the Actor and wait for it to finish
const run = await client.actor("rel8ble/internet-archive-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": ["apollo 11"],
    "mediaTypes": ["movies"],
    "sortBy": "downloads",
    "maxResults": 20,
    "proxyConfiguration": { "useApifyProxy": False },
}

# Run the Actor and wait for it to finish
run = client.actor("rel8ble/internet-archive-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "apollo 11"
  ],
  "mediaTypes": [
    "movies"
  ],
  "sortBy": "downloads",
  "maxResults": 20,
  "proxyConfiguration": {
    "useApifyProxy": false
  }
}' |
apify call rel8ble/internet-archive-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,rel8ble/internet-archive-scraper"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/egbA36IpDxajZfNZQ/builds/dCbSf4f1jSVTDfcCr/openapi.json
