# Internet Archive Search: books, audio, video and software items (`steadydata/internet-archive-search`) Actor

Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.

- **URL**: https://apify.com/steadydata/internet-archive-search.md
- **Developed by:** [Steadydata Team](https://apify.com/steadydata) (community)
- **Categories:** Developer tools, News, Automation
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.65 / 1,000 item listeds

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Internet Archive Search: books, audio, video and software items

Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.

### Why this scraper

- **Only delivered results are charged.** Inputs that fail come back as clear error
  records at no cost.
- Straight from archive.org's own search API, the same index behind the site's search box, with the full Lucene query language: free text, `subject:(jazz)`, `creator:(...)`, `year:[1950 TO 1959]`, AND and OR. Measured on the platform: 240 items for two searches over texts and audio, for an eighth of a cent.
- One row per item with the catalogue fields that matter: title, creator, date and year, description, subject tags, language, publisher, the collections it sits in, download count, total size, declared licence, when it was added, the item page and the file directory for download.
- Sorted by popularity (downloads) by default, or by the date of the work, the date added or relevance; filters on media type (texts, audio, movies, software, image, data, web), a specific collection and a language.
- Lists are folded to one shape whether the archive stored a field as a string or a list, and users' favourite lists are dropped from the collections, so the rows are ready to use.

### Who this is for

Put search terms or Lucene queries in `queries` (up to 50 per run), choose `mediaTypes` (default texts), optionally `collection`, `language`, `sortBy` and `maxItemsPerQuery` (default 100). Built for researchers and librarians building corpora, publishers and rights teams checking what is in the public domain, teachers and archivists collecting material on a subject, and developers who need a dataset of digitised books, recordings, films or old software with metadata.

### Who this is not for

This lists items and their metadata; it does not download the files (the `filesUrl` directory does). The metadata is what uploaders typed: creator, date, language and licence are missing or inconsistent on many items, and `subjects` can be a single free-text string. Download counts are the archive's own and include automated traffic. The archive asks for a polite pace, so a run of many searches takes about a second per page.

### Input fields

| Field | Type | Required or default | What it does |
|---|---|---|---|
| `queries` | list of text | required | One search per row, up to 50: words in title, description and subjects (cookbook), or a Lucene query such as subject:(jazz) AND year:\[1950 TO 1959]. |
| `mediaTypes` | list of text |  | texts, audio, movies, software, image, data, web, collection. Empty means every type. |
| `collection` | text |  | Only items in this collection, by its identifier, for example librivoxaudio or prelinger. Empty means every collection. |
| `language` | text |  | Only items in this language as the archive labels it, for example English, eng or Dutch. Empty means every language. |
| `sortBy` | text (downloads, date, publicdate, relevance) | downloads | downloads (most popular first), date (newest first), publicdate (recently added first) or relevance. |
| `maxItemsPerQuery` | number | 100 | Cost ceiling per search. |

### Input example

```json
{
    "queries": [
        "cookbook",
        "subject:(jazz) AND year:[1950 TO 1959]"
    ],
    "mediaTypes": [
        "texts"
    ],
    "sortBy": "downloads",
    "maxItemsPerQuery": 100
}
```

### Output example

| Field | Type | What it holds |
|---|---|---|
| `identifier` | text | The archive's unique item name, the last part of every archive.org URL of the item. |
| `title` | text | The item title as the uploader gave it. |
| `creator` | text | The author, artist or maker of the work as catalogued; several are joined with semicolons. |
| `mediaType` | text | texts, audio, movies, software, image, data, web or collection. |
| `date` | text | The date of the work itself (publication or recording), when catalogued; not the upload date. |
| `year` | number | The year of the work, when catalogued. |
| `description` | text | The uploader's description of the item; can be long and can contain HTML. |
| `subjects` | list | The subject tags of the item, split on commas and semicolons. |
| `language` | text | The language as catalogued, in words or as a code. |
| `publisher` | text | The publisher of the work, when catalogued. |
| `collections` | list | The collections the item belongs to (favourite lists of users are left out). |
| `downloads` | number | How many times the item was downloaded, the archive's popularity measure. |
| `sizeBytes` | number | The total size of the item's files. |
| `licenseUrl` | text | The licence the uploader declared, when any. |
| `addedAt` | text | When the item was added to the archive. |
| `url` | text | The item page on archive.org. |
| `filesUrl` | text | The directory with the item's files for download. |
| `query` | text | The search term this row was found with. |

Error codes: `INVALID_QUERY`, `NO_RESULTS`, `BLOCKED`.

One delivered row looks like this:

```json
{
  "identifier": "recordchanger11unse",
  "title": "The record changer (Jan-Dec 1952)",
  "creator": "Changer Publications, Inc.",
  "mediaType": "texts",
  "date": "1952-01-01",
  "year": 1952,
  "description": null,
  "subjects": [
    "sound recording periodical",
    "Jazz -- Periodicals",
    "Jazz -- Discography -- Periodicals"
  ],
  "language": "eng",
  "publisher": "New York : Changer Publications, Inc.",
  "collections": [
    "libraryofcongresspackardcampus",
    "mediahistory",
    "fedlink",
    "library_of_congress",
    "americana"
  ],
  "downloads": 34551,
  "sizeBytes": 1482070890,
  "licenseUrl": null,
  "addedAt": "2014-09-04T15:36:51Z",
  "url": "https://archive.org/details/recordchanger11unse",
  "filesUrl": "https://archive.org/download/recordchanger11unse",
  "query": "subject:(jazz) AND year:[1950 TO 1959]",
  "status": "ok"
}
```

### Related actors from steadydata

- [wayback-history](https://apify.com/steadydata/wayback-history): the archived versions of any web page, from the same archive
- [arxiv-papers](https://apify.com/steadydata/arxiv-papers): scientific papers with abstracts and authors
- [huggingface-models](https://apify.com/steadydata/huggingface-models): open models and datasets

### Pricing

Pay per event: one `item-listed` event per delivered result. No charge for inputs
that fail, no separate platform-usage surcharge.

**Free Apify plan:** this actor delivers up to 25 rows per run for accounts on the Apify free
plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the
actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full
size, billed per delivered row, with failed rows never charged.

**Reviews:** if this actor saves you time, a short review on this page is the one thing that
helps most. Ratings are what other buyers look at first, and we have no other way to ask.

### FAQ

**Is personal data collected?**
No. `creator` is the catalogued author or maker of a published work, the same as on a library card; no user profiles are read.

**How do I search a specific field?**
Use the archive's query language in the search term: `title:(cookbook)`, `creator:(Bach)`, `subject:(jazz) AND year:[1950 TO 1959]`, `description:(radio drama)`. A term without fields searches title, description and subjects.

**How do I get only public-domain items?**
Add `licenseurl:(*publicdomain*)` to the query, or filter on `licenseUrl` in the rows; note that many old items carry no licence field at all even when the work is out of copyright.

**Can I list a whole collection?**
Yes: set `collection` to its identifier (for example `librivoxaudio` or `prelinger`) and use `*` or a broad term as the search; raise `maxItemsPerQuery` for the size of the collection.

**Where are the files?**
`filesUrl` is the item's download directory with every file (PDF, MP3, MP4, ZIP); the actor links it and does not download files.

**What does a run cost when a search finds nothing?**
Nothing. `NO_RESULTS`, `INVALID_QUERY` and `BLOCKED` rows are free; only delivered items are charged.

**What happens when the source changes?**
Sources change from time to time; that is the nature of this work. The actor is
monitored daily and fixed fast, and while it is broken you are not charged, because
only delivered results cost anything.

# Changelog

This Actor's version history is a separate document: https://apify.com/steadydata/internet-archive-search/changelog.md

# Actor input Schema

## `queries` (type: `array`):

One search per row, up to 50: words in title, description and subjects (cookbook), or a Lucene query such as subject:(jazz) AND year:\[1950 TO 1959].

## `mediaTypes` (type: `array`):

texts, audio, movies, software, image, data, web, collection. Empty means every type.

## `collection` (type: `string`):

Only items in this collection, by its identifier, for example librivoxaudio or prelinger. Empty means every collection.

## `language` (type: `string`):

Only items in this language as the archive labels it, for example English, eng or Dutch. Empty means every language.

## `sortBy` (type: `string`):

downloads (most popular first), date (newest first), publicdate (recently added first) or relevance.

## `maxItemsPerQuery` (type: `integer`):

Cost ceiling per search.

## Actor input object example

```json
{
  "queries": [
    "cookbook",
    "subject:(jazz) AND year:[1950 TO 1959]"
  ],
  "mediaTypes": [
    "texts"
  ],
  "sortBy": "downloads",
  "maxItemsPerQuery": 100
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "queries": [
        "cookbook",
        "subject:(jazz) AND year:[1950 TO 1959]"
    ],
    "mediaTypes": [
        "texts"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("steadydata/internet-archive-search").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "queries": [
        "cookbook",
        "subject:(jazz) AND year:[1950 TO 1959]",
    ],
    "mediaTypes": ["texts"],
}

# Run the Actor and wait for it to finish
run = client.actor("steadydata/internet-archive-search").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "queries": [
    "cookbook",
    "subject:(jazz) AND year:[1950 TO 1959]"
  ],
  "mediaTypes": [
    "texts"
  ]
}' |
apify call steadydata/internet-archive-search --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,steadydata/internet-archive-search"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/W6FbH25yAMqYduH94/builds/8FnwEyBdHd3eEUp1w/openapi.json
