# Internet Archive (archive.org) Items Scraper (`scrapers_lat/archive-org-scraper`) Actor

Scrape Internet Archive items by keyword, media type or identifier. Get title, creator, downloads, subjects, collections, dates, item size and downloadable files as JSON, CSV or Excel.

- **URL**: https://apify.com/scrapers\_lat/archive-org-scraper.md
- **Developed by:** [Scrapers Lat](https://apify.com/scrapers_lat) (community)
- **Categories:** News, Other
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $6.80 / 1,000 results

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/platform/actors/running/actors-in-store#pay-per-event

## What's an Apify Actor?

Actors are web data automations that power AI and operations. They run on the Apify platform to scrape websites, process data, connect APIs, and automate workflows.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.
Actors are written with capital "A".

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.
The best way to integrate Actors is as follows.

- **AI agents and MCP clients** — the [Apify MCP server](https://docs.apify.com/integrations/mcp.md) at `https://mcp.apify.com` (remote, streamable HTTP, OAuth on first use).
- **Agentic workflows and local Actor development** — [Agent Skills](https://apify.com/.well-known/agent-skills/index.json) with the [Apify CLI](https://docs.apify.com/cli/docs.md): `npm install -g apify-cli`, then `apify login`.
- **JavaScript/TypeScript projects** — the official [JS/TS client](https://docs.apify.com/api/client/js/docs.md): `npm install apify-client`.
- **Python projects** — the official [Python client](https://docs.apify.com/api/client/python/docs.md): `pip install apify-client`.
- **Any other language** — the [REST API](https://docs.apify.com/api/v2.md).

For usage examples, see the [API](#api) section below.

For more details, see Apify documentation as [Markdown index](https://docs.apify.com/llms.txt) and [Markdown full-text](https://docs.apify.com/llms-full.txt).

# README

[![Internet Archive (archive.org) Items Scraper](https://scrapers.lat/banners/archive-org-scraper.png)](https://console.apify.com/actors/fWRnsBqC6KjhkqIET/input)

## Internet Archive (archive.org) Items Scraper

> Extract items from the Internet Archive by keyword, media type or identifier, across a public library of more than 40 million texts, movies, audio recordings, software and images.

![Apify](https://img.shields.io/badge/Platform-Apify-1CE1CE?logo=apify\&logoColor=white)
![Coverage](https://img.shields.io/badge/Coverage-Global-blue)
![Maintained](https://img.shields.io/badge/Maintained-Yes-brightgreen)
![Output](https://img.shields.io/badge/Output-JSON%20%7C%20CSV%20%7C%20Excel-orange)

<table><tr>
<td align="center"><strong>25 fields</strong><br>per record</td>
<td align="center"><strong>Global</strong><br>coverage</td>
<td align="center"><strong>JSON / CSV / Excel</strong><br>output formats</td>
<td align="center"><strong>Updated</strong><br>2026-07-26</td>
</tr></table>

<br>

### What you get

Each record is one Internet Archive item with its full public metadata, its download statistics, and an optional listing of every downloadable file with a direct link.

- **imageUrl**: item thumbnail image
- **title**: item title
- **url**: link to the item page on archive.org
- **identifier**: unique Internet Archive identifier
- **mediatype**: item type (texts, movies, audio, software, image, data, web and more)
- **creator**: author, band, uploader or organization credited with the item
- **downloads**: total number of times the item has been downloaded
- **avgRating**: average user star rating when reviews exist
- **numReviews**: number of user reviews
- **year**: year the work was published or produced
- **date**: publication or production date when available
- **publicdate**: date the item was made public on archive.org
- **language**: language of the work
- **itemSizeBytes**: total item size in bytes
- **itemSizeMB**: total item size in megabytes
- **numberOfFiles\***: exact number of files in the item
- **fileFormats**: list of file formats available in the item
- **collections**: collections the item belongs to
- **subjects**: subjects, tags and keywords describing the item
- **description**: item description text
- **licenseUrl**: usage license URL when the item declares one
- **files\***: every downloadable file with its name, format, size and direct download URL
- **searchQuery**: the query that returned this item
- **observedAt**: when this item was last seen by the scraper

*\*These fields only appear when withFiles is set to true.*

### Who is it for

| Use case | Who benefits |
|---|---|
| Building research and media datasets at scale | Data scientists and machine learning teams |
| Archiving live music, film or radio catalogs | Musicologists, archivists and fan communities |
| Monitoring downloads and popularity of public-domain works | Publishers and rights researchers |
| Sourcing public-domain books, audio and video | Content creators and educators |
| Cataloging software, images and web captures | Digital preservation and library teams |

### Frequently Asked Questions

**What can I search on the Internet Archive with this scraper?**
Anything in the public collection: books and texts, films and video, audio and live music, software, images, data sets and archived web pages. Search by keyword across everything, or narrow to a single media type such as audio or texts.

**How many items can I collect in one run?**
Set the Max Items value to control the volume. A single search can reach many thousands of items, and you can pass several queries at once so one run covers multiple topics. Free Apify plans are capped per run; upgrade for larger pulls.

**Can I get the download links for the files inside an item?**
Yes. Enable the Include downloadable files option and each item returns its complete file list with the file name, format, size and a direct download URL for every file, plus the exact file count.

**Can I fetch specific items instead of searching?**
Yes. Paste Internet Archive identifiers or full archive.org URLs (details, metadata or download links) into the Item identifiers or URLs field and those items are collected directly, with or without a search.

**What happens if an item does not exist or is restricted?**
The item is reported with a clear error note and the run continues with the rest of your input, so one missing or dark item never stops the collection.

### Related scrapers

Need data from the same space? Here are other scrapers we build and maintain:

- [arXiv Research Papers & Abstracts Scraper](https://apify.com/scrapers_lat/arxiv-papers-scraper): Scrape arXiv preprints with authors, abstracts, categories and PDF links.
- [Clinical Trials Scraper](https://apify.com/scrapers_lat/clinicaltrials-scraper): Extract clinical study records with conditions, sponsors, phases and status.
- [Chrome Web Store Extensions Scraper](https://apify.com/scrapers_lat/chrome-web-store-scraper): Collect browser extension listings with ratings, users and versions.
- [YouTube Video & Channel Data Scraper](https://apify.com/scrapers_lat/youtube-scraper): Pull video and channel metadata, views and engagement stats.
- [Reddit Posts & Comments Scraper](https://apify.com/scrapers_lat/reddit-scraper): Gather posts and comment threads with scores and authors.
- [Apple App Store Reviews & Ratings Scraper](https://apify.com/scrapers_lat/app-store-reviews-scraper): Harvest app reviews, ratings and version history.

### More scrapers at scrapers.lat

This actor is built and maintained by [scrapers.lat](https://scrapers.lat), where we publish scrapers for Latin American and US public platforms: real estate, jobs, e-commerce, company registries and government data. Browse the full catalog, see live sample output for each one, or ask us for a custom scraper at [scrapers.lat](https://scrapers.lat).

***

> This actor is an independent tool and has no affiliation with the Internet Archive. It only accesses data that is publicly available on the platform. Use it in accordance with the Internet Archive's terms of service.

# Actor input Schema

## `maxItems` (type: `integer`):

Maximum number of Internet Archive items to collect across all searches. Optional.

## `withFiles` (type: `boolean`):

When enabled, each item also includes its full downloadable file list (name, format, size and direct download URL) plus the exact file count. Adds one metadata request per item.

## `searchQueries` (type: `array`):

One or more words or phrases to search across Internet Archive (for example 'grateful dead', 'apollo 11', 'nasa'). Each query runs a separate search. You can also paste a native Internet Archive query such as 'title:(moon) AND creator:nasa'. Leave empty if you are passing item identifiers or URLs below.

## `itemUrls` (type: `array`):

Fetch specific items directly. Accepts Internet Archive identifiers (for example 'nasa') or full URLs (archive.org/details/... or archive.org/metadata/...). When provided, these items are collected in addition to any search results.

## `mediaType` (type: `string`):

Restrict search results to one media type. Ignored for items passed as identifiers or URLs.

## `sortBy` (type: `string`):

Order search results.

## Actor input object example

```json
{
  "maxItems": 10,
  "withFiles": true,
  "searchQueries": [
    "grateful dead"
  ],
  "mediaType": "all",
  "sortBy": "relevance"
}
```

# Actor output Schema

## `results` (type: `string`):

No description

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "maxItems": 10,
    "searchQueries": [
        "grateful dead"
    ]
};

// Run the Actor and wait for it to finish
const run = await client.actor("scrapers_lat/archive-org-scraper").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "maxItems": 10,
    "searchQueries": ["grateful dead"],
}

# Run the Actor and wait for it to finish
run = client.actor("scrapers_lat/archive-org-scraper").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print("💾 Check your data here: https://console.apify.com/storage/datasets/" + run["defaultDatasetId"])
for item in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "maxItems": 10,
  "searchQueries": [
    "grateful dead"
  ]
}' |
apify call scrapers_lat/archive-org-scraper --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "command": "npx",
            "args": [
                "mcp-remote",
                "https://mcp.apify.com/?tools=scrapers_lat/archive-org-scraper",
                "--header",
                "Authorization: Bearer <YOUR_API_TOKEN>"
            ]
        }
    }
}

```

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/acts/fWRnsBqC6KjhkqIET/builds/kOaqRaD4f4D18t3Gb/openapi.json
