Internet Archive Search: books, audio, video and software items
Pricing
from $0.65 / 1,000 item listeds
Internet Archive Search: books, audio, video and software items
Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.
Pricing
from $0.65 / 1,000 item listeds
Rating
0.0
(0)
Developer
Steadydata Team
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
2 hours ago
Last modified
Categories
Share
Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.
Why this scraper
- Only delivered results are charged. Inputs that fail come back as clear error records at no cost.
- Straight from archive.org's own search API, the same index behind the site's search box, with the full Lucene query language: free text,
subject:(jazz),creator:(...),year:[1950 TO 1959], AND and OR. Measured on the platform: 240 items for two searches over texts and audio, for an eighth of a cent. - One row per item with the catalogue fields that matter: title, creator, date and year, description, subject tags, language, publisher, the collections it sits in, download count, total size, declared licence, when it was added, the item page and the file directory for download.
- Sorted by popularity (downloads) by default, or by the date of the work, the date added or relevance; filters on media type (texts, audio, movies, software, image, data, web), a specific collection and a language.
- Lists are folded to one shape whether the archive stored a field as a string or a list, and users' favourite lists are dropped from the collections, so the rows are ready to use.
Who this is for
Put search terms or Lucene queries in queries (up to 50 per run), choose mediaTypes (default texts), optionally collection, language, sortBy and maxItemsPerQuery (default 100). Built for researchers and librarians building corpora, publishers and rights teams checking what is in the public domain, teachers and archivists collecting material on a subject, and developers who need a dataset of digitised books, recordings, films or old software with metadata.
Who this is not for
This lists items and their metadata; it does not download the files (the filesUrl directory does). The metadata is what uploaders typed: creator, date, language and licence are missing or inconsistent on many items, and subjects can be a single free-text string. Download counts are the archive's own and include automated traffic. The archive asks for a polite pace, so a run of many searches takes about a second per page.
Input fields
| Field | Type | Required or default | What it does |
|---|---|---|---|
queries | list of text | required | One search per row, up to 50: words in title, description and subjects (cookbook), or a Lucene query such as subject:(jazz) AND year:[1950 TO 1959]. |
mediaTypes | list of text | texts, audio, movies, software, image, data, web, collection. Empty means every type. | |
collection | text | Only items in this collection, by its identifier, for example librivoxaudio or prelinger. Empty means every collection. | |
language | text | Only items in this language as the archive labels it, for example English, eng or Dutch. Empty means every language. | |
sortBy | text (downloads, date, publicdate, relevance) | downloads | downloads (most popular first), date (newest first), publicdate (recently added first) or relevance. |
maxItemsPerQuery | number | 100 | Cost ceiling per search. |
Input example
{"queries": ["cookbook","subject:(jazz) AND year:[1950 TO 1959]"],"mediaTypes": ["texts"],"sortBy": "downloads","maxItemsPerQuery": 100}
Output example
| Field | Type | What it holds |
|---|---|---|
identifier | text | The archive's unique item name, the last part of every archive.org URL of the item. |
title | text | The item title as the uploader gave it. |
creator | text | The author, artist or maker of the work as catalogued; several are joined with semicolons. |
mediaType | text | texts, audio, movies, software, image, data, web or collection. |
date | text | The date of the work itself (publication or recording), when catalogued; not the upload date. |
year | number | The year of the work, when catalogued. |
description | text | The uploader's description of the item; can be long and can contain HTML. |
subjects | list | The subject tags of the item, split on commas and semicolons. |
language | text | The language as catalogued, in words or as a code. |
publisher | text | The publisher of the work, when catalogued. |
collections | list | The collections the item belongs to (favourite lists of users are left out). |
downloads | number | How many times the item was downloaded, the archive's popularity measure. |
sizeBytes | number | The total size of the item's files. |
licenseUrl | text | The licence the uploader declared, when any. |
addedAt | text | When the item was added to the archive. |
url | text | The item page on archive.org. |
filesUrl | text | The directory with the item's files for download. |
query | text | The search term this row was found with. |
Error codes: INVALID_QUERY, NO_RESULTS, BLOCKED.
One delivered row looks like this:
{"identifier": "recordchanger11unse","title": "The record changer (Jan-Dec 1952)","creator": "Changer Publications, Inc.","mediaType": "texts","date": "1952-01-01","year": 1952,"description": null,"subjects": ["sound recording periodical","Jazz -- Periodicals","Jazz -- Discography -- Periodicals"],"language": "eng","publisher": "New York : Changer Publications, Inc.","collections": ["libraryofcongresspackardcampus","mediahistory","fedlink","library_of_congress","americana"],"downloads": 34551,"sizeBytes": 1482070890,"licenseUrl": null,"addedAt": "2014-09-04T15:36:51Z","url": "https://archive.org/details/recordchanger11unse","filesUrl": "https://archive.org/download/recordchanger11unse","query": "subject:(jazz) AND year:[1950 TO 1959]","status": "ok"}
Related actors from steadydata
- wayback-history: the archived versions of any web page, from the same archive
- arxiv-papers: scientific papers with abstracts and authors
- huggingface-models: open models and datasets
Pricing
Pay per event: one item-listed event per delivered result. No charge for inputs
that fail, no separate platform-usage surcharge.
Free Apify plan: this actor delivers up to 25 rows per run for accounts on the Apify free plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full size, billed per delivered row, with failed rows never charged.
Reviews: if this actor saves you time, a short review on this page is the one thing that helps most. Ratings are what other buyers look at first, and we have no other way to ask.
FAQ
Is personal data collected?
No. creator is the catalogued author or maker of a published work, the same as on a library card; no user profiles are read.
How do I search a specific field?
Use the archive's query language in the search term: title:(cookbook), creator:(Bach), subject:(jazz) AND year:[1950 TO 1959], description:(radio drama). A term without fields searches title, description and subjects.
How do I get only public-domain items?
Add licenseurl:(*publicdomain*) to the query, or filter on licenseUrl in the rows; note that many old items carry no licence field at all even when the work is out of copyright.
Can I list a whole collection?
Yes: set collection to its identifier (for example librivoxaudio or prelinger) and use * or a broad term as the search; raise maxItemsPerQuery for the size of the collection.
Where are the files?
filesUrl is the item's download directory with every file (PDF, MP3, MP4, ZIP); the actor links it and does not download files.
What does a run cost when a search finds nothing?
Nothing. NO_RESULTS, INVALID_QUERY and BLOCKED rows are free; only delivered items are charged.
What happens when the source changes? Sources change from time to time; that is the nature of this work. The actor is monitored daily and fixed fast, and while it is broken you are not charged, because only delivered results cost anything.