Internet Archive Search: books, audio, video and software items avatar

Internet Archive Search: books, audio, video and software items

Pricing

from $0.65 / 1,000 item listeds

Go to Apify Store
Internet Archive Search: books, audio, video and software items

Internet Archive Search: books, audio, video and software items

Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.

Pricing

from $0.65 / 1,000 item listeds

Rating

0.0

(0)

Developer

Steadydata Team

Steadydata Team

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 hours ago

Last modified

Share

Items from the Internet Archive by search term, media type, collection, subject, creator, language and year: title, creator, date, description, subjects, downloads, size, licence and links to the item and its files, from the official search API. Up to 50 searches a run. Pay per item.

Why this scraper

  • Only delivered results are charged. Inputs that fail come back as clear error records at no cost.
  • Straight from archive.org's own search API, the same index behind the site's search box, with the full Lucene query language: free text, subject:(jazz), creator:(...), year:[1950 TO 1959], AND and OR. Measured on the platform: 240 items for two searches over texts and audio, for an eighth of a cent.
  • One row per item with the catalogue fields that matter: title, creator, date and year, description, subject tags, language, publisher, the collections it sits in, download count, total size, declared licence, when it was added, the item page and the file directory for download.
  • Sorted by popularity (downloads) by default, or by the date of the work, the date added or relevance; filters on media type (texts, audio, movies, software, image, data, web), a specific collection and a language.
  • Lists are folded to one shape whether the archive stored a field as a string or a list, and users' favourite lists are dropped from the collections, so the rows are ready to use.

Who this is for

Put search terms or Lucene queries in queries (up to 50 per run), choose mediaTypes (default texts), optionally collection, language, sortBy and maxItemsPerQuery (default 100). Built for researchers and librarians building corpora, publishers and rights teams checking what is in the public domain, teachers and archivists collecting material on a subject, and developers who need a dataset of digitised books, recordings, films or old software with metadata.

Who this is not for

This lists items and their metadata; it does not download the files (the filesUrl directory does). The metadata is what uploaders typed: creator, date, language and licence are missing or inconsistent on many items, and subjects can be a single free-text string. Download counts are the archive's own and include automated traffic. The archive asks for a polite pace, so a run of many searches takes about a second per page.

Input fields

FieldTypeRequired or defaultWhat it does
querieslist of textrequiredOne search per row, up to 50: words in title, description and subjects (cookbook), or a Lucene query such as subject:(jazz) AND year:[1950 TO 1959].
mediaTypeslist of texttexts, audio, movies, software, image, data, web, collection. Empty means every type.
collectiontextOnly items in this collection, by its identifier, for example librivoxaudio or prelinger. Empty means every collection.
languagetextOnly items in this language as the archive labels it, for example English, eng or Dutch. Empty means every language.
sortBytext (downloads, date, publicdate, relevance)downloadsdownloads (most popular first), date (newest first), publicdate (recently added first) or relevance.
maxItemsPerQuerynumber100Cost ceiling per search.

Input example

{
"queries": [
"cookbook",
"subject:(jazz) AND year:[1950 TO 1959]"
],
"mediaTypes": [
"texts"
],
"sortBy": "downloads",
"maxItemsPerQuery": 100
}

Output example

FieldTypeWhat it holds
identifiertextThe archive's unique item name, the last part of every archive.org URL of the item.
titletextThe item title as the uploader gave it.
creatortextThe author, artist or maker of the work as catalogued; several are joined with semicolons.
mediaTypetexttexts, audio, movies, software, image, data, web or collection.
datetextThe date of the work itself (publication or recording), when catalogued; not the upload date.
yearnumberThe year of the work, when catalogued.
descriptiontextThe uploader's description of the item; can be long and can contain HTML.
subjectslistThe subject tags of the item, split on commas and semicolons.
languagetextThe language as catalogued, in words or as a code.
publishertextThe publisher of the work, when catalogued.
collectionslistThe collections the item belongs to (favourite lists of users are left out).
downloadsnumberHow many times the item was downloaded, the archive's popularity measure.
sizeBytesnumberThe total size of the item's files.
licenseUrltextThe licence the uploader declared, when any.
addedAttextWhen the item was added to the archive.
urltextThe item page on archive.org.
filesUrltextThe directory with the item's files for download.
querytextThe search term this row was found with.

Error codes: INVALID_QUERY, NO_RESULTS, BLOCKED.

One delivered row looks like this:

{
"identifier": "recordchanger11unse",
"title": "The record changer (Jan-Dec 1952)",
"creator": "Changer Publications, Inc.",
"mediaType": "texts",
"date": "1952-01-01",
"year": 1952,
"description": null,
"subjects": [
"sound recording periodical",
"Jazz -- Periodicals",
"Jazz -- Discography -- Periodicals"
],
"language": "eng",
"publisher": "New York : Changer Publications, Inc.",
"collections": [
"libraryofcongresspackardcampus",
"mediahistory",
"fedlink",
"library_of_congress",
"americana"
],
"downloads": 34551,
"sizeBytes": 1482070890,
"licenseUrl": null,
"addedAt": "2014-09-04T15:36:51Z",
"url": "https://archive.org/details/recordchanger11unse",
"filesUrl": "https://archive.org/download/recordchanger11unse",
"query": "subject:(jazz) AND year:[1950 TO 1959]",
"status": "ok"
}

Pricing

Pay per event: one item-listed event per delivered result. No charge for inputs that fail, no separate platform-usage surcharge.

Free Apify plan: this actor delivers up to 25 rows per run for accounts on the Apify free plan, and then stops with a message. That limit is set by us, not by Apify. It exists so the actor keeps paying for itself for the people who do pay. Any paid Apify plan runs it at full size, billed per delivered row, with failed rows never charged.

Reviews: if this actor saves you time, a short review on this page is the one thing that helps most. Ratings are what other buyers look at first, and we have no other way to ask.

FAQ

Is personal data collected? No. creator is the catalogued author or maker of a published work, the same as on a library card; no user profiles are read.

How do I search a specific field? Use the archive's query language in the search term: title:(cookbook), creator:(Bach), subject:(jazz) AND year:[1950 TO 1959], description:(radio drama). A term without fields searches title, description and subjects.

How do I get only public-domain items? Add licenseurl:(*publicdomain*) to the query, or filter on licenseUrl in the rows; note that many old items carry no licence field at all even when the work is out of copyright.

Can I list a whole collection? Yes: set collection to its identifier (for example librivoxaudio or prelinger) and use * or a broad term as the search; raise maxItemsPerQuery for the size of the collection.

Where are the files? filesUrl is the item's download directory with every file (PDF, MP3, MP4, ZIP); the actor links it and does not download files.

What does a run cost when a search finds nothing? Nothing. NO_RESULTS, INVALID_QUERY and BLOCKED rows are free; only delivered items are charged.

What happens when the source changes? Sources change from time to time; that is the nature of this work. The actor is monitored daily and fixed fast, and while it is broken you are not charged, because only delivered results cost anything.