Internet Archive (archive.org) Items Scraper avatar

Internet Archive (archive.org) Items Scraper

Pricing

from $6.80 / 1,000 results

Go to Apify Store
Internet Archive (archive.org) Items Scraper

Internet Archive (archive.org) Items Scraper

Scrape Internet Archive items by keyword, media type or identifier. Get title, creator, downloads, subjects, collections, dates, item size and downloadable files as JSON, CSV or Excel.

Pricing

from $6.80 / 1,000 results

Rating

0.0

(0)

Developer

Scrapers Lat

Scrapers Lat

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

a day ago

Last modified

Categories

Share

Internet Archive (archive.org) Items Scraper

Internet Archive (archive.org) Items Scraper

Extract items from the Internet Archive by keyword, media type or identifier, across a public library of more than 40 million texts, movies, audio recordings, software and images.

Apify Coverage Maintained Output

25 fields
per record
Global
coverage
JSON / CSV / Excel
output formats
Updated
2026-07-26

What you get

Each record is one Internet Archive item with its full public metadata, its download statistics, and an optional listing of every downloadable file with a direct link.

  • imageUrl: item thumbnail image
  • title: item title
  • url: link to the item page on archive.org
  • identifier: unique Internet Archive identifier
  • mediatype: item type (texts, movies, audio, software, image, data, web and more)
  • creator: author, band, uploader or organization credited with the item
  • downloads: total number of times the item has been downloaded
  • avgRating: average user star rating when reviews exist
  • numReviews: number of user reviews
  • year: year the work was published or produced
  • date: publication or production date when available
  • publicdate: date the item was made public on archive.org
  • language: language of the work
  • itemSizeBytes: total item size in bytes
  • itemSizeMB: total item size in megabytes
  • numberOfFiles*: exact number of files in the item
  • fileFormats: list of file formats available in the item
  • collections: collections the item belongs to
  • subjects: subjects, tags and keywords describing the item
  • description: item description text
  • licenseUrl: usage license URL when the item declares one
  • files*: every downloadable file with its name, format, size and direct download URL
  • searchQuery: the query that returned this item
  • observedAt: when this item was last seen by the scraper

*These fields only appear when withFiles is set to true.

Who is it for

Use caseWho benefits
Building research and media datasets at scaleData scientists and machine learning teams
Archiving live music, film or radio catalogsMusicologists, archivists and fan communities
Monitoring downloads and popularity of public-domain worksPublishers and rights researchers
Sourcing public-domain books, audio and videoContent creators and educators
Cataloging software, images and web capturesDigital preservation and library teams

Frequently Asked Questions

What can I search on the Internet Archive with this scraper? Anything in the public collection: books and texts, films and video, audio and live music, software, images, data sets and archived web pages. Search by keyword across everything, or narrow to a single media type such as audio or texts.

How many items can I collect in one run? Set the Max Items value to control the volume. A single search can reach many thousands of items, and you can pass several queries at once so one run covers multiple topics. Free Apify plans are capped per run; upgrade for larger pulls.

Can I get the download links for the files inside an item? Yes. Enable the Include downloadable files option and each item returns its complete file list with the file name, format, size and a direct download URL for every file, plus the exact file count.

Can I fetch specific items instead of searching? Yes. Paste Internet Archive identifiers or full archive.org URLs (details, metadata or download links) into the Item identifiers or URLs field and those items are collected directly, with or without a search.

What happens if an item does not exist or is restricted? The item is reported with a clear error note and the run continues with the rest of your input, so one missing or dark item never stops the collection.

Need data from the same space? Here are other scrapers we build and maintain:

More scrapers at scrapers.lat

This actor is built and maintained by scrapers.lat, where we publish scrapers for Latin American and US public platforms: real estate, jobs, e-commerce, company registries and government data. Browse the full catalog, see live sample output for each one, or ask us for a custom scraper at scrapers.lat.


This actor is an independent tool and has no affiliation with the Internet Archive. It only accesses data that is publicly available on the platform. Use it in accordance with the Internet Archive's terms of service.