Internet Archive (archive.org) Items Scraper
Pricing
from $6.80 / 1,000 results
Internet Archive (archive.org) Items Scraper
Scrape Internet Archive items by keyword, media type or identifier. Get title, creator, downloads, subjects, collections, dates, item size and downloadable files as JSON, CSV or Excel.
Pricing
from $6.80 / 1,000 results
Rating
0.0
(0)
Developer
Scrapers Lat
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
a day ago
Last modified
Share
Internet Archive (archive.org) Items Scraper
Extract items from the Internet Archive by keyword, media type or identifier, across a public library of more than 40 million texts, movies, audio recordings, software and images.
| 25 fields per record | Global coverage | JSON / CSV / Excel output formats | Updated 2026-07-26 |
What you get
Each record is one Internet Archive item with its full public metadata, its download statistics, and an optional listing of every downloadable file with a direct link.
- imageUrl: item thumbnail image
- title: item title
- url: link to the item page on archive.org
- identifier: unique Internet Archive identifier
- mediatype: item type (texts, movies, audio, software, image, data, web and more)
- creator: author, band, uploader or organization credited with the item
- downloads: total number of times the item has been downloaded
- avgRating: average user star rating when reviews exist
- numReviews: number of user reviews
- year: year the work was published or produced
- date: publication or production date when available
- publicdate: date the item was made public on archive.org
- language: language of the work
- itemSizeBytes: total item size in bytes
- itemSizeMB: total item size in megabytes
- numberOfFiles*: exact number of files in the item
- fileFormats: list of file formats available in the item
- collections: collections the item belongs to
- subjects: subjects, tags and keywords describing the item
- description: item description text
- licenseUrl: usage license URL when the item declares one
- files*: every downloadable file with its name, format, size and direct download URL
- searchQuery: the query that returned this item
- observedAt: when this item was last seen by the scraper
*These fields only appear when withFiles is set to true.
Who is it for
| Use case | Who benefits |
|---|---|
| Building research and media datasets at scale | Data scientists and machine learning teams |
| Archiving live music, film or radio catalogs | Musicologists, archivists and fan communities |
| Monitoring downloads and popularity of public-domain works | Publishers and rights researchers |
| Sourcing public-domain books, audio and video | Content creators and educators |
| Cataloging software, images and web captures | Digital preservation and library teams |
Frequently Asked Questions
What can I search on the Internet Archive with this scraper? Anything in the public collection: books and texts, films and video, audio and live music, software, images, data sets and archived web pages. Search by keyword across everything, or narrow to a single media type such as audio or texts.
How many items can I collect in one run? Set the Max Items value to control the volume. A single search can reach many thousands of items, and you can pass several queries at once so one run covers multiple topics. Free Apify plans are capped per run; upgrade for larger pulls.
Can I get the download links for the files inside an item? Yes. Enable the Include downloadable files option and each item returns its complete file list with the file name, format, size and a direct download URL for every file, plus the exact file count.
Can I fetch specific items instead of searching? Yes. Paste Internet Archive identifiers or full archive.org URLs (details, metadata or download links) into the Item identifiers or URLs field and those items are collected directly, with or without a search.
What happens if an item does not exist or is restricted? The item is reported with a clear error note and the run continues with the rest of your input, so one missing or dark item never stops the collection.
Related scrapers
Need data from the same space? Here are other scrapers we build and maintain:
- arXiv Research Papers & Abstracts Scraper: Scrape arXiv preprints with authors, abstracts, categories and PDF links.
- Clinical Trials Scraper: Extract clinical study records with conditions, sponsors, phases and status.
- Chrome Web Store Extensions Scraper: Collect browser extension listings with ratings, users and versions.
- YouTube Video & Channel Data Scraper: Pull video and channel metadata, views and engagement stats.
- Reddit Posts & Comments Scraper: Gather posts and comment threads with scores and authors.
- Apple App Store Reviews & Ratings Scraper: Harvest app reviews, ratings and version history.
More scrapers at scrapers.lat
This actor is built and maintained by scrapers.lat, where we publish scrapers for Latin American and US public platforms: real estate, jobs, e-commerce, company registries and government data. Browse the full catalog, see live sample output for each one, or ask us for a custom scraper at scrapers.lat.
This actor is an independent tool and has no affiliation with the Internet Archive. It only accesses data that is publicly available on the platform. Use it in accordance with the Internet Archive's terms of service.
