Internet Archive Digital Library (archive.org) - Data Scraper avatar

Internet Archive Digital Library (archive.org) - Data Scraper

Pricing

from $50.00 / 1,000 verified record extracteds

Go to Apify Store
Internet Archive Digital Library (archive.org) - Data Scraper

Internet Archive Digital Library (archive.org) - Data Scraper

Extract record data from Internet Archive Digital Library (archive.org) via its official public JSON API. Search by search query, collection, format, media type and export 19 structured fields per record as JSON, CSV or Excel - reliable, with no fragile HTML scraping.

Pricing

from $50.00 / 1,000 verified record extracteds

Rating

0.0

(0)

Developer

Terry Gluff

Terry Gluff

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

3 hours ago

Last modified

Share

Extract record data from Internet Archive Digital Library (archive.org) via its official public JSON API. Search by search query, collection, format, media type and export 19 structured fields per record as JSON, CSV or Excel - reliable, with no fragile HTML scraping.

What you get

Every record is delivered as a clean, structured row you can export to JSON, CSV, or Excel. Fields include:

  • btih
  • collection
  • creator
  • date
  • description
  • downloads
  • format
  • identifier
  • indexflag
  • item_size
  • mediatype
  • month
  • oai_updatedate
  • publicdate
  • stripped_tags
  • subject
  • title
  • week
  • year

Input

Configure the run with these options (all optional -- leave blank to fetch everything):

  • Search query — Search term used to filter results, for example a name, title, or keyword. The actor returns records matching this text. Leave blank to fetch a default sample.
  • Collection — Optional filter on 'collection' (for example 'podcasts_mirror_bobarchives'). Leave blank for no filter.
  • Media type — Optional filter on 'mediatype' (for example 'audio'). Leave blank for no filter. Pick from a dropdown of valid values.
  • Subject — Optional filter on 'subject' (for example 'Podcast'). Leave blank for no filter.
  • Year — Optional filter on 'year' (for example '2022'). Leave blank for no filter. Pick from a dropdown of valid values.
  • Maximum results — cap how many records to return.

How to use

  1. Click Start (or call the Actor via the Apify API or a client SDK).
  2. Set any filters you need and run it.
  3. Download the results as JSON, CSV, or Excel from the run's dataset.

Pricing

This Actor uses pay-per-result pricing: you are billed only for the records it returns, with no monthly subscription.

Notes

This Actor collects publicly available data. Please use it responsibly and in line with the source website's terms of service.