Library of Congress Scraper avatar

Library of Congress Scraper

Pricing

from $14.00 / 1,000 result items

Go to Apify Store
Library of Congress Scraper

Library of Congress Scraper

Export records from the US Library of Congress catalog of 170M+ items. Search books, audio, film, maps, manuscripts, newspapers, photos, sheet music, and web archives. Pull titles, contributors, dates, subjects, languages, image URLs, and direct catalog links.

Pricing

from $14.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Categories

Share

ParseForge Banner

๐Ÿ›๏ธ Library of Congress Scraper

๐Ÿš€ Export the world's largest cultural archive in seconds. Search 170,000,000+ digitized items at the US Library of Congress across 11 format types, including books, audio, film, maps, manuscripts, newspapers, photos, sheet music, and web archives. No login, no manual harvesting.

The Library of Congress Scraper queries the LOC digital catalog and returns 18 structured fields per record, including title, contributors, date, subjects, languages, format, mediums, rights, repository info, and direct links to resource files and image derivatives. The LOC has been digitizing its holdings since the 1990s and exposes the world's most comprehensive open cultural catalog.

The catalog spans books and printed material, audio recordings, films, maps, manuscripts, historical newspapers, photographs, sheet music, notated music, web archives, and curated collections. This Actor returns the data as CSV, Excel, JSON, or XML in under five minutes, with year-range, language, and collection filters applied server-side.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Historians, archivists, journalists, educators, documentary producers, genealogists, digital humanities researchers, museum curatorsSource primary documents, build classroom packs, enrich research databases, locate rights-cleared media, map historical newspapers, source public-domain images

๐Ÿ“‹ What the Library of Congress Scraper does

Five archival workflows in a single run:

  • ๐Ÿ“š Format-scoped search. Pick one of 11 LOC formats (books, audio, film, maps, manuscripts, newspapers, photos, sheet music, web archives, notated music, collections).
  • ๐Ÿ”Ž Keyword search. Free-text search across the chosen format.
  • ๐ŸŒ Language filter. Restrict to a single language slug (e.g. english, spanish, french, chinese, arabic).
  • ๐Ÿ“… Date range. Earliest and latest year inclusive, for time-bounded research.
  • ๐Ÿ—‚๏ธ Collection filter. Restrict to a curated LOC collection slug (e.g. wpa-life-histories, civil-war-maps).

Each record includes the LOC item ID, title, description, contributor list, date, subject tags, language list, format and medium, parent collection, repository, rights statement, every resource URL (manifests, audio, video, IIIF images), and a primary image thumbnail.

๐Ÿ’ก Why it matters: the LOC catalog is the foundational reference for American cultural and political history. Building your own harvester means navigating multiple catalog endpoints, parsing nested metadata, and chasing pagination across millions of records. This Actor turns the entire catalog into a download.

๐Ÿ“Š Data fields

Each record includes: contributors, createdPublished, date, description, digitalIds, format, imageUrl, imageUrls, itemId, languages, mediums, notes, partof, researchCenters, resourceUrls, scrapedAt, sourceCollection, subjects, title, topics, url. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the Library of Congress Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Pick a format, add a keyword, set optional language, year range, or collection.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect catalog records.
  5. ๐Ÿ“ฅ Download. Grab your results from the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the US Library of Congress. All trademarks mentioned are the property of their respective owners. Only publicly available catalog data is collected. Honor each item's individual rights statement.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.