LibriVox Audiobooks Scraper avatar

LibriVox Audiobooks Scraper

Pricing

from $10.00 / 1,000 result items

Go to Apify Store
LibriVox Audiobooks Scraper

LibriVox Audiobooks Scraper

Pull free public domain audiobooks from LibriVox: title, author, narrator, language, runtime, chapter count, genre, copyright year, description, RSS feed, and MP3 download URLs. Export to JSON, CSV, or Excel for educators, podcasters, language learners, and audio content libraries.

Pricing

from $10.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

6 days ago

Last modified

Share

ParseForge Banner

๐Ÿ“š LibriVox Audiobooks Scraper

๐Ÿš€ Export the world's largest public-domain audiobook library in seconds. Browse 20,000+ free audiobooks from LibriVox, filter by title, author, language, or genre, and pull every section's reader credit, runtime, and direct audio URL. No login, no manual catalog scrape.

The LibriVox Audiobooks Scraper queries the LibriVox catalog and returns 23 structured fields per audiobook, including the title, author list, primary author, language, copyright year, runtime in human-readable and seconds form, description, genres, translators, plus direct links to the LibriVox page, RSS podcast feed, ZIP download, Internet Archive page, and original Project Gutenberg text source. Optional extended mode adds the full sections list with per-section reader credits, individual playtimes, and per-file audio URLs.

The catalog includes classic literature read aloud (Project Gutenberg titles), original LibriVox productions, multilingual works, and short-story collections. Most readings are in English but the project covers dozens of other languages including French, German, Spanish, Italian, Dutch, Portuguese, Latin, Japanese, and Mandarin. This Actor turns the catalog into clean CSV, Excel, JSON, or XML in under five minutes.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Audiobook app developers, podcast networks, education and EdTech teams, accessibility specialists, public libraries, audio content curatorsStock audiobook apps with free content, classroom listening assignments, accessibility libraries for visually impaired users, podcast feed generation, language-learning audio decks

๐Ÿ“‹ What the LibriVox Audiobooks Scraper does

Five filtering workflows in a single run:

  • ๐Ÿ”Ž Title substring search. Case-insensitive title match like pride, monte cristo.
  • โœ๏ธ Author substring search. Match by author last name or full name.
  • ๐ŸŒ Language filter. Filter to a specific LibriVox language (English, French, German, Spanish, Japanese, etc.).
  • ๐ŸŽญ Genre filter. Pick one of 28 LibriVox genres (Romance, Crime & Mystery, Philosophy, Poetry, and more).
  • ๐Ÿ“‘ Extended mode toggle. When enabled, each record carries the full sections list with reader credits, runtimes, and audio URLs.

Each record includes the LibriVox ID, title, full author list, primary author, language, copyright year, section count, total runtime (human-readable and seconds), full description (plain text and HTML), genres, translators when applicable, and the canonical LibriVox page, RSS podcast feed, ZIP archive URL, Project Gutenberg source link, Internet Archive page, and any other reference URLs.

๐Ÿ’ก Why it matters: LibriVox is the canonical free audiobook archive, but its catalog page is paginated and its section structure is nested under each book. Building your own crawler means walking thousands of pages and threading section lookups. This Actor returns everything as flat structured rows ready for a database or content platform.

๐Ÿ“Š Data fields

Each record includes: authors, copyrightYear, description, descriptionHtml, genres, language, librivoxId, numSections, primaryAuthor, scrapedAt, sections, sectionsCount, title, totalTime, totalTimeSeconds, translators, urlInternetArchive, urlLibrivox, urlProject, urlRss, urlTextSource, urlZipFile. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the LibriVox Audiobooks Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Add optional title, author, language, or genre filters, choose extended mode if you want per-section data.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect catalog records.
  5. ๐Ÿ“ฅ Download. Grab your results from the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the LibriVox project. All trademarks mentioned are the property of their respective owners. Only publicly available catalog data is collected. LibriVox audiobooks are dedicated to the public domain.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.