LibriVox Audiobooks Scraper
Pricing
from $10.00 / 1,000 result items
LibriVox Audiobooks Scraper
Pull free public domain audiobooks from LibriVox: title, author, narrator, language, runtime, chapter count, genre, copyright year, description, RSS feed, and MP3 download URLs. Export to JSON, CSV, or Excel for educators, podcasters, language learners, and audio content libraries.
Pricing
from $10.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
6 days ago
Last modified
Share

๐ LibriVox Audiobooks Scraper
๐ Export the world's largest public-domain audiobook library in seconds. Browse 20,000+ free audiobooks from LibriVox, filter by title, author, language, or genre, and pull every section's reader credit, runtime, and direct audio URL. No login, no manual catalog scrape.
The LibriVox Audiobooks Scraper queries the LibriVox catalog and returns 23 structured fields per audiobook, including the title, author list, primary author, language, copyright year, runtime in human-readable and seconds form, description, genres, translators, plus direct links to the LibriVox page, RSS podcast feed, ZIP download, Internet Archive page, and original Project Gutenberg text source. Optional extended mode adds the full sections list with per-section reader credits, individual playtimes, and per-file audio URLs.
The catalog includes classic literature read aloud (Project Gutenberg titles), original LibriVox productions, multilingual works, and short-story collections. Most readings are in English but the project covers dozens of other languages including French, German, Spanish, Italian, Dutch, Portuguese, Latin, Japanese, and Mandarin. This Actor turns the catalog into clean CSV, Excel, JSON, or XML in under five minutes.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| Audiobook app developers, podcast networks, education and EdTech teams, accessibility specialists, public libraries, audio content curators | Stock audiobook apps with free content, classroom listening assignments, accessibility libraries for visually impaired users, podcast feed generation, language-learning audio decks |
๐ What the LibriVox Audiobooks Scraper does
Five filtering workflows in a single run:
- ๐ Title substring search. Case-insensitive title match like
pride,monte cristo. - โ๏ธ Author substring search. Match by author last name or full name.
- ๐ Language filter. Filter to a specific LibriVox language (English, French, German, Spanish, Japanese, etc.).
- ๐ญ Genre filter. Pick one of 28 LibriVox genres (Romance, Crime & Mystery, Philosophy, Poetry, and more).
- ๐ Extended mode toggle. When enabled, each record carries the full sections list with reader credits, runtimes, and audio URLs.
Each record includes the LibriVox ID, title, full author list, primary author, language, copyright year, section count, total runtime (human-readable and seconds), full description (plain text and HTML), genres, translators when applicable, and the canonical LibriVox page, RSS podcast feed, ZIP archive URL, Project Gutenberg source link, Internet Archive page, and any other reference URLs.
๐ก Why it matters: LibriVox is the canonical free audiobook archive, but its catalog page is paginated and its section structure is nested under each book. Building your own crawler means walking thousands of pages and threading section lookups. This Actor returns everything as flat structured rows ready for a database or content platform.
๐ Data fields
Each record includes: authors, copyrightYear, description, descriptionHtml, genres, language, librivoxId, numSections, primaryAuthor, scrapedAt, sections, sectionsCount, title, totalTime, totalTimeSeconds, translators, urlInternetArchive, urlLibrivox, urlProject, urlRss, urlTextSource, urlZipFile. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
๐ How to use
- ๐ Sign up. Create a free account with $5 credit (takes 2 minutes).
- ๐ Open the Actor. Go to the LibriVox Audiobooks Scraper page on the Apify Store.
- ๐ฏ Set input. Add optional title, author, language, or genre filters, choose extended mode if you want per-section data.
- ๐ Run it. Click Start and let the Actor collect catalog records.
- ๐ฅ Download. Grab your results from the Dataset tab as CSV, Excel, JSON, or XML.
โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
๐ Recommended Actors
- ๐๏ธ Library of Congress Scraper - 170M+ digitized cultural records
- ๐ฃ๏ธ Tatoeba Sentence Corpus Scraper - 12M+ multilingual example sentences
- ๐ MyMemory Translation Scraper - Bulk text translation across 70+ language codes
- ๐ฐ ArXiv Scraper - Academic preprints with metadata
- ๐จ Met Museum Scraper - Open-access artworks from The Met
๐ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
โ ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by the LibriVox project. All trademarks mentioned are the property of their respective owners. Only publicly available catalog data is collected. LibriVox audiobooks are dedicated to the public domain.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.