Project Gutenberg Books Scraper
Pricing
from $13.00 / 1,000 result items
Project Gutenberg Books Scraper
Search 75,000+ free public-domain books from Project Gutenberg. Returns title, author with birth/death years, cover image, plain-text and EPUB download URLs, Kindle and HTML formats, subjects, bookshelves, language, copyright status, summaries and download counts. Filter by author or language.
Pricing
from $13.00 / 1,000 result items
Rating
0.0
(0)
Developer
ParseForge
Maintained by CommunityActor stats
0
Bookmarked
4
Total users
2
Monthly active users
3 days ago
Last modified
Categories
Share

๐ Project Gutenberg Books Scraper
๐ Search 75,000+ free public-domain books from Project Gutenberg.
The Project Gutenberg Books Scraper searches the Project Gutenberg catalog and returns structured records for any free public-domain ebook. Output includes title, author with birth/death years, cover image, plain-text and EPUB download URLs, Kindle and HTML formats, subjects, bookshelves, language, copyright status, summaries, and download counts.
Project Gutenberg has been digitizing public-domain texts since 1971 and now hosts 75,000+ books across 60+ languages. Filters run server-side, so a single run can isolate every Shakespeare play, all 19th-century French novels, or the most-downloaded books of all time.
| ๐ฏ Target Audience | ๐ก Primary Use Cases |
|---|---|
| Researchers, NLP/ML teams, librarians, educators, content creators, ebook app developers | Building text corpora, NLP training datasets, public-domain ebook libraries, literary research, citation generation |
๐ What the Project Gutenberg Books Scraper does
Five filtering workflows in a single run:
- ๐ Free-text search. Match by title, author, or general keywords.
- ๐ค Author filter. Restrict to one author across all their works.
- ๐ท๏ธ Topic filter. Filter by subject (history, philosophy, science, fiction).
- ๐ Language filter. ISO 639 language codes (en, fr, de, es, zh, ja).
- ๐ Author year filter. Filter authors by birth/death year for period studies.
๐ก Why it matters: clean, server-side filtering removes the parser-and-pagination work from your team and keeps your dataset fresh on every run.
๐ Data fields
Each record includes: authors, copyright, coverUrl, downloadCount, gutenbergId, gutenbergUrl, languages, mediaType, subjectCount, title. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.
๐ How to use
- ๐ Sign up. Create a free account with $5 credit (takes 2 minutes).
- ๐ Open the Actor. Go to the Project Gutenberg Books Scraper page on the Apify Store.
- ๐ฏ Set input. Pick your filters and
maxItems. - ๐ Run it. Click Start and let the Actor collect your data.
- ๐ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.
โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.
๐ Recommended Actors
- ๐ Open Library Books - 30M+ books and editions
- ๐ Wikidata Entity Search - 100M+ open knowledge-graph entities
- ๐จ Openverse Media - 800M+ openly licensed images and audio
- ๐ arXiv Scraper - Academic preprints
- ๐ฌ TVMaze TV Shows - TV show metadata
๐ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.
โ ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Project Gutenberg, the Gutendex project, or any contributing volunteers. All trademarks mentioned are the property of their respective owners. Only publicly available open data is collected.
๐ Need Help?
If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.
For faster answers, join our Discord. It's the best place to get support and suggest new actors.