Project Gutenberg Books Scraper avatar

Project Gutenberg Books Scraper

Pricing

from $13.00 / 1,000 result items

Go to Apify Store
Project Gutenberg Books Scraper

Project Gutenberg Books Scraper

Search 75,000+ free public-domain books from Project Gutenberg. Returns title, author with birth/death years, cover image, plain-text and EPUB download URLs, Kindle and HTML formats, subjects, bookshelves, language, copyright status, summaries and download counts. Filter by author or language.

Pricing

from $13.00 / 1,000 result items

Rating

0.0

(0)

Developer

ParseForge

ParseForge

Maintained by Community

Actor stats

0

Bookmarked

4

Total users

2

Monthly active users

3 days ago

Last modified

Share

ParseForge Banner

๐Ÿ“š Project Gutenberg Books Scraper

๐Ÿš€ Search 75,000+ free public-domain books from Project Gutenberg.

The Project Gutenberg Books Scraper searches the Project Gutenberg catalog and returns structured records for any free public-domain ebook. Output includes title, author with birth/death years, cover image, plain-text and EPUB download URLs, Kindle and HTML formats, subjects, bookshelves, language, copyright status, summaries, and download counts.

Project Gutenberg has been digitizing public-domain texts since 1971 and now hosts 75,000+ books across 60+ languages. Filters run server-side, so a single run can isolate every Shakespeare play, all 19th-century French novels, or the most-downloaded books of all time.

๐ŸŽฏ Target Audience๐Ÿ’ก Primary Use Cases
Researchers, NLP/ML teams, librarians, educators, content creators, ebook app developersBuilding text corpora, NLP training datasets, public-domain ebook libraries, literary research, citation generation

๐Ÿ“‹ What the Project Gutenberg Books Scraper does

Five filtering workflows in a single run:

  • ๐Ÿ” Free-text search. Match by title, author, or general keywords.
  • ๐Ÿ‘ค Author filter. Restrict to one author across all their works.
  • ๐Ÿท๏ธ Topic filter. Filter by subject (history, philosophy, science, fiction).
  • ๐ŸŒ Language filter. ISO 639 language codes (en, fr, de, es, zh, ja).
  • ๐Ÿ“… Author year filter. Filter authors by birth/death year for period studies.

๐Ÿ’ก Why it matters: clean, server-side filtering removes the parser-and-pagination work from your team and keeps your dataset fresh on every run.

๐Ÿ“Š Data fields

Each record includes: authors, copyright, coverUrl, downloadCount, gutenbergId, gutenbergUrl, languages, mediaType, subjectCount, title. These field names come straight from the actor's dataset schema, so what you see here is what lands in your dataset.

๐Ÿš€ How to use

  1. ๐Ÿ“ Sign up. Create a free account with $5 credit (takes 2 minutes).
  2. ๐ŸŒ Open the Actor. Go to the Project Gutenberg Books Scraper page on the Apify Store.
  3. ๐ŸŽฏ Set input. Pick your filters and maxItems.
  4. ๐Ÿš€ Run it. Click Start and let the Actor collect your data.
  5. ๐Ÿ“ฅ Download. Grab your results in the Dataset tab as CSV, Excel, JSON, or XML.

โฑ๏ธ Total time from signup to downloaded dataset: 3-5 minutes. No coding required.

๐Ÿ’ก Pro Tip: browse the complete ParseForge collection for more reference-data scrapers.

โš ๏ธ Disclaimer: this Actor is an independent tool and is not affiliated with, endorsed by, or sponsored by Project Gutenberg, the Gutendex project, or any contributing volunteers. All trademarks mentioned are the property of their respective owners. Only publicly available open data is collected.

๐Ÿ†˜ Need Help?

If you hit a bug, have questions about setup, or need a scraper we haven't built yet, open our contact form or write to parseforge@protonmail.com. We also take on paid custom data projects.

For faster answers, join our Discord. It's the best place to get support and suggest new actors.