Go to example tasks
Multi-Source Chunking Pipeline
Created by
Dennis
Combine Project Gutenberg and Open Library with chapter-level chunking and overlap. Build clean RAG datasets from multiple public-domain catalogs.
Public Domain Ebook RAG Feedcodeclouds/public-domain-ebook-rag-feed
Book ID
Chunk #
Text
Chars
TextNumberBooleanListObject
Input
Data sources:gutenberg+1
Max books to fetch:10
RAG chunking options
Enable chunking:true
Chunking strategy:chapter
Max chunk characters:1500
Overlap characters:150
Output fields
Book ID
Chunk #
Text
Chars
Sign up on Apify01
Create your Apify account to access the Public Domain Ebook RAG Feed.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
