Go to example tasks
Chunk product docs for a support chatbot
Three pages of product documentation go in, chunks ready to embed come out: each chunk keeps its source URL, its heading path and exact token counts, the pricing table stays in one piece, and the page that is mirrored on a second URL is delivered once with both URLs kept. Change the records for your own pages, or point the Actor at a dataset from a crawler.
RAG Dataset Builder - Source-Linked Chunks for Retrievalleadproof/rag-dataset-builder
Source
Section
#
Tokens
+3 fieldsTextNumberBooleanListObject
Input
PageRecords
schemaVersion:1.0+2
inputId:pricing+2
status:succeeded+2
finalUrl:https://docs.example.com/billing/pricing+2
title:Pricing+2
language:en+2
markdown:# Pricing
Every plan includes the API, webhooks and email support.
## Plans
| Plan | Monthly | Seats |
|------|---------|-------|
| Starter | $29 | 3 |
| Team | $99 | 15 |
## Billing questions
### When am I charged?
On the first day of each billing period, for the period ahead.
### Can I change plan mid-month?
Yes. The change takes effect immediately and the next invoice is prorated.+2
Target chunk size:256
Overlap:32
Output fields
Source
Section
#
Tokens
Text
Chunk ID
Document ID
Sign up on Apify01
Create your Apify account to access the RAG Dataset Builder - Source-Linked Chunks for Retrieval.
Start the run02
The Actor will start running based on the input automatically.
Receive the output03
Monitor the progress in real-time. You will be notified as soon as your dataset is complete and ready for review.
Integrate into your workflow04
The final output is delivered in JSON, CSV, or Excel format, ready to be plugged into your workflow.
