Text Splitter & RAG Chunker API — Recursive Embedding Chunks
Pricing
from $0.75 / 1,000 text chunkeds
Text Splitter & RAG Chunker API — Recursive Embedding Chunks
Recursive text splitter and RAG chunker for embeddings, vector databases and LLM context. Split on paragraph, line, sentence, word or hard boundaries, add overlap, stable chunk IDs and previous/next chunk links.
Pricing
from $0.75 / 1,000 text chunkeds
Rating
0.0
(0)
Developer
Rosario Vitale
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
0
Monthly active users
3 days ago
Last modified
Categories
Share
Text Splitter API — RAG, LLM & Embedding Chunks
Why use this Actor?
Split text into bounded overlapping chunks for RAG, embeddings, vector databases and LLM context. Process one or many documents, choose character or token-sized chunks, preserve paragraph or sentence boundaries, clean whitespace and enforce safe batch limits.
Features
- Text — A single document to split into chunks. Use this for one long text.
- Texts (multiple) — Several documents to split in one run. Each array item is one document. Maximum 1,000 documents per run.
- Chunk size — Target size of each chunk (in characters or tokens, see Unit).
- Chunk overlap — How much each chunk overlaps the previous one (in characters or tokens). Helps preserve context across boundaries.
- Unit — Whether chunk size and overlap are measured in characters or approximate tokens (~4 chars/token).
- Split by — Preferred boundary to split on before packing chunks.
- Clean text — Normalize whitespace and collapse excessive blank lines before splitting.
Use cases
- Rag ingestion.
- Embedding preparation.
- Vector database chunking.
- Llm context preparation.
Example input
{"chunkSize": 1000,"chunkOverlap": 100,"unit": "characters","splitBy": "paragraph","clean": true}
Pricing & cost control
Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.
FAQ
What is this Actor for?
It is designed for RAG ingestion, embedding preparation, vector database chunking.
Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.
How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.
Search keywords
text splitter rag, best text splitter for rag, rag, text, chunks, llm, vector, split, splitter, markdown, embeddings, documentation, documents, dataset
Split any text into clean, overlapping chunks that are ready for embeddings, vector databases, RAG pipelines and LLM context windows — without writing your own splitter.
Paste text (or send many documents), pick a chunk size and overlap, and get back tidy chunks with character counts and approximate token counts as JSON or CSV.
Why
Every RAG / LLM pipeline needs chunking, and everyone re-implements the same fiddly logic: respect paragraph and sentence boundaries, keep an overlap so context isn't lost, normalize messy whitespace, and estimate tokens. This Actor does it for you, reliably, in one call.
Features
- ✂️ Smart chunking — packs text up to your target size while respecting paragraph/sentence boundaries.
- 🔁 Bounded overlap — keeps context across boundaries without making chunks exceed the requested target size.
- 📍 Source offsets — every chunk includes
startCharandendCharso it can be mapped back to the cleaned source. - 🔢 Characters or tokens — size and overlap in characters or approximate tokens (~4 chars/token).
- 🧹 Cleaning — normalizes whitespace and collapses excessive blank lines.
- 📦 Batch — split many documents in a single run.
- 📊 Token estimate — every chunk includes
charCountandapproxTokens.
Input
| Field | Type | Description |
|---|---|---|
text | string | A single document to split. |
texts | array | Multiple documents (one per item). |
chunkSize | integer | Target chunk size. Default 1000. |
chunkOverlap | integer | Overlap between chunks. Default 100. |
unit | select | characters or tokens. Default characters. |
splitBy | select | paragraph, sentence or character. Default paragraph. |
clean | boolean | Normalize whitespace. Default true. |
Example input
{"text": "Your long document text goes here...","chunkSize": 1000,"chunkOverlap": 100,"unit": "characters","splitBy": "paragraph","clean": true}
Output
One dataset item per chunk:
{"sourceIndex": 0,"chunkIndex": 0,"totalChunks": 3,"text": "Retrieval-Augmented Generation (RAG) combines a language model ...","charCount": 312,"approxTokens": 78,"startChar": 0,"endChar": 312}
Export as JSON, CSV, or Excel, or pull via the Apify API — then send the chunks straight to your embeddings model or vector DB.
Common use cases
- Prepare documents for embeddings + vector search (Pinecone, Qdrant, Weaviate, pgvector).
- Build RAG context for ChatGPT/Claude apps.
- Fit long content into LLM context windows.
- Pairs perfectly with PDF to Structured Data — extract text from PDFs, then chunk it here.
Notes
- Token counts are an estimate (~4 characters per token); exact tokenization depends on the model.
- The Actor enforces
chunkOverlap < chunkSizeand caps total input at 20 million characters. - Paragraph/sentence modes prefer natural boundaries when a suitable one exists; otherwise they fall back to the hard size boundary.
- Returned chunks never exceed the requested target size after unit conversion.
How to use
Paste one document into Text, or provide several documents in Texts. Choose target chunk size, overlap, measurement unit, and preferred boundary. Run the Actor and use the dataset directly in embeddings, vector databases, or retrieval pipelines.
Pricing and cost estimation
This Actor uses pay-per-event pricing per output chunk, plus the small Actor start event shown in the Store pricing panel. You are charged only for chunks the Actor actually emits. With the current base price of $0.00075 per chunk, 1,000 chunks cost about $0.75 before any Apify platform usage shown by your plan.
FAQ and support
Token mode is intentionally approximate because tokenization differs by language model. Use character mode when exact repeatability matters. Report reproducible edge cases through the Actor Issues tab.