Text Splitter & RAG Chunker API — Recursive Embedding Chunks avatar

Text Splitter & RAG Chunker API — Recursive Embedding Chunks

Pricing

from $0.75 / 1,000 text chunkeds

Go to Apify Store
Text Splitter & RAG Chunker API — Recursive Embedding Chunks

Text Splitter & RAG Chunker API — Recursive Embedding Chunks

Recursive text splitter and RAG chunker for embeddings, vector databases and LLM context. Split on paragraph, line, sentence, word or hard boundaries, add overlap, stable chunk IDs and previous/next chunk links.

Pricing

from $0.75 / 1,000 text chunkeds

Rating

0.0

(0)

Developer

Rosario Vitale

Rosario Vitale

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

0

Monthly active users

3 days ago

Last modified

Share

Text Splitter API — RAG, LLM & Embedding Chunks

Why use this Actor?

Split text into bounded overlapping chunks for RAG, embeddings, vector databases and LLM context. Process one or many documents, choose character or token-sized chunks, preserve paragraph or sentence boundaries, clean whitespace and enforce safe batch limits.

Features

  • Text — A single document to split into chunks. Use this for one long text.
  • Texts (multiple) — Several documents to split in one run. Each array item is one document. Maximum 1,000 documents per run.
  • Chunk size — Target size of each chunk (in characters or tokens, see Unit).
  • Chunk overlap — How much each chunk overlaps the previous one (in characters or tokens). Helps preserve context across boundaries.
  • Unit — Whether chunk size and overlap are measured in characters or approximate tokens (~4 chars/token).
  • Split by — Preferred boundary to split on before packing chunks.
  • Clean text — Normalize whitespace and collapse excessive blank lines before splitting.

Use cases

  • Rag ingestion.
  • Embedding preparation.
  • Vector database chunking.
  • Llm context preparation.

Example input

{
"chunkSize": 1000,
"chunkOverlap": 100,
"unit": "characters",
"splitBy": "paragraph",
"clean": true
}

Pricing & cost control

Use the bounded input limits and filters to keep runs predictable. Pay-per-result Actors only charge primary result rows; summary, status and monitoring metadata are designed to add context without inflating result volume.

FAQ

What is this Actor for?
It is designed for RAG ingestion, embedding preparation, vector database chunking.

Can I run it on a schedule?
Yes. You can schedule Actor runs on Apify and send the resulting dataset into automations, webhooks, storage, or downstream APIs.

How do I control cost and run size?
Use the input limits and filters shown in the Actor input form. The Actor applies bounded defaults and hard caps so large jobs remain predictable.

Search keywords

text splitter rag, best text splitter for rag, rag, text, chunks, llm, vector, split, splitter, markdown, embeddings, documentation, documents, dataset

Split any text into clean, overlapping chunks that are ready for embeddings, vector databases, RAG pipelines and LLM context windows — without writing your own splitter.

Paste text (or send many documents), pick a chunk size and overlap, and get back tidy chunks with character counts and approximate token counts as JSON or CSV.

Why

Every RAG / LLM pipeline needs chunking, and everyone re-implements the same fiddly logic: respect paragraph and sentence boundaries, keep an overlap so context isn't lost, normalize messy whitespace, and estimate tokens. This Actor does it for you, reliably, in one call.

Features

  • ✂️ Smart chunking — packs text up to your target size while respecting paragraph/sentence boundaries.
  • 🔁 Bounded overlap — keeps context across boundaries without making chunks exceed the requested target size.
  • 📍 Source offsets — every chunk includes startChar and endChar so it can be mapped back to the cleaned source.
  • 🔢 Characters or tokens — size and overlap in characters or approximate tokens (~4 chars/token).
  • 🧹 Cleaning — normalizes whitespace and collapses excessive blank lines.
  • 📦 Batch — split many documents in a single run.
  • 📊 Token estimate — every chunk includes charCount and approxTokens.

Input

FieldTypeDescription
textstringA single document to split.
textsarrayMultiple documents (one per item).
chunkSizeintegerTarget chunk size. Default 1000.
chunkOverlapintegerOverlap between chunks. Default 100.
unitselectcharacters or tokens. Default characters.
splitByselectparagraph, sentence or character. Default paragraph.
cleanbooleanNormalize whitespace. Default true.

Example input

{
"text": "Your long document text goes here...",
"chunkSize": 1000,
"chunkOverlap": 100,
"unit": "characters",
"splitBy": "paragraph",
"clean": true
}

Output

One dataset item per chunk:

{
"sourceIndex": 0,
"chunkIndex": 0,
"totalChunks": 3,
"text": "Retrieval-Augmented Generation (RAG) combines a language model ...",
"charCount": 312,
"approxTokens": 78,
"startChar": 0,
"endChar": 312
}

Export as JSON, CSV, or Excel, or pull via the Apify API — then send the chunks straight to your embeddings model or vector DB.

Common use cases

  • Prepare documents for embeddings + vector search (Pinecone, Qdrant, Weaviate, pgvector).
  • Build RAG context for ChatGPT/Claude apps.
  • Fit long content into LLM context windows.
  • Pairs perfectly with PDF to Structured Data — extract text from PDFs, then chunk it here.

Notes

  • Token counts are an estimate (~4 characters per token); exact tokenization depends on the model.
  • The Actor enforces chunkOverlap < chunkSize and caps total input at 20 million characters.
  • Paragraph/sentence modes prefer natural boundaries when a suitable one exists; otherwise they fall back to the hard size boundary.
  • Returned chunks never exceed the requested target size after unit conversion.

How to use

Paste one document into Text, or provide several documents in Texts. Choose target chunk size, overlap, measurement unit, and preferred boundary. Run the Actor and use the dataset directly in embeddings, vector databases, or retrieval pipelines.

Pricing and cost estimation

This Actor uses pay-per-event pricing per output chunk, plus the small Actor start event shown in the Store pricing panel. You are charged only for chunks the Actor actually emits. With the current base price of $0.00075 per chunk, 1,000 chunks cost about $0.75 before any Apify platform usage shown by your plan.

FAQ and support

Token mode is intentionally approximate because tokenization differs by language model. Use character mode when exact repeatability matters. Report reproducible edge cases through the Actor Issues tab.