Text Splitter & Chunker for RAG / LLM Embeddings avatar

Text Splitter & Chunker for RAG / LLM Embeddings

Pricing

from $0.0001 / chunk created

Go to Apify Store
Text Splitter & Chunker for RAG / LLM Embeddings

Text Splitter & Chunker for RAG / LLM Embeddings

Split long text into clean overlapping chunks for RAG, vector embeddings and LLM pipelines. Recursive boundary-aware or fixed-size splitting, by characters, words or approx tokens, with adjustable overlap. Each chunk includes char/word/token counts and offset. Output JSON, CSV or Excel.

Pricing

from $0.0001 / chunk created

Rating

0.0

(0)

Developer

hiper soft

hiper soft

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

2 days ago

Last modified

Categories

Share

Split long text into clean, overlapping chunks ready for RAG (retrieval-augmented generation), vector embeddings and LLM pipelines with the Text Splitter & Chunker. Paste an article, document or transcript and get back well-sized chunks — by characters, words or approximate tokens — with adjustable overlap, exported as JSON, CSV or Excel.

Chunking is the first step of almost every AI knowledge base, semantic search and chatbot build: documents must be broken into pieces that fit an embedding model's context and retrieve cleanly. This tool does exactly that, with no setup and no code.

What it does

  • ✂️ Smart recursive splitting — keeps paragraphs, sentences and words intact wherever possible, so chunks stay readable and meaningful.
  • 📏 Size in your unit — set chunk size in characters, words or approximate tokens (~4 chars each).
  • 🔁 Overlap control — carry a slice of the previous chunk into the next so context isn't lost at the seams (standard practice for RAG).
  • 🧱 Fixed-size mode — need exact, uniform chunks? Switch to a hard fixed-size split.
  • 🗂️ Batch documents — chunk many texts in one run; each is chunked independently and labelled by document.
  • 📊 Useful metadata — every chunk includes character, word and approximate token counts plus its start offset.
  • 📤 Export anywhere — JSON, CSV or Excel, ready to feed into an embedding step, a vector database, or an n8n / Make / Zapier workflow.

Example output

{
"documentIndex": 0,
"index": 2,
"text": "…the second half of the section, kept whole at a sentence boundary…",
"charCount": 812,
"wordCount": 137,
"approxTokens": 203,
"startOffset": 1588,
"unit": "tokens",
"strategy": "recursive"
}

How to use it

  1. Paste your text (or add several texts in the list).
  2. Choose a chunk size and overlap, and the unit (characters, words or tokens).
  3. Pick a strategy — Recursive (recommended) or Fixed.
  4. Run, and download the chunks as JSON, CSV or Excel.

Tip for embeddings: a common starting point is ~800–1,000 tokens per chunk with ~100–200 tokens of overlap. Smaller chunks improve retrieval precision; larger chunks keep more context per chunk.

Input fields

FieldDescription
TextThe text to split.
Multiple textsOptional list of separate documents to chunk in one run.
Chunk sizeTarget maximum size of each chunk, in the chosen unit.
Chunk overlapHow much each chunk overlaps the previous one.
Size unitCharacters, words, or approximate tokens.
Split strategyRecursive (boundary-aware) or Fixed size.
Trim whitespaceTrim each chunk's leading/trailing whitespace.
Max chunksMaximum number of chunks to output (0 = no limit).

Output fields

documentIndex, index, text, charCount, wordCount, approxTokens, startOffset, unit, strategy, collectedAt.

  • RAG knowledge bases — chunk docs before embedding for a chatbot or assistant.
  • Vector search — prepare uniform, overlapping passages for a vector database.
  • LLM context prep — break long inputs into model-sized pieces.
  • Semantic search indexing — split articles and manuals into retrievable passages.
  • Workflow automation — a chunking step inside an n8n, Make or Zapier pipeline.
  • Dataset preparation — turn raw documents into a clean chunk dataset for ML.

FAQ

Do I need an account or key? No. Paste your text, choose the settings, and run.

How are tokens counted? Token counts are an estimate using the common ~4-characters-per-token heuristic for English. They're ideal for sizing chunks to an embedding model without exact tokenizer setup.

What does "recursive" splitting mean? It tries to split on the largest natural boundary first (paragraphs), then lines, sentences and finally words — so chunks rarely cut a sentence in half. Fixed mode cuts at exact sizes instead.

Why use overlap? Overlap repeats a little context between neighbouring chunks so that answers spanning a boundary are still retrievable. It's a standard RAG technique.

Can I export to Excel or Google Sheets? Yes — results download as JSON, CSV or Excel and integrate with Sheets, vector databases and automation tools.


Turn long documents into clean, overlapping, embed-ready chunks — boundary-aware, sized your way, and exported for any RAG or LLM pipeline.