# Text Splitter & Chunker for RAG / LLM Embeddings (`hipersoft/text-splitter-chunker`) Actor

Split long text into clean overlapping chunks for RAG, vector embeddings and LLM pipelines. Recursive boundary-aware or fixed-size splitting, by characters, words or approx tokens, with adjustable overlap. Each chunk includes char/word/token counts and offset. Output JSON, CSV or Excel.

- **URL**: https://apify.com/hipersoft/text-splitter-chunker.md
- **Developed by:** [hiper soft](https://apify.com/hipersoft) (community)
- **Stats:** 2 total users, 1 monthly users, 100.0% runs succeeded, 0 bookmarks
- **User rating**: No ratings yet

## Pricing

from $0.0001 / chunk created

This Actor is paid per event. You are not charged for the Apify platform usage, but only a fixed price for specific events.
Since this Actor supports Apify Store discounts, the price gets lower the higher subscription plan you have.

Learn more: https://docs.apify.com/actors/running/actors-in-store.md#pay-per-event

## What's an Apify Actor?

An Actor is a serverless cloud program that runs on the Apify platform. It has two run modes.
In Batch mode, an Actor accepts a well-defined JSON input, performs an action which can take anything from a few seconds to a few hours,
and optionally produces a well-defined JSON output, datasets with results, or files in key-value store.
In Standby mode, an Actor provides a web server which can be used as a website, API, or an MCP server.

Apify vocabulary and the platform model are defined once, in the agent quickstart at https://apify.com/agents.md.

## How to integrate an Actor?

If asked about integration, you help developers integrate Actors into their projects.
You adapt to their stack and deliver integrations that are safe, well-documented, and production-ready.

Do not guess an integration path. Every one of them is in the agent quickstart at https://apify.com/agents.md: the Apify MCP server, Agent Skills with the Apify CLI, the JavaScript and Python clients, the REST API, and the account-free path for an agent with no human to sign in. It also carries the rule on stating cost before the first paid run.

For examples already wired to this Actor's own input schema, see the [API](#api) section below.

Each client library has reference documentation the quickstart does not restate: [JavaScript/TypeScript](https://docs.apify.com/api/client/js/docs.md) (`npm install apify-client`) and [Python](https://docs.apify.com/api/client/python/docs.md) (`pip install apify-client`).

# README

## Text Splitter & Chunker for RAG / LLM Embeddings

Split long text into clean, overlapping **chunks** ready for **RAG (retrieval-augmented generation), vector embeddings and LLM pipelines** with the **Text Splitter & Chunker**. Paste an article, document or transcript and get back well-sized chunks — by **characters, words or approximate tokens** — with adjustable **overlap**, exported as **JSON, CSV or Excel**.

Chunking is the first step of almost every **AI knowledge base, semantic search and chatbot** build: documents must be broken into pieces that fit an embedding model's context and retrieve cleanly. This tool does exactly that, with no setup and no code.

### What it does

- ✂️ **Smart recursive splitting** — keeps paragraphs, sentences and words intact wherever possible, so chunks stay readable and meaningful.
- 📏 **Size in your unit** — set chunk size in **characters, words or approximate tokens** (~4 chars each).
- 🔁 **Overlap control** — carry a slice of the previous chunk into the next so context isn't lost at the seams (standard practice for RAG).
- 🧱 **Fixed-size mode** — need exact, uniform chunks? Switch to a hard fixed-size split.
- 🗂️ **Batch documents** — chunk many texts in one run; each is chunked independently and labelled by document.
- 📊 **Useful metadata** — every chunk includes character, word and approximate token counts plus its start offset.
- 📤 **Export anywhere** — JSON, CSV or Excel, ready to feed into an embedding step, a vector database, or an n8n / Make / Zapier workflow.

### Example output

```json
{
  "documentIndex": 0,
  "index": 2,
  "text": "…the second half of the section, kept whole at a sentence boundary…",
  "charCount": 812,
  "wordCount": 137,
  "approxTokens": 203,
  "startOffset": 1588,
  "unit": "tokens",
  "strategy": "recursive"
}
```

### How to use it

1. Paste your **text** (or add several texts in the list).
2. Choose a **chunk size** and **overlap**, and the **unit** (characters, words or tokens).
3. Pick a **strategy** — Recursive (recommended) or Fixed.
4. Run, and download the chunks as **JSON, CSV or Excel**.

**Tip for embeddings:** a common starting point is ~800–1,000 tokens per chunk with ~100–200 tokens of overlap. Smaller chunks improve retrieval precision; larger chunks keep more context per chunk.

### Input fields

| Field | Description |
|-------|-------------|
| **Text** | The text to split. |
| **Multiple texts** | Optional list of separate documents to chunk in one run. |
| **Chunk size** | Target maximum size of each chunk, in the chosen unit. |
| **Chunk overlap** | How much each chunk overlaps the previous one. |
| **Size unit** | Characters, words, or approximate tokens. |
| **Split strategy** | Recursive (boundary-aware) or Fixed size. |
| **Trim whitespace** | Trim each chunk's leading/trailing whitespace. |
| **Max chunks** | Maximum number of chunks to output (0 = no limit). |

### Output fields

`documentIndex`, `index`, `text`, `charCount`, `wordCount`, `approxTokens`, `startOffset`, `unit`, `strategy`, `collectedAt`.

### Popular use cases

- **RAG knowledge bases** — chunk docs before embedding for a chatbot or assistant.
- **Vector search** — prepare uniform, overlapping passages for a vector database.
- **LLM context prep** — break long inputs into model-sized pieces.
- **Semantic search indexing** — split articles and manuals into retrievable passages.
- **Workflow automation** — a chunking step inside an n8n, Make or Zapier pipeline.
- **Dataset preparation** — turn raw documents into a clean chunk dataset for ML.

### FAQ

**Do I need an account or key?**
No. Paste your text, choose the settings, and run.

**How are tokens counted?**
Token counts are an estimate using the common ~4-characters-per-token heuristic for English. They're ideal for sizing chunks to an embedding model without exact tokenizer setup.

**What does "recursive" splitting mean?**
It tries to split on the largest natural boundary first (paragraphs), then lines, sentences and finally words — so chunks rarely cut a sentence in half. Fixed mode cuts at exact sizes instead.

**Why use overlap?**
Overlap repeats a little context between neighbouring chunks so that answers spanning a boundary are still retrievable. It's a standard RAG technique.

**Can I export to Excel or Google Sheets?**
Yes — results download as JSON, CSV or Excel and integrate with Sheets, vector databases and automation tools.

***

Turn long documents into clean, overlapping, embed-ready **chunks** — boundary-aware, sized your way, and exported for any RAG or LLM pipeline.

# Actor input Schema

## `text` (type: `string`):

The text to split into chunks. Paste an article, document or transcript here. For multiple documents, use the list below.

## `texts` (type: `array`):

Optional — a list of separate texts/documents to chunk in one run. Each is chunked independently.

## `chunkSize` (type: `integer`):

Target maximum size of each chunk, in the unit chosen below.

## `chunkOverlap` (type: `integer`):

How much each chunk overlaps the previous one (same unit). Overlap preserves context across chunk boundaries — common for RAG.

## `unit` (type: `string`):

Measure chunk size in characters, words, or approximate tokens (~4 chars each).

## `strategy` (type: `string`):

Recursive keeps paragraphs/sentences/words intact where possible (recommended). Fixed cuts at exact sizes.

## `trim` (type: `boolean`):

Trim leading/trailing whitespace on each chunk.

## `maxItems` (type: `integer`):

Maximum number of chunks to output (0 = no limit).

## Actor input object example

```json
{
  "texts": [],
  "chunkSize": 1000,
  "chunkOverlap": 100,
  "unit": "characters",
  "strategy": "recursive",
  "trim": true,
  "maxItems": 0
}
```

# Actor output Schema

## `results` (type: `string`):

The results as dataset items.

# API

You can run this Actor programmatically using our API. Below are code examples in JavaScript, Python, and CLI, as well as the OpenAPI specification and MCP server setup.

## JavaScript example

```javascript
import { ApifyClient } from 'apify-client';

// Initialize the ApifyClient with your Apify API token
// Replace the '<YOUR_API_TOKEN>' with your token
const client = new ApifyClient({
    token: '<YOUR_API_TOKEN>',
});

// Prepare Actor input
const input = {
    "text": "",
    "texts": []
};

// Run the Actor and wait for it to finish
const run = await client.actor("hipersoft/text-splitter-chunker").call(input);

// Fetch and print Actor results from the run's dataset (if any)
console.log('Results from dataset');
console.log(`💾 Check your data here: https://console.apify.com/storage/datasets/${run.defaultDatasetId}`);
const { items } = await client.dataset(run.defaultDatasetId).listItems();
items.forEach((item) => {
    console.dir(item);
});

// 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/js/docs

```

## Python example

```python
from apify_client import ApifyClient

# Initialize the ApifyClient with your Apify API token
# Replace '<YOUR_API_TOKEN>' with your token.
client = ApifyClient("<YOUR_API_TOKEN>")

# Prepare the Actor input
run_input = {
    "text": "",
    "texts": [],
}

# Run the Actor and wait for it to finish
run = client.actor("hipersoft/text-splitter-chunker").call(run_input=run_input)

# Fetch and print Actor results from the run's dataset (if there are any)
print(f"💾 Check your data here: https://console.apify.com/storage/datasets/{run.default_dataset_id}")
for item in client.dataset(run.default_dataset_id).iterate_items():
    print(item)

# 📚 Want to learn more 📖? Go to → https://docs.apify.com/api/client/python/docs/quick-start

```

## CLI example

```bash
echo '{
  "text": "",
  "texts": []
}' |
apify call hipersoft/text-splitter-chunker --silent --output-dataset

```

## MCP server setup

```json
{
    "mcpServers": {
        "apify": {
            "type": "http",
            "url": "https://mcp.apify.com/?tools=fetch-actor-details,hipersoft/text-splitter-chunker"
        }
    }
}
```

The hosted server signs you in with OAuth on first connect, so no API token belongs in this config. Clients without OAuth support can send an `Authorization: Bearer <APIFY_API_TOKEN>` header instead, using a token from API & Integrations in Apify Console (https://console.apify.com/settings/integrations).

## OpenAPI specification

Download the OpenAPI definition: https://api.apify.com/v2/actors/38A8Nv0zQdcRQAvDE/builds/irsvhZCECGqgBw09Z/openapi.json
