RAG Text Chunker (Markdown to Chunks, Token Counts) avatar

RAG Text Chunker (Markdown to Chunks, Token Counts)

Pricing

from $0.12 / 1,000 chunks

Go to Apify Store
RAG Text Chunker (Markdown to Chunks, Token Counts)

RAG Text Chunker (Markdown to Chunks, Token Counts)

Split Markdown or text into RAG-ready chunks with exact OpenAI token counts (o200k/cl100k): heading-aware, code blocks kept whole, sentence-boundary overlap, heading path on every chunk. Chunks any dataset from Website to Markdown, PDF to Markdown or your own texts.

Pricing

from $0.12 / 1,000 chunks

Rating

0.0

(0)

Developer

Murat Uzun

Murat Uzun

Maintained by Community

Actor stats

0

Bookmarked

2

Total users

1

Monthly active users

20 hours ago

Last modified

Share

What is RAG Text Chunker?

RAG Text Chunker splits Markdown or plain text into retrieval-ready chunks for vector databases (Pinecone, Qdrant, Weaviate, pgvector, Chroma) and RAG apps, with exact OpenAI token counts (o200k_base for GPT-4o/GPT-5, or cl100k_base for GPT-4 and text-embedding-3). It is heading-aware: every chunk stays inside one section and carries its heading path (Guide > Install > Linux), code blocks are never cut mid-line, oversized paragraphs split at sentence boundaries, tiny fragments are folded into their neighbours, and inline base64 images are dropped. $0.20 per 1,000 chunks.

Feed it the dataset of a scraping run (Website to Markdown, llms.txt Generator, PDF to Markdown, YouTube transcripts, articles) or paste your own texts, and get one row per chunk ready to embed.

What does RAG Text Chunker output?

FieldDescription
documentId, url, titleThe source document (documentId is its URL, or document-N)
chunkIndex, chunkCount, chunkIdPosition of the chunk in the document; chunkId = documentId#index, a stable key for upserts
headingPath, headingsHeadings the chunk sits under, as an array and as A > B > C
textThe chunk text, with the heading path on top and the overlap from the previous chunk
tokens, overlapTokens, charactersExact token count with the chosen tokenizer, tokens of overlap, length
encodingo200k_base or cl100k_base
any Columns to keepFields copied from the source rows, e.g. lang, crawledAt, site

How to use RAG Text Chunker

  1. Give it text: an Input dataset ID from another run (the text column is found automatically), Documents as JSON, or Texts.
  2. Set Chunk size (default 500 tokens) and Overlap (default 50), and pick the Tokenizer your embedding model uses.
  3. Run, then embed the text column of the dataset and store it with chunkId, url and headings as metadata.

Chaining with a crawler

Run Website to Markdown or PDF to Markdown Converter, copy the run's dataset ID, and pass it as Input dataset ID. With Apify integrations or the API you can start the chunker automatically when the crawl finishes.

Example input

{
"inputDatasetId": "NJMrOuJvOh0X9IP4r",
"chunkSize": 400,
"overlap": 40,
"encoding": "o200k_base",
"keepFields": ["lang"]
}

Example output

{
"documentId": "https://crawlee.dev/js/docs/3.10/quick-start",
"url": "https://crawlee.dev/js/docs/3.10/quick-start",
"title": "Quick Start",
"chunkIndex": 4,
"chunkCount": 26,
"chunkId": "https://crawlee.dev/js/docs/3.10/quick-start#4",
"headingPath": [
"Quick Start",
"Crawling"
],
"headings": "Quick Start > Crawling",
"text": "Quick Start > Crawling\n\nYou need to explicitly install it with NPM. πŸ‘‡\n\n```\nnpm install crawlee puppeteer\n```\n\nRun the following example to perform a recursive crawl of the Crawlee website using the selected crawler.\n\nDon't forget about module imports\n\nTo run ...",
"tokens": 145,
"overlapTokens": 23,
"characters": 589,
"encoding": "o200k_base"
}

In a test on 120 pages of the Crawlee and Apify docs with 400-token chunks and 40-token overlap, the 898 chunks ranged from 23 to 463 tokens (content up to 400, plus heading path and overlap).

How much does it cost?

Pay per result: $0.0002 per chunk ($0.20 per 1,000), with volume discounts on paid Apify plans. A typical documentation page gives 5 to 10 chunks at 400 tokens. A dataset that cannot be read is listed in the ERRORS record and costs nothing. Set Maximum cost per run to cap spend.

Limits

  • Chunk size limits the content; the heading path and the overlap are added on top, so a chunk can be up to about chunk size + overlap + the heading path.
  • Chunking follows Markdown headings (# … ######). Plain text without headings is split by paragraphs and sentences.
  • Tokens are counted with OpenAI's tokenizers; other models (Claude, Gemini, open-source embeddings) count somewhat differently.
  • Datasets are read in batches of 1,000 rows; very large datasets take longer.

Using RAG Text Chunker with AI agents

The Actor is pay-per-event with limited permissions, so AI agents can call it through Apify's MCP server (mcp.apify.com), for example with {"texts": ["# Doc\n\nLong text..."], "chunkSize": 300}, or with inputDatasetId right after a crawl.

FAQ

Why heading-aware chunks? A chunk that mixes two sections embeds to a blurry vector and is hard to cite. Keeping one section per chunk and putting its heading path in the text makes retrieval and citations more precise.

Which chunk size should I use? 300 to 500 tokens with 10 % overlap is a common starting point for question answering; use larger chunks for summarisation.

Part of the webdatatools web-intelligence suite β€” every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:

Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.

Website & domain intelligence

Content for AI, LLMs and RAG

Search, video and social

Leads, jobs and company data

Developer, app and research data