RAG Text Chunker (Markdown to Chunks, Token Counts)
Pricing
from $0.12 / 1,000 chunks
RAG Text Chunker (Markdown to Chunks, Token Counts)
Split Markdown or text into RAG-ready chunks with exact OpenAI token counts (o200k/cl100k): heading-aware, code blocks kept whole, sentence-boundary overlap, heading path on every chunk. Chunks any dataset from Website to Markdown, PDF to Markdown or your own texts.
Pricing
from $0.12 / 1,000 chunks
Rating
0.0
(0)
Developer
Murat Uzun
Maintained by CommunityActor stats
0
Bookmarked
2
Total users
1
Monthly active users
20 hours ago
Last modified
Categories
Share
What is RAG Text Chunker?
RAG Text Chunker splits Markdown or plain text into retrieval-ready chunks for vector databases (Pinecone, Qdrant, Weaviate, pgvector, Chroma) and RAG apps, with exact OpenAI token counts (o200k_base for GPT-4o/GPT-5, or cl100k_base for GPT-4 and text-embedding-3). It is heading-aware: every chunk stays inside one section and carries its heading path (Guide > Install > Linux), code blocks are never cut mid-line, oversized paragraphs split at sentence boundaries, tiny fragments are folded into their neighbours, and inline base64 images are dropped. $0.20 per 1,000 chunks.
Feed it the dataset of a scraping run (Website to Markdown, llms.txt Generator, PDF to Markdown, YouTube transcripts, articles) or paste your own texts, and get one row per chunk ready to embed.
What does RAG Text Chunker output?
| Field | Description |
|---|---|
documentId, url, title | The source document (documentId is its URL, or document-N) |
chunkIndex, chunkCount, chunkId | Position of the chunk in the document; chunkId = documentId#index, a stable key for upserts |
headingPath, headings | Headings the chunk sits under, as an array and as A > B > C |
text | The chunk text, with the heading path on top and the overlap from the previous chunk |
tokens, overlapTokens, characters | Exact token count with the chosen tokenizer, tokens of overlap, length |
encoding | o200k_base or cl100k_base |
| any Columns to keep | Fields copied from the source rows, e.g. lang, crawledAt, site |
How to use RAG Text Chunker
- Give it text: an Input dataset ID from another run (the text column is found automatically), Documents as JSON, or Texts.
- Set Chunk size (default 500 tokens) and Overlap (default 50), and pick the Tokenizer your embedding model uses.
- Run, then embed the
textcolumn of the dataset and store it withchunkId,urlandheadingsas metadata.
Chaining with a crawler
Run Website to Markdown or PDF to Markdown Converter, copy the run's dataset ID, and pass it as Input dataset ID. With Apify integrations or the API you can start the chunker automatically when the crawl finishes.
Example input
{"inputDatasetId": "NJMrOuJvOh0X9IP4r","chunkSize": 400,"overlap": 40,"encoding": "o200k_base","keepFields": ["lang"]}
Example output
{"documentId": "https://crawlee.dev/js/docs/3.10/quick-start","url": "https://crawlee.dev/js/docs/3.10/quick-start","title": "Quick Start","chunkIndex": 4,"chunkCount": 26,"chunkId": "https://crawlee.dev/js/docs/3.10/quick-start#4","headingPath": ["Quick Start","Crawling"],"headings": "Quick Start > Crawling","text": "Quick Start > Crawling\n\nYou need to explicitly install it with NPM. π\n\n```\nnpm install crawlee puppeteer\n```\n\nRun the following example to perform a recursive crawl of the Crawlee website using the selected crawler.\n\nDon't forget about module imports\n\nTo run ...","tokens": 145,"overlapTokens": 23,"characters": 589,"encoding": "o200k_base"}
In a test on 120 pages of the Crawlee and Apify docs with 400-token chunks and 40-token overlap, the 898 chunks ranged from 23 to 463 tokens (content up to 400, plus heading path and overlap).
How much does it cost?
Pay per result: $0.0002 per chunk ($0.20 per 1,000), with volume discounts on paid Apify plans. A typical documentation page gives 5 to 10 chunks at 400 tokens. A dataset that cannot be read is listed in the ERRORS record and costs nothing. Set Maximum cost per run to cap spend.
Limits
- Chunk size limits the content; the heading path and the overlap are added on top, so a chunk can be up to about chunk size + overlap + the heading path.
- Chunking follows Markdown headings (
#β¦######). Plain text without headings is split by paragraphs and sentences. - Tokens are counted with OpenAI's tokenizers; other models (Claude, Gemini, open-source embeddings) count somewhat differently.
- Datasets are read in batches of 1,000 rows; very large datasets take longer.
Using RAG Text Chunker with AI agents
The Actor is pay-per-event with limited permissions, so AI agents can call it through Apify's MCP server (mcp.apify.com), for example with {"texts": ["# Doc\n\nLong text..."], "chunkSize": 300}, or with inputDatasetId right after a crawl.
FAQ
Why heading-aware chunks? A chunk that mixes two sections embeds to a blurry vector and is hard to cite. Keeping one section per chunk and putting its heading path in the text makes retrieval and citations more precise.
Which chunk size should I use? 300 to 500 tokens with 10 % overlap is a common starting point for question answering; use larger chunks for summarisation.
Related Actors
Part of the webdatatools web-intelligence suite β every Actor is pay-per-event, reads public data without a login, and returns one clean row per entity:
Browse the whole suite at webdatatools, or call ten of these Actors straight from Claude, Cursor or Cline with the webdatatools MCP server.
Website & domain intelligence
- Email Extractor β Website Contact & Social Finder β e-mails, phones and social profiles per domain
- Tech Stack Detector β Wappalyzer & BuiltWith Alternative β CMS, e-commerce, analytics, pixels and payments per domain
- Domain DNS & Email Security Checker β SPF, DKIM, DMARC, MX provider, registrar and domain age
- Domain Security Audit (TLS, HTTP headers, redirects, robots) β TLS expiry, security headers, redirect chain, robots and llms.txt
- Subdomain Finder (Certificate Transparency) β every subdomain seen in CT logs, with a live DNS check
- Bulk Core Web Vitals & PageSpeed Audit β Lighthouse scores, LCP, CLS, INP and top fixes per URL
- On-Page SEO Audit β title, meta, headings, links, images and schema issues per page
- Sitemap URL Extractor & Change Monitor β every sitemap URL, or new and removed pages between runs
- Wayback Machine Snapshot & Page Change Tracker β how a page changed over time, or every archived snapshot
- Bulk Domain WHOIS & RDAP Lookup β registrar, dates, status and nameservers per domain
- Web Scraper β CSS Selector & Data Extractor β pull any CSS selector off any page, one row per URL
- Website Screenshot Generator β full-page or viewport PNG/JPEG screenshots of any URL
Content for AI, LLMs and RAG
- AI Web Search & Read: Google results as clean Markdown β a query turned into clean Markdown from the top search results
- Website to Markdown β Content Crawler for LLM & RAG β any site as clean Markdown per page, no browser
- llms.txt Generator (Website to llms.txt & llms-full.txt) β llms.txt and llms-full.txt for any website or docs site
- PDF to Markdown Converter (Text, Headings, Metadata) β PDF files to clean Markdown with headings, metadata and token counts
- Article & News Extractor (clean text, author, date, markdown) β clean article text, author, date and Markdown per URL
- Structured Data & JSON-LD Extractor (Schema.org, Open Graph) β Schema.org and Open Graph data from any page
- Google News Scraper (RSS search by keyword, topic, site) β news results by keyword, topic or site
- Press Release Monitor: PR Newswire, BusinessWire, GlobeNewswire β PR Newswire, Business Wire and GlobeNewswire releases
Search, video and social
- YouTube Shorts Scraper β Shorts from channels, hashtags and searches with view counts
- Pinterest Pins Scraper β latest pins of public Pinterest profiles and boards
- YouTube Transcript Scraper β captions and subtitles as text + timed segments, per video or channel
- Google Search Results Scraper β SERP API β organic SERP results per keyword and country
- YouTube Comments Scraper β Comments & Replies β comments and replies with likes, no API key
- YouTube Channel Latest Videos (RSS, no API key) β the latest 15 videos of any channel from RSS
- YouTube Channel Scraper (videos, shorts, live) β a channel's full video, shorts and stream list
- YouTube Search Results Scraper (videos, channels, no API key) β videos, channels and playlists per query
- YouTube Video Details Scraper (views, likes, description, tags) β views, likes, description, tags and chapters per video
- Apple Podcasts Lookup & Episodes Scraper β podcast metadata and episodes from iTunes and RSS
- Bluesky Post, Search & Profile Scraper β posts, profiles, followers and threads from the AT Protocol API
- X Tweet Scraper (Twitter Posts by URL, No Login) β X (Twitter) posts by URL with likes, replies, author, media and MP4 links
- Telegram Channel Posts Scraper β posts, views and media flags from any public channel
- Substack Publication & Posts Scraper β archive, authors and paywall status per publication
- Google Play Reviews Scraper β reviews, ratings, replies and app versions per app
- App Store Reviews Scraper β iOS reviews and ratings per app and country
- Trustpilot Reviews Scraper (Ratings, Replies, No Login) β Trustpilot reviews with stars, full text, verification and company replies
- Google Trends Scraper β interest over time, by region, and related queries per keyword
- Google Ads Transparency Scraper β ads any advertiser runs on Google, with format and dates
- Keyword Suggestions Scraper (Google, YouTube, Amazon, Bing) β autocomplete keyword ideas from four search engines
- Bilibili Scraper (Videos, Search, Popular) β Chinese video platform: views, likes, coins, danmaku, uploader
- Mastodon Scraper (Hashtags, Accounts, Trending) β public fediverse posts by hashtag, account or trending
- Meetup Events Scraper (Search by Keyword & City) β upcoming events with RSVPs, fees, venues and groups
- Eventbrite Scraper (Events by Keyword & City) β events by keyword and city with venue, dates and organizer
Leads, jobs and company data
- Career Site Jobs API (Greenhouse, Lever, Ashby, Workday +1) β company domains in, their open jobs out, ATS detected automatically
- Workday Jobs Scraper β jobs with full descriptions from any Workday career site
- Google Maps Scraper β businesses with phone, website, address, rating and coordinates per search
- LinkedIn Jobs Scraper β job titles, companies, locations and full descriptions from LinkedIn job search
- Company 360: full company profile from a domain β one row per domain: contacts, tech, security, hiring and company facts
- Hiring Signals Scraper (Greenhouse, Lever, Ashby, Workable) β open jobs and hiring velocity from 10 public ATS boards
- Y Combinator Companies & Founders Scraper β YC startups by batch, industry and hiring status
- Wikidata Entity & Company Enrichment (facts, IDs, links) β HQ, founders, employees, revenue and social IDs per company
- Email Validator & Verifier β Bulk Email Check β syntax, MX, disposable, role and free-provider checks
- OpenStreetMap POI Extractor (Overpass API: shops, amenities) β shops and amenities by radius, bbox or area
- Stock, Crypto & FX Quotes β one row per symbol from Yahoo, Binance and ECB rates
- Remote Jobs Aggregator (RemoteOK, WWR, Hacker News) β one clean row per remote job, de-duplicated across feeds
- Greenhouse Jobs Scraper β jobs with descriptions from any Greenhouse job board
- Lever Jobs Scraper β jobs with descriptions from any Lever careers page
- Ashby Jobs Scraper β jobs, salaries and descriptions from any Ashby job board
- SmartRecruiters Jobs Scraper β jobs with descriptions from any SmartRecruiters company
- Seek Jobs Scraper (Australia & New Zealand) β Seek job ads with salary, work type and location
- Dice Jobs Scraper β US tech jobs from Dice with salary and remote flag
- AutoScout24 Scraper β European car listings with price, mileage and seller
- Rightmove Scraper β UK property for sale or rent with price and agent
- Wellfound Jobs Scraper (AngelList Startup Jobs) β startup jobs with salary and equity ranges, company size and stage
- Yandex Maps Scraper (Places, Ratings, Phones) β businesses in Russia, TΓΌrkiye and the CIS with phones, ratings, hours
- Craigslist Scraper (Listings, Prices, Locations) β listings in any area and category with price, date and coordinates
- JobStreet Scraper (Malaysia, Singapore, PH, ID + JobsDB) β JobStreet and JobsDB jobs in 6 Asian countries with parsed salaries
- InfoJobs Scraper (Spain Jobs, Salaries, Companies) β Spanish jobs with salary range, contract type and full description
- Redfin Scraper (Homes for Sale, Prices, Details) β US homes for sale or sold from any Redfin search, with price and details
- Kleinanzeigen Scraper (Ads, Prices, Locations) β German classifieds with price, VB flag, ZIP, city and seller type
Developer, app and research data
- npm, PyPI & Crates.io Package Health Checker β releases, downloads, deprecation and a health score
- GitHub Repository Health & Activity Report β stars, commits, contributors and risk flags per repo
- VS Code Marketplace Extension Scraper (installs, ratings) β installs, ratings and versions per extension
- Chrome Web Store Extension Scraper (installs, ratings) β users, rating, version and developer per extension
- Google Play Scraper β apps, ratings, installs, developer contact and reviews
- App Store (iOS) App Metadata, Ratings & Top Charts Lookup β ratings, price, version and charts per app
- CrossRef DOI & Citation Metadata Lookup β papers, authors, journals and citation counts
- FDA Recalls & Adverse Events Monitor (openFDA) β food, drug and device recalls from openFDA
- iCal / ICS Calendar Feed to Events Extractor β any public calendar feed as event rows
- Shopify Store Products Scraper β catalog, prices, variants and stock per store
- Hacker News Search & Front Page Scraper β stories, comments and points by query or front page
- GitHub Trending Repositories Scraper β trending repos and developers by language and period
- Stack Overflow & Stack Exchange Q&A Scraper β questions, answers and scores by query, tag or site
- Bulk Image Downloader β download image URLs to storage with size, dimensions and a ZIP
- Google Flights Scraper (Prices, Airlines, Stops) β flight prices, airlines, times, stops and CO2 by route and date
- Google Hotels Scraper (Prices, Ratings, Reviews) β hotel prices per night, stars, rating and reviews by city and dates
- Booking.com Scraper (Hotel Prices, Ratings, Availability) β Booking.com hotels for any city and dates with prices, scores and deals
- AliExpress Scraper (Search Products & Prices) β AliExpress search results with USD price, discount and rank
- Lazada Scraper (Products, Prices, Sold, Ratings) β Lazada products in 6 countries with price, rating, units sold and seller
- Trendyol Scraper (Turkey Products, Prices, Ratings) β Trendyol products with price, basket discount, rating and promotions