RAG Text Chunker

Split text into token-sized chunks for RAG / embeddings prep. Multiple strategies: recursive char, sentence-aware, semantic boundaries. Configurable overlap. All in-browser.

When to use this

Retrieval-Augmented Generation (RAG) and embedding-based search both depend on splitting a corpus into chunks: small pieces that are individually embedded and stored in a vector database. The split happens before any of the AI machinery runs, but the quality of your retrieval quietly depends on it more than most people realise. Too-small chunks lose context; too-large chunks dilute relevance; chunks split mid-sentence retrieve poorly because the embedding lands in a weird semantic spot. This tool gives you a fast in-browser playground to experiment with chunk size, overlap, and strategy before you commit a pipeline to the choice.

The four strategies

Overlap — why and how much

Token estimation

Common gotchas

Worked example

Paste a long document, set chunk size to 512 tokens with 64 tokens of overlap, and pick the recursive strategy. The tool splits on the biggest natural boundary first — blank lines between paragraphs — then falls back to single newlines, then ". ", then spaces, packing pieces greedily until each chunk is just under 512 tokens. Every chunk after the first is prefixed with the last ~64 tokens of the previous one, so a sentence that straddles a boundary still appears whole in at least one chunk. You get per-chunk token/char counts and can export the set as JSON or JSONL for your vector store.

FAQ

How are tokens counted? With a fast heuristic, not a full tokenizer: roughly one token per 3.8 characters of Latin text, but one token per CJK character (Chinese/Japanese/Korean), which better matches how those scripts tokenize. It's an estimate for sizing, labelled "~tokens", not an exact BPE count.

What do the four strategies do? Recursive uses the separator hierarchy above; sentence packs whole sentences; paragraph packs whole paragraphs; semantic additionally breaks at heading-like lines (Markdown # headings or short ALL-CAPS lines) so sections stay together. All of them fall back to smaller splits when a single piece exceeds the size limit.

Why add overlap? Overlap keeps context from leaking away at chunk edges — a retrieved chunk carries a little of what came before, so an answer spanning a boundary isn't cut in half. The overlap is measured in tokens and taken from the tail of the preceding chunk.

Is my text uploaded? No. Chunking, counting and the JSON/JSONL export all happen in your browser — useful when the document is private or proprietary.