Zero-dependency, markdown- and code-aware text chunker for RAG and LLM context windows — with token budgets, overlap, and heading breadcrumbs.
Most RAG pipelines reach for LangChain or a Python toolchain just to split documents. chunklet is a single, dependency-free Node module (and CLI) that does the structural splitting that actually matters for retrieval quality:
- 🧱 Structure-aware — splits on markdown headings → paragraphs → sentences → words, in that order. Never mid-word unless it has to.
- 🔒 Code-block safe — fenced code blocks (
```/~~~) are kept intact and never split on inner blank lines or#lines. - 🧭 Heading breadcrumbs — every chunk carries its heading path (
Guide > Setup), so retrieval knows where each chunk came from. Optionally prepend it into the text for better embeddings. - 🎯 Token budgets + overlap — pack chunks up to a target token count, with configurable overlap for context continuity.
- 🔌 Bring-your-own tokenizer — accurate
~chars/4heuristic by default; plug intiktokenor any counter when you need exact counts. - 🪶 Zero dependencies, ESM, no build step. Works as a library or a CLI.
A naive splitter gives you a chunk that says "Install with npm." — but the retriever has no idea that line lives under Guide › Installation. chunklet attaches that breadcrumb to every chunk (and can prepend it to the embedded text), which measurably improves hybrid retrieval on technical docs.
npm install chunkletOr run the CLI without installing:
npx chunklet README.md -m 400 -o 40 -f jsonl > chunks.jsonlRequires Node.js >= 18.
import { chunk } from 'chunklet';
const text = `# Guide
## Installation
Install with npm, or run it with npx.
## Usage
Pipe markdown in and get JSON chunks out.`;
const chunks = chunk(text, { maxTokens: 400, overlap: 40 });
for (const c of chunks) {
console.log(c.index, c.headingPath.join(' > '), `(${c.tokens} tokens)`);
}Each chunk is a plain object:
{
index: 1,
text: '## Installation\n\nInstall with npm, or run it with npx.',
tokens: 23,
chars: 91,
start: 74, // char offset in the source document
end: 165,
headingPath: ['Guide', 'Installation'],
type: 'text' // 'text' | 'code' | 'mixed'
}const chunks = chunk(text, { maxTokens: 400, prependHeadings: true });
// chunk.text => "Guide > Installation\n\n## Installation\n\nInstall with npm..."import { encoding_for_model } from 'tiktoken';
const enc = encoding_for_model('gpt-4o');
const chunks = chunk(text, {
maxTokens: 512,
tokenizer: (t) => enc.encode(t).length,
});chunklet [file] [options]
cat doc.md | chunklet [options]
-m, --max-tokens <n> target max tokens per chunk (default 512)
-o, --overlap <n> tokens of trailing context to repeat (default 0)
-f, --format <fmt> json | jsonl | text (default json)
--chars-per-token <n> token estimation ratio (default 4)
--prepend-headings prepend the heading breadcrumb into chunk text
--no-split-headings do not start a new chunk at each heading
--no-split-oversize emit oversize blocks as-is instead of splitting
--stats print a summary to stderr
-h, --help show this help
-v, --version show version
Examples:
# JSONL, one chunk per line — ready to stream into an embedder
chunklet docs/guide.md -m 400 -o 40 -f jsonl > chunks.jsonl
# Human-readable preview with a stats summary
cat docs/*.md | chunklet --prepend-headings -f text --stats| Export | Description |
|---|---|
chunk(text, opts) |
Main chunker. Returns an array of chunk objects. |
segment(text) |
Parse markdown into structural blocks (heading/text/code) with breadcrumbs. |
splitText(text, maxTokens, opts) |
Recursive separator-aware splitter for oversize text. |
estimateTokens(text, opts) / makeCounter(opts) |
Token estimation helpers. |
formatChunks(chunks, format) |
Serialize chunks to json / jsonl / text. |
| Option | Default | Description |
|---|---|---|
maxTokens |
512 |
Target maximum tokens per chunk. |
overlap |
0 |
Tokens of trailing context repeated at the start of the next chunk. Must be < maxTokens. |
splitOnHeadings |
true |
Start a new chunk at each heading so every chunk stays within one section. |
splitOversize |
true |
Recursively split any single block larger than maxTokens. |
prependHeadings |
false |
Prepend the heading breadcrumb into each chunk's text. |
charsPerToken |
4 |
Ratio for the built-in token estimate. |
tokenizer |
— | Custom (text) => number counter; overrides charsPerToken. |
markdown ──► segment() ──► structural blocks (headings / paragraphs / code)
│ each tagged with its heading breadcrumb
▼
greedy packing into chunks ≤ maxTokens
│ (oversize blocks → splitText recursively)
▼
overlap + optional breadcrumb prefix ──► chunks[]
npm test # node --test, zero dependencieschunklet is free and open source. If it saves you some time, an optional tip is always welcome (never required):
- USDT — Ethereum (ERC-20) only:
0xad39bdf2df0b8dd6991150fcea0a156150ed19b8 - Verify on Etherscan
Please send only on the Ethereum (ERC-20) network. Thank you! 🙏
MIT © 2026 Ayubjon