Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

chunklet

Zero-dependency, markdown- and code-aware text chunker for RAG and LLM context windows — with token budgets, overlap, and heading breadcrumbs.

chunklet demo

Most RAG pipelines reach for LangChain or a Python toolchain just to split documents. chunklet is a single, dependency-free Node module (and CLI) that does the structural splitting that actually matters for retrieval quality:

  • 🧱 Structure-aware — splits on markdown headings → paragraphs → sentences → words, in that order. Never mid-word unless it has to.
  • 🔒 Code-block safe — fenced code blocks (``` / ~~~) are kept intact and never split on inner blank lines or # lines.
  • 🧭 Heading breadcrumbs — every chunk carries its heading path (Guide > Setup), so retrieval knows where each chunk came from. Optionally prepend it into the text for better embeddings.
  • 🎯 Token budgets + overlap — pack chunks up to a target token count, with configurable overlap for context continuity.
  • 🔌 Bring-your-own tokenizer — accurate ~chars/4 heuristic by default; plug in tiktoken or any counter when you need exact counts.
  • 🪶 Zero dependencies, ESM, no build step. Works as a library or a CLI.

Why heading breadcrumbs matter

A naive splitter gives you a chunk that says "Install with npm." — but the retriever has no idea that line lives under Guide › Installation. chunklet attaches that breadcrumb to every chunk (and can prepend it to the embedded text), which measurably improves hybrid retrieval on technical docs.

Install

npm install chunklet

Or run the CLI without installing:

npx chunklet README.md -m 400 -o 40 -f jsonl > chunks.jsonl

Requires Node.js >= 18.

Library usage

import { chunk } from 'chunklet';

const text = `# Guide

## Installation

Install with npm, or run it with npx.

## Usage

Pipe markdown in and get JSON chunks out.`;

const chunks = chunk(text, { maxTokens: 400, overlap: 40 });

for (const c of chunks) {
  console.log(c.index, c.headingPath.join(' > '), `(${c.tokens} tokens)`);
}

Each chunk is a plain object:

{
  index: 1,
  text: '## Installation\n\nInstall with npm, or run it with npx.',
  tokens: 23,
  chars: 91,
  start: 74,            // char offset in the source document
  end: 165,
  headingPath: ['Guide', 'Installation'],
  type: 'text'          // 'text' | 'code' | 'mixed'
}

Prepend breadcrumbs for embedding

const chunks = chunk(text, { maxTokens: 400, prependHeadings: true });
// chunk.text => "Guide > Installation\n\n## Installation\n\nInstall with npm..."

Bring your own tokenizer

import { encoding_for_model } from 'tiktoken';
const enc = encoding_for_model('gpt-4o');

const chunks = chunk(text, {
  maxTokens: 512,
  tokenizer: (t) => enc.encode(t).length,
});

CLI

chunklet [file] [options]
cat doc.md | chunklet [options]

  -m, --max-tokens <n>      target max tokens per chunk (default 512)
  -o, --overlap <n>         tokens of trailing context to repeat (default 0)
  -f, --format <fmt>        json | jsonl | text (default json)
      --chars-per-token <n> token estimation ratio (default 4)
      --prepend-headings    prepend the heading breadcrumb into chunk text
      --no-split-headings   do not start a new chunk at each heading
      --no-split-oversize   emit oversize blocks as-is instead of splitting
      --stats               print a summary to stderr
  -h, --help                show this help
  -v, --version             show version

Examples:

# JSONL, one chunk per line — ready to stream into an embedder
chunklet docs/guide.md -m 400 -o 40 -f jsonl > chunks.jsonl

# Human-readable preview with a stats summary
cat docs/*.md | chunklet --prepend-headings -f text --stats

API

Export Description
chunk(text, opts) Main chunker. Returns an array of chunk objects.
segment(text) Parse markdown into structural blocks (heading/text/code) with breadcrumbs.
splitText(text, maxTokens, opts) Recursive separator-aware splitter for oversize text.
estimateTokens(text, opts) / makeCounter(opts) Token estimation helpers.
formatChunks(chunks, format) Serialize chunks to json / jsonl / text.

chunk options

Option Default Description
maxTokens 512 Target maximum tokens per chunk.
overlap 0 Tokens of trailing context repeated at the start of the next chunk. Must be < maxTokens.
splitOnHeadings true Start a new chunk at each heading so every chunk stays within one section.
splitOversize true Recursively split any single block larger than maxTokens.
prependHeadings false Prepend the heading breadcrumb into each chunk's text.
charsPerToken 4 Ratio for the built-in token estimate.
tokenizer — Custom (text) => number counter; overrides charsPerToken.

How it works

markdown ──► segment() ──► structural blocks (headings / paragraphs / code)
                               │  each tagged with its heading breadcrumb
                               ▼
                          greedy packing into chunks ≤ maxTokens
                               │  (oversize blocks → splitText recursively)
                               ▼
                          overlap + optional breadcrumb prefix ──► chunks[]

Testing

npm test   # node --test, zero dependencies

Support

chunklet is free and open source. If it saves you some time, an optional tip is always welcome (never required):

  • USDT — Ethereum (ERC-20) only: 0xad39bdf2df0b8dd6991150fcea0a156150ed19b8
  • Verify on Etherscan

Please send only on the Ethereum (ERC-20) network. Thank you! 🙏

License

MIT © 2026 Ayubjon

About

Zero-dependency, markdown- and code-aware text chunker for RAG and LLM context windows: token budgets, overlap, and heading breadcrumbs.

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages