Audit any text splitter for the Unicode corruption your tests don't catch.
Your RAG pipeline chunks text. If it cuts on a code-point index — and almost all of them do — it will eventually cut through the middle of a family emoji, a Devanagari conjunct, or a Hangul syllable. Nothing throws. The chunk still round-trips. It just embeds as garbage, and you find out when retrieval quality drops for one language.
splitlint builds text where the hazard sits exactly on the cut, then checks the invariants a splitter must not break.
$ splitlint langchain_text_splitters:RecursiveCharacterTextSplitter
splitlint · Unicode 17.0.0
splitter: langchain_text_splitters:RecursiveCharacterTextSplitter chunk_size=50
✗ FAIL zwj-family grapheme-integrity
mangling chunk 0: cut at offset 50 splits a grapheme cluster: U+1F469 | U+200D
✗ FAIL combining-marks grapheme-integrity
mangling chunk 0: cut at offset 50 splits a grapheme cluster: U+0065 | U+0301
✗ FAIL devanagari-conjunct grapheme-integrity
mangling chunk 0: cut at offset 50 splits a grapheme cluster: U+094D | U+0937
✗ FAIL hangul-jamo grapheme-integrity
mangling chunk 0: cut at offset 50 splits a grapheme cluster: U+1100 | U+1161
✓ pass crlf 3 chunks
✓ pass astral-pair 3 chunks
...
10 of 12 cases failed. Invariants violated: grapheme-integritypip install splitlint# audit anything importable: module:Attribute
splitlint langchain_text_splitters:RecursiveCharacterTextSplitter
splitlint semantic_text_splitter:TextSplitter
splitlint my_app.rag:split_document --chunk-size 512
splitlint --list-cases # what it tests, and why
splitlint <spec> --json # machine-readable, for CIExit codes:
| Code | Means |
|---|---|
0 |
At least one case was exercised and every invariant held |
1 |
An invariant was violated |
2 |
Nothing could be audited — bad arguments, the splitter wouldn't load, or it never placed a boundary |
Exit 0 is unreachable without a real cut having been made. A splitter that
returns the text unsplit exits 2, not 0, so CI can't go green on a splitter
that was never actually tested.
As a library:
from splitlint import check_all, graphemes
violations = check_all(text, my_splitter.split_text(text))
for v in violations:
print(v.severity, v.invariant, v.detail)| Invariant | Severity | Means |
|---|---|---|
encoding-integrity |
corruption | A chunk holds an unpaired surrogate and cannot be encoded to UTF-8 |
content-preservation |
corruption | Input characters silently vanished |
grapheme-integrity |
mangling | A cut landed inside a grapheme cluster |
no-empty-chunks |
contract | An empty chunk was emitted — embedding APIs reject these |
roundtrip-identity |
contract | --strict: rejoined chunks must equal the input exactly |
size-bound |
contract | --max-size N: no chunk may exceed N characters |
Severity is not decoration. Corruption means data is destroyed. Mangling means the bytes survive but the boundaries are wrong, so the text renders and embeds as nonsense. Contract means a downstream promise was broken.
A family emoji sitting comfortably in the middle of a chunk proves nothing. Each case pads with separator-free filler until the hazard straddles the cut offset — which also forces recursive splitters past their paragraph, line and space separators down to the character splitter, where the boundary bugs actually live.
Twelve cases: ZWJ families, combining marks, regional-indicator flags, skin-tone
modifiers, keycaps, TAG-sequence subdivision flags, Devanagari conjuncts, Hangul jamo,
variation selectors, CRLF, astral pairs, and ZWNJ sequences. Run --list-cases.
Grapheme segmentation is a full implementation of UAX #29 extended grapheme clusters — rules GB1–GB999, including GB9c for Indic conjuncts and GB11 for emoji ZWJ sequences.
It is verified against the official Unicode GraphemeBreakTest.txt: all 766 cases
pass. The property tables are generated from the published UCD by
tools/gen_gbp.py rather than hand-copied, so a new Unicode
version is a regeneration, not a rewrite.
$ python -m pytest tests/test_grapheme_conformance.py
2 passed- A pass is not a proof. splitlint tests the boundary offsets it constructs, not every possible input. It finds the bug class it was built for.
- Inconclusive is reported, not hidden. A splitter that returns one chunk never
placed a boundary, so nothing was tested.
CharacterTextSplitterdoes this whenever its separator is absent. splitlint says "Nothing was tested" and exits2. - Normalization is not destruction. A splitter that emits NFC-normalized text is
reported as
normalization/ mangling, not as data loss — the code points differ but nothing was destroyed. - Plain-ASCII terminals are supported. On a console that can't encode
✓, the markers degrade toPASS/FAILrather than raisingUnicodeEncodeError. grapheme-integrityis a quality bar, not a correctness law. A splitter forced into a tiny chunk size may have no legal cut available. It is gradedmangling, notcorruption, for that reason.content-preservationignores whitespace by default, because most splitters strip legitimately. Use--strictwhen you expect losslessness.- Overlapping chunks are handled: overlap is never reported as data loss.
The bug is real and it ships. I found and fixed an instance of it upstream in langchain4j (#6308): a splitter using UTF-16 code units cut surrogate pairs in half, producing chunks that could not be encoded to UTF-8 at all. It was silent with a character limit and fatal with a token limit. splitlint is that investigation turned into something you can run in ten seconds.
New corpus cases are the most valuable contribution — a hazard splitlint doesn't yet
construct is a bug it can't find. A case needs a name, a one-line rationale, and the
hazard string; the harness does the rest. Run pytest before opening a PR.
MIT