Skip to content
r0h1tbPublic

About

Audit any text splitter for the Unicode corruption your tests don't catch. Full UAX #29 implementation, verified against the official Unicode conformance suite.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

splitlint

Audit any text splitter for the Unicode corruption your tests don't catch.

CI Unicode 17.0.0 License: MIT

Your RAG pipeline chunks text. If it cuts on a code-point index — and almost all of them do — it will eventually cut through the middle of a family emoji, a Devanagari conjunct, or a Hangul syllable. Nothing throws. The chunk still round-trips. It just embeds as garbage, and you find out when retrieval quality drops for one language.

splitlint builds text where the hazard sits exactly on the cut, then checks the invariants a splitter must not break.

$ splitlint langchain_text_splitters:RecursiveCharacterTextSplitter

splitlint · Unicode 17.0.0
splitter: langchain_text_splitters:RecursiveCharacterTextSplitter   chunk_size=50

  ✗ FAIL  zwj-family           grapheme-integrity
      mangling  chunk 0: cut at offset 50 splits a grapheme cluster: U+1F469 | U+200D
  ✗ FAIL  combining-marks      grapheme-integrity
      mangling  chunk 0: cut at offset 50 splits a grapheme cluster: U+0065 | U+0301
  ✗ FAIL  devanagari-conjunct  grapheme-integrity
      mangling  chunk 0: cut at offset 50 splits a grapheme cluster: U+094D | U+0937
  ✗ FAIL  hangul-jamo          grapheme-integrity
      mangling  chunk 0: cut at offset 50 splits a grapheme cluster: U+1100 | U+1161
  ✓ pass  crlf                 3 chunks
  ✓ pass  astral-pair          3 chunks
  ...

10 of 12 cases failed.  Invariants violated: grapheme-integrity

Install

pip install splitlint

Use it

# audit anything importable: module:Attribute
splitlint langchain_text_splitters:RecursiveCharacterTextSplitter
splitlint semantic_text_splitter:TextSplitter
splitlint my_app.rag:split_document --chunk-size 512

splitlint --list-cases          # what it tests, and why
splitlint <spec> --json         # machine-readable, for CI

Exit codes:

Code Means
0 At least one case was exercised and every invariant held
1 An invariant was violated
2 Nothing could be audited — bad arguments, the splitter wouldn't load, or it never placed a boundary

Exit 0 is unreachable without a real cut having been made. A splitter that returns the text unsplit exits 2, not 0, so CI can't go green on a splitter that was never actually tested.

As a library:

from splitlint import check_all, graphemes

violations = check_all(text, my_splitter.split_text(text))
for v in violations:
    print(v.severity, v.invariant, v.detail)

What it checks

Invariant Severity Means
encoding-integrity corruption A chunk holds an unpaired surrogate and cannot be encoded to UTF-8
content-preservation corruption Input characters silently vanished
grapheme-integrity mangling A cut landed inside a grapheme cluster
no-empty-chunks contract An empty chunk was emitted — embedding APIs reject these
roundtrip-identity contract --strict: rejoined chunks must equal the input exactly
size-bound contract --max-size N: no chunk may exceed N characters

Severity is not decoration. Corruption means data is destroyed. Mangling means the bytes survive but the boundaries are wrong, so the text renders and embeds as nonsense. Contract means a downstream promise was broken.

The design idea

A family emoji sitting comfortably in the middle of a chunk proves nothing. Each case pads with separator-free filler until the hazard straddles the cut offset — which also forces recursive splitters past their paragraph, line and space separators down to the character splitter, where the boundary bugs actually live.

Twelve cases: ZWJ families, combining marks, regional-indicator flags, skin-tone modifiers, keycaps, TAG-sequence subdivision flags, Devanagari conjuncts, Hangul jamo, variation selectors, CRLF, astral pairs, and ZWNJ sequences. Run --list-cases.

Correctness

Grapheme segmentation is a full implementation of UAX #29 extended grapheme clusters — rules GB1–GB999, including GB9c for Indic conjuncts and GB11 for emoji ZWJ sequences.

It is verified against the official Unicode GraphemeBreakTest.txt: all 766 cases pass. The property tables are generated from the published UCD by tools/gen_gbp.py rather than hand-copied, so a new Unicode version is a regeneration, not a rewrite.

$ python -m pytest tests/test_grapheme_conformance.py
2 passed

Honest limitations

  • A pass is not a proof. splitlint tests the boundary offsets it constructs, not every possible input. It finds the bug class it was built for.
  • Inconclusive is reported, not hidden. A splitter that returns one chunk never placed a boundary, so nothing was tested. CharacterTextSplitter does this whenever its separator is absent. splitlint says "Nothing was tested" and exits 2.
  • Normalization is not destruction. A splitter that emits NFC-normalized text is reported as normalization / mangling, not as data loss — the code points differ but nothing was destroyed.
  • Plain-ASCII terminals are supported. On a console that can't encode ✓, the markers degrade to PASS/FAIL rather than raising UnicodeEncodeError.
  • grapheme-integrity is a quality bar, not a correctness law. A splitter forced into a tiny chunk size may have no legal cut available. It is graded mangling, not corruption, for that reason.
  • content-preservation ignores whitespace by default, because most splitters strip legitimately. Use --strict when you expect losslessness.
  • Overlapping chunks are handled: overlap is never reported as data loss.

Why this exists

The bug is real and it ships. I found and fixed an instance of it upstream in langchain4j (#6308): a splitter using UTF-16 code units cut surrogate pairs in half, producing chunks that could not be encoded to UTF-8 at all. It was silent with a character limit and fatal with a token limit. splitlint is that investigation turned into something you can run in ten seconds.

Contributing

New corpus cases are the most valuable contribution — a hazard splitlint doesn't yet construct is a bug it can't find. A case needs a name, a one-line rationale, and the hazard string; the harness does the rest. Run pytest before opening a PR.

License

MIT

About

Audit any text splitter for the Unicode corruption your tests don't catch. Full UAX #29 implementation, verified against the official Unicode conformance suite.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages