Repository navigation
fix: match string literals by RDF term identity; normalize language-tag case - #1676
Conversation
…ctionary key
`"bob"`, `"bob"@en` and `"bob"@fr` share one string-dictionary key. SPARQL
lowering attached no datatype/language constraint to ordinary triple objects
(only annotation-block objects carried one), so a constant object reached the
scan as a bare string, took the untyped-string branch, and matched all three.
`VALUES ?o { "bob"@en }` additionally projected "bob"@en for the false
matches, and per-row join/OPTIONAL probes rebuilt the object from a binding
the same way, so joins on a string variable equated every tag.
- SPARQL lowering: plain, language-tagged and explicitly typed string
literals carry their term constraint (`lower_object_with_term_constraint`)
on ordinary triples and property-path endpoints. Bare numerics, booleans
and dates stay unconstrained, so `25` still matches any integer subtype.
- Join / OPTIONAL probe substitution carries the binding's string
constraint; late-materialized (`EncodedLit`) bindings recover theirs via
`BinaryIndexStore::o_type_from_kind`.
- The overlay-only fallback and the raw untranslated-flake filter check the
novelty flake's language tag against the pattern's.
JSON-LD `@language` / `@type` value objects were already exact; a plain JSON
string keeps its lenient cross-tag matching (pinned by test).
W3C: basic#quotes-3/4, open-world#open-eq-02, expr-equals#eq-graph-4,
entailment#lang and entailment#plainLit now pass and leave the skip register.
…ive) RDF 1.1 defines the value space of language tags as lowercase and compares them case-insensitively, so `"x"@EN` and `"x"@en` are one term. We stored the tag as written and looked it up exactly, so the two interned as different lang_ids and `?x :p "string"@EN` could not find `"string"@en`. - `fluree_db_core::normalize_lang_tag` is the single canonical form. `FlakeMeta::with_lang` / new `FlakeMeta::from_parts` apply it, and the three transact sinks build metadata through `from_parts`. - Both `LanguageTagDict`s intern and look up the normalized form; a dictionary persisted before normalization still answers via a case-insensitive fallback scan. - SPARQL and JSON-LD constraints (`"x"@EN`, `@language: "En"`, VALUES cells) normalize before matching; the novelty-side filters compare case-insensitively. - `LANG()` reports the canonical lowercase form. W3C expr-builtin#dawg-lang-3 now passes and leaves the skip register.
aaj3f
left a comment
There was a problem hiding this comment.
@bplatz good catch and seems like the right fix.
My only question is for pre-existing ledgers/data-state involving language tags. It seems the store keeps a verbatim copy of the persisted tags alongside the now-normalized dictionary, and the per-row compare in matches_datatype_constraint reads the verbatim copy, so a ledger indexed before this PR with "chat"@en-US answers ?s ex:name "chat"@en-US with [] at HEAD (I seeded one at commit 1 and re-opened it at HEAD — en-US, en-us and fr-CA all return nothing while FILTER(LANG(?n)="en-us") still finds the rows), and a reindex makes it worse by splitting the tag into two ids and shifting every later lang_id on the read side. Details and two fix options are on global_dict.rs:176.
The second item is the pair of per-row allocations in the EncodedLit probe arm (join.rs:992), which is a few-line hoist. The rest of my inline notes below are docs accuracy (typed non-string literals are now exact too — a good change, just undocumented) and two questions.
Adherence to repo commitments:
- Patterns/abstractions: ✔ Extends
TriplePattern.dtc/DatatypeConstraintand the existing scan checks; one shared lowering helper for triples and path endpoints; no parallel construct.⚠️ Case normalization is applied at ingest and in the dictionaries but not to the store's verbatim tag table or the read-side compares, which is what breaks existing data. - Performance (speed first, memory second):
⚠️ Hot path touched (bind-join probe, novelty filters).Litarms and novelty filters are allocation-free; theEncodedLitarm adds two heap allocations per probe row (registryVec+Sid/Arc<str>) on the literal-object join lane — small, hoistable, should land before merge. No bench covers that lane;bench-compare's advisory numbers are the same ambient noise as on #1667/#1668. - Testing: ✔ New
[[test]]target runs in CI (three tests by name), lowering + dictionary unit tests, W3C 36/36 with seven register removals. ✖ Nothing exercises an index written before this change, which is the case that regresses. - Conventions: ✔ Two well-factored commits with thorough multi-line bodies; fmt clean; clippy
-D warningsclean on all six touched crates; docs updated.⚠️ Body/docs under-state the behaviour change (typed literals) and say nothing about existing ledgers.
Verified locally at branch HEAD (1b0254f): cargo fmt --all -- --check clean; cargo clippy -p fluree-db-query -p fluree-db-core -p fluree-db-binary-index -p fluree-db-sparql -p fluree-db-transact -p fluree-db-indexer --all-targets -- -D warnings clean; cargo nextest run -p fluree-db-api --test it_literal_identity --test it_fastpath_1652_regression 4/4; cargo nextest run -p fluree-db-query 1467 passed; three mutation checks red as expected; legacy-ledger probe (seed at aaaa7d4, query at HEAD) red as described above.
| /// Case-insensitive reverse lookup: the dictionary stores normalized | ||
| /// (lowercase) tags, but a dictionary persisted before normalization may | ||
| /// still hold the tag as written, so an exact miss falls back to a scan. | ||
| fn find_normalized(&self, tag: &str) -> Option<u16> { |
There was a problem hiding this comment.
blocking. Existing ledgers with any non-lowercase language tag stop matching tag-constrained patterns after this upgrade, before and after reindex. The real location is fluree-db-binary-index/src/read/binary_index_store.rs:386 / :1510 / :2669 and fluree-db-query/src/binary_scan.rs:1287 (none in this diff), so I'm parking it on the fallback scan that was meant to cover this case.
The store keeps the persisted tags twice: dicts.language_tags is rebuilt at load by get_or_insert (binary_index_store.rs:2669), which now normalizes, but language_tags: Vec<String> (:386) is a verbatim copy of the FIR6 root, and that verbatim copy is what resolve_lang_tag (:1510) reads. So for a ledger indexed before this PR with "chat"@en-US: the query lowers to LangTag("en-us"); value_to_otype_okey → resolve_lang_id("en-us") finds the right lang_id through the normalized dict, the cursor is pinned to the right o_type and yields the legacy rows — and then matches_datatype_constraint (binary_scan.rs:1287) does resolve_lang_tag(o_type) == Some("en-us") against "en-US", and drops every one. The range-fallback path has the same exact compare at :849. Before this PR the same query matched (exact lookup, exact compare).
I reproduced it end to end rather than trusting the trace: seeded a file-backed ledger at commit 1 (ingest identical to main for tags) with two "chat"@en-US rows and one "chat"@fr-CA, indexed, and confirmed ?s ex:name "chat"@en-US returns both rows there; re-opened the same directory at HEAD and "chat"@en-US, "chat"@en-us and "chat"@fr-CA all return [], while FILTER(LANG(?n) = "en-us") still finds the rows. It does not heal on reindex, and it gets worse: the incremental indexer seeds its dictionary verbatim (resolver.rs:1268) and then get_or_inserts the normalized novelty tag (:711), which is an exact VecBiDict miss, so the root ends up holding both en-US and en-us; the store's load loop then dedups them and every later language's id is off by one (["en-US","en-us","fr"] → find("fr") = 2 against a persisted lang_id of 3). After writing one new "chat"@en-US at HEAD and reindexing, "chat"@en-US → [] (including the row just written) and "chat"@fr-CA → []. en-US, pt-BR, zh-Hant are the ordinary BCP 47 spellings, so this is most real-world tagged data, and there is no migration or release note.
Two ways out, either is fine by me: (a) canonicalize at store load — normalize the Vec at :386 too and fold case-variant ids into one with a small remap, so every read-side lookup sees one lowercase tag per id — or (b) keep persisted ids untouched and make the read side case-insensitive end to end: eq_ignore_ascii_case at binary_scan.rs:849/:1287 (the novelty-side filters this PR adds at :1697/:2404 already do exactly this), and have resolve_lang_id return every id whose tag matches so the bound-object fast path can't pin a single wrong o_type (fall back to the untyped-string seek + per-row check when there is more than one). (a) also needs the indexer's seed path (from_ordered_tags) to normalize so a reindex doesn't re-split. Whichever you pick, the seed-at-old-revision / query-at-new-revision shape above is the test that should exist for it, and the behaviour-changes section of the body should say what happens to existing data.
There was a problem hiding this comment.
Addressed in 8c5f81b.
Thank you for reproducing this rather than just reading the trace — you were right, and the diagnosis was exact. I took the cheap half of your option (a) plus the compare from (b), which turned out to be three one-liners once I looked at where the tags actually enter:
binary_index_store.rs— normalize the root's tags on the way into the store. Sinceresolve_lang_tagreads that copy, this fixes the indexed compare at:1287and the range fallback at:849without touching either. Positions are preserved, solang_idstill indexes the vec 1:1 — case variants stay distinct ids, no folding, no remap.- The indexer's
from_ordered_tagsnow normalizes the ordered seed. That was what madeget_or_insertmiss and append a second id, so a reindex no longer splits the tag — and it heals the root as a side effect. Your "it does not heal on reindex, and it gets worse" is no longer true. binary_scan.rs:849now compares witheq_ignore_ascii_case. I found this one chasing the first two:FlakeMeta.langis a plain deserialized field, so flakes replayed from commits written before normalization reach novelty with the tag as authored — normalizing the store's vec doesn't help there because it's a flake, not a store lookup.
Two unit tests, both mutation-checked; reverting the seed normalization fails the indexer one. I did not do the seed-at-old-revision integration test — the transformations are directly testable without it, and building a legacy FIR6 root in a test is a lot of machinery for what is now a byte scan.
Not covered, and stated in the commit body and the new Existing data section of the PR body: a root that already holds both spellings of one tag still collapses at load and shifts later ids. That needs both spellings to have been written before this change; recovery is dump and re-import. We think the population here is zero, so I did not want to build the multi-id resolve and fast-path fallback for it.
| // a string binding probes only rows with its | ||
| // exact tag / `xsd:string` datatype. | ||
| let store = gv.store(); | ||
| let o_type = store.o_type_from_kind(*o_kind, *dt_id, *lang_id); |
There was a problem hiding this comment.
blocking (performance, per the repo's hot-path rule; small fix). The EncodedLit probe arm now does two fresh heap allocations per driving row on top of the ones already there: o_type_from_kind builds OTypeRegistry::builtin_only() → OTypeRegistry::new(&[]) → Vec::with_capacity(15) (fluree-db-core/src/o_type_registry.rs:31-33), and resolve_datatype_sid hands back Sid::new(..) → Arc::from(&str) (sid.rs:49) for every built-in string (or Arc::from(tag) at :995 for a tagged one). decode_value_from_kind two lines up already constructs the same registry, so this arm now builds it twice per row.
I don't think this changes the order of anything — it's only the literal-object bind-join lane, IRI joins never enter this arm, and the row already clones the pattern and decodes the string — but it is pure added work per probe row and none of it needs to be per row. The encoded triple already says what we need: lang_id != 0 means language-tagged, and dt_id equal to the reserved xsd:string DatatypeDictId means plain string; neither needs a registry or a Sid. Something like:
// no registry, no Sid: decide string-ness from the encoded triple
let dtc = if *lang_id != 0 {
store.resolve_language_tag(*lang_id).map(|t| DatatypeConstraint::LangTag(Arc::from(t)))
} else if DatatypeDictId::from_u16(*dt_id) == DatatypeDictId::STRING {
Some(DatatypeConstraint::Explicit(self.xsd_string_sid.clone())) // built once per operator
} else {
None
};with the tag Arc cached per lang_id on the operator if you want the tagged case allocation-free too. Fold-in-now territory — it is a few lines and it keeps the hot-path rule honest.
There was a problem hiding this comment.
Addressed in 392eeb9.
You were right that none of it needs to be per row, and the diagnosis of the double registry build was correct — decode_value_from_kind constructs one two lines up.
Took your approach with one adjustment. Rather than lang_id != 0, I keyed on dt_id: LANG_STRING for tagged, STRING for plain. OTypeRegistry::resolve only produces a lang-string OType for LEX_ID + LANG_STRING, so lang_id != 0 alone would have been a slightly wider condition than the code it replaces, and I wanted this to be exactly behaviour-preserving.
For the tag I added a borrowed lang_tag_for_id on the store instead of using resolve_language_tag — the latter returns String, so it would have traded the registry for an extra allocation on the tagged path. The new accessor reads the same vec resolve_lang_tag reads, which also keeps the constraint consistent with what the scan-side compare will check; resolve_lang_tag now delegates to it. xsd:string is a OnceLock, so a clone is a refcount bump.
Net: plain-string rows allocate nothing on this path (was registry + Sid), tagged rows allocate only the tag Arc they already did. it_literal_identity and it_fastpath_1652_regression 4/4, fluree-db-query 1886/1886.
| /// booleans and dates stay unconstrained so `25` keeps matching a stored | ||
| /// `"25"^^xsd:int` as before; tightening numeric subtypes is a separate | ||
| /// decision. | ||
| pub(super) fn lower_object_with_term_constraint( |
There was a problem hiding this comment.
optional (docs / release notes). This is more of a heads-up than a suggestion. The Typed { .. } arm pins every explicitly typed literal, not just typed strings, and dt_compatible only relaxes xsd:integer and xsd:double — I checked at HEAD against 25 / "25"^^xsd:int / "25"^^xsd:long rows: "25"^^xsd:int returns only the int row, "25"^^xsd:long only the long row, while "25"^^xsd:integer and bare 25 return all three. I think that's the right behaviour (it's spec-correct and it's what JSON-LD @type already does), but the body's "the blast radius is strings only" and the new paragraph in docs/concepts/datatypes.md:177 both say otherwise, and this is the kind of thing that feeds release notes. We may want a sentence in both saying explicitly typed literals are now exact per datatype, with bare numerics the one lenient form. Minor and non-blocking — but if you agree, I'd rather see it folded in now than lost in the backlog.
There was a problem hiding this comment.
Addressed in 392eeb9 — and you read it right, this was under-stated.
Updated both places. docs/concepts/datatypes.md now says an explicitly typed literal is exact for its datatype in both surfaces, with your xsd:int / xsd:long example, and calls bare numerics the one lenient form. The PR body's changes bullet no longer claims the blast radius is strings only, and Behaviour changes has its own line for it so it reaches release notes.
| DatatypeConstraint::LangTag(lang.clone()), | ||
| DatatypeConstraint::LangTag(normalized_lang_arc(lang)), | ||
| ), | ||
| LiteralValue::Typed { value, datatype } => { |
There was a problem hiding this comment.
fluree-db-sparql/src/lower/term.rs:271 — question. Not introduced here, but this PR makes it load-bearing on ordinary triples: when the datatype IRI isn't encodable the constraint falls back to xsd:string, so ?s ex:p "a"^^ex:NoSuchType now matches plain "a" rows, where the spec answer is no match (no such term exists). Previously it matched everything, so this is strictly narrower, and I'm not sure it matters in practice — but if we're pinning term identity, None-matches-nothing (or an Explicit on a fresh SID that can never match) seems like the more honest fallback. Happy to be told this is deliberate.
There was a problem hiding this comment.
Good question, and I dug into it rather than answering from the comment — which turns out not to justify itself. It points at "the storage-side fallback in term_to_binding", but term_to_binding has no custom-datatype fallback; its xsd:string is the Simple arm. So the stated reason for the fallback is wrong.
And your instinct is right on the substance. encode_iri_strict returns None exactly when the canonical prefix is not a registered namespace on this ledger — so when the fallback fires, no stored row can carry that datatype, and no-match is provably the correct answer rather than merely the spec-nicer one.
I have not changed it here, because None in this Option<DatatypeConstraint> means the opposite of what we want — no constraint, i.e. fully lenient — so expressing "never matches" needs either a new DatatypeConstraint variant or an early empty-result path in lowering. That is a design change rather than a fold-in, and it is pre-existing and strictly narrower than before this PR. I will file it with the above reasoning so the argument is not lost. Shout if you would rather it block here.
There was a problem hiding this comment.
Filed as #1686, with the encode_iri_strict argument and the note that the comment's stated justification does not hold, so the reasoning is not lost.
| // `"bob"@en` and `"bob"@fr` must not probe each | ||
| // other's rows. Numeric/other constraints are left | ||
| // off so cross-subtype matching stays as before. | ||
| if crate::binding::is_string_term_constraint(dtc) { |
There was a problem hiding this comment.
question (JSON-LD parity seam). A bare JSON string in a JSON-LD values cell is typed xsd:string (parse/lower.rs:530), so through this arm it now probes strictly, while a bare JSON string as a constant object stays lenient (pinned by the new test). Both are internally consistent with SPARQL, so I don't think this is wrong — just noting the seam so the "plain JSON string is lenient" product decision is stated as "…as a constant object" rather than in general. No change requested unless you want the docs sentence to say so.
There was a problem hiding this comment.
Addressed in 392eeb9 — docs only, no code change, since as you say both halves are internally consistent with SPARQL.
The datatypes.md sentence now scopes the leniency explicitly: a bare JSON string as a constant object is lenient, and the same string in a values cell is an xsd:string term matching only xsd:string rows. Worth having written down — it is exactly the kind of seam someone hits and reports as a bug.
| // BCP 47 tags compare case-insensitively; match the stored | ||
| // (lowercase) form. | ||
| Some(DatatypeConstraint::LangTag( | ||
| match fluree_db_core::normalize_lang_tag(tag) { |
There was a problem hiding this comment.
nit. UnresolvedTriplePattern::with_lang already normalized the tag at ast.rs:196, so this second normalize_lang_tag is a no-op scan on an already-lowercase Arc<str>. Harmless; either site alone would do.
There was a problem hiding this comment.
You are right that it is redundant on the with_lang path, but I would keep both, and I do not think either site alone would do.
lower.rs:609 is the choke point: UnresolvedDatatypeConstraint::LangTag is also constructed at node_map.rs:1513, values.rs:232, and several r2rml sites, none of which normalize. So that one is load-bearing for everything that does not come through with_lang, and removing ast.rs:196 would leave the parser path relying on it from further away for no measurable gain — an already-lowercase tag returns Cow::Borrowed and the Arc is cloned, so the cost is a byte scan of a 2-5 character string once per query term, not per row.
Happy to drop the ast.rs one if you feel strongly, but defensive at both a constructor and a choke point seems worth more than the scan.
| // either keeps or drops a fact's entries as a whole — so it can never | ||
| // separate an assertion from the retraction that cancels it. | ||
| // A tagged bound object matches only flakes carrying the same tag. | ||
| let bound_lang = self.pattern.dtc.as_ref().and_then(|d| d.lang_tag()); |
There was a problem hiding this comment.
praise. Hoisting bound_lang out of the per-flake closure and comparing with eq_ignore_ascii_case is exactly right — and it's the comparison the indexed path at :849/:1287 is missing (see the first blocking item).
There was a problem hiding this comment.
Thanks — and your instinct was right that the indexed path was missing it. :849 now compares with eq_ignore_ascii_case too, in 8c5f81b. It turned out to matter for a reason neither of us had in the original note: FlakeMeta.lang is a plain deserialized field, so flakes replayed from pre-normalization commits reach novelty with the tag as authored, and normalizing the store's tag vec does not reach them.
| (None, Some(i)) => Some(FlakeMeta::with_index(i)), | ||
| (None, None) => None, | ||
| }; | ||
| let meta = FlakeMeta::from_parts(lang.as_deref(), list_index); |
There was a problem hiding this comment.
praise. The three sink match blocks → FlakeMeta::from_parts refactor is behaviour-identical (I diffed the arms), so the "list-position" in the branch name is not hiding anything in this diff; the body's "not in this PR" note is accurate.
There was a problem hiding this comment.
Appreciated, and thanks for diffing the arms rather than taking the commit message's word for it.
…lang-and-list-position # Conflicts: # fluree-db-api/Cargo.toml
Tag normalization landed on the ingest and dictionary paths, but three read paths still compared against tags as authored, so a ledger written before normalization stopped answering tag-constrained patterns against its own data — `?s ex:name "chat"@en-US` returning nothing on rows it had matched before. `en-US`, `pt-BR` and `zh-Hant` are the ordinary BCP 47 spellings. The store keeps the persisted tags twice: `dicts.language_tags` is rebuilt through `get_or_insert`, which normalizes, while `language_tags` is a verbatim copy of the FIR6 root — and that verbatim copy is what `resolve_lang_tag` reads. A tagged query resolved the right `lang_id`, pinned the right `o_type`, yielded the right rows, then dropped every one of them on the compare. Normalize on the way into the store instead, which fixes the indexed and range-fallback compares at once without touching either. Positions are the `lang_id` mapping, so case variants stay distinct ids rather than folding. The indexer seeded its dictionary from the root verbatim and then `get_or_insert`ed the normalized novelty tag — an exact miss, so a rebuild carried both `en-US` and `en-us` and every later `lang_id` shifted when the store deduped them at load. Normalizing the ordered seed keeps a reindex on one id per language, and heals the root as a side effect. `FlakeMeta.lang` is a plain deserialized field, so flakes replayed from commits written before normalization reach novelty with the tag as authored. The range-fallback filter now compares case-insensitively, matching what the overlay filters already do. Not covered: a root that already holds two case-variant spellings of the same tag — only reachable if both spellings were written before this change. Those still collapse at load and shift later ids; recovery is dump and re-import.
The `EncodedLit` arm of the bind-join probe substitution tagged each driving row's constraint by resolving an `OType` first: `o_type_from_kind` builds an `OTypeRegistry` (a 15-element `Vec`) and `resolve_datatype_sid` returns a fresh `Sid`, whose name is an `Arc<str>` allocation. `decode_value_from_kind` two lines above already builds the same registry, so the arm constructed one twice per row on the literal-object join lane. None of it needs to be per row. The encoded triple already carries the answer: `dt_id == LANG_STRING` is a tagged string and `dt_id == STRING` is a plain one, and nothing else can produce a string term constraint. The tag comes from a new borrowed `lang_tag_for_id` — the same vec `resolve_lang_tag` reads, reachable without an `OType` — and `xsd:string` is built once in a `OnceLock`, so a clone is a refcount bump. Plain-string rows now allocate nothing on this path (was a registry plus a `Sid`); tagged rows allocate only the tag `Arc` they allocated before. `resolve_lang_tag` delegates to the new accessor rather than repeating the 1-based index arithmetic. Also documents in `docs/concepts/datatypes.md` that explicitly typed literals are exact per datatype — bare numerics are the one lenient form — and scopes the "bare JSON string is lenient" note to the constant-object position, since the same string in a `values` cell is an `xsd:string` term.
Problem
"bob","bob"@enand"bob"@frshare one string-dictionary key, and SPARQL lowering attached no datatype/language constraint to ordinary triple objects (only annotation-block objects carried one). A constant object therefore reached the scan as a bare string, took the untyped-string branch, and matched all three;VALUES ?o { "bob"@en }projected"bob"@enfor the false matches; per-row join/OPTIONAL probes rebuilt the object from a binding the same way, so joins on a string variable equated every tag. Pre-existing (byte-identical on v4.1.5) — it was masked init_fastpath_1652_regressiononly by the membership-join lane's(s, o_type, o_key)key, which #1674's gate change stops applying at small driving sets.Separately, language tags were stored as written and looked up exactly, so
"x"@ENand"x"@eninterned as two terms (W3Cdawg-lang-3).Changes
Term identity (commit 1)
lower_object_with_term_constraint) on ordinary triples and property-path endpoints. Bare numerics, booleans and dates stay unconstrained —25still matches any integer subtype. Note this pins every explicitly typed literal, not only typed strings:"25"^^xsd:intnow matchesxsd:introws only, not the same number stored asxsd:long.binding::is_string_term_constraint); late-materializedEncodedLitbindings recover theirs viaBinaryIndexStore::o_type_from_kind.binary_scanhonour the pattern's language tag.@language/@typevalue objects were already exact. A plain JSON string keeps its lenient cross-tag matching; that's pinned by test as a deliberate product decision (note: lenient is also the sloweruntyped_stringscan path on multilingual predicates).Language-tag case (commit 2)
fluree_db_core::normalize_lang_tag(ASCII-lowercase) is the canonical form;FlakeMeta::with_lang/ newFlakeMeta::from_partsapply it and the three transact sinks use it, so every ingest path stores lowercase. BothLanguageTagDicts intern/look up normalized (with a case-insensitive fallback scan for pre-normalization dictionaries). SPARQL / JSON-LD / VALUES constraints normalize;LANG()returns lowercase.Behaviour changes (release notes)
?s p "bob"now matches onlyxsd:stringvalues, not"bob"@en— spec-correct."25"^^xsd:intno longer matches anxsd:longrow. Bare numerics (25) stay lenient.LANG()on data written with an uppercase tag returns lowercase (RDF 1.1 value space).Existing data
Tag normalization is applied on the read side too, so a ledger written before
this change keeps answering tag-constrained patterns against its own data:
the store normalizes the root's tags at load (the copy
resolve_lang_tagreads), the indexer normalizes its ordered seed so a reindex cannot split
en-USfromen-usand shift laterlang_ids, and the range-fallbackfilter compares case-insensitively for flakes replayed from pre-normalization
commits. No migration or reindex is required.
The one uncovered shape is a root that already holds both spellings of a
tag, which needs both to have been written before this change; those collapse
at load and shift later ids. Recovery is dump and re-import.
Tests
fluree-db-api/tests/it_literal_identity.rs: 11 SPARQL shapes × {novelty, indexed} (tagged, plain, typed, VALUES one/multi-row, self-join, cross-predicate join, OPTIONAL, property path, FILTER), JSON-LD pins, numeric-leniency pin, mixed-case tags incl.LANG().LanguageTagDictcase test.basic#quotes-3/4,open-world#open-eq-02,expr-equals#eq-graph-4,expr-builtin#dawg-lang-3,entailment#lang,entailment#plainLit.--all-features --all-targets -D warningsclean.Docs:
docs/concepts/datatypes.mddocuments string-literal matching and tag case.Not in this PR
eq-graph-1/2stay skipped by design (numeric leniency).o_ihalf of the 1652 finding: an inconsistency between join lanes over whether list position is part of a literal's identity — a semantics call to settle explicitly, not patch here.