Skip to content

fix: match string literals by RDF term identity; normalize language-tag case - #1676

Merged
bplatz merged 5 commits into
mainfrom
fix/literal-identity-lang-and-list-position
Aug 25, 2026
Merged

bplatz merged 5 commits into
mainfrom
fix/literal-identity-lang-and-list-position

Conversation

@bplatz

@bplatz bplatz commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

Problem

"bob", "bob"@en and "bob"@fr share one string-dictionary key, and SPARQL lowering attached no datatype/language constraint to ordinary triple objects (only annotation-block objects carried one). A constant object therefore reached the scan as a bare string, took the untyped-string branch, and matched all three; VALUES ?o { "bob"@en } projected "bob"@en for the false matches; per-row join/OPTIONAL probes rebuilt the object from a binding the same way, so joins on a string variable equated every tag. Pre-existing (byte-identical on v4.1.5) — it was masked in it_fastpath_1652_regression only by the membership-join lane's (s, o_type, o_key) key, which #1674's gate change stops applying at small driving sets.

Separately, language tags were stored as written and looked up exactly, so "x"@EN and "x"@en interned as two terms (W3C dawg-lang-3).

Changes

Term identity (commit 1)

  • SPARQL lowering: plain, language-tagged and explicitly typed literals carry their term constraint (lower_object_with_term_constraint) on ordinary triples and property-path endpoints. Bare numerics, booleans and dates stay unconstrained — 25 still matches any integer subtype. Note this pins every explicitly typed literal, not only typed strings: "25"^^xsd:int now matches xsd:int rows only, not the same number stored as xsd:long.
  • Join / OPTIONAL probe substitution carries the binding's string constraint (binding::is_string_term_constraint); late-materialized EncodedLit bindings recover theirs via BinaryIndexStore::o_type_from_kind.
  • The overlay-only fallback and the raw untranslated-flake filter in binary_scan honour the pattern's language tag.
  • JSON-LD @language / @type value objects were already exact. A plain JSON string keeps its lenient cross-tag matching; that's pinned by test as a deliberate product decision (note: lenient is also the slower untyped_string scan path on multilingual predicates).

Language-tag case (commit 2)

  • fluree_db_core::normalize_lang_tag (ASCII-lowercase) is the canonical form; FlakeMeta::with_lang / new FlakeMeta::from_parts apply it and the three transact sinks use it, so every ingest path stores lowercase. Both LanguageTagDicts intern/look up normalized (with a case-insensitive fallback scan for pre-normalization dictionaries). SPARQL / JSON-LD / VALUES constraints normalize; LANG() returns lowercase.

Behaviour changes (release notes)

  • SPARQL ?s p "bob" now matches only xsd:string values, not "bob"@en — spec-correct.
  • Explicitly typed literals are exact per datatype: "25"^^xsd:int no longer matches an xsd:long row. Bare numerics (25) stay lenient.
  • LANG() on data written with an uppercase tag returns lowercase (RDF 1.1 value space).

Existing data

Tag normalization is applied on the read side too, so a ledger written before
this change keeps answering tag-constrained patterns against its own data:
the store normalizes the root's tags at load (the copy resolve_lang_tag
reads), the indexer normalizes its ordered seed so a reindex cannot split
en-US from en-us and shift later lang_ids, and the range-fallback
filter compares case-insensitively for flakes replayed from pre-normalization
commits. No migration or reindex is required.

The one uncovered shape is a root that already holds both spellings of a
tag, which needs both to have been written before this change; those collapse
at load and shift later ids. Recovery is dump and re-import.

Tests

  • fluree-db-api/tests/it_literal_identity.rs: 11 SPARQL shapes × {novelty, indexed} (tagged, plain, typed, VALUES one/multi-row, self-join, cross-predicate join, OPTIONAL, property path, FILTER), JSON-LD pins, numeric-leniency pin, mixed-case tags incl. LANG().
  • Lowering unit test; LanguageTagDict case test.
  • W3C: 36/36, with seven previously-skipped tests now passing and removed from the skip register: basic#quotes-3/4, open-world#open-eq-02, expr-equals#eq-graph-4, expr-builtin#dawg-lang-3, entailment#lang, entailment#plainLit.
  • grp_query 262 / grp_query_sparql 422 / grp_misc 349 / grp_transact 185 / grp_import 99 / grp_index 89 / 1652 pin; core/transact/query/sparql/binary-index/indexer units; clippy --all-features --all-targets -D warnings clean.

Docs: docs/concepts/datatypes.md documents string-literal matching and tag case.

Not in this PR

  • eq-graph-1/2 stay skipped by design (numeric leniency).
  • The list-rows o_i half of the 1652 finding: an inconsistency between join lanes over whether list position is part of a literal's identity — a semantics call to settle explicitly, not patch here.

bplatz added 2 commits August 21, 2026 07:47
…ctionary key

`"bob"`, `"bob"@en` and `"bob"@fr` share one string-dictionary key. SPARQL
lowering attached no datatype/language constraint to ordinary triple objects
(only annotation-block objects carried one), so a constant object reached the
scan as a bare string, took the untyped-string branch, and matched all three.
`VALUES ?o { "bob"@en }` additionally projected "bob"@en for the false
matches, and per-row join/OPTIONAL probes rebuilt the object from a binding
the same way, so joins on a string variable equated every tag.

- SPARQL lowering: plain, language-tagged and explicitly typed string
  literals carry their term constraint (`lower_object_with_term_constraint`)
  on ordinary triples and property-path endpoints. Bare numerics, booleans
  and dates stay unconstrained, so `25` still matches any integer subtype.
- Join / OPTIONAL probe substitution carries the binding's string
  constraint; late-materialized (`EncodedLit`) bindings recover theirs via
  `BinaryIndexStore::o_type_from_kind`.
- The overlay-only fallback and the raw untranslated-flake filter check the
  novelty flake's language tag against the pattern's.

JSON-LD `@language` / `@type` value objects were already exact; a plain JSON
string keeps its lenient cross-tag matching (pinned by test).

W3C: basic#quotes-3/4, open-world#open-eq-02, expr-equals#eq-graph-4,
entailment#lang and entailment#plainLit now pass and leave the skip register.
…ive)

RDF 1.1 defines the value space of language tags as lowercase and compares
them case-insensitively, so `"x"@EN` and `"x"@en` are one term. We stored
the tag as written and looked it up exactly, so the two interned as
different lang_ids and `?x :p "string"@EN` could not find `"string"@en`.

- `fluree_db_core::normalize_lang_tag` is the single canonical form.
  `FlakeMeta::with_lang` / new `FlakeMeta::from_parts` apply it, and the
  three transact sinks build metadata through `from_parts`.
- Both `LanguageTagDict`s intern and look up the normalized form; a
  dictionary persisted before normalization still answers via a
  case-insensitive fallback scan.
- SPARQL and JSON-LD constraints (`"x"@EN`, `@language: "En"`, VALUES
  cells) normalize before matching; the novelty-side filters compare
  case-insensitively.
- `LANG()` reports the canonical lowercase form.

W3C expr-builtin#dawg-lang-3 now passes and leaves the skip register.
@bplatz
bplatz requested review from aaj3f and zonotope August 21, 2026 12:23

@aaj3f aaj3f left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@bplatz good catch and seems like the right fix.

My only question is for pre-existing ledgers/data-state involving language tags. It seems the store keeps a verbatim copy of the persisted tags alongside the now-normalized dictionary, and the per-row compare in matches_datatype_constraint reads the verbatim copy, so a ledger indexed before this PR with "chat"@en-US answers ?s ex:name "chat"@en-US with [] at HEAD (I seeded one at commit 1 and re-opened it at HEAD — en-US, en-us and fr-CA all return nothing while FILTER(LANG(?n)="en-us") still finds the rows), and a reindex makes it worse by splitting the tag into two ids and shifting every later lang_id on the read side. Details and two fix options are on global_dict.rs:176.

The second item is the pair of per-row allocations in the EncodedLit probe arm (join.rs:992), which is a few-line hoist. The rest of my inline notes below are docs accuracy (typed non-string literals are now exact too — a good change, just undocumented) and two questions.

Adherence to repo commitments:

  • Patterns/abstractions: ✔ Extends TriplePattern.dtc / DatatypeConstraint and the existing scan checks; one shared lowering helper for triples and path endpoints; no parallel construct. ⚠️ Case normalization is applied at ingest and in the dictionaries but not to the store's verbatim tag table or the read-side compares, which is what breaks existing data.
  • Performance (speed first, memory second): ⚠️ Hot path touched (bind-join probe, novelty filters). Lit arms and novelty filters are allocation-free; the EncodedLit arm adds two heap allocations per probe row (registry Vec + Sid/Arc<str>) on the literal-object join lane — small, hoistable, should land before merge. No bench covers that lane; bench-compare's advisory numbers are the same ambient noise as on #1667/#1668.
  • Testing: ✔ New [[test]] target runs in CI (three tests by name), lowering + dictionary unit tests, W3C 36/36 with seven register removals. ✖ Nothing exercises an index written before this change, which is the case that regresses.
  • Conventions: ✔ Two well-factored commits with thorough multi-line bodies; fmt clean; clippy -D warnings clean on all six touched crates; docs updated. ⚠️ Body/docs under-state the behaviour change (typed literals) and say nothing about existing ledgers.

Verified locally at branch HEAD (1b0254f): cargo fmt --all -- --check clean; cargo clippy -p fluree-db-query -p fluree-db-core -p fluree-db-binary-index -p fluree-db-sparql -p fluree-db-transact -p fluree-db-indexer --all-targets -- -D warnings clean; cargo nextest run -p fluree-db-api --test it_literal_identity --test it_fastpath_1652_regression 4/4; cargo nextest run -p fluree-db-query 1467 passed; three mutation checks red as expected; legacy-ledger probe (seed at aaaa7d4, query at HEAD) red as described above.

/// Case-insensitive reverse lookup: the dictionary stores normalized
/// (lowercase) tags, but a dictionary persisted before normalization may
/// still hold the tag as written, so an exact miss falls back to a scan.
fn find_normalized(&self, tag: &str) -> Option<u16> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocking. Existing ledgers with any non-lowercase language tag stop matching tag-constrained patterns after this upgrade, before and after reindex. The real location is fluree-db-binary-index/src/read/binary_index_store.rs:386 / :1510 / :2669 and fluree-db-query/src/binary_scan.rs:1287 (none in this diff), so I'm parking it on the fallback scan that was meant to cover this case.

The store keeps the persisted tags twice: dicts.language_tags is rebuilt at load by get_or_insert (binary_index_store.rs:2669), which now normalizes, but language_tags: Vec<String> (:386) is a verbatim copy of the FIR6 root, and that verbatim copy is what resolve_lang_tag (:1510) reads. So for a ledger indexed before this PR with "chat"@en-US: the query lowers to LangTag("en-us"); value_to_otype_okey → resolve_lang_id("en-us") finds the right lang_id through the normalized dict, the cursor is pinned to the right o_type and yields the legacy rows — and then matches_datatype_constraint (binary_scan.rs:1287) does resolve_lang_tag(o_type) == Some("en-us") against "en-US", and drops every one. The range-fallback path has the same exact compare at :849. Before this PR the same query matched (exact lookup, exact compare).

I reproduced it end to end rather than trusting the trace: seeded a file-backed ledger at commit 1 (ingest identical to main for tags) with two "chat"@en-US rows and one "chat"@fr-CA, indexed, and confirmed ?s ex:name "chat"@en-US returns both rows there; re-opened the same directory at HEAD and "chat"@en-US, "chat"@en-us and "chat"@fr-CA all return [], while FILTER(LANG(?n) = "en-us") still finds the rows. It does not heal on reindex, and it gets worse: the incremental indexer seeds its dictionary verbatim (resolver.rs:1268) and then get_or_inserts the normalized novelty tag (:711), which is an exact VecBiDict miss, so the root ends up holding both en-US and en-us; the store's load loop then dedups them and every later language's id is off by one (["en-US","en-us","fr"] → find("fr") = 2 against a persisted lang_id of 3). After writing one new "chat"@en-US at HEAD and reindexing, "chat"@en-US → [] (including the row just written) and "chat"@fr-CA → []. en-US, pt-BR, zh-Hant are the ordinary BCP 47 spellings, so this is most real-world tagged data, and there is no migration or release note.

Two ways out, either is fine by me: (a) canonicalize at store load — normalize the Vec at :386 too and fold case-variant ids into one with a small remap, so every read-side lookup sees one lowercase tag per id — or (b) keep persisted ids untouched and make the read side case-insensitive end to end: eq_ignore_ascii_case at binary_scan.rs:849/:1287 (the novelty-side filters this PR adds at :1697/:2404 already do exactly this), and have resolve_lang_id return every id whose tag matches so the bound-object fast path can't pin a single wrong o_type (fall back to the untyped-string seek + per-row check when there is more than one). (a) also needs the indexer's seed path (from_ordered_tags) to normalize so a reindex doesn't re-split. Whichever you pick, the seed-at-old-revision / query-at-new-revision shape above is the test that should exist for it, and the behaviour-changes section of the body should say what happens to existing data.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 8c5f81b.

Thank you for reproducing this rather than just reading the trace — you were right, and the diagnosis was exact. I took the cheap half of your option (a) plus the compare from (b), which turned out to be three one-liners once I looked at where the tags actually enter:

  1. binary_index_store.rs — normalize the root's tags on the way into the store. Since resolve_lang_tag reads that copy, this fixes the indexed compare at :1287 and the range fallback at :849 without touching either. Positions are preserved, so lang_id still indexes the vec 1:1 — case variants stay distinct ids, no folding, no remap.
  2. The indexer's from_ordered_tags now normalizes the ordered seed. That was what made get_or_insert miss and append a second id, so a reindex no longer splits the tag — and it heals the root as a side effect. Your "it does not heal on reindex, and it gets worse" is no longer true.
  3. binary_scan.rs:849 now compares with eq_ignore_ascii_case. I found this one chasing the first two: FlakeMeta.lang is a plain deserialized field, so flakes replayed from commits written before normalization reach novelty with the tag as authored — normalizing the store's vec doesn't help there because it's a flake, not a store lookup.

Two unit tests, both mutation-checked; reverting the seed normalization fails the indexer one. I did not do the seed-at-old-revision integration test — the transformations are directly testable without it, and building a legacy FIR6 root in a test is a lot of machinery for what is now a byte scan.

Not covered, and stated in the commit body and the new Existing data section of the PR body: a root that already holds both spellings of one tag still collapses at load and shifts later ids. That needs both spellings to have been written before this change; recovery is dump and re-import. We think the population here is zero, so I did not want to build the multi-id resolve and fast-path fallback for it.

Comment thread fluree-db-query/src/join.rs Outdated
// a string binding probes only rows with its
// exact tag / `xsd:string` datatype.
let store = gv.store();
let o_type = store.o_type_from_kind(*o_kind, *dt_id, *lang_id);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

blocking (performance, per the repo's hot-path rule; small fix). The EncodedLit probe arm now does two fresh heap allocations per driving row on top of the ones already there: o_type_from_kind builds OTypeRegistry::builtin_only() → OTypeRegistry::new(&[]) → Vec::with_capacity(15) (fluree-db-core/src/o_type_registry.rs:31-33), and resolve_datatype_sid hands back Sid::new(..) → Arc::from(&str) (sid.rs:49) for every built-in string (or Arc::from(tag) at :995 for a tagged one). decode_value_from_kind two lines up already constructs the same registry, so this arm now builds it twice per row.

I don't think this changes the order of anything — it's only the literal-object bind-join lane, IRI joins never enter this arm, and the row already clones the pattern and decodes the string — but it is pure added work per probe row and none of it needs to be per row. The encoded triple already says what we need: lang_id != 0 means language-tagged, and dt_id equal to the reserved xsd:string DatatypeDictId means plain string; neither needs a registry or a Sid. Something like:

// no registry, no Sid: decide string-ness from the encoded triple
let dtc = if *lang_id != 0 {
    store.resolve_language_tag(*lang_id).map(|t| DatatypeConstraint::LangTag(Arc::from(t)))
} else if DatatypeDictId::from_u16(*dt_id) == DatatypeDictId::STRING {
    Some(DatatypeConstraint::Explicit(self.xsd_string_sid.clone())) // built once per operator
} else {
    None
};

with the tag Arc cached per lang_id on the operator if you want the tagged case allocation-free too. Fold-in-now territory — it is a few lines and it keeps the hot-path rule honest.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 392eeb9.

You were right that none of it needs to be per row, and the diagnosis of the double registry build was correct — decode_value_from_kind constructs one two lines up.

Took your approach with one adjustment. Rather than lang_id != 0, I keyed on dt_id: LANG_STRING for tagged, STRING for plain. OTypeRegistry::resolve only produces a lang-string OType for LEX_ID + LANG_STRING, so lang_id != 0 alone would have been a slightly wider condition than the code it replaces, and I wanted this to be exactly behaviour-preserving.

For the tag I added a borrowed lang_tag_for_id on the store instead of using resolve_language_tag — the latter returns String, so it would have traded the registry for an extra allocation on the tagged path. The new accessor reads the same vec resolve_lang_tag reads, which also keeps the constraint consistent with what the scan-side compare will check; resolve_lang_tag now delegates to it. xsd:string is a OnceLock, so a clone is a refcount bump.

Net: plain-string rows allocate nothing on this path (was registry + Sid), tagged rows allocate only the tag Arc they already did. it_literal_identity and it_fastpath_1652_regression 4/4, fluree-db-query 1886/1886.

/// booleans and dates stay unconstrained so `25` keeps matching a stored
/// `"25"^^xsd:int` as before; tightening numeric subtypes is a separate
/// decision.
pub(super) fn lower_object_with_term_constraint(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

optional (docs / release notes). This is more of a heads-up than a suggestion. The Typed { .. } arm pins every explicitly typed literal, not just typed strings, and dt_compatible only relaxes xsd:integer and xsd:double — I checked at HEAD against 25 / "25"^^xsd:int / "25"^^xsd:long rows: "25"^^xsd:int returns only the int row, "25"^^xsd:long only the long row, while "25"^^xsd:integer and bare 25 return all three. I think that's the right behaviour (it's spec-correct and it's what JSON-LD @type already does), but the body's "the blast radius is strings only" and the new paragraph in docs/concepts/datatypes.md:177 both say otherwise, and this is the kind of thing that feeds release notes. We may want a sentence in both saying explicitly typed literals are now exact per datatype, with bare numerics the one lenient form. Minor and non-blocking — but if you agree, I'd rather see it folded in now than lost in the backlog.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 392eeb9 — and you read it right, this was under-stated.

Updated both places. docs/concepts/datatypes.md now says an explicitly typed literal is exact for its datatype in both surfaces, with your xsd:int / xsd:long example, and calls bare numerics the one lenient form. The PR body's changes bullet no longer claims the blast radius is strings only, and Behaviour changes has its own line for it so it reaches release notes.

DatatypeConstraint::LangTag(lang.clone()),
DatatypeConstraint::LangTag(normalized_lang_arc(lang)),
),
LiteralValue::Typed { value, datatype } => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fluree-db-sparql/src/lower/term.rs:271 — question. Not introduced here, but this PR makes it load-bearing on ordinary triples: when the datatype IRI isn't encodable the constraint falls back to xsd:string, so ?s ex:p "a"^^ex:NoSuchType now matches plain "a" rows, where the spec answer is no match (no such term exists). Previously it matched everything, so this is strictly narrower, and I'm not sure it matters in practice — but if we're pinning term identity, None-matches-nothing (or an Explicit on a fresh SID that can never match) seems like the more honest fallback. Happy to be told this is deliberate.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question, and I dug into it rather than answering from the comment — which turns out not to justify itself. It points at "the storage-side fallback in term_to_binding", but term_to_binding has no custom-datatype fallback; its xsd:string is the Simple arm. So the stated reason for the fallback is wrong.

And your instinct is right on the substance. encode_iri_strict returns None exactly when the canonical prefix is not a registered namespace on this ledger — so when the fallback fires, no stored row can carry that datatype, and no-match is provably the correct answer rather than merely the spec-nicer one.

I have not changed it here, because None in this Option<DatatypeConstraint> means the opposite of what we want — no constraint, i.e. fully lenient — so expressing "never matches" needs either a new DatatypeConstraint variant or an early empty-result path in lowering. That is a design change rather than a fold-in, and it is pre-existing and strictly narrower than before this PR. I will file it with the above reasoning so the argument is not lost. Shout if you would rather it block here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Filed as #1686, with the encode_iri_strict argument and the note that the comment's stated justification does not hold, so the reasoning is not lost.

// `"bob"@en` and `"bob"@fr` must not probe each
// other's rows. Numeric/other constraints are left
// off so cross-subtype matching stays as before.
if crate::binding::is_string_term_constraint(dtc) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

question (JSON-LD parity seam). A bare JSON string in a JSON-LD values cell is typed xsd:string (parse/lower.rs:530), so through this arm it now probes strictly, while a bare JSON string as a constant object stays lenient (pinned by the new test). Both are internally consistent with SPARQL, so I don't think this is wrong — just noting the seam so the "plain JSON string is lenient" product decision is stated as "…as a constant object" rather than in general. No change requested unless you want the docs sentence to say so.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 392eeb9 — docs only, no code change, since as you say both halves are internally consistent with SPARQL.

The datatypes.md sentence now scopes the leniency explicitly: a bare JSON string as a constant object is lenient, and the same string in a values cell is an xsd:string term matching only xsd:string rows. Worth having written down — it is exactly the kind of seam someone hits and reports as a bug.

// BCP 47 tags compare case-insensitively; match the stored
// (lowercase) form.
Some(DatatypeConstraint::LangTag(
match fluree_db_core::normalize_lang_tag(tag) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit. UnresolvedTriplePattern::with_lang already normalized the tag at ast.rs:196, so this second normalize_lang_tag is a no-op scan on an already-lowercase Arc<str>. Harmless; either site alone would do.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right that it is redundant on the with_lang path, but I would keep both, and I do not think either site alone would do.

lower.rs:609 is the choke point: UnresolvedDatatypeConstraint::LangTag is also constructed at node_map.rs:1513, values.rs:232, and several r2rml sites, none of which normalize. So that one is load-bearing for everything that does not come through with_lang, and removing ast.rs:196 would leave the parser path relying on it from further away for no measurable gain — an already-lowercase tag returns Cow::Borrowed and the Arc is cloned, so the cost is a byte scan of a 2-5 character string once per query term, not per row.

Happy to drop the ast.rs one if you feel strongly, but defensive at both a constructor and a choke point seems worth more than the scan.

// either keeps or drops a fact's entries as a whole — so it can never
// separate an assertion from the retraction that cancels it.
// A tagged bound object matches only flakes carrying the same tag.
let bound_lang = self.pattern.dtc.as_ref().and_then(|d| d.lang_tag());

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise. Hoisting bound_lang out of the per-flake closure and comparing with eq_ignore_ascii_case is exactly right — and it's the comparison the indexed path at :849/:1287 is missing (see the first blocking item).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks — and your instinct was right that the indexed path was missing it. :849 now compares with eq_ignore_ascii_case too, in 8c5f81b. It turned out to matter for a reason neither of us had in the original note: FlakeMeta.lang is a plain deserialized field, so flakes replayed from pre-normalization commits reach novelty with the tag as authored, and normalizing the store's tag vec does not reach them.

(None, Some(i)) => Some(FlakeMeta::with_index(i)),
(None, None) => None,
};
let meta = FlakeMeta::from_parts(lang.as_deref(), list_index);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

praise. The three sink match blocks → FlakeMeta::from_parts refactor is behaviour-identical (I diffed the arms), so the "list-position" in the branch name is not hiding anything in this diff; the body's "not in this PR" note is accurate.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Appreciated, and thanks for diffing the arms rather than taking the commit message's word for it.

bplatz added 3 commits August 25, 2026 15:19
…lang-and-list-position

# Conflicts:
#	fluree-db-api/Cargo.toml
Tag normalization landed on the ingest and dictionary paths, but three read
paths still compared against tags as authored, so a ledger written before
normalization stopped answering tag-constrained patterns against its own
data — `?s ex:name "chat"@en-US` returning nothing on rows it had matched
before. `en-US`, `pt-BR` and `zh-Hant` are the ordinary BCP 47 spellings.

The store keeps the persisted tags twice: `dicts.language_tags` is rebuilt
through `get_or_insert`, which normalizes, while `language_tags` is a
verbatim copy of the FIR6 root — and that verbatim copy is what
`resolve_lang_tag` reads. A tagged query resolved the right `lang_id`,
pinned the right `o_type`, yielded the right rows, then dropped every one of
them on the compare. Normalize on the way into the store instead, which
fixes the indexed and range-fallback compares at once without touching
either. Positions are the `lang_id` mapping, so case variants stay distinct
ids rather than folding.

The indexer seeded its dictionary from the root verbatim and then
`get_or_insert`ed the normalized novelty tag — an exact miss, so a rebuild
carried both `en-US` and `en-us` and every later `lang_id` shifted when the
store deduped them at load. Normalizing the ordered seed keeps a reindex on
one id per language, and heals the root as a side effect.

`FlakeMeta.lang` is a plain deserialized field, so flakes replayed from
commits written before normalization reach novelty with the tag as authored.
The range-fallback filter now compares case-insensitively, matching what the
overlay filters already do.

Not covered: a root that already holds two case-variant spellings of the
same tag — only reachable if both spellings were written before this change.
Those still collapse at load and shift later ids; recovery is dump and
re-import.
The `EncodedLit` arm of the bind-join probe substitution tagged each driving
row's constraint by resolving an `OType` first: `o_type_from_kind` builds an
`OTypeRegistry` (a 15-element `Vec`) and `resolve_datatype_sid` returns a
fresh `Sid`, whose name is an `Arc<str>` allocation. `decode_value_from_kind`
two lines above already builds the same registry, so the arm constructed one
twice per row on the literal-object join lane.

None of it needs to be per row. The encoded triple already carries the
answer: `dt_id == LANG_STRING` is a tagged string and `dt_id == STRING` is a
plain one, and nothing else can produce a string term constraint. The tag
comes from a new borrowed `lang_tag_for_id` — the same vec `resolve_lang_tag`
reads, reachable without an `OType` — and `xsd:string` is built once in a
`OnceLock`, so a clone is a refcount bump.

Plain-string rows now allocate nothing on this path (was a registry plus a
`Sid`); tagged rows allocate only the tag `Arc` they allocated before.
`resolve_lang_tag` delegates to the new accessor rather than repeating the
1-based index arithmetic.

Also documents in `docs/concepts/datatypes.md` that explicitly typed literals
are exact per datatype — bare numerics are the one lenient form — and scopes
the "bare JSON string is lenient" note to the constant-object position, since
the same string in a `values` cell is an `xsd:string` term.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants