Skip to content

fix(query): typed numeric constants match exactly, bare numbers match any numeric datatype - #1990

Merged
bplatz merged 4 commits into
mainfrom
fix/numeric-constant-matching
Sep 30, 2026
Merged

bplatz merged 4 commits into
mainfrom
fix/numeric-constant-matching

Conversation

@bplatz

@bplatz bplatz commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #1737

Summary

A numeric constant in a triple pattern now means the same thing on every lane and both query surfaces:

  • Typed ("25"^^xsd:long, {"@value": "25", "@type": "xsd:long"}): matches that datatype only.
  • Bare (25, 25.0, a JSON number in a JSON-LD query): matches every numeric datatype holding an equal value — xsd:integer, xsd:long, xsd:int, xsd:double, …

This resolves the product question #1737 was left open on (triage:needs-decision): bare numbers stay lenient, typed numbers are exact. It is what docs/concepts/datatypes.md already documented; the engine now agrees with it.

What was wrong

Measured on main with one subject per datatype, all holding 25:

constant JSON-LD novelty JSON-LD indexed SPARQL novelty SPARQL indexed
"25"^^xsd:long integer, int, long integer long long
bare 25 (integer + long + int + double on the predicate) all four integer int, integer, long integer
bare 25.0 (predicate holds only xsd:long) long nothing not measured nothing
COUNT of bare 25, indexed fast path 1, generic 4

Three causes:

  1. JSON-LD rewrote typed constants. parse_value_object ran @type through normalize_numeric_datatype (xsd:long/int/… → xsd:integer, xsd:float → xsd:double), a v4-baseline assumption that storage normalized numeric subtypes. It doesn't: writes keep the declared datatype. The rewritten constant then sought the wrong o_type once indexed.
  2. dt_compatible made typed xsd:integer/xsd:double lenient. It existed to let those rewritten constants match their family, and so also made SPARQL "25"^^xsd:integer match xsd:long rows in novelty (but not once indexed).
  3. The index had no multi-datatype seek for bare numbers. infer_exact_datatype_sid_from_stats narrowed to the predicate's datatype only when it had exactly one; otherwise value_to_otype_okey_simple pinned XSD_INTEGER/XSD_DOUBLE, so one exact seek missed every other numeric datatype. The single-datatype narrowing was also wrong across families: a bare double on an integer-typed predicate paired XSD_INTEGER with an f64 key. The predicate_object_count fast path had the same encoding.
  4. The policy lane pinned one numeric family. A scan whose predicate a policy might touch falls back to the range provider (binary_range.rs), whose bare-number prefilter left o_type open but pinned o_key to one family's encoding: under a policy 25 dropped "25.0"^^xsd:double, and 25.0 kept only it.

Changes

  • fluree-db-query/src/parse/node_map.rs: a constant @type is kept as written. The now-unused fluree_vocab::xsd::normalize_* helpers are removed.
  • dt_compatible is removed; its five callers compare datatypes exactly. Flake identity already includes the exact datatype, so the two write-side callers in stage.rs (retraction meta copy, delete-verb classification) were lenient about pairs that never cancel each other.
  • untyped_numeric_slices (binary_scan.rs): for a bare Long/Double, one (o_type, o_key) per numeric datatype the predicate carries, keyed by that family's encoding (i64 for integer types, f64 for float types; conversions follow FlakeValue::numeric_cmp). The datatypes come from the predicate's observed-datatype stats; with no predicate (?s ?p 25) or no observed set, every integer and float o_type (fifteen point seeks).
    • Decimals and big integers are keyed by per-predicate NumBig arena handles. For a current read the value's handles are looked up in the predicate's arena, or in every arena of the graph when the predicate is unbound. History and below-max_t reads scan the NumBig o_type under the decoded-value filter instead. A value only in novelty has no handle and arrives through the raw-flake lane.
    • One slice is a single seek, as before; several become a chain of cursors drained in order (pending_cursors), used only when every key component ahead of o_type in the sort order is bound, so the output stays in index order. Each cursor sets its own key range. No row can hold the value → overlay-only. A non-finite value → the unnarrowed scan under the decoded-value filter.
    • A bare_number_seek routing stamp records whether the scan seeks or walks.
  • The range provider seeks the same slices, so a policy-governed read matches the rows the root view does.
  • The predicate_object_count fast path sums a count per slice, and declines to the generic count when a slice needs the decoded-value filter.
  • Commit metadata in novelty (commit_flakes.rs) typed f:t, f:asserts and f:retracts as xsd:long, while the indexer and bulk import write them as xsd:integer. With exact typed matching that divergence became visible (a typed f:t lookup answered differently before and after indexing), so novelty now uses xsd:integer too. These flakes are generated from the commit envelope on load, so no existing data changes; docs/reference/vocabulary.md listed a third type for f:t (xsd:int) and is corrected.
  • Docs: concepts/datatypes.md spells out typed vs bare numeric matching; reference/vocabulary.md corrected as above.

Tests

New standalone binary it_numeric_constant_matching (it toggles the fast-path kill switch): thirteen cases — bare integer, double and decimal, each typed datatype, a single-datatype predicate, a bare number with no predicate, and COUNT — each on SPARQL and JSON-LD (JSON-LD can't express a constant object under a variable predicate), with fast paths on and off. The predicate holds an xsd:decimal 25 among its numeric rows, and a decimal on another predicate shares its arena handle number. Six lanes: novelty, indexed, indexed under a policy, a new numeric datatype written over the index (novelty-aware datatype stats), the same under a policy, and a time-travel read below the index's max_t (no arena handles). Routing is pinned: the indexed COUNT asserts the predicate_object_count stamp proceeds, every indexed bare-number scan asserts bare_number_seek proceeds (a walk returns the same rows, only slower), and the policy lanes assert the scan fell back to the range provider.

Updated tests that pinned the old leniency, each now asserting both sides:

  • it_query_datatype::datatype_query_explicit_typed_value_object_matches: xsd:integer constant no longer matches the xsd:int row.
  • it_query_misc::{untyped_value_matching_parity, indexed_untyped_value_matching_parity}: commit f:t is xsd:integer; an xsd:int constant matches nothing, xsd:integer matches it, bare still matches.

Each fix was reintroduced on its own to confirm the new test catches it: the JSON-LD @type rewrite (typed JSON-LD cases, all lanes), lenient typed xsd:integer (typed integer, novelty lane), the scan pinning XSD_INTEGER (bare cases, indexed lane), the COUNT fast path pinning it (indexed COUNT, 1 vs 4), novelty f:t as xsd:long (untyped_value_matching_parity), no seek without a predicate (bare_number_seek stamps), the range provider pinning one family's o_key (policy lanes, 40 checks), no NumBig slices (decimal row, 80 checks), no arena lookup without a predicate, and no NumBig slice for a below-max_t read (time-travel lane).

?s ?p 25 on a 200k-flake ledger (debug build, best of five, same rows): 90 ms at 63a0f31cf, 0.53 ms now. A bare number on an all-decimal predicate: 42 ms → 0.40 ms.

Unit tests for untyped_numeric_slices replace the numeric cases of the old inference tests (observed set over counts, an empty set seeks every numeric datatype, NumBig handles vs the NumBig range, non-finite declines, 2^53 + 1 has no float slice and 2^60 has one).

Behavior changes

  • JSON-LD constants typed with a numeric subtype no longer match other subtypes in novelty. They used to — and returned the wrong subtype once indexed.
  • SPARQL "25"^^xsd:integer and "1.5"^^xsd:double written out are exact. Bare 25 / 1.5e0 remain lenient.
  • Bare numbers now also match across the integer/float divide once indexed (25 finds "25.0"^^xsd:double), and match xsd:decimal rows, as they did in novelty. The same holds with no predicate and under a policy.

Notes for review

  • fix: keep typed literal values through the index (reindex affected ledgers) #1988 has merged; its indexed-lane skip of the JSON-LD xsd:long/xsd:float constants in it_typed_literal_index.rs is removed here, and the test passes.
  • The genesis-path datatype check in fluree-db-core/src/range.rs became the same exact comparison but isn't reached by these tests (unindexed file-backed ledgers answer through the scan operator's range fallback, which is covered).
  • it_bounded_overlay_translation::warm_whole_product_short_circuits_and_respects_epoch failed once in a parallel grp_query run on this branch, then passed alone three times and in the next full grp_query run; its failing assertion is a tracing-span check (which overlay-translation lane a read took), not a result.

Not in this PR

… any numeric datatype (#1737)

A numeric constant in a triple pattern now means the same thing on every
lane and both query surfaces: a typed literal ("25"^^xsd:long, or a JSON-LD
value object with @type) matches its own datatype only, and a bare number
(25, 25.0, a JSON number) matches every numeric datatype holding an equal
value.

- The JSON-LD query parser no longer rewrites a constant's @type to
  xsd:integer / xsd:double (normalize_numeric_datatype). The rewrite made
  typed constants lenient in novelty and sought the wrong o_type once
  indexed: {"@value":"25","@type":"xsd:long"} returned the xsd:integer row.
  The unused fluree_vocab::xsd::normalize_* helpers are removed.
- dt_compatible is removed; its callers compare datatypes exactly. It
  existed to let the rewritten constants match their numeric family, and
  also made SPARQL "25"^^xsd:integer lenient in novelty. Flake identity
  already includes the exact datatype, so the two write-side callers were
  lenient about pairs that never cancel.
- Bare Long/Double bound objects seek one (o_type, o_key) slice per numeric
  datatype in the predicate's observed-datatype stats, keyed by that
  family's encoding (untyped_numeric_slices). Several slices chain cursors
  (pending_cursors) where the sort order keeps them in index order; no
  provable slices fall back to the unnarrowed scan. Before, a predicate with
  more than one datatype pinned XSD_INTEGER / XSD_DOUBLE, and a bare double
  on an integer-typed predicate paired XSD_INTEGER with an f64 key.
- The predicate_object_count fast path sums a count per slice (it counted
  xsd:integer only).
- Novelty commit metadata types f:t, f:asserts and f:retracts as
  xsd:integer, as the indexer and bulk import already write them, so a
  typed constant finds a commit before and after indexing.

Tests: it_numeric_constant_matching (own binary) runs typed and bare cases
on SPARQL and JSON-LD, fast paths on and off, across novelty, indexed and
novelty-over-index lanes, and pins the COUNT fast-path stamp. Tests that
pinned the old leniency now assert the exact match and the non-match.

@aaj3f aaj3f left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@bplatz settling #1737 as "typed is exact, bare is lenient" is the right call, and taking out the @type rewrite and dt_compatible rather than working around them is great. Two things I'd want in before merging: a bare number with an unbound predicate (?s ?p 25) lost its seek and now walks the whole graph (about 99× slower at 1M flakes, growing with the ledger), and the policy lane still answers the old way. Claude-assisted review below:


✅ Approving to unblock you — with 2 things that need to be addressed before this merges.

Settling #1737 as "typed is exact, bare is lenient" and making the engine agree with the docs is the right call, and removing the JSON-LD @type rewrite and dt_compatible outright, instead of working around them, leaves both surfaces on one rule. The multi-slice seek is clean: o_type sorts ahead of o_key in all four V2 comparators, so the chained cursors do come out in index order, and the two stage.rs callers were never cancelling those lenient pairs anyway, since Flake equality includes dt (flake.rs:576-584). Two things need to land first:

  1. fluree-db-query/src/binary_scan.rs:2275, CRITICAL performance. With an unbound predicate (?s ?p 25) there are no slices, and this arm leaves the object unnarrowed, so the forced OPST scan walks the whole graph with a decode per row. I measured 0.78 ms at the merge base against 67 ms here on a 200k-flake ledger, and 3.5 ms against 343 ms at 1M flakes, same rows each time. Bound predicates holding decimals or big integers, and predicates without an observed set, fall into the same full walk. Seeking every integer and float o_type (at most fifteen point seeks, and object_slices_stay_ordered already allows it for OPST) keeps it a seek.
  2. fluree-db-query/src/binary_range.rs:524-536, the policy lane. When a policy might touch the predicate, the scan declines to the range fallback, whose bare-number prefilter pins one family's o_key. I reproduced 25 returning three of the four rows and 25.0 one of four, while the root view returns all four. The body scopes this provider out as internal-only, but it's the lane a policy-governed query falls to whenever the policy might touch the predicate. Not introduced here, but it's the one place the invariant this PR states still fails. (Inline on untyped_numeric_slices, since binary_range.rs isn't in the diff.)

On the other "Not in this PR" item: I agree the f:size divergence (sum of flake sizes in novelty, commit blob size in the index) is separable. It needs a decision about which of the two is the canonical value, and it has nothing to do with constant matching.

On #1988: its bound-constant test skips the indexed JSON-LD xsd:long/xsd:float constants (fluree-db-api/tests/it_typed_literal_index.rs:273-278) because of the bug this PR fixes. With both merged onto main, that test passes with the skip removed, so whichever of the two merges second should drop it. I'd land #1988 first, so the skip would come out here when this rebases.

Adherence to repo commitments:

  • Patterns/abstractions: ✔ One matching rule for both surfaces (no @type rewrite, no dt_compatible), and the multi-slice seek reuses BinaryCursor and per-cursor overlay windows rather than adding an operator. ⚠️ The range provider keeps its own bare-number rule (encode_bound_object_prefilter) instead of untyped_numeric_slices, which is how the policy lane diverges.
  • Performance (speed first, memory second): ✖ CRITICAL: bare numbers with an unbound predicate lost their seek (0.78 ms to 67 ms at 200k flakes, 3.5 ms to 343 ms at 1M, growing with the ledger rather than the matches), and predicates holding decimals or big integers, or without an observed set, fall to full-predicate walks. The multi-slice path itself adds k seeks and no per-row work, and the COUNT fast path's per-slice counts are fine.
  • Deployment targets: ⚠️ A query-engine change every host runs. I found nothing in fluree/solo that sends typed numeric constants or calls the removed dt_compatible and fluree_vocab::xsd::normalize_*, and its commit-history view reads f:t through variables, so the semantic change is safe for it to repin. The two findings above are what reach it: a hand-written or LLM-generated ?s ?p 25 on solo's Lambdas pays the whole-ledger walk under a 15-minute ceiling, and requests that carry a policy run on the lane that still disagrees.
  • Testing: ✔ it_numeric_constant_matching is its own [[test]] target and runs under CI's nextest; I ran it by name, and seeking only the first slice turned 31 checks red. The updated grp_query tests and the new unit tests pass. ⚠️ There's no unbound-predicate case and no policy-lane case, which are the two places the promise doesn't hold yet.
  • Conventions: ✔ Self-describing title and a thorough body, and datatypes.md and vocabulary.md now match the engine.

Verified locally at 63a0f31cf: it_numeric_constant_matching (1 passed), the seven untyped_numeric_slices unit tests and the three updated grp_query tests; the single-slice mutation (red, then restored); a throwaway timing probe at this head and at ba984c7f5; a throwaway policy-lane probe.

Approving now so you can merge without waiting on another pass from me — just be sure the unbound-predicate seek and the policy-lane fix are in before you do.

}
_ => object_slices = slices,
}
} else if untyped_number {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Must address before merge (CRITICAL, performance): a bare number with an unbound predicate now walks the whole graph — ?s ?p 25 went from one seek to a full OPST scan.

untyped_numeric_slices never runs for this shape: numeric_slices is filter.p_id.and_then(…) (:2214-2224), so with no predicate it's None, and this arm leaves o_type and o_key unset. open() forces OPST whenever only the object is bound (binary_scan.rs:707-716), and with nothing in the filter use_range is false, so the cursor is BinaryCursor::scan_all over the graph's whole OPST branch, decoding every object (a dictionary lookup for each string and IRI) to compare it with 25. Before, inferred_dt_sid needed the predicate too, so the constant fell through to value_to_otype_okey_simple, which pinned (XSD_INTEGER, 25): one seek, though it missed xsd:long, xsd:int and xsd:double, which is #1737 for this shape. The comment by the OPST choice already warns that an object open() can't encode "would devolve into a wide scan", and bare numbers now take that path.

I measured it with a throwaway test in grp_query: an in-memory ledger with three strings and an integer per subject, reindexed, best of five, debug build. At 200k flakes, SELECT ?s ?p WHERE { ?s ?p 25 } takes 0.78 ms at the merge base and 67 ms at this head; at 1M flakes it takes 3.5 ms at the base and 343 ms here, with the same rows each time, so the cost follows the ledger rather than the matches. The same arm also catches a bound predicate whose observed set holds DECIMAL or UNKNOWN (untyped_numeric_slices returns None at :4224) or that has no observed set at all; each of those was one seek and is now a full-predicate walk.

Everything needed to keep it a seek is already at hand. object_slices_stay_ordered is true for OPST whatever else is bound, so the chained cursors work with no predicate; seeking every integer and float o_type (at most fifteen point seeks) needs no stats at all, and if you'd rather narrow it, the novelty-aware stats_view computed just above has every predicate's observed datatypes through get_graph_properties(g_id). Decimals and big integers live in per-predicate arenas, so they stay out of reach without a predicate, as they were before this PR; for a bound predicate that holds decimals, decimal_object_key could contribute the arena slice rather than declining the whole seek. A ?s ?p 25 case in cases() would pin both the rows and the routing.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5c61f37. With no predicate (or no observed-datatype set), a bare number now seeks every integer and float o_type: fifteen point seeks, chained in OPST order. Decimals and big integers are sought by NumBig arena handle. With a predicate that's its own arena; with none it's every arena in the graph, and colliding handle numbers are left to the decoded-value filter. That covers the DECIMAL/UNKNOWN walk too. History and below-max_t reads scan just the NumBig o_type. The cursor for each slice now sets its own key range: use_range was computed before slicing, so unbound-predicate slices would each have been scan_all.

?s ?p 25 at 200k flakes, debug build, same rows: 90 ms at 63a0f31cf, 0.53 ms now. A bare number on an all-decimal predicate: 42 ms → 0.40 ms. it_numeric_constant_matching has ?s ?p 25 / ?s ?p 7.5e0 cases, and a bare_number_seek stamp that every indexed bare-number scan must proceed on. Dropping the unbound-predicate seek turns 12 checks red.

/// value, or decimal / big-number values, whose keys are per-predicate arena
/// handles — so the caller scans unnarrowed under the decoded-value filter.
/// `Some(vec![])` means no numeric datatype is present at all.
pub(crate) fn untyped_numeric_slices(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Must address before merge: under a policy, a bare number still misses the other numeric family, so "the same thing on every lane" doesn't hold on the lane policy-governed reads use.

fluree-db-query/src/binary_range.rs:524-536 builds its bare-number prefilter with encode_bound_object_prefilter, which leaves o_type open but pins o_key to one family's encoding (binary_scan.rs:3992-4004): 25 keeps the integer-family rows and drops "25.0"^^xsd:double, and 25.0 keeps only the doubles. The PR body leaves this provider out as used by internal lookups only, but BinaryScanOperator::open routes to open_range_fallback whenever a policy might touch the scanned predicate (binary_scan.rs:2057-2087), and whenever there's no binary store (:2050).

I checked it with a throwaway test in grp_policy built from it_policy_predicate_fast_lanes.rs's helpers: ex:v holding 25 as xsd:integer, xsd:int and xsd:long plus "25.0"^^xsd:double, reindexed, and a view with an allow rule f:onProperty ex:v under default allow, so every row stays visible. The root view returns all four subjects for both 25 and 25.0. The policed view stamps policy_predicate_scan as fallback:gate_declined and returns three for 25 (no ex:d) and one for 25.0 (only ex:d).

This isn't introduced here, but it's the one lane where the answer still depends on the lane, and in fluree/solo's query Lambda a request that carries a policy hits this gate for every predicate its policy might touch. Minimally the range provider could drop the o_key prefilter for an untyped Long/Double and let its value filter decide; better, it could seek these same slices so it stays a seek. A policy-lane pass in it_numeric_constant_matching (the view setup from it_policy_predicate_fast_lanes.rs is a few lines) would pin it.

Commenting here because fluree-db-query/src/binary_range.rs is not in this diff.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 5c61f37. The range provider now seeks the same untyped_numeric_slices, one cursor per slice, instead of pinning o_key to one family. it_numeric_constant_matching runs every case under a policy that may touch each scanned predicate but hides nothing, on the indexed lane and with novelty over the index. It asserts policy_predicate_scan declined. Restoring the old single-family prefilter turns 40 checks red, reproducing your three-of-four / one-of-four.

}
// The COUNT must be served by its fast path wherever that path
// applies (a fully indexed ledger), or this case pins nothing.
if c.jsonld.is_none() && fast_paths && lane == "indexed" && !count_proceeded {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

👍 Pinning the COUNT case to the predicate_object_count stamp is what makes the fast-path half of this test real.

I cut the scan down to its first slice locally and 31 checks went red across both surfaces, both fast-path settings and the two indexed lanes, while the fast COUNT stayed at 4, because it has its own slice loop. The stamp is what proves that loop, not the generic count, answered it.

…s, and under policy

A bare number now keeps a seek on every shape, and the policy lane matches
the same rows as the root lane.

- `?s ?p 25` seeks every integer and float o_type (OPST) instead of walking
  the whole graph. A predicate with no observed-datatype set seeks every
  numeric o_type the same way.
- Decimal and big-integer rows (DECIMAL / UNKNOWN in the observed set) are
  sought by NumBig arena handle: the bound predicate's arena, or every
  arena in the graph when the predicate is unbound. The handles reflect the
  index at max_t, so history and below-max_t reads scan the NumBig o_type
  under the decoded-value filter instead. Novelty-only values have no
  handle and arrive through the raw-flake lane.
- The range provider, where a policy-governed scan falls back, seeks the
  same slices. It previously pinned o_key to one family's encoding, so
  under a policy `25` dropped "25.0"^^xsd:double and `25.0` kept only it.
- Each slice cursor decides its own key range. The shared `use_range` was
  computed before slicing, so slices on an unbound predicate would each
  have scanned the whole graph.
- An integer gets a float slice whenever f64 holds it exactly, matching
  numeric_cmp (2^60 matched "1.152921504606847e18"^^xsd:double in novelty
  but not once indexed).
- `bare_number_seek` stamps whether a bare-number scan seeks or walks.
#1988 skipped these constants on its indexed lane because a JSON-LD
constant typed with a numeric subtype matched nothing once indexed. The
typed-constant fix on this branch closes that gap, so the checks run on
every lane.
@bplatz
bplatz merged commit 60a37e9 into main Sep 30, 2026
16 checks passed
@bplatz
bplatz deleted the fix/numeric-constant-matching branch September 30, 2026 10:51
@aaj3f aaj3f mentioned this pull request Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working as expected

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Numeric subtype leniency holds on novelty but not once indexed — the same constant-object query answers differently across an index publish

2 participants