Repository navigation
fix(query): typed numeric constants match exactly, bare numbers match any numeric datatype - #1990
Conversation
… any numeric datatype (#1737) A numeric constant in a triple pattern now means the same thing on every lane and both query surfaces: a typed literal ("25"^^xsd:long, or a JSON-LD value object with @type) matches its own datatype only, and a bare number (25, 25.0, a JSON number) matches every numeric datatype holding an equal value. - The JSON-LD query parser no longer rewrites a constant's @type to xsd:integer / xsd:double (normalize_numeric_datatype). The rewrite made typed constants lenient in novelty and sought the wrong o_type once indexed: {"@value":"25","@type":"xsd:long"} returned the xsd:integer row. The unused fluree_vocab::xsd::normalize_* helpers are removed. - dt_compatible is removed; its callers compare datatypes exactly. It existed to let the rewritten constants match their numeric family, and also made SPARQL "25"^^xsd:integer lenient in novelty. Flake identity already includes the exact datatype, so the two write-side callers were lenient about pairs that never cancel. - Bare Long/Double bound objects seek one (o_type, o_key) slice per numeric datatype in the predicate's observed-datatype stats, keyed by that family's encoding (untyped_numeric_slices). Several slices chain cursors (pending_cursors) where the sort order keeps them in index order; no provable slices fall back to the unnarrowed scan. Before, a predicate with more than one datatype pinned XSD_INTEGER / XSD_DOUBLE, and a bare double on an integer-typed predicate paired XSD_INTEGER with an f64 key. - The predicate_object_count fast path sums a count per slice (it counted xsd:integer only). - Novelty commit metadata types f:t, f:asserts and f:retracts as xsd:integer, as the indexer and bulk import already write them, so a typed constant finds a commit before and after indexing. Tests: it_numeric_constant_matching (own binary) runs typed and bare cases on SPARQL and JSON-LD, fast paths on and off, across novelty, indexed and novelty-over-index lanes, and pins the COUNT fast-path stamp. Tests that pinned the old leniency now assert the exact match and the non-match.
aaj3f
left a comment
There was a problem hiding this comment.
@bplatz settling #1737 as "typed is exact, bare is lenient" is the right call, and taking out the @type rewrite and dt_compatible rather than working around them is great. Two things I'd want in before merging: a bare number with an unbound predicate (?s ?p 25) lost its seek and now walks the whole graph (about 99× slower at 1M flakes, growing with the ledger), and the policy lane still answers the old way. Claude-assisted review below:
✅ Approving to unblock you — with 2 things that need to be addressed before this merges.
Settling #1737 as "typed is exact, bare is lenient" and making the engine agree with the docs is the right call, and removing the JSON-LD @type rewrite and dt_compatible outright, instead of working around them, leaves both surfaces on one rule. The multi-slice seek is clean: o_type sorts ahead of o_key in all four V2 comparators, so the chained cursors do come out in index order, and the two stage.rs callers were never cancelling those lenient pairs anyway, since Flake equality includes dt (flake.rs:576-584). Two things need to land first:
fluree-db-query/src/binary_scan.rs:2275, CRITICAL performance. With an unbound predicate (?s ?p 25) there are no slices, and this arm leaves the object unnarrowed, so the forced OPST scan walks the whole graph with a decode per row. I measured 0.78 ms at the merge base against 67 ms here on a 200k-flake ledger, and 3.5 ms against 343 ms at 1M flakes, same rows each time. Bound predicates holding decimals or big integers, and predicates without an observed set, fall into the same full walk. Seeking every integer and float o_type (at most fifteen point seeks, andobject_slices_stay_orderedalready allows it for OPST) keeps it a seek.fluree-db-query/src/binary_range.rs:524-536, the policy lane. When a policy might touch the predicate, the scan declines to the range fallback, whose bare-number prefilter pins one family'so_key. I reproduced25returning three of the four rows and25.0one of four, while the root view returns all four. The body scopes this provider out as internal-only, but it's the lane a policy-governed query falls to whenever the policy might touch the predicate. Not introduced here, but it's the one place the invariant this PR states still fails. (Inline onuntyped_numeric_slices, sincebinary_range.rsisn't in the diff.)
On the other "Not in this PR" item: I agree the f:size divergence (sum of flake sizes in novelty, commit blob size in the index) is separable. It needs a decision about which of the two is the canonical value, and it has nothing to do with constant matching.
On #1988: its bound-constant test skips the indexed JSON-LD xsd:long/xsd:float constants (fluree-db-api/tests/it_typed_literal_index.rs:273-278) because of the bug this PR fixes. With both merged onto main, that test passes with the skip removed, so whichever of the two merges second should drop it. I'd land #1988 first, so the skip would come out here when this rebases.
Adherence to repo commitments:
- Patterns/abstractions: ✔ One matching rule for both surfaces (no
@typerewrite, nodt_compatible), and the multi-slice seek reusesBinaryCursorand per-cursor overlay windows rather than adding an operator.⚠️ The range provider keeps its own bare-number rule (encode_bound_object_prefilter) instead ofuntyped_numeric_slices, which is how the policy lane diverges. - Performance (speed first, memory second): ✖ CRITICAL: bare numbers with an unbound predicate lost their seek (0.78 ms to 67 ms at 200k flakes, 3.5 ms to 343 ms at 1M, growing with the ledger rather than the matches), and predicates holding decimals or big integers, or without an observed set, fall to full-predicate walks. The multi-slice path itself adds k seeks and no per-row work, and the COUNT fast path's per-slice counts are fine.
- Deployment targets:
⚠️ A query-engine change every host runs. I found nothing in fluree/solo that sends typed numeric constants or calls the removeddt_compatibleandfluree_vocab::xsd::normalize_*, and its commit-history view readsf:tthrough variables, so the semantic change is safe for it to repin. The two findings above are what reach it: a hand-written or LLM-generated?s ?p 25on solo's Lambdas pays the whole-ledger walk under a 15-minute ceiling, and requests that carry a policy run on the lane that still disagrees. - Testing: ✔
it_numeric_constant_matchingis its own[[test]]target and runs under CI's nextest; I ran it by name, and seeking only the first slice turned 31 checks red. The updatedgrp_querytests and the new unit tests pass.⚠️ There's no unbound-predicate case and no policy-lane case, which are the two places the promise doesn't hold yet. - Conventions: ✔ Self-describing title and a thorough body, and
datatypes.mdandvocabulary.mdnow match the engine.
Verified locally at 63a0f31cf: it_numeric_constant_matching (1 passed), the seven untyped_numeric_slices unit tests and the three updated grp_query tests; the single-slice mutation (red, then restored); a throwaway timing probe at this head and at ba984c7f5; a throwaway policy-lane probe.
Approving now so you can merge without waiting on another pass from me — just be sure the unbound-predicate seek and the policy-lane fix are in before you do.
| } | ||
| _ => object_slices = slices, | ||
| } | ||
| } else if untyped_number { |
There was a problem hiding this comment.
🔴 Must address before merge (CRITICAL, performance): a bare number with an unbound predicate now walks the whole graph — ?s ?p 25 went from one seek to a full OPST scan.
untyped_numeric_slices never runs for this shape: numeric_slices is filter.p_id.and_then(…) (:2214-2224), so with no predicate it's None, and this arm leaves o_type and o_key unset. open() forces OPST whenever only the object is bound (binary_scan.rs:707-716), and with nothing in the filter use_range is false, so the cursor is BinaryCursor::scan_all over the graph's whole OPST branch, decoding every object (a dictionary lookup for each string and IRI) to compare it with 25. Before, inferred_dt_sid needed the predicate too, so the constant fell through to value_to_otype_okey_simple, which pinned (XSD_INTEGER, 25): one seek, though it missed xsd:long, xsd:int and xsd:double, which is #1737 for this shape. The comment by the OPST choice already warns that an object open() can't encode "would devolve into a wide scan", and bare numbers now take that path.
I measured it with a throwaway test in grp_query: an in-memory ledger with three strings and an integer per subject, reindexed, best of five, debug build. At 200k flakes, SELECT ?s ?p WHERE { ?s ?p 25 } takes 0.78 ms at the merge base and 67 ms at this head; at 1M flakes it takes 3.5 ms at the base and 343 ms here, with the same rows each time, so the cost follows the ledger rather than the matches. The same arm also catches a bound predicate whose observed set holds DECIMAL or UNKNOWN (untyped_numeric_slices returns None at :4224) or that has no observed set at all; each of those was one seek and is now a full-predicate walk.
Everything needed to keep it a seek is already at hand. object_slices_stay_ordered is true for OPST whatever else is bound, so the chained cursors work with no predicate; seeking every integer and float o_type (at most fifteen point seeks) needs no stats at all, and if you'd rather narrow it, the novelty-aware stats_view computed just above has every predicate's observed datatypes through get_graph_properties(g_id). Decimals and big integers live in per-predicate arenas, so they stay out of reach without a predicate, as they were before this PR; for a bound predicate that holds decimals, decimal_object_key could contribute the arena slice rather than declining the whole seek. A ?s ?p 25 case in cases() would pin both the rows and the routing.
There was a problem hiding this comment.
Fixed in 5c61f37. With no predicate (or no observed-datatype set), a bare number now seeks every integer and float o_type: fifteen point seeks, chained in OPST order. Decimals and big integers are sought by NumBig arena handle. With a predicate that's its own arena; with none it's every arena in the graph, and colliding handle numbers are left to the decoded-value filter. That covers the DECIMAL/UNKNOWN walk too. History and below-max_t reads scan just the NumBig o_type. The cursor for each slice now sets its own key range: use_range was computed before slicing, so unbound-predicate slices would each have been scan_all.
?s ?p 25 at 200k flakes, debug build, same rows: 90 ms at 63a0f31cf, 0.53 ms now. A bare number on an all-decimal predicate: 42 ms → 0.40 ms. it_numeric_constant_matching has ?s ?p 25 / ?s ?p 7.5e0 cases, and a bare_number_seek stamp that every indexed bare-number scan must proceed on. Dropping the unbound-predicate seek turns 12 checks red.
| /// value, or decimal / big-number values, whose keys are per-predicate arena | ||
| /// handles — so the caller scans unnarrowed under the decoded-value filter. | ||
| /// `Some(vec![])` means no numeric datatype is present at all. | ||
| pub(crate) fn untyped_numeric_slices( |
There was a problem hiding this comment.
🔴 Must address before merge: under a policy, a bare number still misses the other numeric family, so "the same thing on every lane" doesn't hold on the lane policy-governed reads use.
fluree-db-query/src/binary_range.rs:524-536 builds its bare-number prefilter with encode_bound_object_prefilter, which leaves o_type open but pins o_key to one family's encoding (binary_scan.rs:3992-4004): 25 keeps the integer-family rows and drops "25.0"^^xsd:double, and 25.0 keeps only the doubles. The PR body leaves this provider out as used by internal lookups only, but BinaryScanOperator::open routes to open_range_fallback whenever a policy might touch the scanned predicate (binary_scan.rs:2057-2087), and whenever there's no binary store (:2050).
I checked it with a throwaway test in grp_policy built from it_policy_predicate_fast_lanes.rs's helpers: ex:v holding 25 as xsd:integer, xsd:int and xsd:long plus "25.0"^^xsd:double, reindexed, and a view with an allow rule f:onProperty ex:v under default allow, so every row stays visible. The root view returns all four subjects for both 25 and 25.0. The policed view stamps policy_predicate_scan as fallback:gate_declined and returns three for 25 (no ex:d) and one for 25.0 (only ex:d).
This isn't introduced here, but it's the one lane where the answer still depends on the lane, and in fluree/solo's query Lambda a request that carries a policy hits this gate for every predicate its policy might touch. Minimally the range provider could drop the o_key prefilter for an untyped Long/Double and let its value filter decide; better, it could seek these same slices so it stays a seek. A policy-lane pass in it_numeric_constant_matching (the view setup from it_policy_predicate_fast_lanes.rs is a few lines) would pin it.
Commenting here because fluree-db-query/src/binary_range.rs is not in this diff.
There was a problem hiding this comment.
Fixed in 5c61f37. The range provider now seeks the same untyped_numeric_slices, one cursor per slice, instead of pinning o_key to one family. it_numeric_constant_matching runs every case under a policy that may touch each scanned predicate but hides nothing, on the indexed lane and with novelty over the index. It asserts policy_predicate_scan declined. Restoring the old single-family prefilter turns 40 checks red, reproducing your three-of-four / one-of-four.
| } | ||
| // The COUNT must be served by its fast path wherever that path | ||
| // applies (a fully indexed ledger), or this case pins nothing. | ||
| if c.jsonld.is_none() && fast_paths && lane == "indexed" && !count_proceeded { |
There was a problem hiding this comment.
👍 Pinning the COUNT case to the predicate_object_count stamp is what makes the fast-path half of this test real.
I cut the scan down to its first slice locally and 31 checks went red across both surfaces, both fast-path settings and the two indexed lanes, while the fast COUNT stayed at 4, because it has its own slice loop. The stamp is what proves that loop, not the generic count, answered it.
…s, and under policy A bare number now keeps a seek on every shape, and the policy lane matches the same rows as the root lane. - `?s ?p 25` seeks every integer and float o_type (OPST) instead of walking the whole graph. A predicate with no observed-datatype set seeks every numeric o_type the same way. - Decimal and big-integer rows (DECIMAL / UNKNOWN in the observed set) are sought by NumBig arena handle: the bound predicate's arena, or every arena in the graph when the predicate is unbound. The handles reflect the index at max_t, so history and below-max_t reads scan the NumBig o_type under the decoded-value filter instead. Novelty-only values have no handle and arrive through the raw-flake lane. - The range provider, where a policy-governed scan falls back, seeks the same slices. It previously pinned o_key to one family's encoding, so under a policy `25` dropped "25.0"^^xsd:double and `25.0` kept only it. - Each slice cursor decides its own key range. The shared `use_range` was computed before slicing, so slices on an unbound predicate would each have scanned the whole graph. - An integer gets a float slice whenever f64 holds it exactly, matching numeric_cmp (2^60 matched "1.152921504606847e18"^^xsd:double in novelty but not once indexed). - `bare_number_seek` stamps whether a bare-number scan seeks or walks.
#1988 skipped these constants on its indexed lane because a JSON-LD constant typed with a numeric subtype matched nothing once indexed. The typed-constant fix on this branch closes that gap, so the checks run on every lane.
Fixes #1737
Summary
A numeric constant in a triple pattern now means the same thing on every lane and both query surfaces:
"25"^^xsd:long,{"@value": "25", "@type": "xsd:long"}): matches that datatype only.25,25.0, a JSON number in a JSON-LD query): matches every numeric datatype holding an equal value —xsd:integer,xsd:long,xsd:int,xsd:double, …This resolves the product question #1737 was left open on (
triage:needs-decision): bare numbers stay lenient, typed numbers are exact. It is whatdocs/concepts/datatypes.mdalready documented; the engine now agrees with it.What was wrong
Measured on
mainwith one subject per datatype, all holding 25:"25"^^xsd:long25(integer + long + int + double on the predicate)25.0(predicate holds onlyxsd:long)COUNTof bare25, indexedThree causes:
parse_value_objectran@typethroughnormalize_numeric_datatype(xsd:long/int/… →xsd:integer,xsd:float→xsd:double), a v4-baseline assumption that storage normalized numeric subtypes. It doesn't: writes keep the declared datatype. The rewritten constant then sought the wrong o_type once indexed.dt_compatiblemade typedxsd:integer/xsd:doublelenient. It existed to let those rewritten constants match their family, and so also made SPARQL"25"^^xsd:integermatchxsd:longrows in novelty (but not once indexed).infer_exact_datatype_sid_from_statsnarrowed to the predicate's datatype only when it had exactly one; otherwisevalue_to_otype_okey_simplepinnedXSD_INTEGER/XSD_DOUBLE, so one exact seek missed every other numeric datatype. The single-datatype narrowing was also wrong across families: a bare double on an integer-typed predicate pairedXSD_INTEGERwith anf64key. Thepredicate_object_countfast path had the same encoding.binary_range.rs), whose bare-number prefilter lefto_typeopen but pinnedo_keyto one family's encoding: under a policy25dropped"25.0"^^xsd:double, and25.0kept only it.Changes
fluree-db-query/src/parse/node_map.rs: a constant@typeis kept as written. The now-unusedfluree_vocab::xsd::normalize_*helpers are removed.dt_compatibleis removed; its five callers compare datatypes exactly. Flake identity already includes the exact datatype, so the two write-side callers instage.rs(retraction meta copy, delete-verb classification) were lenient about pairs that never cancel each other.untyped_numeric_slices(binary_scan.rs): for a bareLong/Double, one(o_type, o_key)per numeric datatype the predicate carries, keyed by that family's encoding (i64for integer types,f64for float types; conversions followFlakeValue::numeric_cmp). The datatypes come from the predicate's observed-datatype stats; with no predicate (?s ?p 25) or no observed set, every integer and float o_type (fifteen point seeks).max_treads scan the NumBig o_type under the decoded-value filter instead. A value only in novelty has no handle and arrives through the raw-flake lane.pending_cursors), used only when every key component ahead ofo_typein the sort order is bound, so the output stays in index order. Each cursor sets its own key range. No row can hold the value → overlay-only. A non-finite value → the unnarrowed scan under the decoded-value filter.bare_number_seekrouting stamp records whether the scan seeks or walks.predicate_object_countfast path sums a count per slice, and declines to the generic count when a slice needs the decoded-value filter.commit_flakes.rs) typedf:t,f:assertsandf:retractsasxsd:long, while the indexer and bulk import write them asxsd:integer. With exact typed matching that divergence became visible (a typedf:tlookup answered differently before and after indexing), so novelty now usesxsd:integertoo. These flakes are generated from the commit envelope on load, so no existing data changes;docs/reference/vocabulary.mdlisted a third type forf:t(xsd:int) and is corrected.concepts/datatypes.mdspells out typed vs bare numeric matching;reference/vocabulary.mdcorrected as above.Tests
New standalone binary
it_numeric_constant_matching(it toggles the fast-path kill switch): thirteen cases — bare integer, double and decimal, each typed datatype, a single-datatype predicate, a bare number with no predicate, andCOUNT— each on SPARQL and JSON-LD (JSON-LD can't express a constant object under a variable predicate), with fast paths on and off. The predicate holds anxsd:decimal25 among its numeric rows, and a decimal on another predicate shares its arena handle number. Six lanes: novelty, indexed, indexed under a policy, a new numeric datatype written over the index (novelty-aware datatype stats), the same under a policy, and a time-travel read below the index'smax_t(no arena handles). Routing is pinned: the indexedCOUNTasserts thepredicate_object_countstamp proceeds, every indexed bare-number scan assertsbare_number_seekproceeds (a walk returns the same rows, only slower), and the policy lanes assert the scan fell back to the range provider.Updated tests that pinned the old leniency, each now asserting both sides:
it_query_datatype::datatype_query_explicit_typed_value_object_matches:xsd:integerconstant no longer matches thexsd:introw.it_query_misc::{untyped_value_matching_parity, indexed_untyped_value_matching_parity}: commitf:tisxsd:integer; anxsd:intconstant matches nothing,xsd:integermatches it, bare still matches.Each fix was reintroduced on its own to confirm the new test catches it: the JSON-LD
@typerewrite (typed JSON-LD cases, all lanes), lenient typedxsd:integer(typed integer, novelty lane), the scan pinningXSD_INTEGER(bare cases, indexed lane), the COUNT fast path pinning it (indexed COUNT, 1 vs 4), noveltyf:tasxsd:long(untyped_value_matching_parity), no seek without a predicate (bare_number_seekstamps), the range provider pinning one family'so_key(policy lanes, 40 checks), no NumBig slices (decimal row, 80 checks), no arena lookup without a predicate, and no NumBig slice for a below-max_tread (time-travel lane).?s ?p 25on a 200k-flake ledger (debug build, best of five, same rows): 90 ms at63a0f31cf, 0.53 ms now. A bare number on an all-decimal predicate: 42 ms → 0.40 ms.Unit tests for
untyped_numeric_slicesreplace the numeric cases of the old inference tests (observed set over counts, an empty set seeks every numeric datatype, NumBig handles vs the NumBig range, non-finite declines,2^53 + 1has no float slice and2^60has one).Behavior changes
"25"^^xsd:integerand"1.5"^^xsd:doublewritten out are exact. Bare25/1.5e0remain lenient.25finds"25.0"^^xsd:double), and matchxsd:decimalrows, as they did in novelty. The same holds with no predicate and under a policy.Notes for review
xsd:long/xsd:floatconstants init_typed_literal_index.rsis removed here, and the test passes.fluree-db-core/src/range.rsbecame the same exact comparison but isn't reached by these tests (unindexed file-backed ledgers answer through the scan operator's range fallback, which is covered).it_bounded_overlay_translation::warm_whole_product_short_circuits_and_respects_epochfailed once in a parallelgrp_queryrun on this branch, then passed alone three times and in the next fullgrp_queryrun; its failing assertion is a tracing-span check (which overlay-translation lane a read took), not a result.Not in this PR
f:sizeholds different values in novelty (sum of flake sizes) and in the index (commit blob size). Follow-up: Commitf:sizeholds a different value in novelty than in the index #1998