Skip to content

fix(query): evaluate grouped SELECT expressions once per group - #2006

Open
aaj3f wants to merge 80 commits into
mainfrom
fix/sparql-grouped-projection
Open

aaj3f wants to merge 80 commits into
mainfrom
fix/sparql-grouped-projection

Conversation

@aaj3f

@aaj3f aaj3f commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

This is more than #1978 needs on its own, and that's deliberate. What an expression means under GROUP BY, HAVING or implicit aggregation was decided in several places, so each shape broke its own way: #1362 with the expression as the group key, then #1978 with it in the SELECT. So rather than patch the formatter, grouping becomes one concept that owns its expressions: each is evaluated once per group, HAVING is just a filter, and a value the grouping can't supply is a plan-time error, not silently wrong rows. That leaves one place to reason about and tune. Fwiw, it's one of five PRs taking this approach, with #2007, #2008, #2009 and #2010.

Fixes #1978

Since the review

@bplatz approved c67dc3516. This push is that branch rebased onto 745736a04, with every review item addressed, plus a few pre-existing bugs in the same class that I decided to fold in rather than file. Briefly:

  • The eight inline comments (each thread has a reply): EXISTS-body variables are free over the group row, in HAVING and in a SELECT expression; a Cypher aggregating WITH / RETURN reads a projected node's property after the aggregation; a sh:sparql constraint that can't run follows its shape's severity and the graph's mode; at the top level, the generated binds now run after a trailing VALUES, as they do in a sub-SELECT; update WHERE errors name the variable; the stream refuses a bind-stage grouped read before it starts; and the second having doc example is fixed.
  • The four suggested follow-ups are in this PR too: ASK and CONSTRUCT group to spec instead of refusing; ORDER BY ?nosuch orders nothing instead of failing with a 500; every fast path that gates on the grouping's binds has a routing-stamp canary pair; and HAVING reads a SELECT expression's alias (my call, over a diagnostic; it's an extension, documented in the compatibility reference).
  • Pre-existing bugs, folded in. The aggregate fast paths ignored a top-level trailing VALUES; ASK dropped OFFSET and LIMIT; ASK, CONSTRUCT and DESCRIBE didn't parse a trailing VALUES; a Cypher WITH's WHERE filtered before its SKIP / LIMIT; a WITH inside CALL (p) { … } lost its correlation with p; and three SHACL bugs: an sh:sparql-only node shape lost its sh:severity, sh:and / sh:or / sh:xone lists written by SPARQL UPDATE didn't resolve, and a nested shape's Warning results didn't count toward conformance. There are also a handful of smaller ones (error texts, JSON-LD ask options, per-group lists on indexed ledgers). Each of these changes what main returns today, so folding them in was a deliberate call rather than a review fix.

Every one of these that changes behavior has a row in the table below.

What was wrong, and the fix

#1978 shows up in the SPARQL-results formatter, but the defect is upstream of it, in how we lower grouped SELECT expressions. A SELECT expression at a grouped query level is an Extend over the group rows, evaluated once per group after HAVING, in SELECT order (SPARQL 1.1 §18.2.4.4). Our SPARQL lowering instead made every aggregate-free SELECT expression a per-solution BIND before grouping, so the grouping stage carried the alias as a per-group list (Binding::Grouped) — and each reader of that list invented its own meaning for it:

  • the SPARQL-results formatters (JSON, XML, NDJSON, CLI table) expanded it into one row per solution — 5 rows for 2 groups in the issue, a cartesian product for two expression columns, and a lost row for an implicit group over no solutions;
  • FILTER and HAVING panicked a debug build and read it as unbound in release;
  • a join on it matched anything or nothing, depending on the operator — substituted into a nested-loop join's pattern, the list left the slot a variable, which matched anything; compared by value, as in a hash join or a correlated subquery's merge, a list equals no scalar, so it matched nothing (the JSON-LD subquery and SHACL $this rows below). A GROUP BY over it merged every group into one, and an UPDATE template refused it.

So the class is: after the group step, something reads a WHERE variable that is neither a GROUP BY key nor an aggregate output. Rather than teach the formatter to cope, this PR closes that class structurally:

  1. Lowering (fluree-db-sparql). A level's SELECT expressions are now placed after its modifiers are lowered, from one shared rule (ir::SelectExprPlacer). In a grouping level (GROUP BY, or an aggregate anywhere in SELECT, HAVING or ORDER BY) every expression is a per-group Extend, in one list in SELECT order — except an expression whose alias is itself a group key (the GROUP BY (LCASE(?a)) shortcut), which stays a WHERE bind. SELECT * of a grouping level projects its GROUP BY keys (nothing, under implicit grouping). The top level and sub-SELECTs share one function, lower_select_level.
  2. IR (fluree-db-query). Per-group binds moved from Aggregation to Grouping, so a dedup-only GROUP BY can carry them (and removing the field turned every reader into a compile error). QueryOutput::Select now carries an UngroupedProjection: Reject by default, PerGroupList only for the JSON-LD user-query entry. A sub-query's projection is a plain variable list, so it can't carry a per-group list at all.
  3. Plan-time check. Grouping::first_ungrouped_read, enforced in apply_solution_modifiers (which the top-level and sub-query pipelines share), rejects any HAVING, bind, ORDER BY or Reject-projection read of such a variable with a 4xx, in every build, and replaces two ad-hoc ORDER BY checks. The lowerers rewrite HAVING and ORDER BY reads of such a variable to SAMPLE(?v) first (sample_ungrouped_reads, item 5), so for those the check is a backstop: it catches SELECT-expression and projection reads under Reject, and IR from a lowerer that doesn't apply the rewrite (Cypher: a property of a node the clause doesn't project, WITH e.area AS a, count(*) AS c ORDER BY e.name). A Grouped value reaching scalar evaluation, a join or an OPTIONAL substitution is now an internal error instead of a debug panic / silent unbound.
  4. Formatters. The SPARQL-results writers stop disaggregating: a list cell is a FormatError, as a Cypher path or list already was, and disaggregate_row, the XML write_grouped_rows and the CLI table fallbacks are deleted. The streaming endpoint refuses a JSON-LD per-group list before the 200.

5 and 6 below are calls I made on spec edges and on JSON-LD, and 7 is the rest of the class, which an adversarial review pass turned up — one piece of it, trailing VALUES, is also a call of mine. So that's probably where review time is best spent; happy to talk through any of them.

  1. Spec edges, on SPARQL and JSON-LD through two shared IR helpers:
    • a HAVING or ORDER BY read of a non-key variable means SAMPLE(?v) (§18.2.4.1);
    • HAVING without grouping is a Filter over the solutions (§18.2.4.2), which can't see the SELECT expressions;
    • the validator's GROUP BY projection-scope check (V4) uses the same definition of "grouped" as lowering, so a HAVING/ORDER BY-only aggregate triggers it;
    • an aggregate over a same-level SELECT alias is a named validator error (V009, AggregateOverSelectAlias); per the spec it would see the alias unbound.
  2. JSON-LD. A select expression that reads only group keys, aggregate outputs or earlier per-group aliases (or a constant) is now evaluated once per group, like SPARQL, instead of becoming a per-group list. An expression over a non-key variable keeps the documented per-group list, and so does a key-only alias that such an expression reads (it moves before grouping with its reader). A subquery projecting a non-key variable of its grouping is now an error — it used to return the list to the enclosing query, where a join on it silently matched nothing.
  3. The rest of the class, from that review pass:
    • a trailing VALUES clause still joins before grouping, as at base, so "VALUES as a parameter" keeps restricting what the aggregates count (I decided to keep base placement here, a deviation from §18.2.4.3 now documented in docs/reference/compatibility.md). HAVING still reads a VALUES variable as unbound, as it would in the spec's order, unless the WHERE binds it or it's a GROUP BY key or an aggregate output — at the top level and in a sub-SELECT (the implicit SAMPLE had kept every group for HAVING (?v = 1) VALUES ?v { 1 });
    • HAVING evaluates with FILTER's evaluator, so HAVING (EXISTS …) works (it was never resolved: always false), and a Cypher metadata read in an aggregating WITH … WHERE sees policy-filtered flakes (it read as empty under a view policy);
    • SPARQL ASK / CONSTRUCT and JSON-LD ask group to spec instead of dropping GROUP BY / HAVING (the first version of this PR refused them; the review asked for the spec semantics);
    • SHACL: a pre-bound variable ($this, $PATH) groups any level that projects it, so the sub-SELECT family that must project $this evaluates, and a shape query the planner rejects is a sh:sparql constraint failure naming the constraint, its shape and the variable, decided by the shape's severity and the graph's mode;
    • the stream refuses select * under groupBy before the response starts when /query would return per-group lists for it (the grouping's GroupByOperator lane). Main streamed those lists as cartesian rows, one per list element;
    • a grouped projection of a variable nothing binds is one named 400 on /query and on the stream ("projected variable ?x is unbound"). Main failed it on both with "Variable not found: Selected variable VarId(n) not found in query schema" — a 500 on the tracked /query path, and on the stream an error record after the 200;
    • grouped-read plan errors name variables (?x), never internal ids, and describe a synthetic variable (a Cypher property access) instead of printing its internal name.

Behavior changes

Class key: RESULTS CHANGE = the query still runs and returns different rows or values; NOW ERRORS = a query that runs today fails (4xx); ERROR → ROWS = a query that fails today returns rows; ERROR → COMMIT / COMMIT → ERROR = a write SHACL rejects today now commits, or the reverse. The shapes that don't change (each pinned by a test) and the Rust API changes are folded under the table.

Surface Shape Before After Class
SPARQL grouped aggregate-free SELECT expression (#1978) one row per solution; Σk² rows for two expression columns one row per group RESULTS CHANGE
SPARQL implicit group over no solutions + SELECT expression, e.g. ("x" AS ?c) (COUNT(*) AS ?n) 0 rows 1 row (x, 0) RESULTS CHANGE
SPARQL ORDER BY a SELECT alias of a grouped query error sorts ERROR → ROWS
SPARQL sub-SELECT expression read by an outer FILTER / GROUP BY / aggregate / DISTINCT / join debug panic, merged groups, or silent wrong rows correct RESULTS CHANGE
SPARQL UPDATE template reading a grouped sub-SELECT expression transaction error commits ERROR → ROWS
SPARQL, JSON-LD HAVING reading a non-key variable debug panic; release / streaming lane: 0 rows SAMPLE(?v) semantics RESULTS CHANGE (groups can appear)
SPARQL, JSON-LD ORDER BY a non-key variable 400 ("Sort variable … not found") sorted by SAMPLE(?v) ERROR → ROWS
SPARQL, JSON-LD HAVING without GROUP BY or an aggregate silently ignored filter over the solutions RESULTS CHANGE (fewer rows)
SPARQL SELECT * with an aggregate only in HAVING / ORDER BY 1 empty row, one row per solution, or a join panic (sub-SELECT) 1 row, no variables RESULTS CHANGE (traditional lane only)
SPARQL projected variable or expression with a HAVING/ORDER BY-only aggregate expanded rows, or silently grouped by the projected variable V4 error ("… is projected but is neither a GROUP BY key nor aggregated") NOW ERRORS
SPARQL aggregate over a same-level SELECT alias, e.g. (IF(…) AS ?seg) (COUNT(?seg) AS ?c) counted solutions; alias expanded V009 error naming the alias NOW ERRORS
SPARQL grouped HAVING (EXISTS …) / (NOT EXISTS …), and an EXISTS SELECT expression EXISTS never resolved: no group kept / every group kept evaluated per group row: keys and aggregates bound, any other body variable free RESULTS CHANGE
Cypher aggregating WITH … WHERE (lowered to HAVING) that reads node metadata (keys, labels, properties, …) under a non-root view policy read as empty: HAVING used the synchronous readers, which are fail-closed under a policy resolved through the policy filter, as in a MATCH … WHERE RESULTS CHANGE
SPARQL ASK / CONSTRUCT with GROUP BY, HAVING or an aggregate ORDER BY clause dropped: ASK { … } HAVING (?a = "Nope") was true; CONSTRUCT built every solution's triples grouped to spec: ASK is true when some group passes HAVING; CONSTRUCT instantiates its template once per group, and a template variable that is not a GROUP BY key is unbound, so its triples are skipped. DESCRIBE still refuses them (400) RESULTS CHANGE
JSON-LD key-only (or constant) select expression under groupBy per-group list scalar RESULTS CHANGE (cell shape)
JSON-LD ask with groupBy / having option dropped (having that rejects every group answered true) grouped: true when some group passes having RESULTS CHANGE
JSON-LD subquery projecting a non-key variable under groupBy list crossed into the outer query; joins on it matched nothing error NOW ERRORS
JSON-LD per-group list (explicit, or through select * under groupBy when /query returns lists, i.e. on the GroupByOperator lane) on the NDJSON stream / SPARQL JSON / SPARQL XML cartesian rows 4xx before the stream starts (naming the columns; use /query or project the keys) / FormatError NOW ERRORS
JSON-LD projected variable that nothing binds, under groupBy /query: "Variable not found: Selected variable VarId(n) not found in query schema" (500 on the tracked path); stream: the same, after the 200 400 "projected variable ?x is unbound: nothing in the query binds it" on both, the stream before it starts error text / status
JSON-LD update template reading a grouped subquery's key-only select expression transaction error ("Object cannot be a grouped value") commits, one value per group ERROR → ROWS
SHACL sh:select (lowered like SPARQL, but the SPARQL validator does not run) with HAVING and no grouping, or HAVING over a non-key variable HAVING dropped: every solution a violation; non-key read: debug panic, release no violation the SPARQL semantics above: a filter; SAMPLE(?v) RESULTS CHANGE (a transaction's verdict can flip either way)
SHACL sh:select projecting a non-key variable of a grouped level, e.g. SELECT $this ?value … GROUP BY $this (V4 would reject it in a SPARQL query) violation reported at the focus node a sh:sparql constraint failure naming the constraint, its shape and ?value, decided by severity and mode (row below) NOW ERRORS on a Violation shape in a reject-mode graph; logged otherwise
SHACL a sub-SELECT grouped by another variable, projecting $this (every sub-SELECT must), e.g. { SELECT $this ?v (COUNT(?x) AS ?c) … GROUP BY ?v } FILTER(?c > 1) never fired ($this came back as a list, which joined nothing) evaluates: $this groups the level (it is constant per evaluation) RESULTS CHANGE (this constraint family now fires; it never fired before)
SPARQL, JSON-LD, Cypher HAVING (Cypher: an aggregating WITH … WHERE) reading a SELECT expression's alias, e.g. (COUNT(?e) + 0 AS ?n) … HAVING (?n > 1), WITH p, count(f) + 0 AS c WHERE c > 1 alias unbound in HAVING: no group kept the expression runs once per group before HAVING, and HAVING reads its value. A Fluree extension, documented in the compatibility reference RESULTS CHANGE
Cypher aggregating WITH … WHERE / ORDER BY or aggregating RETURN … ORDER BY reading a property of a node the clause projects, e.g. WITH p, count(f) AS c WHERE p.age > 30, ORDER BY p.age + 1 WHERE: []; ORDER BY: 400 the property is read after the aggregation, as a following WITH p, c WHERE p.age > 30 reads it, so a read in WHERE or ORDER BY doesn't change the clause's aggregates. A property with several values joins every value, as MATCH … WHERE does: in WHERE the row comes once per value that passes; in ORDER BY once per value, and WITH DISTINCT keeps those copies (the sort key is projected with them) while RETURN DISTINCT removes them; an aggregate in a later clause counts each copy. A read inside an aggregate's argument (collect(p.age)) is still joined before grouping, so it repeats the group's rows for every aggregate of that clause (unchanged; follow-up) RESULTS CHANGE / ERROR → ROWS
Cypher aggregating WITH … ORDER BY / RETURN … ORDER BY on an expression over its aggregates, e.g. ORDER BY c + 1, ORDER BY -c 400 "an ORDER BY key is neither a GROUP BY key nor an aggregate result" sorted ERROR → ROWS
Cypher ORDER BY a value read from a collect() list, e.g. WITH p, collect(f) AS fs ORDER BY size(fs), ORDER BY any(x IN fs WHERE …) 400 "ORDER BY on a collect() list is not supported in v1" sorted after the aggregation; a key whose value is the list itself, or a list built from one (fs, tail(fs)), is still a 400 ERROR → ROWS
JSON-LD orderBy holding an expression or an aggregate, e.g. "(desc (count ?e))", ["desc", "(count ?e)"] "orderBy must be an array of objects with 'var' field" / "Invalid variable syntax" still a 400; the message lists the forms orderBy takes and says to sort on an alias, (as <expr> ?k) error text
SPARQL, JSON-LD ORDER BY a variable nothing binds, e.g. ORDER BY ?nosuch, grouped or not, on /query and the stream 500 "Sort variable VarId(n) not found in query schema" the key orders nothing; ordered by the other keys ERROR → ROWS
SPARQL top-level trailing VALUES read by an aggregate input, a GROUP BY expression or an ungrouped SELECT expression, e.g. SELECT (SUM(?n * ?v) AS ?s) … VALUES ?v { 2 } the expression ran before the join: ?v unbound, SUM = 0 (a sub-SELECT already saw ?v) runs after the join, as in a sub-SELECT RESULTS CHANGE
SPARQL trailing VALUES on an indexed ledger, on a shape a fast path serves (COUNT(*), the GROUP BY count top-k, the star top-k, SUM, the per-predicate directory count) the fast path ignored the VALUES, e.g. COUNT(*) … VALUES ?a { "Net" } gave 6, not 3 the generic lane applies it RESULTS CHANGE
SPARQL trailing VALUES row naming an f:reifies* IRI accepted (the join matched nothing) refused like an in-WHERE VALUES row NOW ERRORS
SHACL sh:sparql constraint whose own query fails (it does not parse, lower or plan, $PATH has nothing to bind, or it fails on its own terms, e.g. SUM(?nosuch) or an unknown function) on a Warning / Info shape, or in a warn-mode graph rejected every write to the shape's targets (the grouped-projection plan failure is new in this PR) logged once per distinct failure, with the number of focus nodes; the write commits. A Violation shape in a reject-mode graph still fails the write, with a message that names the shape. Budgets, storage / catalog / policy access, data state and internal faults still fail the write whatever the severity ERROR → COMMIT
SHACL an sh:sparql constraint that cannot run, in a shape checked as a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, sh:qualifiedValueShape) took the nested shape's severity (Violation by default): rejected the write under a Warning outer shape, and was only logged under a Violation outer shape when the nested shape was a Warning takes the severity of the outermost shape that reports the result: logged under a Warning or Info outer shape, or in a warn-mode graph; fails the write under a Violation outer shape in a reject-mode graph, whatever the nested shape's severity ERROR → COMMIT / COMMIT → ERROR
JSON-LD update, Cypher writes an update WHERE whose subquery is a grouped-read error (SPARQL UPDATE's validator refuses that query first, with a named message) message said VarId(0) names the variable error text
JSON-LD grouped SELECT expression reading a non-key variable, on the stream 400 on /query; on the stream, the error arrived after the 200 the same 400 on the stream, before it starts error timing
SPARQL, JSON-LD, sh:select an aggregate over a variable nothing binds before the grouping, e.g. SUM(?nosuch) "Aggregate input variable VarId(n) not found in schema", a 500 on the tracked path "an aggregate reads variable ?nosuch, which is unbound: nothing before the grouping binds it", a 400 on both paths error text / status
JSON-LD, sh:select an aggregate whose output the WHERE binds, or two aggregates with one output (SPARQL's validator refuses both first) "Aggregate output variable VarId(0) already exists in schema", "Duplicate aggregate output variable VarId(1)" "aggregate output variable ?a is already bound in the WHERE pattern", "variable ?n is the output of more than one aggregate" error text
Cypher an aggregate of a sibling aggregate's output, WITH p, count(f) AS c, sum(c) AS s "Aggregate input variable VarId(2) not found in schema" the named unbound-aggregate-input error, a 400 error text / status
SPARQL ASK … OFFSET n and ASK … LIMIT 0 the modifiers were dropped: OFFSET 2 over two solutions and LIMIT 0 answered true true only when a solution remains after them RESULTS CHANGE
SPARQL a trailing VALUES after ASK, CONSTRUCT or DESCRIBE (the grammar allows one after every query form) parse error joined with the WHERE's solutions, as after SELECT; DESCRIBE joins it right after its WHERE, so DESCRIBE ?x VALUES ?x { ex:a } describes ex:a ERROR → ROWS
JSON-LD ask with offset, limit or values; construct with values ignored applied, as for select RESULTS CHANGE
JSON-LD a per-group list (projection of a non-key variable) of IRIs or literals on an indexed ledger "format_binding called without QueryResult for encoded IRI binding" (or EncodedLit), from both JSON-LD formatters rendered (typed JSON already rendered it) ERROR → ROWS
Cypher a WITH with both a WHERE and a SKIP or LIMIT, e.g. UNWIND range(1, 10) AS x WITH x ORDER BY x LIMIT 5 WHERE x > 3 the WHERE filtered before the sort and the slice: 4 to 8 the WHERE filters the clause's sliced rows, as openCypher places it: 4 and 5. Plain, aggregating and DISTINCT WITHs alike; a WITH without a slice is unchanged RESULTS CHANGE
Cypher after a sliced aggregating or DISTINCT WITH, a WHERE reading a variable the clause doesn't project, e.g. WITH DISTINCT p.age AS a ORDER BY a LIMIT 2 WHERE p.name = 'x' aggregating: a grouped-read plan error; DISTINCT: filtered before the DISTINCT an error: "WITH … WHERE reads p, which is out of scope: after an aggregating or DISTINCT WITH, its WHERE sees only the variables the WITH projects" NOW ERRORS (DISTINCT) / error text (aggregating)
Cypher inside CALL (p) { … }, a WITH (or an aggregating RETURN that sorts on an expression) whose body reads the import, e.g. MATCH (p:P) CALL (p) { MATCH (p)-[:knows]->(f) WITH f RETURN f.name AS fn } RETURN p.name, fn every outer p crossed with the body's whole result: 24 rows instead of 6 on a four-person fixture. A slice, DISTINCT or aggregate in the body ran once for all of them, so WITH f ORDER BY f.age LIMIT 1 gave every p the same friend the body runs per import: p's own friends, and a slice, DISTINCT or aggregate applies per p. As for an aggregating CALL RETURN (unchanged), an import whose body matches nothing gets no row from an aggregate, even one with no grouping key; OPTIONAL MATCH keeps it as 0 (follow-up) RESULTS CHANGE
SHACL a node shape whose only constraint is sh:sparql (nothing but its rdf:type sh:NodeShape registers it), with sh:severity sh:Warning or sh:Info the severity was dropped and the shape compiled as a Violation shape, so a result rejected the write the shape's severity applies: logged, and the write commits. The same holds for its sh:message, sh:name and sh:description, and for a shape whose registering statement is in a later shapes graph ERROR → COMMIT
SHACL a nested shape (through sh:node, sh:not, sh:and, sh:or, sh:xone or sh:qualifiedValueShape, at node level or on a property's values) that reports a Warning or Info result for the node only Violation results counted, so the node conformed: sh:node and the logical combinators passed, sh:not failed, and a qualified count counted the value the node does not conform, whatever the severity of the result (SHACL §3.5). The referencing shape reports at its own severity: a Violation shape rejects the write, a Warning shape logs it, and sh:not is satisfied. The sh:sparql-only severity fix (row above) made this reachable for a nested sh:sparql-only Warning shape, which used to compile as Violation COMMIT → ERROR / ERROR → COMMIT
SHACL an sh:and / sh:or / sh:xone list written by a SPARQL UPDATE (( ex:A ex:B ), stored as rdf:first / rdf:rest) "Referenced shape …#coll0 could not be resolved", on every focus node resolved, like the same list written in JSON-LD or Turtle RESULTS CHANGE
JSON-LD on an indexed ledger, a per-group list beside a key and its count, in the shape the GROUP BY ?o count top-k or the per-predicate directory count serves, e.g. "select": ["?a", "?e", "(as (count ?e) ?n)"] with orderBy (desc ?n) / limit the top-k answered null for the list; the directory count failed ("Projected variable not in child schema") the generic lane answers, list included RESULTS CHANGE / ERROR → ROWS
Shapes that don't change (pinned), and the Rust API changes
Surface Shape Before After Class
SPARQL grouped query or sub-SELECT with a trailing VALUES clause, e.g. … GROUP BY ?a VALUES ?e { ex:e1 } (VALUES as a parameter) joined before grouping: restricts what the aggregates count ((Net, 1)); VALUES ?x { 1 2 } doubles COUNT the same (the generated binds now run after the join at the top level too) unchanged
SPARQL HAVING reading a trailing VALUES variable that the WHERE does not bind and that is not a GROUP BY key, e.g. GROUP BY ?a HAVING (?v = 1) VALUES ?v { 1 } read as unbound: no group (a debug build on the traditional grouping lane panicked) read as unbound: no group, top level and sub-SELECT (the implicit SAMPLE does not reach it). A VALUES variable that is a key is read as the key: GROUP BY ?v HAVING (?v = 1) VALUES ?v { 1 2 } gives (1, 6), as before unchanged
JSON-LD key-only alias read by a per-solution expression, e.g. (as (strlen ?a) ?len) + (as (+ ?len (strlen (str ?e))) ?x) both per-group lists both per-group lists (the alias moves before grouping with its reader) unchanged
JSON-LD having with an ["exists", …] form not accepted (no EXISTS form in having) not accepted (pinned) unchanged
JSON-LD bare aggregate column name ((count ?t) → ?count) ?count ?count (now pinned) unchanged
Rust API Grouping::assemble Option<Grouping> Result<Option<Grouping>, GroupingError> (HAVING / binds with no key and no aggregate are an error, not dropped) —
Rust API Aggregation.binds field moved to the Grouping variants; Grouping::binds() unchanged —
Rust API QueryOutput::Select { projection, restriction } + ungrouped: UngroupedProjection —
Rust API QueryError — new variant UngroupedRead(ir::UngroupedRead); QueryError::name_variables(&VarRegistry) turns it into InvalidQuery —
Rust API ShaclError::SparqlConstraint { constraint, .. } constraint: Sid constraint: String, the constraint's IRI —
Rust API HavingOperator::new (child, expr) (child, expr, planning) —
Rust API Query.post_values Option<Pattern> Option<PostValues> (vars, rows, then: the binds after the join) —
Rust API fluree_db_shacl — new ConstraintFailure, ConstraintFailures; ShaclEngine::with_constraint_failures —
Rust API ir::ConstructTemplate — new substitute_var —
Rust API fluree_db_sparql::ast::{AskQuery, ConstructQuery, DescribeQuery} — new values: Option<Box<GraphPattern>> (the trailing VALUES clause, as on SelectQuery); a struct literal outside the crate must set it —
Rust API ir::ReadStage (new in this PR) — also UnboundAggregateInput, BoundAggregateOutput, RepeatedAggregateOutput: the grouping-stage plan errors about one variable, named by QueryError::name_variables —

Worth calling out: stored artifacts. Saved queries that use any shape in the table above behave differently after upgrade — the RESULTS CHANGE rows change silently, and the NOW ERRORS rows start failing. And since SHACL shapes are stored in the ledger, a shape whose sh:select has one of the SHACL shapes above changes which transactions commit. What that means for solo is just below.

Downstream: fluree/solo

The short version for solo is that its own queries are fine, but stored artifacts may not be, and we can't enumerate those from the repo.

Per a survey of fluree/solo's checked-in and generated queries (at solo main 0262cf8, and re-diffed since with no relevant change), no static solo query changes result: every grouped query in the repo projects only keys, aggregates or expressions of aggregates, and solo's reads of the implicit ?count column name keep working (pinned here). The must-not-change guards in this PR cover solo's static grouped shapes.

Stored artifacts do change, though: saved dashboards, decks and publications. Publications in particular write SPARQL-JSON rows into customers' Iceberg tables on every run and have no update endpoint. So before we bump the db pin, we should scan the stored dashboard/deck specs and the _system TablePublication queries for the shapes in the table above, in particular:

  • grouped aggregate-free SELECT expressions: warehouse row counts drop to one per group on the next run (replace mode rewrites; append mode appends de-duplicated rows next to exploded history);
  • constant-label KPIs over possibly-empty input: a run refused as empty now publishes one row;
  • HAVING without GROUP BY or an aggregate: now filters;
  • HAVING on a non-key variable: now evaluates SAMPLE, so groups can appear;
  • ORDER BY a non-key variable: now succeeds;
  • a projected variable/expression with a HAVING/ORDER BY-only aggregate, and an aggregate over a same-level SELECT alias: now a named error — recreate with an explicit GROUP BY, or aggregate the expression;
  • grouped HAVING (EXISTS …) / (NOT EXISTS …): now evaluated per group;
  • ASK or CONSTRUCT with GROUP BY or HAVING: now grouped to spec;
  • JSON-LD key-only select expressions: lists become scalars;
  • HAVING reading a SELECT expression's alias: now filters on its value (was: no rows);
  • a top-level trailing VALUES read by an aggregate input or a SELECT expression: now read;
  • a trailing VALUES on a fast-path shape (indexed ledgers): now applied;
  • ASK with OFFSET or LIMIT 0: now applied (was: ignored);
  • a trailing VALUES after ASK, CONSTRUCT or DESCRIBE: now parsed and joined (was: a parse error);
  • JSON-LD ask with offset, limit or values, and construct with values: now applied (was: ignored). A malformed offset or limit on ask (a string, a negative or fractional number, one beyond u64) is now a 400, as for select (was: ignored);
  • JSON-LD per-group lists on indexed ledgers: now rendered (was: a format error); beside a count fast-path shape, now a list (was: null, or an error);
  • a Cypher WITH with both a WHERE and a SKIP or LIMIT: the WHERE now filters the sliced rows;
  • a Cypher WITH inside a CALL (p) { … } body now stays correlated with p.

Solo generates neither of those two Cypher shapes. Its Cypher surfaces (the client's cypherQuery / cypherWrite, the query and transact lambdas, the Cypher import) pass a caller's statement through, and a grep of its Rust, TypeScript, Markdown and JSON at solo main c1efea7c6b finds no WITH … LIMIT, no WITH … SKIP and no Cypher CALL in its code, prompts or docs. So those reach solo only through user-written statements.

SHACL is the other exposure (the SHACL and Rust-API greps below are at solo main 3bf84bc, whose db pin is still ba984c7f5). Solo's query Lambda runs validate over user-authored shapes and bounds sh:sparql evaluation with a fuel ceiling (fluree-lambda-query/src/handler.rs:1358, :1511); solo itself authors no sh:select shapes (grep of 3bf84bc: the only sh:sparql mentions are those two comments). So it's user-authored constraints that move:

  • a SPARQL constraint that projects a non-key variable under grouping now fails validation on a Violation shape in a reject-mode graph (a SHACL failure naming the shape and the variable), and is logged otherwise, instead of reporting a violation at the focus node;
  • a constraint whose sub-SELECT groups by another variable and projects $this now fires — it never fired before, so writes it should have refused start being refused;
  • constraints using HAVING without grouping, HAVING over a non-key variable, or HAVING with EXISTS change their verdicts per the SHACL rows of the behavior table;
  • a constraint that cannot run follows its shape's sh:severity and the graph's mode. On a Warning or Info shape, or in a warn-mode graph, it's logged and writes commit; at base, parse and lower failures rejected every write; in a shape checked as a nested shape (sh:node and the logical constraints), the outermost reporting shape's severity applies;
  • a Warning or Info node shape whose only constraint is sh:sparql now commits with a warning (was: rejected). Solo authors no sh:sparql shape (its query lambda only bounds user-authored ones), and the node shapes it generates carry sh:targetClass, which already registered them;
  • sh:and / sh:or / sh:xone lists written by SPARQL UPDATE now resolve. Solo writes shapes as JSON-LD and authors no sh:and, sh:or or sh:xone;
  • a nested shape's Warning or Info results now count toward conformance. Solo's only nested shape is the vocabulary membership shape (web/src/lib/ontology/writes/shacl.ts, sh:node to …-scheme). Its property shape …-in never carries sh:severity, so its results are Violations, which already counted. Solo authors no sh:not, logical combinators or qualified shapes. A user-authored nested Warning shape follows the new rule.

On the Rust API (grep of solo 3bf84bc), solo references none of the changed items (Grouping::assemble, Aggregation, QueryOutput::Select, HavingOperator, ShaclError::SparqlConstraint, ReadStage, UngroupedRead, VariableNotFound), and its QueryError matches all have a catch-all arm; fluree-lambda-transact constructs ShaclError::{CoreError, QueryError, InvalidPattern}, which are unchanged. Solo main f5d7518f72 (2026-10-05) references none of post_values, ConstructTemplate, ConstraintFailure, with_constraint_failures or QueryOutput::Ask, and no solo Rust or TypeScript builds a grouped ASK or CONSTRUCT. Re-run at solo main 177bf9ac6a (2026-10-06): no reference to AskQuery, DescribeQuery, ReadStage or UngroupedRead (the ConstructQuery hits are an unrelated sparqlConstructQuery field), none of the replaced error strings ("Aggregate input variable", "already exists in schema", "Duplicate aggregate output", "format_binding called without", "Projected variable not in child schema"), no ASK with OFFSET, LIMIT or VALUES, and no JSON-LD ask with options. That's all from grep, fwiw — I haven't compiled this against solo. Dedicated query pools take the change only when their image tag moves.

Lockstep for solo: duplicate_pref_labels should switch to the expression form, "having": "(> (count ?concept) 1)", a one-line change in fluree-lambda-model/src/jsonld/queries/reports.rs:75. The query sends "having": [["filter", "(> (count ?concept) 1)"]], the form the old db doc example taught (that example is fixed here). db refuses it at parse ("filter operator must be a string") on base and on this head alike, so that report has presumably failed on every run. Solo's unit test only checks that the having key exists. The fix doesn't depend on this PR, so it can land before or after the db pin bump.

Nothing persisted changes: no commit, index, nameservice or graph-registry format. Stored queries and shapes run without migration, and behave per the table above.

Performance

The #1978 shapes get a lot cheaper. New bench query_hot_grouped_projection (first commit, so its base-tree run is the baseline from the same bench code): N entities, N/200 groups, indexed file ledger; each scenario runs the query and renders the response (SPARQL JSON; JSON-LD for the JSON-LD scenario).

Scenario tiny: 2k entities, 10 groups small: 20k entities, 100 groups small: peak / allocated per query small: response size
key_count (control) 107 µs → 107 µs 0.99 ms → 1.00 ms 6.0 / 13.1 MB (unchanged) 13 KB (unchanged)
expr_count (#1978) 1.92 ms → 114 µs 20.0 ms → 0.99 ms (20×) 14.7 / 83.3 MB → 5.7 / 12.2 MB 2.68 MB → 13 KB
expr_min 2.74 ms → 254 µs 30.1 ms → 2.47 ms (12×) 15.9 / 80.9 MB → 5.7 / 13.4 MB 1.64 MB → 8 KB
dedup_expr 1.38 ms → 90 µs 15.0 ms → 0.74 ms (20×) 10.1 / 51.5 MB → 2.9 / 6.5 MB 860 KB → 4 KB
implicit_const_count 1.50 ms → 21 µs 17.8 ms → 59 µs (300×) 30.2 / 59.8 MB → 0.6 / 1.6 MB 2.60 MB → 182 B
subselect_expr 2.16 ms → 222 µs 25.9 ms → 1.97 ms (13×) 13.6 / 97.3 MB → 5.9 / 19.3 MB 2.68 MB → 13 KB
jsonld_keyonly_expr 1.24 ms → 111 µs 13.8 ms → 1.00 ms (14×) 14.7 / 50.9 MB → 5.7 / 12.2 MB 161 KB → 1.4 KB

The "before" response is the exploded result (one row per solution). After the fix, a grouped SELECT expression costs what the key-only query costs, because it takes the same streaming GroupAggregateOperator plan (the expression runs once per group above it) instead of GroupByOperator + AggregateOperator, which stored every input row and built a list per column. A dedup-only GROUP BY with an expression now gets WHERE-level early dedup too (plan pins in it_query_explain), which is why dedup_expr lands below the key_count control.

Quiet box. EC2 c7i.4xlarge, the fat-LTO bench profile, base and head in interleaved rounds; a change counts as a win or a loss only when the base and head ranges don't overlap and the median delta exceeds the bench's budget (5% at small). These ran at the head that was reviewed, c67dc3516, and before it:

  • The target path wins clearly. expr_count went 48.9 → 1.93 ms (25×) and jsonld_keyonly_expr 27.8 → 2.00 ms (14×) in session 1, and at c67dc3516 expr_count is 47.1 → 1.77 ms (27×). The key_count control is stable.

  • The count queries, settled. Session 1 found count_novelty +9.4% and count_cached +5.1%, both losses by the rule. There were two causes:

    • real per-query work: every query level collected the variables its WHERE binds, and a grouping level built a SELECT-expression placer and a copy of those variables for the SAMPLE rewrite, even when it had nothing to sample or place. The fix (c67dc3516 as reviewed, 51c153337 after the rebase) collects them only for a stage that reads them;
    • code placement: under perf stat, the old merge ran count_novelty at +0.09% instructions but +5.44% cycles.

    Session 2, at c67dc3516: count_novelty +2.5%, count_cached +1.8% and count_base +2.2%, all stable by the rule. The new merge is at instruction parity on count_novelty (−0.05% instructions, +1.57% cycles), and brings count_base and count_cached instructions back to base level (−0.5% and +0.4%).

  • What remains: a 2–3% timing residual on the count queries, with the head slower in every round. It's layout, not logic: the instruction counts match base.

  • Everything else measured on the box is stable, negation_count, bsbm q9 and the other overlay rows included.

The local runs, the per-commit analysis, and both sessions' quiet-box tables are in a comment on this PR, moved out to keep this description under GitHub's size limit.

Hot paths touched, and what they cost now:

  • SPARQL-results serializers (format_string, NDJSON rows, XML): the per-row scan of every cell for Binding::Grouped is gone.
  • Plan build (apply_solution_modifiers): one pass over HAVING, binds, order binds, sort keys and projection, scanning a few variables with no allocation, replacing two ad-hoc checks.
  • FILTER and HAVING (RowPredicate): the per-row code is unchanged — the same filter_batch call when there is no EXISTS or metadata read, plus one extra .await per batch. Equivalent by inspection; not re-measured at the final head.
  • Implicit SAMPLE: one extra streamable aggregate per sampled variable. HAVING without grouping: one FILTER over the solutions.
  • Queries with a top-level trailing VALUES clause no longer take any fast path (8cfebf749); each such fast path returned a wrong answer for them. I don't have a timing, since a wrong answer gives no baseline to compare against. To keep the scan narrow, put the VALUES block inside the WHERE clause: it binds the variable before the scan, so the scan reads only the matching rows (and it is portable, per the compatibility note).
  • Sort-key pruning (bindable_sort_keys): borrows the ordering unchanged when every key is projected, grouped or an ORDER BY expression; otherwise one walk of the WHERE's produced variables, at plan time, once per sub-query operator (not per parent row).
  • sh:sparql failures: nothing per row. Collecting one costs a mutex push, only when a constraint fails.

Since the review, measured. All three use the same fixture: a release-build probe on an indexed 5,000-node ledger with 20,000 knows edges.

  • The review fixes. 1a20ac6bc (the head after the first batch of review fixes, rebased on 7ea640093) against 56ea09988 (the head after the second batch, on 7ea640093), interleaved, three runs each of seven reps, medians of the per-run medians (µs per query, query plus formatting): count top-k SPARQL 27.6 → 27.4 and JSON-LD 21.6 → 21.4, directory count 14.8 → 14.7, 20,000 JSON-LD rows streamed 16,184 → 16,248 (+0.4%) and DOM 18,020 → 17,971, Cypher aggregating WITH … WHERE c > 3 7,524 → 7,526 and aggregating RETURN … ORDER BY 16,868 → 16,629, ASK 16.8 → 16.7. All within run-to-run spread on a shared machine (load 38–45 on 16 cores). The base binary can't format a JSON-LD per-group list of literals on this ledger (the bug c44d22aca fixes), so that path has no baseline. The Cypher two-level form runs only for queries that used to fail or count wrong, and it is the plan an explicit second WITH already got. After the rebases the same patches are d6b2edfe6 and 4e30eaf0e. The fixes after those change no hot path (their one routing change, the two-level gate, takes only shapes that were a 400), except the two Cypher changes below.
  • A Cypher WITH's WHERE after its slice. 7288a2f54 against a4aa18827 on 8632e33bc (the same patches after the rebase are d466ca9df and b2ef72fba), which differ only in the Cypher lowering, interleaved, two runs of three rounds. A WITH without a slice, or with a slice and no WHERE, lowers to byte-identical IR (20 shapes dumped at both), and its timings are within noise (−4.7% to +5.7%, with overlapping ranges). A WITH with a slice and a WHERE now costs what the same WITH costs without its WHERE (0–5%, 11% on one shape inside its spread). Against the old plan it costs more where the WHERE had shrunk the input before the sort: 1.01–1.07× on three shapes; 1.2× for ORDER BY p.age DESC LIMIT 100 WHERE p.age > 75; 1.6–1.9× for the two-level aggregate (ORDER BY p.age DESC LIMIT 50 WHERE c > 4) and for ORDER BY p.name SKIP 100 WHERE p.age > 70; and 6× for ORDER BY p.name LIMIT 10 WHERE p.area = 'A7', which the old plan answered from an index seek on 100 of the 5,000 nodes. That is the price of the openCypher answer: the old plan answered a different question, and the gap grows with the number of rows the slice ranks. Only Cypher, and only a WITH with both.
  • A WITH inside a CALL body. b6f867305 against the same tree with 07d0892ad's Cypher lowering and planner (before the fix), interleaved, three rounds.
    • CALL shapes the bug doesn't reach got no slower: rows, a count, a selective parent, OPTIONAL MATCH with a count, and a sliced RETURN ran −3% to −12%. A CALL with no import ran −0.5%, and a non-CALL control −2.5%.
    • Outside a CALL body nothing changes. The lowering adds nothing there, and the planner rule reads only pinned_vars, which only Cypher's CALL sets.

Deviations from the design

Several of the items above are on this list too: trailing VALUES (an earlier commit here moved it after HAVING, per the spec, and af22f8e49 reverts that; docs/reference/compatibility.md now documents the deviation), HAVING on FILTER's evaluator, ASK / CONSTRUCT grouping, SHACL sh:select, and the stream's lane-dependent select * refusal. The review fixes added a few more: HAVING reading SELECT aliases (an extension), EXISTS-body variables free over the group row, a Cypher WITH's WHERE after its slice, a WITH inside a CALL body, Cypher reads after the aggregation, nested SHACL conformance counting every result, and sh:sparql failures following severity and mode. Beyond those, the placement rule has three refinements over the design's literal version of the JSON-LD rule (JSON-LD keeps per-group lists for non-key variables), grouped-read errors get their names through QueryError::name_variables, produced_vars_of now lives in ir::pattern with all five callers on it, and the JSON-LD alias guard is standalone until #1867 lands — plus a handful of smaller structural ones. Each is written up in full in a comment on this PR, moved out to keep this description under GitHub's size limit.

Follow-ups (not filed yet)

  • SHACL sh:select skips SPARQL validation, so the SPARQL validator's checks (V4 and the rest) never run on a shape's query.
  • Spec alignment for trailing VALUES (§18.2.4.3: join after HAVING), at the top level and in sub-SELECTs, as its own change with release notes, since it changes the answers of the "VALUES as a parameter" idiom. A sub-SELECT would also need a post-HAVING VALUES slot on SubqueryPattern, honored by every grouped-sub-query fast path and rewrite.
  • /query's JSON-LD select * under groupBy depends on the grouping lane: the keys and aggregates on the streaming lane, plus per-group lists on the GroupByOperator lane (pre-existing). The stream follows it.
  • Plan errors outside the grouped tail still print variable ids: a sort key the plan lost (a key nothing binds is now dropped instead) and the geo / S2 / index / vector search operators' unbound-input checks. The grouped-read and aggregate-variable errors in this PR show the pattern: a typed error carrying the id, named by QueryError::name_variables.
  • The stream reports a plan-time error after the head record (the grouped reads are now checked before it). Building the plan before the 200 would close the class, but it moves reasoning and planning ahead of the response, which is a latency trade-off I'd rather decide on its own.
  • Cypher properties with several values. A property read joins every value of the property. So WHERE gives the row once per value that passes, ORDER BY once per value (WITH DISTINCT keeps those copies, since the sort key is projected with them; RETURN DISTINCT removes them), and an aggregate in a later clause counts each copy: at base, MATCH (p:P) WHERE p.age > 30 RETURN count(p) is 4 over three people when one of them has two ages above 30. A read inside an aggregate's argument (collect(p.age), avg(n.age)) is joined before grouping, so a property with several values repeats the group's rows for every aggregate of that clause (WITH p, count(f) AS c, collect(p.age) AS ages counts each friend once per age). openCypher has no multi-valued properties, so it never repeats a row; matching it needs a model-wide change (WHERE as a semi-join, a sort key of one value per row, aggregate arguments read per row), in the explicit-WITH and MATCH paths at once, which is a lot more scope than this PR. The tests pin today's rows so that change flips them on purpose.
  • Cypher + on strings returns null (RETURN p.name + 'x', at base too); openCypher concatenates strings (and lists) with +. So ORDER BY n + 'x' sorts on nulls.
  • Cypher: an aggregate in a CALL body drops an import whose body matches nothing. MATCH (p:P) WHERE p.name = 'Carol' CALL (p) { MATCH (p)-[:knows]->(f) RETURN count(f) AS c } RETURN p.name, c gives no row for a friendless Carol, where the standalone MATCH … RETURN count(f) gives [[0]]. collect drops her [] the same way.
    • Cause. A CALL body's aggregate is grouped by the import (lower_call_branch for RETURN, and in this PR carry_imports for WITH), and an import with no body rows has no group.
    • Scope. It was there at base for RETURN. This PR's CALL fix brings WITH in a CALL body to the same model: at base every outer row got the whole body's wrong answer.
    • Workaround, documented in cypher.md: OPTIONAL MATCH.
    • Fix sketch. Left-join the imports with the grouped body, filling each aggregate's empty value: count 0, collect [], sum 0.
  • DESCRIBE with a trailing VALUES on an indexed ledger: once fix(query): BIND and UNWIND do not wait for the patterns that bind their target #2013 has landed, add the indexed-ledger case to construct_and_describe_read_a_trailing_values's DESCRIBE checks (SPARQL, both lanes), asserting the described graph (see Sequencing).

Sequencing

Rebase history

The first commits rebased onto 2941d470c with three one-line conflicts (#2004's multi-default-graph planning around this PR's name_variables maps); cargo check --workspace --all-targets --all-features was clean right after, before any fix. The branch then rebased onto ac26b463e (#2003, the storage compare-and-swap fix), onto 7ea640093 (#2012, overflow integers keep xsd:integer; 6a03e83a4, the transaction WHERE decodes through its own graph; docs and CI) and onto 6b9d5b619 (#2011, async WAL recovery; it shares two files with this PR, the CLI's query command and the benches doc, in disjoint hunks) with no conflicts, and the check was clean after each. #2012 changes how encoded overflow integers are decoded, so the indexed lanes suite now runs them through this PR's grouping: a per-group DATATYPE of the key, HAVING, MIN / MAX / SAMPLE and the implicit SAMPLE, a JSON-LD per-group list, the count top-k, and a group a decoded copy joins. Reverting either of #2012's DATATYPE or group-key changes turns it red.

Next, the branch rebased onto 8632e33bc (#2018, #2020 and #2022). #2022 touched three files of this PR: view/stream_query.rs (its row_compactor beside this PR's ensure_streamable; both kept), view/dataset_query.rs (its restructure, with this PR's two name_variables maps applied to it), and it_iceberg_local_fs.rs, which ends as main's file, since this PR's canary moved to its own binary. git range-diff shows seven commits changed in context only and the rest patch-identical. The check was clean right after.

Last, it rebased onto 745736a04 (#2023, #2024 and #2025). The one conflict was in shacl_tests.rs, where #2024's violation_carries_resolved_results sits beside this PR's sh:sparql helpers; both are kept. Every other commit is patch-identical. #2024's SPARQL parameters walk each query form's AST, and they did not walk the trailing VALUES this PR adds after ASK, CONSTRUCT and DESCRIBE; bd3797dcc closes that seam. The 21 commits that touch a file #2023–#2025 also changed each pass cargo check (the api lib in test mode and grp_query_sparql, plus the sparql crate where they touch it), and cargo check --workspace --all-targets --all-features was clean after the rebase.

Tests

New coverage is the SPARQL matrix (it_query_sparql_grouped_projection.rs, each case asserting SPARQL-JSON rows and JSON-LD values on a non-injective fixture), its JSON-LD twin, UPDATE twins, stream refusals, explain pins, a Cypher-under-policy case, SHACL lib tests, unit tests, and a new it_grouped_projection_lanes binary with MustFire / MustNotFire routing stamps — and, since the review, a test for each item above: grouped ASK / CONSTRUCT, EXISTS free over the group row, Cypher reads after the aggregation and the multi-value model pinned in row order, WITH … WHERE after the slice on every lowering path, WITH inside CALL (p), the trailing-VALUES cases at both levels and after every query form, the SHACL severity × mode cells, nested conformance through seven nestings on both write surfaces, and canary pairs on every binds-gated fast path. Every row of the non-vacuity table was committed first, then only the fix reverted, rebuilt, watched fail, restored, and the needle grepped.

The full test list is in a comment on this PR. The non-vacuity logs, what was reverted and what went red, are in the first comment (the PR as first opened) and a follow-up comment (every fix since). Two things the follow-up log calls out: the mutations that remove only a fast path's binds gate stay green where the path also declines on its projection check (removing both turns it red; both runs are in the log), and the EXISTS seed mapping, which no shape reaches, is now a debug_assert! (d6b2edfe6) instead of an untested branch. Release builds keep the spec reading (the list seeds as unbound).

Gates

The gates ran at 31e238a7e on 745736a04. The one commit after it, 87576fdb1, changes only docs/query/cypher.md; at 87576fdb1 fmt is clean (the workspace and testsuite-sparql) and mdbook build succeeds. At 31e238a7e, every one exited 0:

  • cargo fmt --all -- --check, the workspace and testsuite-sparql;
  • cargo clippy --workspace --all-targets --locked -- -D warnings, with --all-features and at default features, and clippy in testsuite-sparql: 0 warnings each;
  • cargo nextest run --all-features on fluree-db-query, fluree-db-sparql, fluree-db-cypher, fluree-db-shacl and fluree-db-transact: 3,002 passed, 1 skipped;
  • the same on the fluree-db-api lib plus 25 test binaries: the five grp_* groups this PR touches (grp_query, grp_query_sparql, grp_misc, grp_transact, grp_policy), six Cypher binaries, and 14 others on the paths it changes (explain, the indexed SPARQL suite, the SQL and Iceberg lanes, the lanes and fused-aggregate canaries, and the fast-path and count regression binaries): 3,436 passed, 13 skipped;
  • testsuite-sparql (W3C): 36/36;
  • doc tests for the sparql, query, cypher and shacl crates: 16 passed, 3 ignored.

testsuite-shacl is local only — no CI workflow runs it — and it doesn't compile on main today; d3df023d7 adds its two missing ValidateOptions fields as its own commit. SHACL Core 81/98 and SHACL-SPARQL 17/22, with the same tests passing with and without the nested-conformance fix, and with and without the sh:sparql-only severity and list fixes; at the PR's first head it matched base test for test. The suite loads lists as indexed items and registers its shapes by targets or properties, so the list and severity fixes don't reach it, and none of its SHACL-SPARQL tests uses GROUP BY, HAVING or an aggregate. The shacl_tests pin those instead.

Not run at this head: the workspace-wide nextest (CI's test job), and with it the other fluree-db-api test binaries and the server, CLI and python crate tests (all of which compile and pass clippy in the workspace runs above); the api doc tests; the wasm check and the wasm-smoke browser suites; the SQL live bridge; the bench.yml compare against the committed baseline; and an openCypher reference engine for the WITH … WHERE change (none was available, so that evidence is read, not executed). The workspace-wide run from when this PR was opened (13,765 run, 13,760 passed, at 2a540f61e) is in the same comment as the perf detail, with its gate log.

@aaj3f aaj3f added bug Something isn't working as expected area:sparql SPARQL/Turtle/TriG/JSON-LD parsing, lowering, UPDATE semantics, W3C conformance area:query Query execution, planning, fast paths, overlay, result formatting labels Sep 30, 2026
@aaj3f
aaj3f marked this pull request as ready for review October 2, 2026 02:06
@aaj3f
aaj3f requested review from bplatz and zonotope and removed request for bplatz October 2, 2026 02:06
@aaj3f

aaj3f commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Non-vacuity log for this PR, moved out of the description to keep it under GitHub's size limit: for each fix, what was reverted and what went red. Every row was committed first, then only the fix reverted, rebuilt, watched fail, restored, and the needle grepped.

The rows through 1936c5d32 were run before the rebase onto bf523e24e; the rebase applied without conflicts and changed none of the fix code. The rows from 2c4aa850c on were run after it. Mutations sharing a build letter (A through F) were applied together, and each test went red with its own mutation's failure message. The one test that went red outside its own row is noted in the lanes row.

Commit Fix reverted Went red
20b3be1db SPARQL placement SelectExprPlacer::place always before grouping (the old lowering) 12: one row per group, every results format, alias chains, ORDER BY/LIMIT, HAVING reads an alias, empty implicit group, STRUUID, sub-SELECT scalars, both explain pins, UPDATE twin, lanes
20b3be1db SELECT * of a grouping level projects every variable (top level; then sub-SELECT only) select_star_under_implicit_grouping_projects_nothing (sub-SELECT variant: debug panic in join.rs)
20b3be1db the expression alias becomes a GROUP BY key (the Cypher reading) one row per group (2 rows, network 3 / other 3)
20b3be1db compound SELECT items placed after the aliases alias chains (?t unbound)
f985bec71 + 20b3be1db Grouping::bind_list drops a dedup-only level's binds one row per group, explain dedup-only, UPDATE twin, lanes
f985bec71 + 20b3be1db dependency trimming skips the grouping's binds one row per group, ORDER BY/LIMIT, explain dedup-only, lanes
19bd22e0b spec edges no implicit SAMPLE (both surfaces) 5: SPARQL HAVING/ORDER BY sample tests, JSON-LD twins, the updated it_query_sparql test
19bd22e0b HAVING without grouping dropped (both surfaces) HAVING-as-filter, SPARQL and JSON-LD
19bd22e0b no SELECT-alias rename in having_as_filter HAVING-as-filter, SPARQL and JSON-LD
19bd22e0b SAMPLE every non-key read (no WHERE-binding test) ORDER BY/LIMIT, HAVING reads an alias as unbound, others
4a414cba6 V4 / V009 V4 keeps the old definition of "grouped" implicit_grouping_via_having_checks_the_projection
4a414cba6 V009 disabled aggregate_over_a_same_level_select_alias_is_rejected
7db6de5d2 plan-time check first_ungrouped_read not called operator_tree unit test, hand-built ORDER BY executor test, JSON-LD subquery error
7db6de5d2 eval of a Grouped binding returns unbound eval unit test
780fdae82 JSON-LD key-only expressions old JSON-LD placement rule jsonld_key_only_select_expression_is_one_value_per_group
780fdae82 alias-shadow guard disabled jsonld_select_alias_cannot_shadow_a_where_variable
f61a9fd88 SHACL HAVING filter dropped / SAMPLE rewrite skipped / plan-time check skipped (one build) each SHACL test, for its own piece: a violation at ex:low; "grouped (list-valued) binding reached scalar evaluation"; a violation at the focus node
1936c5d32 formatters SPARQL JSON / XML writers render a list cell (nothing / unbound) instead of refusing it; the stream endpoint neither refuses a per-group list nor plans with Reject (one build) both grouped_cell_is_a_format_error unit tests, jsonld_per_group_list_is_refused_by_sparql_results_formats, jsonld_per_group_list_is_rejected_before_streaming
base sources the whole fix Cypher parity fails at base (the SPARQL side returns lists). The solo must-not-change guards pass at base, as they should.
2c4aa850c error names QueryError::name_variables leaves the id in the message jsonld_grouped_read_errors_name_the_variable ("projected variable VarId(1) …"), ungrouped_read_message_names_variables
067dd65f8 key-only alias moves with its reader (C) SelectExprPlacer::place_all stops after the forward pass jsonld_key_only_alias_read_per_solution_stays_a_list, select_expr_placement_rule
5916aedb8 trailing VALUES after HAVING (C) a grouped query's VALUES joins before grouping again trailing_values_joins_after_having. The commit was later reverted by f4ec91a13, and its test became trailing_values_join_before_grouping
2c0746ce9 HAVING vs trailing VALUES (F, after f4ec91a13) HAVING's reads of trailing VALUES variables are not renamed trailing_values_join_before_grouping and sub_select_having_reads_trailing_values_as_unbound: HAVING (?v = 1) VALUES ?v { 1 } kept all 3 groups (SAMPLE(?v)), top level and sub-SELECT; expected 0
f4ec91a13 placement kept (D) a grouped query's trailing VALUES no longer restricts the aggregates trailing_values_join_before_grouping: GROUP BY ?a VALUES ?e { ex:e1 } gave full counts for all three areas; expected (Net, 1)
f4ec91a13 / 2a540f61e (E) trailing VALUES variables are not pre-group variables for the lowering trailing_values_join_before_grouping: ORDER BY ?v failed the plan ("ORDER BY variable ?v is neither …") instead of sampling
527936dfd stream wildcard by lane (D) the stream refuses every wildcard with a non-key WHERE variable jsonld_wildcard_with_list_columns_is_rejected_before_streaming: the count-in-having query was refused
6e7be00ac internal names (D) named_message prints internal names cypher_grouped_read_error_names_no_internal_variable ("ORDER BY variable ?#__prop_e_area …"), ungrouped_read_message_names_variables
d18c572b7 SHACL pre-bound grouping (C) levels are not grouped by pre-bound variables; plan failures stay QueryError shacl_sparql_grouped_constraints_fire_on_their_groups (every write refused), shacl_sparql_grouped_projection_of_a_non_key_fails_closed (not a SparqlConstraint)
924545d82 ASK / CONSTRUCT refusal (A) the four SPARQL refusals and JSON-LD ask's ask_and_construct_refuse_grouping (all four forms answered), jsonld_ask_refuses_grouping (both options answered)
d8f6b176b HAVING with FILTER's evaluator (A) HAVING evaluates with the synchronous batch filter again having_exists_is_evaluated_per_group (HAVING (EXISTS …) kept no group)
dcb896772 (its policy test) the same, alone cypher_with_where_metadata_after_aggregation_under_policy_sees_filtered_flakes (no rows: keys(n) read as empty)
d4cee5fe2 stream and unbound projection (A) the stream's select * check never refuses; /query returns the old VarId error jsonld_wildcard_with_list_columns_is_rejected_before_streaming; jsonld_unbound_projection_is_the_same_4xx_on_query_and_stream ("Selected variable VarId(1) not found")
d4cee5fe2 (B) the stream calls an unbound projection list-valued jsonld_unbound_projection_is_the_same_4xx_on_query_and_stream ("list-valued column ?nosuch")
2b921ce08 lanes stamps (A) the top-k detector ignores the grouping's binds it_grouped_projection_lanes (MustNotFire), and grouped_select_expression_order_by_and_limit on a memory ledger (?seg missing: the top-k path serves that shape too)
2b921ce08 (B) the top-k detector never matches it_grouped_projection_lanes (MustFire on the canary)

@bplatz bplatz left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The grouping restructure looks right to me. Please go through the inline comments before merging, in particular the EXISTS-under-grouping one, the Cypher WITH … WHERE one and the SHACL one.

Comment thread fluree-db-query/src/ir/grouping.rs Outdated
output_var
});
if let Some(having) = having.as_deref_mut() {
having.substitute_var(v, sampled);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The SAMPLE rewrite reaches into EXISTS patterns through substitute_var, so a non-key variable inside an EXISTS becomes the sampled value, and the result depends on which row SAMPLE picks. The group row doesn't bind ?e, so it's free inside the pattern (Sample(?e) can't stand in a triple pattern anyway).

SELECT ?a (COUNT(?e) AS ?n) WHERE { ?e ex:area ?a } GROUP BY ?a
HAVING (EXISTS { ?e ex:flag true })

With the flag on e1 the Net group is kept; move it to e2 (also Net) and the group is dropped. By the spec every group is kept. NOT EXISTS flips the same way. having_exists_is_evaluated_per_group only pins a shape where every possible sample agrees.

Suggest leaving EXISTS-body variables out of the rewrite, and out of the first_ungrouped_read HAVING check, which would otherwise reject them.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, and agreed. It's fixed in 5aab1c166. §18.2.4.1 rewrites an expression's unaggregated variables, and a pattern variable inside an EXISTS isn't one of them: the group solution doesn't bind it, so substitute(P, μ) leaves it free.

Expression now has a row-read walk (row_reads / substitute_row_read) that skips EXISTS bodies, and reads a pattern comprehension's projection only where its own pattern doesn't bind the variable. sample_ungrouped_reads, first_ungrouped_read and the SELECT-expression placer all go through it. Dependency tracing follows the post-group stages by row reads too, plus only those correlated variables that grouping produces, so a pattern-only variable no longer rides through the GroupByOperator lane as a per-group list.

Your repro now keeps every group with the flag on e1 or on e2, and NOT EXISTS keeps none (exists_body_variables_are_free_over_the_group_row, GROUP_CONCAT lane included). The third case of having_exists_is_evaluated_per_group, which had pinned the SAMPLE reading, now expects all three groups.

|| self.aggregate_inputs.contains(alias)
|| refs
.iter()
.any(|v| where_vars.contains(v) && !self.keys.contains(v))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This has the same root cause: referenced_vars() includes EXISTS pattern variables, so an EXISTS over a non-key variable puts the expression before grouping, and the plan check then rejects it:

SELECT ?a (EXISTS { ?e ex:flag true } AS ?f) (COUNT(?e) AS ?n)
WHERE { ?e ex:area ?a } GROUP BY ?a

This returns 400 "projected variable ?f is neither a GROUP BY key nor an aggregate result". V4 accepts the query (the AST check skips EXISTS bodies), and the spec gives one row per group.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same root cause, so it's the same commit (5aab1c166). The placer goes by row reads now, so (EXISTS { ?e ex:flag true } AS ?f) becomes a per-group Extend. Your query returns one row per group, with ?f true in each, and since the placer is shared, JSON-LD gets the same placement. It's pinned in the same test, exists_body_variables_are_free_over_the_group_row.

@@ -1269,7 +1271,8 @@ fn lower_with<E: IriEncoder>(
projection.aggregates,
projection.post_binds,
having,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An aggregating WITH … WHERE that reads a property of a key node now returns a 400, and the behavior table doesn't list it:

MATCH (p:E)-[:knows]->(f) WITH p, count(f) AS c WHERE c >= 1 AND p.age > 30 RETURN p.age, c

This returns "HAVING reads a variable that is neither a GROUP BY key nor an aggregate result". Base returned [], which was also wrong. WITH e, count(*) AS c WHERE e.age > 30 behaves the same way. The p.age accessor triple lands in the WITH body before grouping, and Cypher lowering doesn't apply sample_ungrouped_reads. Applying it here would give the exact answer, since a property of a key node has a single value.

WHERE exists { (p)-[:knows]->(f) } after the WITH also returns a 400 ("HAVING reads variable f"). f is out of scope there, so it should be a fresh variable. That's the EXISTS issue in grouping.rs.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for this one. It's fixed in 422d986c4, though not quite the way you suggested, because the SAMPLE reading turned out to be wrong for a property with several values.

An aggregating WITH's WHERE and ORDER BY, and an aggregating RETURN's ORDER BY, can now read a property of a node the clause projects. When they do, the clause lowers as two levels: the aggregation, then a stage that joins the property, filters on the WHERE, sorts and slices. That's what WITH p, count(f) AS c WITH p, c WHERE p.age > 30 already lowered to, and the two forms now return the same rows. (When the clause also has a SKIP or LIMIT, the WHERE now filters after the slice, as openCypher places it; there's more on that in my reply to your follow-ups.) A property of a node the clause doesn't project is still an error, which matches Cypher's scoping.

An earlier commit in this PR (23c3f1ee2) did read the property as a SAMPLE inside the aggregation. The catch is that it joins the property before grouping, so each value repeats every row of the group: with Alice aged 40 and 41, count gave 4 instead of 2, sum doubled, and collect listed every friend twice. Reading the property after the aggregation keeps the clause's own aggregates exact, whatever the property holds.

Past that, the second stage keeps Fluree's model for a property with several values, which joins every value, as MATCH … WHERE already does:

  • in WHERE, the row comes once per value that passes;
  • in ORDER BY, it comes once per value, and WITH DISTINCT keeps those copies (the sort key is projected with them) while RETURN DISTINCT removes them;
  • an aggregate in a later clause counts each copy (… WHERE p.age > 30 RETURN count(*) counts Alice twice).

A read inside an aggregate's argument, such as collect(p.age), is still joined before grouping, so it repeats the group's rows for every aggregate of that clause; that predates this PR. openCypher has no multi-valued properties and never repeats a row. Matching it means changing the model everywhere at once, the explicit-WITH and MATCH paths included, which is a much bigger scope than this PR, so it's a follow-up in the description. The reference now describes the model (5422809d4), and the tests pin it in row order, so that change will flip them on purpose. They cover an unsliced ORDER BY in a WITH and a RETURN, WITH DISTINCT … ORDER BY, the later count(*) and sum(c), and collect(p.age) beside count(f).

Your first query returns [[40, 2], [50, 3]] on a knows/age fixture. WITH e, count(*) AS c WHERE e.age > 30 and the ORDER BY forms return rows too (cypher_reads_a_key_nodes_property_after_grouping), and so does ORDER BY p.age + 1. A sort on an expression over the clause's aggregates, ORDER BY c + 1 or ORDER BY -c, was still a 400; it now takes the same two-level form (ad89f13b9). cypher_output_node_properties_are_read_after_aggregation pins count, sum and collect with a two-valued age under WHERE, a WITH's ORDER BY and a RETURN's ORDER BY, plus the explicit two-WITH form. The error-naming test moved to a node the WITH doesn't project (WITH e.area AS a, count(*) AS c ORDER BY e.name).

WHERE exists { (p)-[:knows]->(f) } after the WITH was the EXISTS issue from grouping.rs. With 5aab1c166, f is free there, and the same test pins it with a likes edge.

The composite-alias case, WITH p, count(f) + 0 AS c WHERE c > 1, dropped every row at base too. It's fixed by 19512aec1 (the HAVING-alias change in my follow-ups reply) and pinned in the same test.

Comment thread fluree-db-shacl/src/sparql.rs Outdated
)
.await?;
.await
.map_err(|e| match e.name_variables(&vars) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This failure leaves validation as an Err, so it bypasses sh:severity and the graph's Warn mode (the severity split in stage.rs only sees results). Because it's raised at plan time, it refuses conforming writes too.

Repro: a shape with sh:severity sh:Warning, sh:targetClass ex:Player and sh:select "SELECT $this ?value WHERE { $this ex:score ?value } GROUP BY $this", then insert an ex:Player with no ex:score. Base commits; here it returns 400 Invalid sh:sparql constraint …: projected variable ?value is neither….

So after an upgrade, a stored shape like this blocks every write to its target class. Reporting a failure is spec-correct, but the behavior table should say exactly that, or severity and Warn mode should apply in this case.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I went with your second option, in 2845a168c and 764519b36: severity and mode now apply. They cover every failure of the constraint's own query:

  • it doesn't parse, lower or plan;
  • a $PATH has nothing to bind it;
  • the query fails on its own terms, for example an aggregate over a variable nothing binds, or a call to an unknown function.

The request's budgets (fuel, deadline, memory), its storage, catalog and policy access, the state of the data and internal faults still fail the write whatever the severity. That split is an exhaustive match over the query error type, so a new error variant has to be placed on one side.

On the transaction path the engine records the failure, with the owning shape's severity and the graph, and goes on to the shape's other constraints. validate_view_with_shacl then decides it the way it decides a result:

  • a Violation shape in a reject-mode graph fails the transaction with the same ShaclError::SparqlConstraint as before;
  • a Warning or Info shape, or a warn-mode graph, logs it and commits. A failure that repeats across focus nodes is logged once, with the number of focus nodes.

The other engine callers (validate_all, reports) still raise the failure.

A failure in a shape checked as a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, sh:qualifiedValueShape) takes the severity of the outermost shape that reports the result, as that shape's own results do (7999b625e). So a broken constraint in a nested Violation shape is logged under a Warning outer shape, and one in a nested Warning shape fails the write under a Violation outer shape in a reject-mode graph.

The message now names the shape every time (on shape …:), and names the variable rather than its internal id. The aggregate case used to print Aggregate input variable VarId(3) not found in schema. It's now a typed plan error, an aggregate reads variable ?nosuch, which is unbound: nothing before the grouping binds it, and that's also what a query gets. It's a 400 on both query paths, where the tracked path used to return a 500 (aggregate_over_a_variable_nothing_binds_is_a_named_error and its JSON-LD twin). The aggregate-output plan errors are named the same way, in de26852a2.

Worth saying plainly: this changes main's behavior, not only this PR's. A parse or lower failure in a Warning or Info shape, or in a warn-mode graph, used to reject the write. It now commits with a warning. It's in the behavior table and the solo notes.

The tests (shacl_sparql_constraint_failure_*) cover every severity × mode cell for:

  • a plan, a parse and a lower failure;
  • an aggregate over a variable nothing binds;
  • two aggregates with one output;
  • an unknown function.

That includes the fail-closed cell with its message, which names the shape and has no VarId. They also cover a Warning shape's failure next to a Violation shape that still rejects. Your repro (a Warning shape, GROUP BY $this projecting ?value) commits.

The split itself has a table test over every error variant (795e31f96, c0e831164). Fuel, cancellation and timeout, memory, storage, catalog, graph-source and policy errors must be raised whatever the severity, and the constraint's own failures must not. Making every error the constraint's own turns it red. The table and a match with no wildcard come from one list, so a new error variant doesn't compile until it has a row.

Comment thread docs/reference/compatibility.md Outdated

**Grouped queries:** A query groups when it has a `GROUP BY` or an aggregate anywhere in SELECT, HAVING or ORDER BY (§18.2.4.1). Its SELECT expressions are evaluated once per group, after HAVING, in SELECT order (§18.2.4.4), and `SELECT *` projects its GROUP BY keys. In HAVING and ORDER BY, a variable that is neither a key nor aggregated means `SAMPLE(?v)` (§18.2.4.1); HAVING without grouping filters the solutions (§18.2.4.2). Two shapes are rejected rather than evaluated literally: a projected variable that is neither a GROUP BY key nor aggregated (§11.4, including when the only aggregate is in HAVING or ORDER BY), and an aggregate over an alias of the same SELECT clause, which the spec would evaluate as unbound (`COUNT` = 0 for every group).

**Trailing VALUES in a grouped query (deviation from §18.2.4.3):** A `VALUES` block after the WHERE clause, at the top level or in a sub-SELECT, is joined with the WHERE solutions before grouping, so it restricts what the aggregates count: `SELECT ?a (COUNT(?e) AS ?n) WHERE { ?e ex:area ?a } GROUP BY ?a VALUES ?e { ex:e1 }` counts `ex:e1` alone. The spec joins it after HAVING instead, pairing each group row with each VALUES row. HAVING still reads a VALUES variable as unbound, as it would in the spec's order, unless the WHERE binds it or it is a GROUP BY key. For queries that behave the same on any SPARQL processor, put the `VALUES` block inside the WHERE clause to restrict the aggregates' input, or move the grouped query into a sub-SELECT and put the `VALUES` block after the outer WHERE to join it with the grouped rows.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

At the top level this isn't quite what happens. post_values joins after the WHERE tree, including the binds lowering generates for GROUP BY expressions and aggregate inputs (lower/mod.rs ~518). A sub-SELECT splices VALUES right after its WHERE (select.rs ~875).

SELECT (SUM(?n * ?v) AS ?s) WHERE { ?e ex:n ?n } VALUES ?v { 2 }

This returns 0 at the top level and 42 as a sub-SELECT; SUM(?v) in the same top-level query does see the value. The behavior predates this PR, but this paragraph now describes the two levels as the same. Either align the top level with the sub-SELECT or narrow the doc.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I went with aligning them, in e6e2061e5. The VALUES join stays where it was, after the WHERE tree and before grouping. The binds the level generates now run after it at the top level, as they do in a sub-SELECT: aggregate inputs, GROUP BY expressions, and the SELECT expressions the level places before grouping (all of them, when the level doesn't group). Query.post_values is now a PostValues: the VALUES vars and rows, plus then, the binds that follow it, in order.

Your query now gives the same answer at both levels. On the test fixture SUM(?n * ?v) is 44 at both, and 66 with VALUES ?v { 1 2 }. The ungrouped (?n * ?v AS ?p) and GROUP BY (?n * ?v AS ?k) read ?v too (generated_binds_read_trailing_values_at_both_levels). The compatibility paragraph now says where those expressions run.

While tracing this, it turned out most of the aggregate fast paths ignored a top-level trailing VALUES altogether. On an indexed ledger, SELECT (COUNT(*) AS ?n) WHERE { ?e ex:area ?a } VALUES ?a { "Net" } counted all 6 rows. The count top-k, the star top-k, SUM(?o) and the per-predicate directory count had the same problem; the whole-graph aggregates already declined. 8cfebf749 sends any query with a trailing VALUES to the generic lane. It's pinned with MustNotFire and hand-derived rows in both lanes, the whole-graph aggregates included (3b45c92dd).

And the grammar allows a trailing VALUES after every query form, while only SELECT parsed one. 5efca0d81 reads it after ASK, CONSTRUCT and DESCRIBE too. ASK and CONSTRUCT carry it as post_values, like SELECT. DESCRIBE joins it in its target subquery right after the WHERE, so DESCRIBE ?x VALUES ?x { ex:a ex:b } describes the listed resources.

/// [`Self::UngroupedRead`] becomes an [`Self::InvalidQuery`] naming them
/// from `vars`. Every other error is returned unchanged.
#[must_use]
pub fn name_variables(self, vars: &crate::var_registry::VarRegistry) -> Self {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The JSON-LD update WHERE path doesn't call this. stage.rs runs execute_where_streaming, where a subquery is planned at run time, so a where containing ["query", {"select": ["?a", "?x"], "where": {"@id": "?x", "ex:area": "?a"}, "groupBy": ["?a"]}] returns … projected variable VarId(0) is neither a GROUP BY key…. A name_variables(&txn.vars) map at the two execute_where_streaming call sites in stage.rs would fix it, or one inside execute_where_streaming / WhereCursor::next_batch, which already hold vars.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I took your second option and fixed it in the cursor itself (cd021b64e). execute_where_streaming names the variables in its plan and open errors. WhereCursor::next_batch names them in errors from the pull, which is where a subquery gets planned. Both go through the registry the cursor already holds, so this covers every caller: JSON-LD, SPARQL and Cypher updates.

Your JSON-LD update now says projected variable ?x is neither a GROUP BY key … (update_grouped_read_error_names_the_variable). The validator rejects the SPARQL UPDATE twin before it gets here, so the test pins that one only for parity.

/// list element — a cartesian product across list columns. A projected
/// variable nothing binds is refused as unbound. The plan's output is then set
/// to refuse per-group lists outright.
fn ensure_streamable(query: &mut fluree_db_query::ir::Query, vars: &VarRegistry) -> Result<()> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This catches list columns and unbound projections before the 200, but not a bind-stage ungrouped read. That one is only found when the plan is built inside run_stream_query, after the head record. Example: JSON-LD select ["?a", "(as (+ (count ?e) (strlen ?n)) ?x)"] with groupBy ["?a"] returns a 400 on /query, but on the stream the error arrives as a record after the 200. Calling first_ungrouped_read here, with produced_vars_of(patterns) as the WHERE variables, would match /query.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done as you suggested, in 97b5b9ed1. ensure_streamable runs first_ungrouped_read over the variables the WHERE and any trailing VALUES bind, with the query's own projection policy. So your JSON-LD query is the same named 400 on /query and on the stream, before the stream starts (jsonld_grouped_read_is_the_same_4xx_on_query_and_stream).

Every other plan-time error still arrives after the head record. Building the plan before the 200 would close that, but it moves reasoning and planning ahead of the response. That's a latency trade-off I'd rather decide on its own, so it's in the description's follow-ups.

"select": ["?category", "(count ?product)"],
"groupBy": ["?category"],
"having": [["filter", "(> (count ?product) 10)"]],
"having": "(> (count ?product) 10)",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The second having example (~line 1645) still uses the [["filter", …]] form.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, fixed in 6ca526c24: "having": "(> (count ?product) 5)". Fwiw, that [["filter", …]] form is also what solo's duplicate_pref_labels report sends, and db refuses it at parse, on main too. So the description's solo section now has a one-line lockstep fix for it.

@bplatz

bplatz commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Suggested follow-ups (not blocking):

  • ASK and CONSTRUCT with GROUP BY or HAVING are now refused. CONSTRUCT … GROUP BY … HAVING is meaningful, and lower_select_level should make supporting it cheap.
  • ORDER BY ?nosuch (grouped or not) is still a 500: "Sort variable VarId(n) not found in query schema".
  • Routing stamps for the other detectors that now gate on binds: implicit_single_aggregate, stats_count_by_predicate, star_topk, whole_graph_scalar_aggs, and the fused R2RML / SQL lanes.
  • HAVING that reads a non-aggregate SELECT alias (e.g. (COUNT(?e)+0 AS ?n) … HAVING (?n > 1)) silently drops every group. A diagnostic would help.

aaj3f added 12 commits October 8, 2026 07:39
Grouped queries that project an aggregate-free SELECT expression
(#1978): a control (key + COUNT), the expression with COUNT, with MIN,
with no aggregate (dedup-only GROUP BY), a constant under implicit
grouping, the expression inside a sub-SELECT, and the JSON-LD twin.
Each scenario measures execution plus the rendered response, with
allocation peak and churn recorded alongside.
A SELECT expression of a grouped level is an Extend over the group rows
(SPARQL 1.1 §18.2.4.4). It belongs to the grouping phase whether or not
an aggregation stage exists: a dedup-only GROUP BY ?a projecting
(IF(?a ...) AS ?seg) needs one, and Aggregation had no room for it
(Grouping::assemble dropped binds when there were no aggregates).

Removing the field makes every reader a compile error, so each one was
moved deliberately. Detectors that decline on binds now ask the grouping
(any binds, with or without aggregates) rather than the aggregation
stage. No behavior change: nothing lowers binds onto a dedup-only
grouping yet.

Also corrects stale docs: binds run after HAVING (the executor's order),
not before, and collect() produces a Binding::List, not Binding::Grouped.
A SELECT expression of a grouped query level is an Extend over the group
rows, evaluated after HAVING, in SELECT order (SPARQL 1.1 §18.2.4.4).
The lowering made every aggregate-free expression a per-solution BIND
before grouping. The grouping stage then carried it as a per-group list,
and the SPARQL-results formatters expanded the list into one row per
solution (#1978): 5 rows for 2 groups, a cartesian product for two
expression columns, a lost row for an implicit group over no solutions.
The same list reached an enclosing query through a sub-SELECT (outer
FILTER, GROUP BY, aggregates, joins) and an UPDATE template.

The level's SELECT expressions are now placed after its solution
modifiers are lowered, so whether it groups is read from the lowered
keys and aggregates: GROUP BY, or an aggregate in SELECT, HAVING or
ORDER BY. In a grouping level every expression is a per-group Extend,
compound aggregate items included, in one list in SELECT order; that
also fixes the alias chains that read an alias before it was bound. An
expression whose alias is a group key (the GROUP BY (LCASE(?a))
shortcut) stays a WHERE bind. The placement rule is one helper in the
query IR (SelectExprPlacer), so JSON-LD can share it.

SELECT * of a grouping level projects its GROUP BY keys, which is
nothing for implicit grouping, instead of every WHERE variable as a
per-group list; a sub-SELECT likewise exports nothing.

Grouped queries with a SELECT expression now take the streaming
GroupAggregateOperator plan the key-only query takes, and a dedup-only
GROUP BY gets WHERE-level early dedup.
Two shapes that returned silent wrong answers, errors or debug panics
now follow SPARQL 1.1 on both surfaces, through two shared IR helpers:

- In a grouping level, a HAVING or ORDER BY read of a variable that the
  WHERE binds and that is not a group key means SAMPLE(?v)
  (§18.2.4.1). The read is rewritten to a hoisted SAMPLE aggregate,
  reusing one the level already has. Keys, aggregate outputs, SELECT
  Extend outputs and variables nothing binds are left alone. It used to
  return no rows on the streaming lane, panic a debug build on the
  traditional lane, and reject ORDER BY with "Sort variable ... not
  found".
- HAVING on a level that does not group is a Filter over its solutions
  (§18.2.4.2). It was dropped. The Filter reads the level's SELECT
  aliases as fresh, unbound variables, since HAVING cannot see SELECT
  expressions.

SPARQL lowering of a SELECT level (top level and sub-SELECT) is now one
shared function in §18.2.4 order: modifiers and aggregates, HAVING,
then SELECT expressions. JSON-LD applies both rules to having and
orderBy in queries and subqueries.

The test that pinned an error for ORDER BY over a non-key variable now
pins the sampled result.
A query level groups when it has a GROUP BY or an aggregate anywhere in
its SELECT, HAVING or ORDER BY (SPARQL 1.1 §18.2.4.1). The projection
check (V4) only counted SELECT aggregates, while lowering counted all
three, so a query whose only aggregate was in HAVING or ORDER BY was
lowered as grouped but never checked: a projected expression expanded
into one row per solution, and a projected variable silently became an
implicit GROUP BY key. Both now get the existing V4 error. The
definition is one AST helper (SolutionModifiers::level_groups), and
lowering asserts in debug builds that it agrees.

An aggregate in SELECT, HAVING or ORDER BY that reads an alias assigned
by the same SELECT clause is now a named error (V009). The alias is
bound by the SELECT's Extend, after aggregation, so per the spec the
aggregate sees it unbound (COUNT = 0 for every group). The error points
at aggregating the expression itself.
After grouping, a variable that the WHERE binds but the grouping neither
keys, aggregates nor binds is a per-group list (Binding::Grouped). Only
a JSON-LD top-level projection may carry one, as documented. Every other
reader used to invent its own meaning: FILTER and HAVING panicked a
debug build and read unbound in release, a join matched anything, a
GROUP BY merged every list into one group, and a JSON-LD subquery
returned the list to its enclosing query, where a join on it silently
matched nothing.

- QueryOutput::Select carries UngroupedProjection: Reject by default,
  PerGroupList set only by the JSON-LD user-query entry (parse_query).
  A sub-query's projection is a plain variable list, so it cannot carry
  a per-group list at all.
- Grouping::first_ungrouped_read is the one rule, enforced in
  apply_solution_modifiers, which the top-level and sub-query pipelines
  share: HAVING, the grouping's binds, ORDER BY binds and keys, and a
  Reject projection may not read such a variable. It fails the plan
  with a 4xx in every build and replaces the two ad-hoc ORDER BY checks.
  SPARQL and JSON-LD never reach it on HAVING or ORDER BY: lowering
  reads those as SAMPLE. It is the backstop for Cypher, hand-built IR
  and any future lowerer.
- Without dependency sets (the sub-query pipeline), the traditional
  grouping lane now carries only keys and aggregate inputs, instead of
  every WHERE column as a list.
- Grouping::assemble returns an error instead of dropping a HAVING or
  binds that have neither a key nor an aggregate.
- A Grouped binding reaching scalar evaluation, a join or an OPTIONAL
  substitution is an Internal error in every build; the sort comparator
  and the group-key conversion, which cannot fail, keep a debug
  assertion.
The SPARQL JSON, NDJSON and XML writers expanded a Binding::Grouped
cell into one row per list element: a cartesian product across list
columns, and a dropped row for an empty list. That was the visible
symptom of #1978, and it still reached JSON-LD queries rendered as
SPARQL results (97 streamed rows from 2 groups on a two-list query).
SPARQL queries can no longer produce such a cell, and SPARQL results
have no list type, so the writers now refuse it with a FormatError, as
they already refused a Cypher path or list.

Deleted: disaggregate_row and the per-row list scans in format_string
and stream_ndjson_rows (every row now streams cell by cell), the XML
write_grouped_rows, and the CLI table's disaggregation fallbacks (the
table renders a stray list ';'-joined, like CSV).

The streaming endpoint refuses a JSON-LD query that projects a
per-group list with a 4xx before the stream starts, naming the column,
and plans the rest with per-group lists disallowed, so a wildcard
projection streams keys and aggregates.
JSON-LD placed a select expression before grouping unless it read an
aggregate output, so a key-only expression under groupBy -- the #1978
shape, e.g. (as (if (= ?a "Net") "network" "other") ?seg) -- came
back as a per-group list of one value repeated, where SPARQL and
Cypher give one value per group.

JSON-LD now places its select expressions with the SPARQL rule
(SelectExprPlacer), in one shared function for queries and subqueries.
In a grouping query, an expression over keys, aggregate outputs or
earlier per-group aliases, or a constant, runs once per group after
having. Unchanged: an expression over a non-key variable keeps the
documented per-group list, as does a chain over it and an alias an
aggregate reads, and an alias that is itself a groupBy key is computed
before grouping.

A per-group alias may not reuse a variable the where clause binds: it
would silently replace that variable's value after grouping, the
groupBy key included (SPARQL rejects the same alias).

The implicit output name of a bare aggregate ((count ?e) -> ?count),
which downstream code reads by name, is now pinned by a test.

Docs: the select-expression, groupBy and having sections of
jsonld-query.md describe the grouping rules; the having example used a
form that does not parse.
Cypher makes a non-aggregate RETURN expression a grouping key, so
RETURN CASE ... AS seg, count(e) groups by the label, which is SPARQL's
GROUP BY (IF(...) AS ?seg). Grouping by the key first and mapping after
(WITH e.area AS a, count(e) AS n RETURN CASE ... AS seg, n) is what a
SPARQL SELECT expression over GROUP BY ?a now means. The parity test
pins both pairs on a non-injective mapping (three areas, two labels),
where the two readings differ.

Must-not-change guards for the grouped shapes fluree/solo runs today:
a nested grouped sub-SELECT, SUM(IF(...)) buckets per key, aggregate-
only sub-SELECTs under SELECT *, a compound expression of SAMPLEs
beside a distinct count, and the grouped paging subquery ordered by
its aggregate.

SHACL sh:select constraints are lowered like SPARQL queries but skip
the SPARQL validator, so the grouped semantics reach stored shapes
directly: HAVING without grouping filters, HAVING over a non-key
variable reads a sample, and a projected non-key variable fails the
transaction with the plan error instead of reporting a violation at
the focus node.

Docs: sparql.md and compatibility.md describe grouped SELECT
expressions (one per group, after HAVING, in SELECT order), SELECT *
under grouping, implicit SAMPLE in HAVING and ORDER BY, HAVING without
grouping, and the two shapes rejected instead of evaluated literally.
…ed lowering

The planner's produced_vars_of gives the variables some pattern may
bind (must_bind_vars gives the ones bound on every row). The grouped
lowering computed the same set in four more places: the SPARQL level
(pre_group_vars) and its SAMPLE rewrite, and the JSON-LD select
computations and SAMPLE rewrite. It now lives in ir::pattern beside the
other pattern-set helpers, and all five callers use it.

No behavior change.
The plan-time grouped-read check (Grouping::first_ungrouped_read) runs
in the planner, which has no variable registry, so its 4xx named the
variable by internal id: "projected variable VarId(3) is neither a
GROUP BY key nor an aggregate result". A JSON-LD subquery projecting a
non-key variable reaches it, as does a JSON-LD select expression that
mixes an aggregate with a non-key variable.

The planner now returns a typed QueryError::UngroupedRead carrying the
ids, and QueryError::name_variables turns it into InvalidQuery naming
them wherever the query's registry is at hand: the runner's execute
and execute_prepared (which also covers a sub-query, planned while the
query runs), the API's view, dataset and stream prepare sites, and
EXPLAIN. An id the registry does not hold prints as before. Where the
typed error surfaces unnamed it keeps InvalidQuery's 400 and stream
error code.
…ader

J3 evaluates a JSON-LD select expression over group keys once per group.
When a later expression reads that alias together with a non-key
variable, as in `(as (strlen ?a) ?len)` then
`(as (+ ?len (strlen (str ?e))) ?x)` under `groupBy ?a`, the reader ran
per group as well, and the plan-time check rejected its non-key read.
Before J3 the query returned per-group lists.

An alias that a before-grouping expression reads now moves before
grouping with it: the alias is constant within its group, so the reader
stays per solution and both come back as per-group lists again.
SelectExprPlacer places a level's expressions together (place_all) so
it can see the readers. The SPARQL twin stays a V4 error, because a
SELECT expression of a grouped level cannot read a non-key variable.
aaj3f added 16 commits October 8, 2026 07:40
…gates

`WITH p, count(f) AS c ORDER BY c + 1` and `RETURN …, count(f) AS c
ORDER BY -c` failed: "an ORDER BY key is neither a GROUP BY key nor an
aggregate result". The sort key lowers as a bind, and one level placed
it in the aggregation's body, before grouping, where `c` does not exist.

The clause now takes the two-level form whenever its WHERE or ORDER BY
lowered anything the second stage can see, not only an output node's
property: the aggregation, then the binds, the filter, the sort and the
slice, as an explicit second WITH does. A clause that lowers nothing
there keeps one level, and a read the second stage cannot see stays a
plan error.
A property read joins every value, as `MATCH (p) WHERE p.age > 30`
does. Reading it after the aggregation keeps the clause's own
aggregates exact, but the rest of that model stands, and the reference
now says so: in WHERE the row comes once per value that passes; in
ORDER BY once per value, and DISTINCT keeps the copies (the sort key is
projected with them); an aggregate in a later clause counts every copy.
A read inside an aggregate's argument is still joined before grouping,
so it repeats the group's rows for every aggregate of the clause.

The test pins each in row order, so a change to the model flips them
on purpose: an unsliced ORDER BY in a WITH and in a RETURN, a WITH
DISTINCT … ORDER BY, count(*) and sum(c) after the WHERE, and
collect(p.age) beside count(f).
A table test over raised_whatever_the_severity, one representative per
QueryError variant (R2rml excepted: its error type lives in a crate
this one does not depend on). The request's budgets, cancellation,
memory, storage, catalog and policy access, data state and internal
faults must fail the write whatever the severity; the constraint's own
failures follow severity and mode. An exhaustive variant-name match
makes a new variant fail to compile until the table places it.

Three arms can't be reached today, and each now says why: the caller
names variables before classifying, which turns an UngroupedRead into
an InvalidQuery; NoSuitableIndex and UnsupportedMode are constructed
nowhere.
A MustFire that failed because its site declined at open read "must
proceed [proceeded: [..., group_by_object_count_topk]]". It now reads
"must serve it" and lists the sites that declined at open beside those
that proceeded.
As for select, a malformed value is now a 400; ask ignored both before.
Each entry of the table names a QueryError variant's pattern, the class
it must get and its representatives. A macro builds the rows and a
match over the same patterns with no wildcard, so a new variant fails to
compile until it has an entry, which brings a class and a representative
with it; each representative must match its own pattern. The earlier
version forced only the variant's name, and its count was a manual step.

R2rml has a row now (R2rmlError::Parse), through a dev-dependency on
fluree-db-r2rml, which fluree-db-query already depends on; normal,
no-default-features and wasm dependency trees are unchanged.
`WITH p, collect(f) AS fs ORDER BY size(fs)` was refused as an ORDER BY
on a collect() list: the check refused any key that read the list. The
reference says only a key on the list itself is refused, and a key
that reads the list and yields one value now sorts after the
aggregation, as other expression keys do: size(fs), a list predicate
such as any(x IN fs WHERE …), a comparison. A key whose value is the
list, or a list built from one (fs, tail(fs), fs + ['x']), is still
refused, since sorting a list value is deferred.

The reference also now says that `WITH DISTINCT` keeps the copies a
multi-valued sort key makes and `RETURN DISTINCT` removes them.
`"orderBy": "(desc (count ?e))"` failed with "orderBy must be an array of
objects with 'var' field", which is not what orderBy takes. Sort keys
are variables; the error now lists the forms and says to select an
expression or aggregate under an alias and order by the alias. An
expression in a term's variable position (`["desc", "(count ?e)"]`,
`{"var": "(count ?e)"}`) gets the same error instead of "Invalid
variable syntax".
…y alias

An aggregate can't appear in orderBy, so a grouped query is one with a
groupBy or an aggregate in select or having. The orderBy section shows
sorting on an aggregate through its select alias.
An sh:sparql constraint that cannot run was recorded with the severity
of the shape that owns it. Checked as a nested shape (sh:node, sh:not,
sh:and, sh:or, sh:xone, or sh:node and sh:qualifiedValueShape on a
property shape), that shape's own results surface only through the
outer shape's, at its severity, but the failure kept the nested
shape's: Violation by default, so it rejected the write under a
Warning outer shape, and a Warning nested shape's failure was only
logged under a Violation outer one.

The failure now takes the severity of the outermost shape that reports
the result: node-level combinators report under the node shape, the
nested checks of a property shape under the property shape, and an
enclosing nesting keeps its own. The nested shape's results keep their
severity, which the conformance check reads. Errors that fail the write
whatever the severity (budgets, cancellation, storage, catalogs, policy,
internal faults) are raised before any failure is recorded, on the
nested route as on any other.

Tests cover every nesting, including one through a Violation middle
shape, for Warning, Info and Violation outer shapes, a Warning inner
shape, and a warn-mode graph, with shapes written in JSON-LD and with
SPARQL UPDATE.
A sub-SELECT projecting an expression over its group keys, with an
outer ORDER BY on the alias, ran the expression per solution before
this branch: the alias was a per-group list, and the outer sort
compared two lists (a debug_assert in the sort comparator, and no
order in release builds). Evaluated once per group, it sorts. Pinned
for SPARQL (STRLEN of a key, CONCAT over two keys, and COUNT * 2 as a
control that already worked) and for a JSON-LD subquery.
openCypher's WITH applies its WHERE to the clause's results
(`WITH … [ORDER BY] [SKIP] [LIMIT] [WHERE]`), so
`WITH x ORDER BY x LIMIT 5 WHERE x > 3` keeps 4 and 5. Every WITH
lowering filtered before the slice and kept 4 to 8.

A WITH with SKIP or LIMIT and a WHERE now lowers without its WHERE,
and a stage over the sliced rows joins the WHERE's property accessors
and filters. The rows keep the clause's order, and the stage repeats
DISTINCT. This covers the plain, one-level aggregating and two-level
aggregating forms.

After a plain WITH, the sliced rows carry the earlier variables the
WHERE reads, including those of an exists { … } body, and the stage
drops them. After an aggregating or DISTINCT WITH, a WHERE read of
anything the clause does not project is an error rather than a free
variable.

A WITH without a slice lowers exactly as before.
The shape compiler scans SHACL predicates in a fixed order, and the
sh:severity, sh:message, sh:name and sh:description arms wrote only to
a subject a shape map already held. A node shape registered later lost
them: one whose only constraint is sh:sparql is registered by its
rdf:type sh:NodeShape after the scan, so it compiled as a Violation
shape with no message, and a Warning or Info shape rejected the write.
So did a shape whose registering predicate is in a later graph.

The arms now record metadata for an unregistered subject, and
claim_metadata gives it to the shape once the scan is done. It routes
the metadata as the arms do (property shape first) and keeps any value
the shape was given after it was registered.

The nested-severity test no longer needs its inner shape to target an
unused class.
…ction

A JSON-LD @list or a Turtle collection stores one statement per
member, and the sh:and / sh:or / sh:xone arms took each object for a
shape. A SPARQL UPDATE stores `( … )` as an rdf:first / rdf:rest
collection, so the arm got the collection's head, and validation
failed with "Referenced shape … could not be resolved".

List processing now replaces a member that heads a list in the graph
with the list's members, as it already does for sh:ignoredProperties;
sh:in, sh:languageIn and sh:path already read both forms. The
and_list / or_list / xone_list fields, which nothing ever set, and
their expansion loops are gone.

The nested-severity test now writes every nesting with SPARQL UPDATE
too.
…very query form

Parameter substitution refuses a parameter that the query itself
assigns (BIND, VALUES, AS), and for SELECT that includes the trailing
VALUES clause. ASK, CONSTRUCT and DESCRIBE now take a trailing VALUES
too, and their arms didn't walk it, so a parameter named there was
substituted in the WHERE while the VALUES still bound the variable.
All four forms now walk their trailing VALUES through one helper.
@aaj3f
aaj3f removed the request for review from zonotope October 8, 2026 13:41
aaj3f added 7 commits October 8, 2026 10:59
…riable it carries

The item-31 test's exists cases gave the same rows whether the body's
f was carried or free. Two cases tell them apart: exists over a carried
f keeps Alice's one row whose friend she likes, and a body variable of
its own (g) is free, so both of her rows pass.
`CALL (p) { MATCH (p)-[:knows]->(f) WITH f RETURN f.name }` crossed every
outer p with every knows edge. The WITH lowers to a subquery that
selects f alone, and a subquery correlates with its parent only on what
it selects, so the CALL's seeded p never reached it. A slice in the
body ran once for all of them.

Inside a CALL body, a subquery the Cypher lowering builds (a WITH, its
two-level and after-the-slice stages, an aggregating RETURN's first
level) now selects the imports it reads, groups by them when it groups,
and pins them. The imports it carries without projecting stay out of
scope after the WITH: a nested CALL can't import them, and WITH * doesn't
bring them back.

The planner treated a CALL whose body binds its import as uncorrelated,
because the nested WITH now selects it, and could place the CALL before
the import's producer. A pinned import is now always a correlation input
for placement. The executor still evaluates once and hash-joins when
the body allows it, and runs per row when a slice needs it.
…r their severity

Whether a node conforms to a nested shape (sh:node, sh:not, sh:and,
sh:or, sh:xone, sh:qualifiedValueShape, at node level and per value)
counted only Violation results. SHACL §3.5 counts every result: a node
conforms to a shape when validating it there reports none, and §2.1.4
has severity only categorize results. So a Warning nested shape that
reported a result still conformed: sh:node passed and sh:not failed.

The six sites now share one `conforms` check: the node-level
combinators and check_value_against_nested_shape, which serves every
per-value form and the literal-focus checks. The severity that decides a
write is the outermost reporting shape's, on the result it reports. A
constraint that cannot run is not a result, and is decided at that same
outermost severity, as before.
@aaj3f
aaj3f force-pushed the fix/sparql-grouped-projection branch from c67dc35 to 87576fd Compare October 8, 2026 20:14
@aaj3f

aaj3f commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

The deviations from the design, each written up in full: moved out of the description to keep it under GitHub's size limit. They describe the branch at 87576fdb1.

  • Placement rule, three refinements (ir::SelectExprPlacer, shared by SPARQL and JSON-LD) over the design's literal version of the JSON-LD rule (JSON-LD keeps per-group lists for non-key variables):
    • A read of a variable that nothing binds before grouping doesn't keep an expression before grouping (it's unbound either way — the same rule the plan-time check uses). Without this, the design's literal rule would place SPARQL (EXISTS { ?key ex:p ?inner } AS ?x) before grouping (?inner is not a key) and then fail the plan.
    • An expression that reads an aggregate output or an earlier per-group alias is always per-group (today's JSON-LD rule, kept as the first clause). A mixed read of an aggregate output and a non-key variable is therefore rejected by the plan-time check, instead of silently becoming a list.
    • A key-only alias that a per-solution expression reads is evaluated per solution too (from review). place_all places forward, then moves such aliases back before grouping until nothing changes, so the reader sees a value per solution and both stay per-group lists in JSON-LD, as at base. In SPARQL the reader is a grouped SELECT expression over a non-key variable, which V4 rejects.
  • The shared placement helper arrives with the SPARQL commit, not the JSON-LD one, so the JSON-LD commit only touches JSON-LD lowering.
  • One SPARQL function for a SELECT level (lower_select_level), used by the top level and sub-SELECTs, instead of the design's per-call-site sequence. lower_select_extends takes the level's pre-group variables (the helper needs them).
  • JSON-LD alias guard implemented standalone. Reject Cypher alias and JSON-LD bind collisions that silently returned zero rows #1867 isn't merged, so this PR doesn't extend its check_bind_targets; it adds a narrow check that a per-group alias is not a variable the WHERE binds (which closes the shape where a post-group alias silently replaced a group key, e.g. (as (str ?n) ?a) under groupBy ?a). It should fold into check_bind_targets when this rebases onto Reject Cypher alias and JSON-LD bind collisions that silently returned zero rows #1867.
  • Streaming lane, wildcard (the design didn't cover it; review caught an earlier head of this branch silently narrowing the columns on the stream, so the wildcard now gets the same refuse-before-the-stream treatment as an explicit list). What a JSON-LD select "*" under groupBy returns on /query depends on the grouping lane: with aggregates that all stream, the streaming GroupAggregateOperator returns the keys and aggregates only; otherwise the GroupByOperator lane also returns the other WHERE variables as per-group lists. The stream refuses the wildcard, before the response starts, only on the GroupByOperator lane when the WHERE binds a variable that is not a key — naming the list columns and pointing at /query or an explicit projection of the keys and aggregates. Grouping::aggregates_stream gives the lane from the same predicate as the planner (AggregateFn::is_streamable, which GroupAggregateOperator::all_streamable now delegates to). The stream still plans with Reject.
  • Trailing VALUES placement is kept, a deviation from §18.2.4.3 now documented in docs/reference/compatibility.md (0509e806e), which also gives the portable forms. I decided to keep base placement: a trailing VALUES clause joins right after the WHERE, before grouping, at the top level and in a sub-SELECT, as at base. It therefore restricts what the aggregates count (the "VALUES as a parameter" idiom: … GROUP BY ?a VALUES ?e { ex:e1 } gives (Net, 1)), and a VALUES variable the WHERE doesn't bind multiplies their input (VALUES ?x { 1 2 } doubles COUNT). The spec joins it after HAVING. An earlier commit on this branch (97eb988d0) moved it there for grouped queries; that silently changed the idiom's answers, and af22f8e49 reverts it. What the branch keeps is HAVING's view (089e6c96e): the SELECT-level lowering takes the trailing VALUES pattern at both levels, counts its variables as bound before grouping (so ORDER BY can read one as a SAMPLE), and, before the SAMPLE rewrite, has HAVING read each one the WHERE doesn't bind, and that is not a key or an aggregate output, as a fresh variable nothing binds: unbound, as after HAVING. That's the same renaming that already kept a HAVING-as-filter from seeing the SELECT expressions, factored out as ir::read_as_unbound. The generated binds (aggregate inputs, GROUP BY expressions, SELECT expressions placed before grouping) now follow the join at the top level too, as they already did in a sub-SELECT (e6e2061e5): Query.post_values carries them as PostValues::then.
  • HAVING shares FILTER's evaluator (from review; not in the design). FilterOperator's predicate is now a RowPredicate that HavingOperator uses too, so an EXISTS in HAVING is resolved per group row (uncorrelated once per batch, correlated per row with the semijoin cache). A Cypher metadata read in HAVING (an aggregating WITH … WHERE) under a non-root view policy now resolves through the policy filter too; it read as empty. A HAVING with neither runs the same synchronous batch filter as before. HavingOperator::new takes the planning context for the EXISTS subplans. JSON-LD having has no EXISTS form, so that surface has no twin to fix; a test pins the refusal.
  • ASK and CONSTRUCT group (from bplatz's review). Both lower their solution modifiers through lower_select_level with a SELECT that projects nothing. ASK keeps LIMIT 1, which now stops at the first surviving group. CONSTRUCT renames each template variable the grouping doesn't produce to a variable nothing binds, so its triples are skipped (§16.2: an unbound template variable produces no triple). JSON-LD ask parses groupBy / having. DESCRIBE still refuses them. A CONSTRUCT with an aggregate ORDER BY is now one implicit group, where it used to be refused (sparql_construct_aggregate_order_by_groups).
  • The Cypher parity test lives in it_query_sparql_grouped_projection.rs (grp_query_sparql) rather than it_query_cypher.rs, to stay clear of Reject Cypher alias and JSON-LD bind collisions that silently returned zero rows #1867's changes to that file.
  • The indexed-ledger lanes test (fast paths on and off) is its own test binary, it_grouped_projection_lanes, because it toggles the process-global fast-path switch (the reason it_distinct_object_numbig_gate is its own binary). It also covers the dedup-only shape.
  • Grouping::assemble errors have two variants (HavingWithoutGrouping, BindsWithoutGrouping); SPARQL maps them to a new LowerError::InvalidGrouping, JSON-LD to ParseError::InvalidOption, Cypher to its generic lowering error. All three are unreachable from the current lowerers.
  • Trim under Reject also keeps COUNT(DISTINCT *)'s visible variables into the traditional grouping lane, which reads them.
  • SELECT * with an explicit GROUP BY (reachable only from unvalidated entry points) projects the user-visible keys.
  • Docs: the JSON-LD having example used a form that doesn't parse ([["filter", …]]); it now uses the string form.
  • SHACL sh:select isn't in the design. It's lowered like SPARQL but never validated, so it inherits the grouped semantics (HAVING filter, implicit SAMPLE). A pre-bound variable ($this, $PATH) is constant within one evaluation, so a (sub-)SELECT level that projects one is also grouped by it, which changes no group (from review). That's what lets the family of sub-SELECTs that must project $this (the pre-binding restriction) evaluate: at base they never fired, and at the first head that review ran against, they refused every write. A shape query the planner still rejects (a projected non-key variable, or one nothing binds) fails as ShaclError::SparqlConstraint, whose message names the shape and the variable, and is decided by severity and mode (below); its constraint field is now the constraint's IRI (String, was its Sid). Pinned by the shacl_tests below; listed in the behavior table.
  • An existing test (sparql_order_by_expression_over_grouped_var_errors_cleanly) pinned the error for ORDER BY over a non-key variable; per my earlier call to follow the spec where it's clean, that shape now returns rows, so it pins the sampled result instead.
  • Error names (not in the design). The plan-time check has no variable names, so it returns a typed QueryError::UngroupedRead carrying ids. QueryError::name_variables turns it into InvalidQuery naming them wherever the query's registry is at hand: the runner's execute and execute_prepared (which also covers a sub-query, planned at run time), the API's view, dataset and stream prepare sites, and EXPLAIN. Nothing is plumbed into the planner. An error that surfaces unnamed keeps InvalidQuery's 400 and stream error code. A grouped projection of a variable nothing binds is the same typed error (ReadStage::UnboundProjection), raised by the post-group schema check where VariableNotFound("Selected variable VarId(n) …") was, and the stream's pre-flight renders the same message. A synthetic variable no user wrote (?__…, ?#…, _:…; for example the variable a Cypher property access e.area lowers to) is described instead of named ("an ORDER BY key is neither …"). That internal-name rule now has one definition, var_registry::is_internal_var_name, which the formatters' wildcard filter re-exports and SPARQL's SELECT * lowering uses; both had their own copy. Plan errors outside the grouped tail that still print ids are a follow-up.
  • produced_vars_of shared (after the rebase). main's planner already had produced_vars_of (the variables a pattern list may bind). The grouped lowering computed the same set in four more places; it now lives in ir::pattern, and all five callers use it. SelectExprPlacer doesn't duplicate main's placement machinery (Settled in where_plan): Settled decides where in the WHERE join chain a BIND or FILTER can run, while the placer decides whether a SELECT expression runs per solution or per group, from the GROUP BY keys, the aggregates and the WHERE's may-bind set. A per-solution expression becomes a WHERE BIND, and main's Settled then schedules it.
  • HAVING reads SELECT aliases (an extension of §18.2.4; my call). Grouping::binds_before_having splits the grouping's binds: the ones HAVING reads, directly or through another bind, run before it, and the rest after it, each part in SELECT order and each once per group. The executor, the plan-time ungrouped-read check and dependency tracing walk the same order. SPARQL 1.1 leaves such an alias unbound in HAVING. Many engines read it, and the W3C aggregate tests repeat the aggregate in HAVING, so they pin neither reading. Documented once in the compatibility reference, for simple and compound aliases alike.
  • EXISTS-body variables are free over the group row (from bplatz's review). Expression::row_reads / substitute_row_read skip EXISTS bodies. The implicit SAMPLE, the plan-time check, the SELECT-expression placer and dependency tracing all use them. This replaces this PR's earlier reading, where a non-key variable in an EXISTS body read a SAMPLE.
  • A Cypher WITH's WHERE filters after its slice (my call, aligning with openCypher).
    • The evidence was read, not executed. openCypher's grammar places the WHERE last: <with statement> ::= WITH <return statement body> [ <order by and page clause> ] [ <where clause> ]. The openCypher 9 reference says that after WITH, "WHERE simply filters the results". The TCK has no WITH with both, and no openCypher engine was run.
    • The lowering. A WITH with both lowers without its WHERE. A stage over the sliced rows then joins the WHERE's property accessors and filters, and repeats DISTINCT, because an accessor over a property with several values can repeat a row.
    • Scope. After a plain WITH, Fluree's WHERE also sees the variables before it, as it always did with no slice, where it filtered in the clause's body. So the sliced rows carry the ones it reads (an exists { … } body's included), and the stage drops them. After an aggregating or DISTINCT WITH, it sees only the projection, and any other read is an error rather than a free variable.
    • Unsliced shapes. A WITH without a slice lowers exactly as before.
  • A WITH inside a CALL body carries the imports (pre-existing).
    • The bug. A subquery correlates with its parent only on what it selects. A WITH f after MATCH (p)-[:knows]->(f) selected f alone, so the CALL's per-row p never reached it, and the CALL crossed every outer row with the body's whole result.
    • The lowering. Inside a CALL body, every subquery the Cypher lowering builds now selects the imports it reads, groups by them when it groups (as the CALL's own grouping RETURN already did), and pins them. That covers a WITH, its two-level and after-the-slice stages, and an aggregating RETURN's first level.
    • Scope. An import carried this way but not projected stays out of scope after the WITH. A nested CALL can't import it, and WITH * doesn't bring it back.
    • The planner, one rule (subquery_correlation_vars). A pinned import counts as a correlation input even when the body binds it. Before, a CALL whose body binds its import looked uncorrelated and could be placed before the import's producer. The executor still picks evaluate-once plus hash join when the body allows it, and per-row when a nested slice needs it. Only Cypher sets pinned_vars, so SPARQL and JSON-LD planning is unchanged.
  • Cypher reads after the aggregation. When an aggregating WITH's WHERE / ORDER BY or an aggregating RETURN's ORDER BY lowers anything (a property accessor of a node the clause projects, or the bind of an expression sort key such as c + 1), the clause lowers as two levels: the aggregation, then a stage that joins the accessors and computes the binds, filters on the WHERE, sorts and slices, which is what an explicit second WITH lowers to (with a SKIP or LIMIT, the WHERE moves to a stage after the slice; see the bullet above). Every read must be visible to the second stage; otherwise the clause stays one level and a property of a node the clause doesn't project stays a plan error, matching Cypher's scope rule. This replaces an earlier SAMPLE rewrite on this branch (23c3f1ee2), which joined the property before grouping and so repeated a group's rows once per value of a multi-valued property (422d986c4, ad89f13b9). The second stage keeps Fluree's model for a property with several values: it joins every value, as MATCH … WHERE does, and the reference says what that does in WHERE, ORDER BY, DISTINCT and a later aggregate (5422809d4).
  • Nested conformance counts every result (my call, per SHACL §3.5). Whether a node conforms to a nested shape counted only Violation results. It now counts every result: one conforms check covers all six sites (the node-level sh:node, sh:not, sh:and, sh:or and sh:xone, and check_value_against_nested_shape, which serves every per-value form, qualified value shapes and the literal-focus checks). Severity still decides the write, on the result the outermost referencing shape reports. It composes with the next bullet: a constraint that cannot run is not a result, so it doesn't change conformance, and it is decided at that same outermost severity.
  • sh:sparql failures follow severity and mode (my call). The engine records them when the caller collects them (the transaction path) and raises them otherwise. A Violation shape in a reject-mode graph fails the write with the same ShaclError::SparqlConstraint. Which query errors count as the constraint's own failure is an exhaustive match over QueryError (raised_whatever_the_severity): the request's budgets, storage / catalog / policy access, data state and internal faults are raised whatever the severity; everything else (an invalid query, filter or expression, a named plan error, an unsupported feature, arithmetic and comparison errors) is the constraint's failure. A new variant has to be placed. A failure in a shape checked as a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, and sh:node or sh:qualifiedValueShape on a property shape) takes the severity of the outermost shape that reports the result, as that shape's own results do (also my call): node-level nesting reports under the node shape, a property shape's under the property shape, and an enclosing nesting keeps its own. The nested shape's results keep their severity, which the conformance check reads (7999b625e).
  • ASK applies OFFSET and LIMIT. ASK is true when a solution remains after them: OFFSET applies and LIMIT becomes at most 1. ORDER BY cannot change the answer and is dropped. With an OFFSET, ASK no longer takes the WHERE planner's multiplicity-blind license.
  • A trailing VALUES after every query form. One parser helper reads it after SELECT, ASK, CONSTRUCT, DESCRIBE and a sub-SELECT. ASK and CONSTRUCT lower it into post_values as SELECT does; DESCRIBE appends it to its target subquery's WHERE, as a sub-SELECT does. SPARQL parameter substitution (feat: SPARQL parameters, fulltext() in SPARQL, structured SHACL rejections #2024) walks it after every form, as it did after SELECT, so a parameter the trailing VALUES assigns is refused rather than substituted in the WHERE while the VALUES still binds it (bd3797dcc).
  • The count top-k and the bound-object count stamp a decline at open. Both stamp Proceed at plan time; a fallback at open (novelty, policy, history) now also stamps Fallback(GateDeclined), as the star top-k and the directory count already did.
  • Sort keys nothing binds are dropped at plan time, for the top level and each sub-query operator. A key some stage binds is kept, so a plan that lost one still fails validation.
  • The stream's pre-flight runs Grouping::first_ungrouped_read with the query's own projection policy.
  • The WHERE cursor names variables at plan, open and pull, so every update surface gets named grouped-read errors.
  • The fused R2RML aggregate now records a routing stamp (fused_r2rml_aggregate) at open, for its canary.

@aaj3f

aaj3f commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Moved out of the description to keep it under GitHub's size limit: the full test list (current at 87576fdb1), and the perf detail and gate log from when this PR was first opened, at the pre-rebase SHAs they name.

Full test list

  • grp_query_sparql — it_query_sparql_grouped_projection.rs (new): the SPARQL matrix on a non-injective fixture (3 areas → 2 labels), each case asserting SPARQL-JSON rows and JSON-LD values; alias chains; ORDER BY / LIMIT order; sub-SELECT scalars under outer FILTER / GROUP BY / aggregates; HAVING reading an aggregate or expression alias, through a chain, and RAND() / STRUUID() tested as returned; STRUUID once per group; empty implicit group; SELECT * under implicit grouping (top level and as a sub-SELECT joined outside); every results format (XML, CSV, TSV, typed JSON); implicit SAMPLE for HAVING (both lanes) and ORDER BY (bare and expression); HAVING as a filter; V4 under HAVING-only grouping; V4 for an expression over a non-key variable read by another expression; V009; trailing VALUES before grouping, as at base (VALUES as a parameter restricts COUNT, a VALUES-only variable multiplies it, ORDER BY samples one), and HAVING reading a VALUES-only variable as unbound (grouped and ungrouped, top level and sub-SELECT; a VALUES variable that is a key is still read); grouped ASK (true / false through grouped and implicit HAVING, one implicit group over no solutions, EXISTS free over the group row) and CONSTRUCT (per group, a skipped non-key triple), DESCRIBE still refused; HAVING EXISTS / NOT EXISTS per group, EXISTS-body variables free (flag on e1 or e2, GROUP_CONCAT lane, EXISTS SELECT expression); Cypher parity; a Cypher grouped-read error that prints no internal variable (now on a node the WITH doesn't project); solo must-not-change guards; ORDER BY a variable nothing binds (grouped, ungrouped, sub-SELECT); generated binds read a trailing VALUES at both levels; Cypher key-node properties after an aggregating WITH / RETURN (WHERE, ORDER BY, composite alias, exists {}); the trailing-VALUES internal-predicate check (it_query_sparql_annotations); CONSTRUCT with an aggregate ORDER BY (it_query_construct).
  • grp_query — it_query_jsonld_grouped_projection.rs (new): the JSON-LD twins (SAMPLE, having-as-filter), key-only select expressions as scalars, unchanged list shapes (including a key-only alias read by a per-solution expression), the subquery error, the grouped-read error naming ?x with no VarId( (subquery, planned at run time, and a top-level expression, planned before), SPARQL-results refusal of lists, ask with groupBy / having, having refusing an EXISTS form, the ?count name pin, solo guards; having reading a select expression alias; orderBy a variable nothing binds; top-level values reaching select expressions; the update grouped-read error naming ?x (and the SPARQL UPDATE twin, refused at validation).
  • grp_transact — SPARQL UPDATE over a grouped sub-SELECT expression (with and without an aggregation stage), and the JSON-LD update twin over a grouped subquery.
  • grp_misc — it_stream_query.rs: SPARQL grouped expression streams 3 rows; a JSON-LD per-group list is refused before the stream; select * under groupBy is refused on the stream when /query returns per-group lists, and streams when every WHERE variable is a key or when every aggregate streams (/query then returns the keys only); a projected variable nothing binds is the same named 400 on /query and the stream. A grouped SELECT expression's read is the same 400 on /query and the stream, before it starts; ORDER BY a variable nothing binds streams every row.
  • it_query_explain — plan pins: streaming GroupAggregateOperator under the Extend (COUNT and MIN); dedup-only grouping with early dedup.
  • it_policy_cypher — an aggregating WITH … WHERE size(keys(n)) = 1 under a view policy that hides one property sees the policy-filtered keys (with a no-policy control).
  • it_grouped_projection_lanes (new binary; toggles the process-global fast-path switch) — bulk-imported indexed ledger, fast paths on and off, hand-derived answers. Routing stamps as canary pairs for every aggregate fast path that gates on the grouping's binds: the GROUP BY ?o count top-k, the star top-k, the per-predicate directory count, COUNT(*), SUM(?o) and the whole-graph scalar aggregates. Each serves the key-only query (MustFire) and declines it with a SELECT expression (MustNotFire, with no fast-path stamp but the fused chain's). Each also declines a trailing VALUES.
  • fluree-db-api lib shacl_tests (feature shacl) — sh:select constraints: HAVING without grouping filters; HAVING over a non-key variable samples; a projected non-key variable fails validation with a sh:sparql constraint failure naming the constraint, its shape and ?value; the seven probe shapes from review (a sub-SELECT grouped by another variable that projects $this, sub-SELECTs grouped explicitly and implicitly, HAVING with and without GROUP BY, ?message and ?value expressions over $this and an aggregate) fire on the violating node and pass the conforming one. Severity × mode cells for plan, parse and lower failures, including the fail-closed cell, and a failing Warning shape next to an enforced Violation shape.
  • Unit: SPARQL lowering placement (every grouping form, SELECT order, sub-SELECT, SELECT *); validator (V4 counting a HAVING/ORDER BY aggregate as grouping, and V009); SelectExprPlacer rule; sample_ungrouped_reads (keys / aggregate outputs / Extend outputs / never-bound left alone, JSON-LD pre alias sampled, EXISTS renamed, SortSpec moved, idempotent, reuses an existing SAMPLE); having_as_filter; Grouping::assemble errors; first_ungrouped_read stage × variable-kind × policy table; rewrite-then-predicate pairing; apply_solution_modifiers without dependency sets; eval on a Grouped binding; formatter refusal (JSON, NDJSON, XML); binds_before_having; EXISTS bodies are not group-row reads; sort keys nothing binds dropped, and a lost key still an error.
  • it_sql_pushdown_lane (features sql,native): the grouped-statement canary pair (with a SELECT expression the query stays in the engine).
  • it_fused_aggregate_routing (new binary, features iceberg,native; its own binary because it asserts the absence of a span-capture event): the fused R2RML aggregate canary pair on the local people Iceberg fixture.
  • Updated: sparql_order_by_expression_over_grouped_var_errors_cleanly pinned the old error for ORDER BY over a non-key variable; it now pins the sampled result. Two hand-built-IR executor tests updated for the new error text / projection. cypher_grouped_read_error_names_no_internal_variable now uses a node the WITH doesn't project; having_exists_is_evaluated_per_group (third case: every group kept); test_build_operator_tree_validates_sort_vars became …_drops_sort_keys_nothing_binds (in both unit-test modules); sparql_construct_aggregate_order_by_is_rejected became …_groups.

Added since the review, by binary:

  • grp_query_sparql:
    • cypher_output_node_properties_are_read_after_aggregation: Alice with two ages, with count, sum and collect under WHERE, a WITH's ORDER BY and a RETURN's ORDER BY, hand-derived, and the explicit two-WITH form. It also pins the multi-value model in row order: an unsliced ORDER BY in a WITH and in a RETURN, WITH DISTINCT … ORDER BY, count(*) and sum(c) after the WHERE, and collect(p.age) beside count(f).
    • ORDER BY c + 1 and ORDER BY -c after an aggregating WITH, and ORDER BY c + 1 after an aggregating RETURN (cypher_reads_a_key_nodes_property_after_grouping).
    • aggregate_over_a_variable_nothing_binds_is_a_named_error (both query paths) and cypher_aggregate_of_a_sibling_aggregate_is_a_named_error.
    • ask_applies_offset_limit_and_values (SPARQL and JSON-LD, grouped included, in it_query_ask) and construct_and_describe_read_a_trailing_values (both CONSTRUCT forms, grouped, DESCRIBE with and without a WHERE, and the JSON-LD construct values twin, in it_query_construct).
    • subselect_grouped_expression_sorts_in_the_outer_query, with its JSON-LD twin in grp_query (d466ca9df): a sub-SELECT's grouped SELECT expression sorts in the outer ORDER BY. At base it was a per-group list there, and the outer sort compared lists: a debug build panics at sort.rs:208 ("Grouped bindings should not appear in sort comparisons"), a release build leaves the rows unsorted.
    • cypher_call_body_with_keeps_the_import, inside CALL (p) { … }: WITH f; WITH f ORDER BY f.age LIMIT 1, and the same with a WHERE after the slice; WITH DISTINCT; OPTIONAL MATCH … WITH count(f) (a friendless p counts 0); a WHERE reading p; and an aggregating RETURN … ORDER BY c + 1. Each gives the per-p answer. After WITH f, and after WITH f WITH *, a nested CALL (p) is refused.
    • cypher_with_where_filters_after_the_slice: a WITH's WHERE after LIMIT, after SKIP alone and without a slice (unchanged), on each lowering path: plain, including a WHERE reading a variable from before the WITH and an exists { … } that correlates with them; aggregating in one level and in two; and DISTINCT. The rows are pinned in order. It also pins that the variables the WHERE read leave scope (a later MATCH (p:P) binds p afresh), and that a property with several values gives a row once per passing value, or once after DISTINCT. Three reads of a variable an aggregating or DISTINCT WITH doesn't project are errors. Two exists cases tell a carried f from a body variable of its own, g.
  • grp_query: jsonld_aggregate_variable_errors_name_the_variable.
  • it_query_cypher: ORDER BY size(friends) after an aggregating WITH and a RETURN, ORDER BY any(x IN friends WHERE …), and friends, tail(friends), friends + ['x'] still refused (cypher_collect_through_with).
  • shacl_tests:
    • shacl_sparql_constraint_failure_follows_severity_and_mode gains an aggregate over a variable nothing binds, two aggregates with one output and an unknown function, and asserts the message names the shape and prints no VarId.
    • shacl_nested_constraint_failure_takes_the_outer_severity: every nesting (sh:node, sh:not, sh:and, sh:or, sh:xone, property sh:node, sh:qualifiedValueShape, and a chain through a Violation middle shape) × a Warning / Info / Violation outer shape, a Warning inner shape, and a warn-mode graph; shapes written in JSON-LD and with SPARQL UPDATE.
    • shacl_sparql_only_shape_keeps_its_severity: a Warning, Info and Violation shape whose only constraint is sh:sparql, targeting its own instances as a class, written in JSON-LD and with SPARQL UPDATE.
    • shacl_nested_warning_results_count_toward_conformance: a Warning inner shape whose check reports a result for the node, either an sh:sparql constraint or a Warning property shape. It's reached through each of seven nestings (sh:node, sh:not, sh:and, sh:or, sh:xone, and sh:node and sh:qualifiedValueShape on a property), under a Violation and a Warning outer shape, written in JSON-LD and with SPARQL UPDATE. sh:not commits; the rest reject under Violation and commit under Warning. All 28 Violation-outer cells failed before the fix.
    • shacl_logical_lists_resolve_from_either_surface: sh:and, sh:or and sh:xone over two shapes each, a conforming and a violating node for each, both surfaces.
  • it_grouped_projection_lanes: JSON-LD canary pairs for the count top-k and the directory count with a per-group list; the per-group IRI list through both JSON-LD formatters; the whole-graph aggregates' trailing-VALUES case; a MustFire also requires no decline stamped at open, and a canary failure lists the sites that declined at open; and a second indexed ledger with overflow integers (SPARQL and JSON-LD, both lanes, typed JSON for the datatypes).
  • fluree-db-shacl unit: constraint_query_errors_are_classified_by_variant, one representative per QueryError variant. The classifier table is built by a macro whose patterns are also a match with no wildcard, so a new QueryError variant fails to compile until it has an entry; R2rml has its row (a dev-dependency on fluree-db-r2rml, already in the graph).
  • fluree-db-query unit: test_parse_order_by_expression_names_the_alias_route, every orderBy term form.
  • fluree-db-sparql unit: test_ask_keeps_offset_and_caps_limit and test_trailing_values_after_every_query_form (lowering); a_parameter_the_query_assigns_is_an_error adds a trailing VALUES after ASK, CONSTRUCT and DESCRIBE (with and without a WHERE).
  • Updated: shacl_nested_constraint_failure_takes_the_outer_severity now writes every nesting with SPARQL UPDATE too, and its inner shape no longer targets an unused class.

Perf detail, when this PR was first opened

New bench. Local run: bench profile (fat LTO), quick profile, 16-core dev Mac shared with other jobs (load 6–16). Before and after binaries were built from the same bench code: base sources, then this branch at its pre-rebase tip. Runs alternate before/after, and the two tiny reps use opposite order. Tiny figures are rep 2; rep 1 agrees within 8% on every figure. After the fix, expr_count and jsonld_keyonly_expr are within noise of key_count, as the design expected. dedup_expr runs below it (0.74 vs 1.00 ms at small, 90 vs 107 µs at tiny) because it has no aggregate: a dedup-only GROUP BY gets WHERE-level early dedup (pinned in it_query_explain), so it groups the distinct keys instead of counting every entity. implicit_const_count is far below it too (one group and no key). subselect_expr is expr_count wrapped in a sub-SELECT, so expr_count is its control: the gap between them (1.97 vs 0.99 ms at small after; 25.9 vs 20.0 ms before) is the sub-query wrapper, present on both sides. expr_min aggregates with MIN over IRIs rather than COUNT, so key_count is not its control.

Watched benches, which the design expected not to change: query_hot_bsbm (Q9: GROUP BY + COUNT + HAVING + ORDER BY), query_overlay_matrix (incl. groupby_{base,cached,novelty,overlay}), query_hot_negation_count (incl. grouped_join_count), query_hot_whole_graph_agg and query_hot_bsbm_bi. All runs are interleaved before/after on a shared 16-core machine. A single rep can swing ±15% at tiny and up to ±40% at small under a load spike, so every conclusion here rests on several reps. The tiny and small sweeps were taken at the pre-rebase tip, before the fixes from the review pass; the check at dcb896772 follows them.

Tiny (CI's scale, 10% budget), four reps, two in each order. All 29 scenarios have a mean change within ±5.3%: Q9 +1.3%, groupby_base +4.2%, grouped_join_count −0.4%, whole-graph count +4.2%. Two moved the same way in all four reps:

  • join_count was faster by 1.0–1.6%.
  • not_exists_after_optional was slower by 0.5–5.0% (mean +2.5%). The query has no grouping.

Small (5% budget). I ran the whole suite twice in opposite orders, the overlay matrix twice more, and then the outliers in isolation using criterion's filter, three alternations each. One overlay rep ran under a load spike of 44 and is excluded.

  • query_hot_bsbm: Q9 −3.2%, Q5 −2.6%. query_hot_whole_graph_agg and query_hot_bsbm_bi are within ±5%.
  • grouped_join_count: +17.4% and +4.0% in the full runs, then +4.8%, +2.0% and +2.8% isolated (mean +3.2%). That's within budget but always positive. On its per-row path this PR changes only Binding::Grouped arms that never run there: a debug_assert in the group-key function, which release builds compile out, and an error return in join substitution.
  • Quiet box, session 1 (2026-10-01, three AB/BA/AB rounds of the previous head's merge bd885e365 against main): two query_overlay_matrix losses, both on the COUNT query, count_cached +5.1% (+0.9 µs) and count_novelty +9.4% (312.6 → 342.0 µs); count_base was +3.9%, stable.
    • count_base and count_cached were real per-query work this PR added. Every query level collected the variables its WHERE binds, and a grouping level also built a SELECT-expression placer and a copy of those variables for the SAMPLE rewrite. V009 and the plan-time ungrouped-read check built sets too. All of it ran even when the level had nothing to sample or place. c67dc3516 collects them only for a stage that reads them. Per query on the COUNT query (n=1000), main / before / after: allocations 190 / 197 / 189 indexed and 1238 / 1245 / 1237 novelty; instructions 168.8k / 173.6k / 168.1k indexed. In the bench itself (the box's trees, an instrumented copy of query_overlay_matrix, M4), count_base and count_cached run 186–187k instructions per query at c67dc3516 and on main, against 191–192k before it.
    • count_novelty runs no code of this PR per row or per flake. Per-row allocations are identical (the previous head added 7 allocations per query at every n and on both ledgers). On an M4 the box's exact trees matched main, alone and in the full matrix: 3.40–3.47M instructions and 671–685k cycles per query on both sides. Cross-built for x86-64 (rustc 1.97, fat LTO, config L), every function on that path is instruction-for-instruction the same with immediates masked; the 178 functions that differ are the per-query code this PR touched. What moved is placement: the hottest loop, SipHash-1-3's write (about a third of the query's CPU, hashing fact keys in remove_stale_flakes), sits at offset 0 of a 64-byte line on main and at offset 32 in bd885e365. At c67dc3516 the SipHash loop is back at offset 0, and its callers are at new offsets. Session 2 confirmed the placement reading under perf stat (part C below): equal instructions with different cycles.

Quiet box, session 2, at c67dc3516 (merge 2caedb047; same box type, harness and rule; 3 rounds, plus count_novelty run alone). The rows for this PR's target and the count queries; every other row is stable:

bench scenario scale base median [range] head median [range] Δ head>base (paired rounds) budget verdict
query_hot_grouped_projection expr_count/small small 47.116 ms [46.887 ms–47.413 ms] 1.771 ms [1.750 ms–1.869 ms] -96.2% 0/3 5% win
query_hot_grouped_projection key_count/small small 1.716 ms [1.702 ms–1.753 ms] 1.798 ms [1.739 ms–1.811 ms] +4.8% 2/3 5% stable
query_overlay_matrix count_base/small small 16.65 µs [16.53 µs–16.68 µs] 17.01 µs [16.82 µs–17.47 µs] +2.2% 3/3 5% stable
query_overlay_matrix count_cached/small small 16.85 µs [16.73 µs–16.89 µs] 17.14 µs [17.09 µs–17.54 µs] +1.8% 3/3 5% stable
query_overlay_matrix count_novelty/small small 310.91 µs [310.53 µs–314.86 µs] 318.84 µs [315.30 µs–333.20 µs] +2.5% 3/3 5% stable
query_overlay_matrix@count_novelty_alone count_novelty/small small 311.53 µs [308.57 µs–313.87 µs] 320.15 µs [315.17 µs–320.93 µs] +2.8% 3/3 5% stable

Part C, perf stat. Each count scenario alone per binary, per-iteration medians over 3 alternations; the old merge is bd885e365, the new one 2caedb047:

scenario binary instructions/iter, median [range] cycles/iter, median [range] IPC Δ instr vs base Δ cycles vs base n
count_base base 133,151 [132,429–133,174] 61,454 [60,899–61,613] 2.167 +0.00% +0.00% 3
count_base old #2006 merge 136,895 [136,131–136,973] 62,557 [62,181–65,078] 2.188 +2.81% +1.79% 3
count_base new #2006 merge 132,440 [132,357–132,905] 60,576 [60,241–60,850] 2.186 -0.53% -1.43% 3
count_cached base 134,554 [133,595–134,958] 61,177 [61,008–62,218] 2.199 +0.00% +0.00% 3
count_cached old #2006 merge 138,012 [137,559–139,235] 63,051 [61,894–63,944] 2.189 +2.57% +3.06% 3
count_cached new #2006 merge 135,041 [134,562–135,275] 61,945 [61,778–62,257] 2.180 +0.36% +1.26% 3
count_novelty base 3,330,670 [3,329,111–3,331,395] 1,154,091 [1,126,520–1,156,722] 2.886 +0.00% +0.00% 3
count_novelty old #2006 merge 3,333,677 [3,329,666–3,334,173] 1,216,845 [1,198,847–1,219,133] 2.740 +0.09% +5.44% 3
count_novelty new #2006 merge 3,328,882 [3,327,291–3,329,659] 1,172,266 [1,154,495–1,193,057] 2.840 -0.05% +1.57% 3

Formatter micro-bench format::tests::bench_stream_vs_dom: 10k rows, release builds of base and this branch, two alternated reps. SPARQL JSON moved −2.6% to −4.6% (streaming) and −0.6% to −5.4% (DOM) across the three shapes, inside the −7%..+5% (one +18% outlier) that the formats this PR doesn't touch moved — so no measurable change either way from dropping the per-row list scan.

Gates, when this PR was first opened

The full gates ran at 2a540f61e (on bf523e24e). 4f992fa4a only adds the compatibility-reference paragraph on trailing VALUES, and c67dc3516, the per-query-cost fix, was gated on its own on top of it, with 0 failures (the log below has the runs). At 2a540f61e, cargo check (right after the rebase), fmt (workspace and testsuite-sparql), clippy with -D warnings, the doc tests, testsuite-sparql (36/36 suites, registers checked both ways) and the wasm check are all clean. The full nextest run (CI's test job) is 13,765 run / 13,760 passed, and the other five need a word:

  • 1 failure, fluree-raft-core forward::tests::propose_connect_timeout_is_retryable: ECONNRESET on macOS, in a crate that depends on no fluree-db crate, and it failed the same way at the pre-rebase tip.
  • 4 timeouts at nextest's 360 s limit, all bulk-import / incremental-indexing tests that passed in the pre-rebase full run and pass when run alone at 2a540f61e (per-test times in the gate log below). The machine is just slower than it was for the pre-rebase run (load 7–26 here).

Across the four interleaved A/B pairs, the head is faster on two (8%, 15%) and slower on two (6%, 21%), and the 21% — the last compaction run — ran at load 13–15 against 10–12 for its base run. Both compaction tests run no code this PR changes, and the same slowness shows on main, so it's the box, not this PR.

Not run: the wasm-smoke browser suites and npm package build (headless Chrome), testsuite-sparql clippy (no file there changed), the sql-bridge workspace (it depends on no fluree crate), and the bench.yml compare against the committed baseline (host_class differs). testsuite-shacl is local only — no CI workflow runs it — and it doesn't compile on main today; e112dddf3 (d3df023d7 after the rebases) adds its two missing ValidateOptions fields as its own commit, and base and head then have the same status for every test (details below).

c67dc3516 (the per-query lowering cost) was gated on its own, on top of 4f992fa4a:

  • cargo fmt --all -- --check: clean. cargo clippy --all --all-features --all-targets --locked -- -D warnings, after touching the five changed files: clean.
  • cargo test -p fluree-db-query -p fluree-db-sparql: green (lib 1,658 passed with 1 ignored, and 658 passed, plus integration and doc tests).
  • cargo nextest run --all-features --no-fail-fast -p fluree-db-query -p fluree-db-sparql -p fluree-db-api -p fluree-db-shacl -p fluree-db-cypher: 7,198 run, 7,192 passed, 6 timed out at the 360 s limit, 28 skipped, 0 failed (5,568 s at load 7–11). The six are indexing-heavy: the four in the table below, plus it_indexing_stats::large_dataset_statistics_accuracy and it_indexing_workflow::file_based_indexing_then_new_connection_loads_and_queries. Rerun alone with no limit, all six pass (the four grp_index tests together in 1,204 s, numbig in 440 s, rdf_type_star in 267 s; load spiked to 116 during the reruns).
  • testsuite-sparql cargo test: 36/36 suites pass.
  • Mutation check: making sample_ungrouped_reads collect the pre-group variables eagerly fails the new sample_ungrouped_reads_collects_where_vars_only_for_a_non_key_read.
  • Under plain cargo test (every test in one process), it_annotation_filter_pushdown::annotation_body_threshold_reduces_scan_work_on_both_surfaces missed its lane event. It passes alone and under nextest, which runs each test in its own process, as CI does.

Run locally (CARGO_BUILD_JOBS=4, --test-threads 4) at 2a540f61e on bf523e24e, unless noted (4f992fa4a adds only the compatibility-reference paragraph on trailing VALUES):

  • cargo check --workspace --all-targets --all-features --locked right after the rebase: clean.

  • cargo fmt --all -- --check: clean. testsuite-sparql: cargo fmt --all -- --check clean.

  • cargo clippy --all --all-features --all-targets --locked -- -D warnings, after touching every .rs file the PR changes so no target is linted from cache: clean (0 warnings).

  • cargo nextest run --workspace --all-features --no-fail-fast --locked (CI's test job): 13,765 run, 13,760 passed, 1 failed, 4 timed out, 41 skipped (4,655 s at load 8–14; the pre-rebase run of the same command took 840 s). The failure is fluree-raft-core forward::tests::propose_connect_timeout_is_retryable (ECONNRESET on macOS; the crate depends on no fluree-db crate, and it failed the same way at the pre-rebase tip). The four timeouts hit nextest's 360 s limit. All four are bulk-import / incremental-indexing tests. Each passes alone with libtest (no per-test limit) at 2a540f61e:

    Test alone at 2a540f61e pre-rebase full run interleaved A/B, same box: base bf523e24e / this head
    it_query_sparql_indexed::indexed_rdf_type_star_count_exact_across_incremental_builds 219 s 28 s 204 s, 219 s / 185 s, 206 s
    it_distinct_object_numbig_gate::distinct_object_count_counts_numbig_object_keys_exactly 391 s 46 s 344 s / 365 s
    it_fwd_pack_compaction::incremental_cycles_compact_the_forward_pack_tail 651 s 100 s 731 s / 625 s
    it_fwd_pack_compaction::a_namespace_that_goes_quiet_stays_bounded_and_readable 850 s 127 s 987 s / 1194 s

    The machine is slower than for the pre-rebase run (load 7–26 here), and on it base and this head are close. The A/B columns are one or two runs per side, interleaved: the head was faster on two tests (by 8% and 15%) and slower on two (by 6% and 21%); the 21% (the last compaction run) ran at load 13–15 against 10–12 for its base run. The two compaction tests run no code this PR changes, and the same slowness shows on main: a_namespace_that_goes_quiet… was still running at 8 min there (12.5 min alone in a separate measurement).

  • cargo test --workspace --all-features --doc --locked (CI's doc-test step): 48 passed, 164 ignored, 0 failed.

  • testsuite-sparql: cargo test --no-fail-fast (CI's command): 36/36 suites pass with registers checked both ways (no register churn); lib 45, w3c_rdf 3.

  • wasm: cargo check -p fluree-db-wasm -p fluree-db-browser --target wasm32-unknown-unknown --lib --tests --locked with CI's RUSTFLAGS: clean.

  • testsuite-shacl, local only — not run by CI (no workflow references it; it is an excluded workspace). It does not compile on main today: ValidateOptions gained max_fuel and cancellation, and its runner still builds the struct without them. e112dddf3 (d3df023d7 after the rebases) adds the two fields (None, as before) as its own commit. Base bf523e24e with the same two lines, and this PR's head: SHACL Core 81/98 and SHACL-SPARQL 17/22 on both, with the same status for every test. The only difference in failure details, after masking blank-node labels, is pre-binding/shapesGraph-001 (failing on both): its message now names the constraint by IRI instead of by SID (d18c572b7, now b998482c5). The five SHACL-SPARQL failures are the same on both sides: three conforms expectations, the unsupported $shapesGraph, and unsupported-sparql-006. No W3C SHACL-SPARQL test uses GROUP BY, HAVING or an aggregate (grep of tests/sparql), which is why this PR's sh:select change doesn't show up there; the four shacl_tests above pin it.

@aaj3f

aaj3f commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

This continues the non-vacuity log in the first comment, which covers the PR as first opened. It has every fix since then, from the review onward.

Same method throughout: each fix was committed first. Then only the fix was reverted, using a needle taken from the post-fmt source and required to match an exact number of times. The tree was rebuilt, the named tests were run and watched fail on an assertion (never a compile error, except the one row below that is meant not to compile), and the files were restored with git checkout --, with git diff empty afterwards. One honest miss: the first run of the severity-and-mode mutation used a needle that matched nothing, so that run was green on unmutated code. It was discarded and re-run against the formatted line. Commits are named by their SHAs on the branch at 87576fdb1. Each mutation ran on the branch as it stood at the time, before the later rebases, and those rebases left every fix commit patch-identical apart from context. A SHA marked "measured on" names a build used for a comparison, not a commit on the branch.

Grouping: EXISTS, HAVING and sort keys

Fix Mutation (only the fix reverted) Went red, and why
5aab1c166: EXISTS-body variables are free over the group row row reads become every referenced variable, and the row-read substitution descends into EXISTS bodies (the walk before the fix) exists_body_variables_are_free_over_the_group_row, having_exists_is_evaluated_per_group; run again later, ask_and_construct_group (its NOT EXISTS case)
5aab1c166 (filter::correlation_seed): the seed mapping the per-row EXISTS / pattern-comprehension seed keeps a Binding::Grouped cell Nothing, because no shape reaches it. I tried grouped ASK with a dedup-only GROUP BY; GROUP_CONCAT, COUNT and SAMPLE beside an EXISTS; grouped CONSTRUCT; an EXISTS SELECT expression; and JSON-LD per-group lists. JSON-LD and Cypher have no post-group EXISTS over a non-key variable. Resolved in d6b2edfe6: the case is now a debug_assert!, so a change that makes it reachable fails in debug and gets its own test, and release builds keep the spec reading. With the assert in place, the probe shapes and the debug suites pass: 2,468 API tests (grp_query_sparql, grp_query, grp_misc, it_query_cypher, it_policy_cypher, the lanes binary, and the lib including shacl_tests) and 2,500 query, Cypher and SPARQL crate tests
19512aec1: HAVING reads the SELECT expressions it names binds_before_having returns all false having_reads_a_select_alias, having_tests_the_value_an_alias_returns, jsonld_having_reads_a_select_expression_alias, and cypher_reads_a_key_nodes_property_after_grouping (the composite alias WITH p, count(f) + 0 AS c WHERE c > 1 returned [])
a4c2aa764: a sort key nothing binds orders nothing bindable_sort_keys keeps every key order_by_a_variable_nothing_binds_orders_nothing, jsonld_order_by_a_variable_nothing_binds_orders_nothing, order_by_a_variable_nothing_binds_streams_every_row, and both test_build_operator_tree_drops_sort_keys_nothing_binds unit tests. test_apply_solution_modifiers_validates_sort_vars stays green, as it should: it pins the error for a key the plan lost

Cypher

Fix Mutation (only the fix reverted) Went red, and why
23c3f1ee2: a key node's property read as a SAMPLE (since replaced by the next row) sample_key_properties returns at once cypher_reads_a_key_nodes_property_after_grouping (a plan error on the key-property reads)
422d986c4: read an output node's property after the aggregation fluree-db-cypher/src/lower/stmt.rs restored whole to its version before the fix, the SAMPLE rewrite (measured on 1a20ac6bc, the branch head then) cypher_output_node_properties_are_read_after_aggregation, first case WITH p, count(f) AS c WHERE p.age > 40: [["Dave", 3]] instead of [["Alice", 2], ["Dave", 3]]. The code before the fix reads Alice's age as a SAMPLE inside the aggregation, and the sample was 40, which fails > 40. The test stops at that first assertion, so its count, sum and collect cases didn't run under this mutation (the table below shows them). cypher_reads_a_key_nodes_property_after_grouping stayed green under the revert: it uses single-valued properties, which is the gap review found
ad89f13b9: an expression sort key after aggregating reads the aggregates reads_after_aggregation takes two levels only when the clause reads a property (post.iter().any(|p| matches!(p, Pattern::Optional(_))) in place of !post.is_empty()), the rule before the fix cypher_reads_a_key_nodes_property_after_grouping: WITH p, count(f) AS c ORDER BY c + 1 DESC LIMIT 2 fails with "an ORDER BY key is neither a GROUP BY key nor an aggregate result", the 400 from before the fix
92f4913be: sort on a value read from a collected list the size / length exemption removed, so a key that reads the list is refused again cypher_collect_through_with: "ORDER BY on a collect() list is not supported in v1" on … ORDER BY size(friends)
b2ef72fba: a WITH's WHERE filters after its slice the guard that sends a sliced WITH's WHERE to the stage becomes false && …, so the WHERE stays in the clause cypher_with_where_filters_after_the_slice: UNWIND range(1, 10) AS x WITH x ORDER BY x LIMIT 5 WHERE x > 3 gives [4..8], not [4, 5]
b2ef72fba: carried variables a plain WITH's sliced rows carry nothing for the WHERE same test: MATCH (p:P) WITH p.age AS a ORDER BY a LIMIT 2 WHERE p.age > 30 gives six rows (25 ×3, 35 ×3), not [35]. The stage's accessor reads p, which is unbound there, so it matches every node's age
b2ef72fba: DISTINCT repeated the stage doesn't repeat DISTINCT same test: WITH DISTINCT p ORDER BY p.name LIMIT 2 WHERE p.age > 30 gives Alice twice (both of her ages pass)
b2ef72fba: scope after an aggregating or DISTINCT WITH the scope check is always off same test: WITH p, count(f) AS c ORDER BY c LIMIT 2 WHERE f.age > 30 fails with the planner's "projected variable f is neither a GROUP BY key nor an aggregate result", not the "reads f, which is out of scope" error. The DISTINCT cases come after it in the same loop, so they didn't run under this mutation
07d0892ad (test): an exists body's variables are carried a plain WITH carries only what its WHERE reads directly (row_reads), not an exists body's variables cypher_with_where_filters_after_the_slice fails on the first exists case, fa DESC … LIMIT 3 ([40, 35, 35] against [35]), before it reaches the new ones. The new pair shows the carry by itself: the two queries differ only in whether the body variable is carried (f) or its own (g), and they give [[40]] and [[40], [40]]
f553cd86c: a WITH inside a CALL body carries the imports carry_imports returns at once cypher_call_body_with_keeps_the_import: CALL (p) { MATCH (p)-[:knows]->(f) WITH f … } gives the 24 crossed rows
f553cd86c: the planner keeps a pinned import as a correlation || sq.pinned_vars.contains(&v) dropped from subquery_correlation_vars same test: WITH f ORDER BY f.age LIMIT 1 gives one row, [40, 25], so the body's slice ran once for all p. Placement wasn't instrumented; the planner treating p as uncorrelated is the reading consistent with the result
f553cd86c: carried imports stay out of scope a nested CALL sees imports a WITH carried same test: a nested CALL (p) after WITH f is accepted (the is_err assertion)

What the earlier SAMPLE reading did with a two-valued property. The 422d986c4 row stops at the test's first assertion, so a throwaway test (added, run, then removed with git checkout --) printed three WHERE p.age > 30 queries on the same fixture, under the revert and with the fix. Alice (aged 40 and 41) knows Bob (score 10) and Carol (100); Dave knows Alice (1), Bob and Carol:

Query (MATCH (p:P)-[:knows]->(f) …) before the fix (1a20ac6bc's stmt.rs) fixed
WITH p, count(f) AS c WHERE p.age > 30 RETURN p.name, c Alice 4, Dave 3 Alice 2, Alice 2, Dave 3
WITH p, sum(f.score) AS s WHERE p.age > 30 RETURN p.name, s Alice 220, Dave 111 Alice 110, Alice 110, Dave 111
WITH p, collect(f.name) AS fs WHERE p.age > 30 RETURN p.name, fs Alice [Bob, Bob, Carol, Carol], Dave [Alice, Bob, Carol] Alice [Bob, Carol] twice, Dave [Alice, Bob, Carol]

Alice appears twice in the fixed column because both of her ages pass the filter: the second stage reads every value. That matches the existing Cypher behavior for a multi-valued property, checked on the same fixture. MATCH (p:P) WHERE p.age > 30 RETURN p.name returns Alice twice, Carol and Dave; with > 40, it returns Alice once and Dave; and RETURN p.name ORDER BY p.age DESC lists Alice twice.

Trailing VALUES, and ASK / CONSTRUCT / DESCRIBE

Fix Mutation (only the fix reverted) Went red, and why
71af73a38: the internal-predicate check covers trailing VALUES rows the post_values check removed sparql_values_row_naming_reifies_is_refused_in_both_positions
8cfebf749: no fast path answers a query with a trailing VALUES trailing_values is always false lanes: all five trailing-VALUES MustNotFire checks (COUNT rows, group_by_object_count_topk, group_by_object_star_topk, COUNT by predicate (directory), SUM(?o))
e6e2061e5: generated binds run after the top-level trailing VALUES the top level puts its generated binds back in the WHERE, before the VALUES join generated_binds_read_trailing_values_at_both_levels (top-level SUM(?n * ?v) = 0)
b383c709d: ASK and CONSTRUCT group, and so does JSON-LD ask SPARQL ASK drops its grouping; JSON-LD ask refuses groupBy / having again ask_and_construct_group, jsonld_ask_groups
5efca0d81: ASK OFFSET lower_ask sets offset: None ask_applies_offset_limit_and_values: ASK { ?p ex:name ?n } OFFSET 2 answered true, expected false
5efca0d81: a trailing VALUES after CONSTRUCT lower_construct lowers no trailing VALUES construct_and_describe_read_a_trailing_values: the first CONSTRUCT (VALUES ?s { ex:jdoe ex:bbob }) built all four people's triples
5efca0d81: a trailing VALUES after ASK lower_ask lowers no trailing VALUES ask_applies_offset_limit_and_values: VALUES ?n { "Charlie" } answered true
5efca0d81: a trailing VALUES after DESCRIBE lower_describe ignores the trailing VALUES construct_and_describe_read_a_trailing_values: DESCRIBE ?s VALUES ?s { ex:jdoe } described nothing, expected jdoe's seven values
5efca0d81: JSON-LD ask offset the ask branch doesn't parse offset ask_applies_offset_limit_and_values: {"ask": …, "offset": 2} answered true
5efca0d81: JSON-LD construct values construct doesn't parse values construct_and_describe_read_a_trailing_values, the JSON-LD twin: fbueller and jbob in the graph too
5efca0d81: JSON-LD ask values the ask branch doesn't parse values ask_applies_offset_limit_and_values: "values": ["?n", ["Charlie"]] answered true
bd3797dcc: SPARQL parameters (#2024) and a trailing VALUES after ASK / CONSTRUCT / DESCRIBE ASK's arm doesn't walk its trailing VALUES a_parameter_the_query_assigns_is_an_error: ASK { ?s <p> ?x } VALUES ?x { 1 } is accepted, and the AST shows the WHERE's ?x replaced by 1 while the VALUES still binds ?x. A first version of the test named ?x only in the VALUES, and under the mutation that version failed for another reason ("not a variable in the query"). The cases now name it in the WHERE too

Fast paths and routing canaries

The canary pairs check that each fast path that gates on the grouping's binds serves the key-only query (MustFire) and declines the same query with a SELECT expression (MustNotFire).

Fix Mutation Went red, and why
count top-k's binds gate (existing) binds gate removed lanes: group_by_object_count_topk must not proceed
directory count's binds gate (existing) binds gate removed lanes: the fast path answered, and the query failed (?ps not in its schema)
star top-k, COUNT(*), whole-graph: binds gate only the three binds gates removed at once green: their projection checks still decline the expression's shape
star top-k binds gate and projection check both removed lanes: the fast path answered (an internal error)
COUNT(*) binds gate and projection check both removed lanes: the query went to the count planner (count-plan), not COUNT rows. With only the canaries of 923c428e1, just the row check caught it ([6, null] against [6, 60]). With 1d2c0d6a3 the routing check fails too: "no fast path may answer [… count-plan]"
whole-graph scalar aggregates binds gate and projection check both removed lanes: the fast path answered (an internal error)
SQL lane and fused R2RML: binds gate only binds gates removed green: shadowed by the projection checks
SQL lane and fused R2RML binds gate and projection check both removed SQL: sql_aggregate_pushdown proceeded on a declined shape, m was empty in the rows, and a grouped statement was sent. Fused canary (then in it_iceberg_local_fs, since moved unchanged to its own binary): fused_r2rml_aggregate must not proceed
1d106590b: the fused R2RML canary in its own binary binds gate and projection check both removed it_fused_aggregate_routing: fused_r2rml_aggregate must not proceed
SUM(?o), the expression case — not opened separately (shadowed like COUNT(*)); its trailing-VALUES case is covered above
a99f444d1: the count fast paths decline a projected per-group list the count top-k detector accepts any projection lanes: the JSON-LD per-group list query planned the count top-k ("no fast path may answer [proceeded: fused_chain, group_by_object_count_topk]")
270dcb51d, 4e30eaf0e: a MustFire also requires no decline at open the count top-k declines at open on every query (if true || should_fallback(ctx)) lanes: both key-only MustFire queries, SPARQL and JSON-LD, "must proceed": the plan stamped Proceed and the open stamped the decline. Before 4e30eaf0e these passed on the plan-time stamp alone (the test collects every misroute before asserting)
#2012 (9d3ca1840), exercised by 3b45c92dd: DATATYPE of an overflow key in eval/rdf.rs, the NUM_BIG arm of datatype_of_binding disabled (if false && …), so DATATYPE reads the dt_id lanes, overflow ledger: SELECT ?z (COUNT(?b) AS ?n) (DATATYPE(?z) AS ?dt) … GROUP BY ?z ORDER BY DESC(?n) LIMIT 2 gave xsd:decimal for the overflow key
#2012, exercised by 4e30eaf0e: a group a decoded copy joins group_aggregate::materialize_encoded takes the datatype from the dt_id again lanes, overflow ledger: … { ?b ex:size ?z } UNION { BIND(1234…890 AS ?z) } … GROUP BY ?z split into groups of 1 and 3, expected one group of 4. An earlier run of the same mutation left the lanes test green, because its cases then had only scanned copies, which key by their encoding and never reach that function. That's why 4e30eaf0e adds the UNION case and its JSON-LD values twin

SHACL

Fix Mutation Went red, and why
2845a168c: a sh:sparql constraint that cannot run follows severity and mode the engine never collects a failure (it always raises it) shacl_sparql_constraint_failure_follows_severity_and_mode, shacl_sparql_constraint_failure_keeps_other_shapes_enforced
764519b36: every sh:sparql query failure follows severity and mode the classification back to the rule before the fix: only InvalidQuery is the constraint's failure (if !matches!(e, QueryError::InvalidQuery(_)) in place of if raised_whatever_the_severity(&e)) shacl_sparql_constraint_failure_follows_severity_and_mode: "expression failure, severity None, warn mode false: expected the constraint's error, got Err(Transact(Shacl(QueryError(InvalidFilter("Unknown function: http://example.org/fn")))))". The unknown-function failure is raised as a query error instead of a constraint failure. …_keeps_other_shapes_enforced stayed green; it has no reclassified error
795e31f96 (test): the classifier's raised half raised_whatever_the_severity becomes false && match e { … }, so every error is the constraint's own constraint_query_errors_are_classified_by_variant: "FuelLimitExceeded must be raised whatever the severity", "Cancelled must be raised …" and so on, for every raised-side variant. Before this test, the same mutation left 216 tests green, the 164 SHACL tests among them
c0e831164: the classifier table the same mutation constraint_query_errors_are_classified_by_variant lists all 21 raised entries, now including QueryError::R2rml(_)
c0e831164: a new error variant needs an entry the Policy entry deleted from the table (a stand-in for a variant with no entry) doesn't compile, by design: E0004 "non-exhaustive patterns: &QueryError::Policy(_) not covered", at the macro's match
7999b625e: a nested shape's constraint failure takes the outer severity a recorded failure keeps the owning shape's severity (exec.failure_severity.unwrap_or(severity) back to severity) shacl_nested_constraint_failure_takes_the_outer_severity, 39 cells. 26 should commit (a Warning or Info outer shape over a Violation inner one, 13 nestings across both write routes) but got the inner constraint's error; 13 should reject (a Violation outer shape over a Warning inner one) but committed
7999b625e: the outermost reporting shape wins reported_under lets the innermost reporting shape win (Some(severity) in place of self.failure_severity.or(Some(severity))) the 6 chain cells: a Violation middle shape between a Warning or Info outer shape and the inner shape made the write reject
2db734978: a shape registered after its metadata keeps its severity the claim_metadata() call dropped shacl_sparql_only_shape_keeps_its_severity: Warning and Info, JSON-LD and SPARQL UPDATE, all four reject with "score is negative". shacl_nested_constraint_failure_takes_the_outer_severity stays green without it (see below)
8913e8c49: sh:and, sh:or and sh:xone lists resolve from an RDF collection a member that heads an RDF list is kept as a shape reference shacl_logical_lists_resolve_from_either_surface: the three commit cases written by SPARQL UPDATE fail, e.g. "sh:and constraint - Referenced shape fdb-…#coll2 could not be resolved". shacl_nested_constraint_failure_takes_the_outer_severity: the six Violation-outer And / Or / Xone SPARQL cells give that error, not the inner constraint's
4594e0dfb: a nested shape's results count toward conformance whatever their severity conforms counts only Violation results again shacl_nested_warning_results_count_toward_conformance: all 28 Violation-outer cells, across all seven nestings, both inner checks and both surfaces. The sh:not cells reject with "Node conforms to shape InnerShape which is not allowed"; the others commit. The other 95 SHACL tests stay green under it, the outer-severity test included, so the two rules compose both ways. A first pass left sh:and collecting messages outside conforms, and its four cells stayed green; sh:and now goes through conforms too. The same 28 cells failed before the fix was written

The outer-severity test without its workaround. shacl_nested_constraint_failure_takes_the_outer_severity passes with or without 2db734978. The outer shape's severity governs every cell, so the inner shape's own severity, which 2db734978 restores, never decides one. That's why its workaround (an inner sh:targetClass ex:Unused) could go, and why shacl_sparql_only_shape_keeps_its_severity is the test that pins 2db734978.

Error names, formatting and the stream

Fix Mutation Went red, and why
cd021b64e: the WHERE cursor names the variables in its errors WhereCursor::next_batch returns the raw error update_grouped_read_error_names_the_variable
97b5b9ed1: the stream refuses a grouped read before it starts ensure_streamable skips first_ungrouped_read jsonld_grouped_read_is_the_same_4xx_on_query_and_stream
7de911c5e: a JSON-LD orderBy on an expression says how to sort on one validate_order_var no longer maps an expression to InvalidOrderBy test_parse_order_by_expression_names_the_alias_route: [["desc","(count ?e)"]] gives InvalidVariable("(count ?e)")
c44d22aca: JSON-LD materializes an encoded binding inside a per-group list lists, sets and maps format their elements without the QueryResult (6 call sites, both formatters) grouped_select_expression_on_an_indexed_ledger_in_both_lanes: "to_jsonld: … format_binding called without QueryResult for encoded IRI binding" on the per-group list query, the error from before the fix

The sub-SELECT sort panic, checked rather than mutated. A sub-SELECT's grouped SELECT expression, sorted in the outer ORDER BY, is fixed by this PR's per-group evaluation itself, not by one commit. So I ran the same queries on debug CLIs built at main (2941d470c), at the branch head then (measured on ba598107f), and at the branch as first reviewed, rebased (measured on 8f8cf77ff). Main panics at fluree-db-query/src/sort.rs:208:13, "Grouped bindings should not appear in sort comparisons", on the SPARQL STRLEN query and on the JSON-LD subquery. Both branch builds return Remote 6, Local 5, Net 3 on both. The COUNT * 2 control sorts at all three. subselect_grouped_expression_sorts_in_the_outer_query and its JSON-LD twin (d466ca9df) pin it.

Not mutated, and why

  • de26852a2 and 320ae3e98 (named aggregate-variable errors): their tests assert the named text and the absence of VarId(. The text before the fix is exactly a VarId( print, so a revert fails them by construction. Not run.
  • a99f444d1's directory-count check: only the count top-k half was mutated.
  • 5efca0d81's LIMIT cap (LIMIT 0 is false): not mutated separately.
  • The bound-object count's decline stamp (270dcb51d): diagnostic only; no canary covers that site.
  • 7ce891589 is a test.
  • The multi-value model pins (5422809d4 and its tests): they pin today's model, not a fix. The per-value rows, DISTINCT keeping the copies, a later count and sum, and collect(p.age) beside count(f) are meant to flip on purpose if the model changes.
  • The canary failure message (62fa6798a): it's text. The routing check behind it is served(), which the MustFire row above shows is load-bearing.
  • claim_metadata's rule that a severity given to a shape after it was registered stays. It only decides between two conflicting sh:severity values on one shape in different graphs, and no test covers that.
  • The claim of sh:message, sh:name and sh:description. For a shape whose only constraint is sh:sparql, validation reads none of them: its results take the constraint node's message. They reach the compiled shape, which the GraphQL schema reads for name and description; no test covers that.
  • Docs and comments: 6ca526c24, 290964aa1, 492683294, 2acbe23e6, 241d3fe41, 622bb386c, b6f867305, 87576fdb1.

@aaj3f

aaj3f commented Oct 8, 2026

Copy link
Copy Markdown
Contributor Author

Thanks for these, @bplatz. All four ended up in this PR rather than as follow-ups, and the branch is rebased onto 745736a04 with them:

  • CONSTRUCT and ASK with GROUP BY / HAVING (b383c709d).
    • Both lower through lower_select_level with a SELECT that projects nothing.
    • ASK is true when some group passes HAVING. CONSTRUCT instantiates its template once per group.
    • On the one judgment call, I went with the spec rather than a 400: a template variable that isn't a GROUP BY key is unbound in the group solution, so the triples that read it are skipped (§16.2).
    • JSON-LD ask takes groupBy / having, and DESCRIBE still refuses them.
    • A CONSTRUCT with an aggregate ORDER BY is now one implicit group where it used to be refused, and its test moved with it.
    • ASK also dropped OFFSET and LIMIT: ASK { … } OFFSET 2 over two solutions was true, and so was LIMIT 0. 5efca0d81 makes ASK true only when a solution remains after them; ORDER BY can't change that and is still dropped. JSON-LD ask now applies offset, limit and values, and construct applies values; both ignored them before.
  • ORDER BY ?nosuch (a4c2aa764).
    • A sort key that no stage of the level can bind is unbound in every row, so it orders nothing and is dropped. The stages are the WHERE, a trailing VALUES, grouping, an ORDER BY expression and the projection.
    • Sub-queries drop it once per operator, not per parent row.
    • A key some stage binds is kept, so a plan that loses one still fails the sort validation.
    • Tested on SPARQL (grouped, ungrouped, sub-SELECT), JSON-LD and the stream.
  • Routing stamps for the other binds-gated detectors (923c428e1, tightened in 1d2c0d6a3 and 4e30eaf0e).
    • Each has a canary pair: the key-only query must take the fast path, and the same query with a SELECT expression must not. Both are checked with hand-derived rows in both lanes.
    • On the indexed ledger, the pairs cover the star top-k, the per-predicate directory count, COUNT(*), SUM(?o) and the whole-graph scalar aggregates. The SQL lane's grouped statement has a pair in it_sql_pushdown_lane.
    • The fused R2RML aggregate has a pair on a local Iceberg table, with its rows checked with fast paths on only. It had no routing stamp before, so it now records Proceed or Fallback at open.
    • Only the count top-k and the directory count decline the expression on their binds gate alone. The rest also decline it on their projection check. So removing just their binds gate leaves the test green, and removing both turns it red. The non-vacuity log has both runs, in the follow-up comment that continues the first comment's.
    • The mutation runs turned up two more checks:
      • With the single-aggregate gate removed, the COUNT(*) query went to a sibling lane, the count planner. So the MustNotFire queries now also assert that the fused-chain stamp is the only one.
      • The count top-k stamped Proceed when it was planned, but could still fall back at open without saying so. It now stamps the decline at open (270dcb51d, as does the bound-object count), and a MustFire requires that no decline was stamped.
    • The count top-k and the directory count also never read the projection, so a JSON-LD query projecting a per-group list next to the key and the count took them. The top-k answered null for the list, and the directory count failed with "Projected variable not in child schema". Both now decline any projection other than the key and the count (a99f444d1), pinned as JSON-LD canary pairs.
  • HAVING reading a non-aggregate SELECT alias (19512aec1).
    • I decided to extend visibility rather than add a diagnostic. In a grouped query, HAVING reads a SELECT expression's alias as well as an aggregate's.
    • The expressions HAVING reads, and the aliases they read in turn, run before HAVING. The rest run after it, each once per group, so HAVING tests the value the query returns. That's pinned with RAND() and STRUUID() over 40 groups.
    • (COUNT(?e) + 0 AS ?n) … HAVING (?n > 1) keeps Local and Net, and the ?seg case keeps the Net group.
    • It's a Fluree extension, though many engines share it. The W3C aggregate tests repeat the aggregate in HAVING, so they pin neither reading. It's documented once in the compatibility reference.
    • The same rule holds for JSON-LD having and Cypher WITH … WHERE, and it fixes the Cypher composite alias (WITH p, count(f) + 0 AS c WHERE c > 1), which dropped every row at base too.

A few more things turned up while tracing these. They're all pre-existing, and I folded them in rather than file them, so each has a row in the description's behavior table:

  • the internal-predicate check in SPARQL lowering didn't walk the trailing VALUES; now it does (71af73a38);
  • two comments still described the old HAVING/EXISTS order (290964aa1);
  • the aggregate-output plan errors printed raw variable ids, Aggregate output variable VarId(0) already exists in schema and Duplicate aggregate output variable VarId(1). SPARQL's validator catches both first, but JSON-LD and sh:select reach the plan check. They're now aggregate output variable ?a is already bound in the WHERE pattern and variable ?n is the output of more than one aggregate (de26852a2). Cypher's WITH p, count(f) AS c, sum(c) AS s gets the named unbound-aggregate-input error (320ae3e98);
  • a Cypher ORDER BY on a value read from a collected list, such as WITH p, collect(f) AS fs ORDER BY size(fs), was refused as an ORDER BY on the list. It sorts after the aggregation now; a key whose value is the list itself is still refused (92f4913be);
  • a JSON-LD orderBy on an expression said "orderBy must be an array of objects with 'var' field". It now lists the forms orderBy takes, and says to sort on an alias (7de911c5e);
  • on an indexed ledger, the JSON-LD formatters failed on a per-group list of IRIs: "format_binding called without QueryResult for encoded IRI binding". They now materialize encoded bindings inside lists and maps, as typed JSON already did (c44d22aca).

The last ones change the answers, or the write verdicts, of queries and shapes that run today without an error, so I'd especially like your eyes on these. Happy to talk through any of them:

  • A Cypher WITH's WHERE filtered before the clause's SKIP / LIMIT. openCypher's grammar puts it after them, and it filters the clause's results. So UNWIND range(1, 10) AS x WITH x ORDER BY x LIMIT 5 WHERE x > 3 gave 4 to 8, where openCypher gives 4 and 5. It now filters the sliced rows (b2ef72fba), and a WITH without a slice lowers as before. That evidence is from the grammar and the reference; no openCypher engine was run.
  • A WITH inside a CALL (p) { … } body lost its correlation with p. MATCH (p:P) CALL (p) { MATCH (p)-[:knows]->(f) WITH f RETURN f.name AS fn } RETURN p.name, fn crossed every p with every knows edge, and a slice in the body ran once for all of them. The body's subqueries now carry the imports they read, and the planner keeps a CALL's import as a correlation input (f553cd86c). This was there at base too.
  • Two SHACL compile bugs.
    • The shape compiler dropped the sh:severity of a node shape whose only constraint is sh:sparql. Such a shape is registered by its rdf:type after the compiler has read the severity, so a Warning shape rejected writes. It keeps its severity now (2db734978).
    • sh:and / sh:or / sh:xone lists written by SPARQL UPDATE are stored as an rdf:first / rdf:rest collection, and the compiler took the collection's head for a shape ("Referenced shape … could not be resolved"). They resolve now, like JSON-LD and Turtle lists (8913e8c49).
  • Whether a node conforms to a nested shape counted only Violation results. That covers sh:node, sh:not, sh:and, sh:or, sh:xone and sh:qualifiedValueShape. SHACL §3.5 counts every result, and severity only categorizes them, so a Warning nested shape that reported a result let sh:node pass and made sh:not fail. Every result counts now, through one check (4594e0dfb). The referencing shape still reports at its own severity, and a constraint that cannot run is still decided at the outermost one.

@aaj3f
aaj3f requested a review from bplatz October 8, 2026 20:16

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:query Query execution, planning, fast paths, overlay, result formatting area:sparql SPARQL/Turtle/TriG/JSON-LD parsing, lowering, UPDATE semantics, W3C conformance bug Something isn't working as expected

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPARQL: an expression in the SELECT of a grouped query emits one row per input binding instead of one per group

2 participants