Repository navigation
Conversation
|
Non-vacuity log for this PR, moved out of the description to keep it under GitHub's size limit: for each fix, what was reverted and what went red. Every row was committed first, then only the fix reverted, rebuilt, watched fail, restored, and the needle grepped. The rows through
|
bplatz
left a comment
There was a problem hiding this comment.
Approving. The grouping restructure looks right to me. Please go through the inline comments before merging, in particular the EXISTS-under-grouping one, the Cypher WITH … WHERE one and the SHACL one.
| output_var | ||
| }); | ||
| if let Some(having) = having.as_deref_mut() { | ||
| having.substitute_var(v, sampled); |
There was a problem hiding this comment.
The SAMPLE rewrite reaches into EXISTS patterns through substitute_var, so a non-key variable inside an EXISTS becomes the sampled value, and the result depends on which row SAMPLE picks. The group row doesn't bind ?e, so it's free inside the pattern (Sample(?e) can't stand in a triple pattern anyway).
SELECT ?a (COUNT(?e) AS ?n) WHERE { ?e ex:area ?a } GROUP BY ?a
HAVING (EXISTS { ?e ex:flag true })With the flag on e1 the Net group is kept; move it to e2 (also Net) and the group is dropped. By the spec every group is kept. NOT EXISTS flips the same way. having_exists_is_evaluated_per_group only pins a shape where every possible sample agrees.
Suggest leaving EXISTS-body variables out of the rewrite, and out of the first_ungrouped_read HAVING check, which would otherwise reject them.
There was a problem hiding this comment.
Good catch, and agreed. It's fixed in 5aab1c166. §18.2.4.1 rewrites an expression's unaggregated variables, and a pattern variable inside an EXISTS isn't one of them: the group solution doesn't bind it, so substitute(P, μ) leaves it free.
Expression now has a row-read walk (row_reads / substitute_row_read) that skips EXISTS bodies, and reads a pattern comprehension's projection only where its own pattern doesn't bind the variable. sample_ungrouped_reads, first_ungrouped_read and the SELECT-expression placer all go through it. Dependency tracing follows the post-group stages by row reads too, plus only those correlated variables that grouping produces, so a pattern-only variable no longer rides through the GroupByOperator lane as a per-group list.
Your repro now keeps every group with the flag on e1 or on e2, and NOT EXISTS keeps none (exists_body_variables_are_free_over_the_group_row, GROUP_CONCAT lane included). The third case of having_exists_is_evaluated_per_group, which had pinned the SAMPLE reading, now expects all three groups.
| || self.aggregate_inputs.contains(alias) | ||
| || refs | ||
| .iter() | ||
| .any(|v| where_vars.contains(v) && !self.keys.contains(v)) |
There was a problem hiding this comment.
This has the same root cause: referenced_vars() includes EXISTS pattern variables, so an EXISTS over a non-key variable puts the expression before grouping, and the plan check then rejects it:
SELECT ?a (EXISTS { ?e ex:flag true } AS ?f) (COUNT(?e) AS ?n)
WHERE { ?e ex:area ?a } GROUP BY ?aThis returns 400 "projected variable ?f is neither a GROUP BY key nor an aggregate result". V4 accepts the query (the AST check skips EXISTS bodies), and the spec gives one row per group.
There was a problem hiding this comment.
Same root cause, so it's the same commit (5aab1c166). The placer goes by row reads now, so (EXISTS { ?e ex:flag true } AS ?f) becomes a per-group Extend. Your query returns one row per group, with ?f true in each, and since the placer is shared, JSON-LD gets the same placement. It's pinned in the same test, exists_body_variables_are_free_over_the_group_row.
| @@ -1269,7 +1271,8 @@ fn lower_with<E: IriEncoder>( | |||
| projection.aggregates, | |||
| projection.post_binds, | |||
| having, | |||
There was a problem hiding this comment.
An aggregating WITH … WHERE that reads a property of a key node now returns a 400, and the behavior table doesn't list it:
MATCH (p:E)-[:knows]->(f) WITH p, count(f) AS c WHERE c >= 1 AND p.age > 30 RETURN p.age, cThis returns "HAVING reads a variable that is neither a GROUP BY key nor an aggregate result". Base returned [], which was also wrong. WITH e, count(*) AS c WHERE e.age > 30 behaves the same way. The p.age accessor triple lands in the WITH body before grouping, and Cypher lowering doesn't apply sample_ungrouped_reads. Applying it here would give the exact answer, since a property of a key node has a single value.
WHERE exists { (p)-[:knows]->(f) } after the WITH also returns a 400 ("HAVING reads variable f"). f is out of scope there, so it should be a fresh variable. That's the EXISTS issue in grouping.rs.
There was a problem hiding this comment.
Thanks for this one. It's fixed in 422d986c4, though not quite the way you suggested, because the SAMPLE reading turned out to be wrong for a property with several values.
An aggregating WITH's WHERE and ORDER BY, and an aggregating RETURN's ORDER BY, can now read a property of a node the clause projects. When they do, the clause lowers as two levels: the aggregation, then a stage that joins the property, filters on the WHERE, sorts and slices. That's what WITH p, count(f) AS c WITH p, c WHERE p.age > 30 already lowered to, and the two forms now return the same rows. (When the clause also has a SKIP or LIMIT, the WHERE now filters after the slice, as openCypher places it; there's more on that in my reply to your follow-ups.) A property of a node the clause doesn't project is still an error, which matches Cypher's scoping.
An earlier commit in this PR (23c3f1ee2) did read the property as a SAMPLE inside the aggregation. The catch is that it joins the property before grouping, so each value repeats every row of the group: with Alice aged 40 and 41, count gave 4 instead of 2, sum doubled, and collect listed every friend twice. Reading the property after the aggregation keeps the clause's own aggregates exact, whatever the property holds.
Past that, the second stage keeps Fluree's model for a property with several values, which joins every value, as MATCH … WHERE already does:
- in
WHERE, the row comes once per value that passes; - in
ORDER BY, it comes once per value, andWITH DISTINCTkeeps those copies (the sort key is projected with them) whileRETURN DISTINCTremoves them; - an aggregate in a later clause counts each copy (
… WHERE p.age > 30 RETURN count(*)counts Alice twice).
A read inside an aggregate's argument, such as collect(p.age), is still joined before grouping, so it repeats the group's rows for every aggregate of that clause; that predates this PR. openCypher has no multi-valued properties and never repeats a row. Matching it means changing the model everywhere at once, the explicit-WITH and MATCH paths included, which is a much bigger scope than this PR, so it's a follow-up in the description. The reference now describes the model (5422809d4), and the tests pin it in row order, so that change will flip them on purpose. They cover an unsliced ORDER BY in a WITH and a RETURN, WITH DISTINCT … ORDER BY, the later count(*) and sum(c), and collect(p.age) beside count(f).
Your first query returns [[40, 2], [50, 3]] on a knows/age fixture. WITH e, count(*) AS c WHERE e.age > 30 and the ORDER BY forms return rows too (cypher_reads_a_key_nodes_property_after_grouping), and so does ORDER BY p.age + 1. A sort on an expression over the clause's aggregates, ORDER BY c + 1 or ORDER BY -c, was still a 400; it now takes the same two-level form (ad89f13b9). cypher_output_node_properties_are_read_after_aggregation pins count, sum and collect with a two-valued age under WHERE, a WITH's ORDER BY and a RETURN's ORDER BY, plus the explicit two-WITH form. The error-naming test moved to a node the WITH doesn't project (WITH e.area AS a, count(*) AS c ORDER BY e.name).
WHERE exists { (p)-[:knows]->(f) } after the WITH was the EXISTS issue from grouping.rs. With 5aab1c166, f is free there, and the same test pins it with a likes edge.
The composite-alias case, WITH p, count(f) + 0 AS c WHERE c > 1, dropped every row at base too. It's fixed by 19512aec1 (the HAVING-alias change in my follow-ups reply) and pinned in the same test.
| ) | ||
| .await?; | ||
| .await | ||
| .map_err(|e| match e.name_variables(&vars) { |
There was a problem hiding this comment.
This failure leaves validation as an Err, so it bypasses sh:severity and the graph's Warn mode (the severity split in stage.rs only sees results). Because it's raised at plan time, it refuses conforming writes too.
Repro: a shape with sh:severity sh:Warning, sh:targetClass ex:Player and sh:select "SELECT $this ?value WHERE { $this ex:score ?value } GROUP BY $this", then insert an ex:Player with no ex:score. Base commits; here it returns 400 Invalid sh:sparql constraint …: projected variable ?value is neither….
So after an upgrade, a stored shape like this blocks every write to its target class. Reporting a failure is spec-correct, but the behavior table should say exactly that, or severity and Warn mode should apply in this case.
There was a problem hiding this comment.
I went with your second option, in 2845a168c and 764519b36: severity and mode now apply. They cover every failure of the constraint's own query:
- it doesn't parse, lower or plan;
- a
$PATHhas nothing to bind it; - the query fails on its own terms, for example an aggregate over a variable nothing binds, or a call to an unknown function.
The request's budgets (fuel, deadline, memory), its storage, catalog and policy access, the state of the data and internal faults still fail the write whatever the severity. That split is an exhaustive match over the query error type, so a new error variant has to be placed on one side.
On the transaction path the engine records the failure, with the owning shape's severity and the graph, and goes on to the shape's other constraints. validate_view_with_shacl then decides it the way it decides a result:
- a Violation shape in a reject-mode graph fails the transaction with the same
ShaclError::SparqlConstraintas before; - a Warning or Info shape, or a warn-mode graph, logs it and commits. A failure that repeats across focus nodes is logged once, with the number of focus nodes.
The other engine callers (validate_all, reports) still raise the failure.
A failure in a shape checked as a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, sh:qualifiedValueShape) takes the severity of the outermost shape that reports the result, as that shape's own results do (7999b625e). So a broken constraint in a nested Violation shape is logged under a Warning outer shape, and one in a nested Warning shape fails the write under a Violation outer shape in a reject-mode graph.
The message now names the shape every time (on shape …:), and names the variable rather than its internal id. The aggregate case used to print Aggregate input variable VarId(3) not found in schema. It's now a typed plan error, an aggregate reads variable ?nosuch, which is unbound: nothing before the grouping binds it, and that's also what a query gets. It's a 400 on both query paths, where the tracked path used to return a 500 (aggregate_over_a_variable_nothing_binds_is_a_named_error and its JSON-LD twin). The aggregate-output plan errors are named the same way, in de26852a2.
Worth saying plainly: this changes main's behavior, not only this PR's. A parse or lower failure in a Warning or Info shape, or in a warn-mode graph, used to reject the write. It now commits with a warning. It's in the behavior table and the solo notes.
The tests (shacl_sparql_constraint_failure_*) cover every severity × mode cell for:
- a plan, a parse and a lower failure;
- an aggregate over a variable nothing binds;
- two aggregates with one output;
- an unknown function.
That includes the fail-closed cell with its message, which names the shape and has no VarId. They also cover a Warning shape's failure next to a Violation shape that still rejects. Your repro (a Warning shape, GROUP BY $this projecting ?value) commits.
The split itself has a table test over every error variant (795e31f96, c0e831164). Fuel, cancellation and timeout, memory, storage, catalog, graph-source and policy errors must be raised whatever the severity, and the constraint's own failures must not. Making every error the constraint's own turns it red. The table and a match with no wildcard come from one list, so a new error variant doesn't compile until it has a row.
|
|
||
| **Grouped queries:** A query groups when it has a `GROUP BY` or an aggregate anywhere in SELECT, HAVING or ORDER BY (§18.2.4.1). Its SELECT expressions are evaluated once per group, after HAVING, in SELECT order (§18.2.4.4), and `SELECT *` projects its GROUP BY keys. In HAVING and ORDER BY, a variable that is neither a key nor aggregated means `SAMPLE(?v)` (§18.2.4.1); HAVING without grouping filters the solutions (§18.2.4.2). Two shapes are rejected rather than evaluated literally: a projected variable that is neither a GROUP BY key nor aggregated (§11.4, including when the only aggregate is in HAVING or ORDER BY), and an aggregate over an alias of the same SELECT clause, which the spec would evaluate as unbound (`COUNT` = 0 for every group). | ||
|
|
||
| **Trailing VALUES in a grouped query (deviation from §18.2.4.3):** A `VALUES` block after the WHERE clause, at the top level or in a sub-SELECT, is joined with the WHERE solutions before grouping, so it restricts what the aggregates count: `SELECT ?a (COUNT(?e) AS ?n) WHERE { ?e ex:area ?a } GROUP BY ?a VALUES ?e { ex:e1 }` counts `ex:e1` alone. The spec joins it after HAVING instead, pairing each group row with each VALUES row. HAVING still reads a VALUES variable as unbound, as it would in the spec's order, unless the WHERE binds it or it is a GROUP BY key. For queries that behave the same on any SPARQL processor, put the `VALUES` block inside the WHERE clause to restrict the aggregates' input, or move the grouped query into a sub-SELECT and put the `VALUES` block after the outer WHERE to join it with the grouped rows. |
There was a problem hiding this comment.
At the top level this isn't quite what happens. post_values joins after the WHERE tree, including the binds lowering generates for GROUP BY expressions and aggregate inputs (lower/mod.rs ~518). A sub-SELECT splices VALUES right after its WHERE (select.rs ~875).
SELECT (SUM(?n * ?v) AS ?s) WHERE { ?e ex:n ?n } VALUES ?v { 2 }This returns 0 at the top level and 42 as a sub-SELECT; SUM(?v) in the same top-level query does see the value. The behavior predates this PR, but this paragraph now describes the two levels as the same. Either align the top level with the sub-SELECT or narrow the doc.
There was a problem hiding this comment.
I went with aligning them, in e6e2061e5. The VALUES join stays where it was, after the WHERE tree and before grouping. The binds the level generates now run after it at the top level, as they do in a sub-SELECT: aggregate inputs, GROUP BY expressions, and the SELECT expressions the level places before grouping (all of them, when the level doesn't group). Query.post_values is now a PostValues: the VALUES vars and rows, plus then, the binds that follow it, in order.
Your query now gives the same answer at both levels. On the test fixture SUM(?n * ?v) is 44 at both, and 66 with VALUES ?v { 1 2 }. The ungrouped (?n * ?v AS ?p) and GROUP BY (?n * ?v AS ?k) read ?v too (generated_binds_read_trailing_values_at_both_levels). The compatibility paragraph now says where those expressions run.
While tracing this, it turned out most of the aggregate fast paths ignored a top-level trailing VALUES altogether. On an indexed ledger, SELECT (COUNT(*) AS ?n) WHERE { ?e ex:area ?a } VALUES ?a { "Net" } counted all 6 rows. The count top-k, the star top-k, SUM(?o) and the per-predicate directory count had the same problem; the whole-graph aggregates already declined. 8cfebf749 sends any query with a trailing VALUES to the generic lane. It's pinned with MustNotFire and hand-derived rows in both lanes, the whole-graph aggregates included (3b45c92dd).
And the grammar allows a trailing VALUES after every query form, while only SELECT parsed one. 5efca0d81 reads it after ASK, CONSTRUCT and DESCRIBE too. ASK and CONSTRUCT carry it as post_values, like SELECT. DESCRIBE joins it in its target subquery right after the WHERE, so DESCRIBE ?x VALUES ?x { ex:a ex:b } describes the listed resources.
| /// [`Self::UngroupedRead`] becomes an [`Self::InvalidQuery`] naming them | ||
| /// from `vars`. Every other error is returned unchanged. | ||
| #[must_use] | ||
| pub fn name_variables(self, vars: &crate::var_registry::VarRegistry) -> Self { |
There was a problem hiding this comment.
The JSON-LD update WHERE path doesn't call this. stage.rs runs execute_where_streaming, where a subquery is planned at run time, so a where containing ["query", {"select": ["?a", "?x"], "where": {"@id": "?x", "ex:area": "?a"}, "groupBy": ["?a"]}] returns … projected variable VarId(0) is neither a GROUP BY key…. A name_variables(&txn.vars) map at the two execute_where_streaming call sites in stage.rs would fix it, or one inside execute_where_streaming / WhereCursor::next_batch, which already hold vars.
There was a problem hiding this comment.
I took your second option and fixed it in the cursor itself (cd021b64e). execute_where_streaming names the variables in its plan and open errors. WhereCursor::next_batch names them in errors from the pull, which is where a subquery gets planned. Both go through the registry the cursor already holds, so this covers every caller: JSON-LD, SPARQL and Cypher updates.
Your JSON-LD update now says projected variable ?x is neither a GROUP BY key … (update_grouped_read_error_names_the_variable). The validator rejects the SPARQL UPDATE twin before it gets here, so the test pins that one only for parity.
| /// list element — a cartesian product across list columns. A projected | ||
| /// variable nothing binds is refused as unbound. The plan's output is then set | ||
| /// to refuse per-group lists outright. | ||
| fn ensure_streamable(query: &mut fluree_db_query::ir::Query, vars: &VarRegistry) -> Result<()> { |
There was a problem hiding this comment.
This catches list columns and unbound projections before the 200, but not a bind-stage ungrouped read. That one is only found when the plan is built inside run_stream_query, after the head record. Example: JSON-LD select ["?a", "(as (+ (count ?e) (strlen ?n)) ?x)"] with groupBy ["?a"] returns a 400 on /query, but on the stream the error arrives as a record after the 200. Calling first_ungrouped_read here, with produced_vars_of(patterns) as the WHERE variables, would match /query.
There was a problem hiding this comment.
Done as you suggested, in 97b5b9ed1. ensure_streamable runs first_ungrouped_read over the variables the WHERE and any trailing VALUES bind, with the query's own projection policy. So your JSON-LD query is the same named 400 on /query and on the stream, before the stream starts (jsonld_grouped_read_is_the_same_4xx_on_query_and_stream).
Every other plan-time error still arrives after the head record. Building the plan before the 200 would close that, but it moves reasoning and planning ahead of the response. That's a latency trade-off I'd rather decide on its own, so it's in the description's follow-ups.
| "select": ["?category", "(count ?product)"], | ||
| "groupBy": ["?category"], | ||
| "having": [["filter", "(> (count ?product) 10)"]], | ||
| "having": "(> (count ?product) 10)", |
There was a problem hiding this comment.
The second having example (~line 1645) still uses the [["filter", …]] form.
There was a problem hiding this comment.
Thanks, fixed in 6ca526c24: "having": "(> (count ?product) 5)". Fwiw, that [["filter", …]] form is also what solo's duplicate_pref_labels report sends, and db refuses it at parse, on main too. So the description's solo section now has a one-line lockstep fix for it.
|
Suggested follow-ups (not blocking):
|
Grouped queries that project an aggregate-free SELECT expression (#1978): a control (key + COUNT), the expression with COUNT, with MIN, with no aggregate (dedup-only GROUP BY), a constant under implicit grouping, the expression inside a sub-SELECT, and the JSON-LD twin. Each scenario measures execution plus the rendered response, with allocation peak and churn recorded alongside.
A SELECT expression of a grouped level is an Extend over the group rows (SPARQL 1.1 §18.2.4.4). It belongs to the grouping phase whether or not an aggregation stage exists: a dedup-only GROUP BY ?a projecting (IF(?a ...) AS ?seg) needs one, and Aggregation had no room for it (Grouping::assemble dropped binds when there were no aggregates). Removing the field makes every reader a compile error, so each one was moved deliberately. Detectors that decline on binds now ask the grouping (any binds, with or without aggregates) rather than the aggregation stage. No behavior change: nothing lowers binds onto a dedup-only grouping yet. Also corrects stale docs: binds run after HAVING (the executor's order), not before, and collect() produces a Binding::List, not Binding::Grouped.
A SELECT expression of a grouped query level is an Extend over the group rows, evaluated after HAVING, in SELECT order (SPARQL 1.1 §18.2.4.4). The lowering made every aggregate-free expression a per-solution BIND before grouping. The grouping stage then carried it as a per-group list, and the SPARQL-results formatters expanded the list into one row per solution (#1978): 5 rows for 2 groups, a cartesian product for two expression columns, a lost row for an implicit group over no solutions. The same list reached an enclosing query through a sub-SELECT (outer FILTER, GROUP BY, aggregates, joins) and an UPDATE template. The level's SELECT expressions are now placed after its solution modifiers are lowered, so whether it groups is read from the lowered keys and aggregates: GROUP BY, or an aggregate in SELECT, HAVING or ORDER BY. In a grouping level every expression is a per-group Extend, compound aggregate items included, in one list in SELECT order; that also fixes the alias chains that read an alias before it was bound. An expression whose alias is a group key (the GROUP BY (LCASE(?a)) shortcut) stays a WHERE bind. The placement rule is one helper in the query IR (SelectExprPlacer), so JSON-LD can share it. SELECT * of a grouping level projects its GROUP BY keys, which is nothing for implicit grouping, instead of every WHERE variable as a per-group list; a sub-SELECT likewise exports nothing. Grouped queries with a SELECT expression now take the streaming GroupAggregateOperator plan the key-only query takes, and a dedup-only GROUP BY gets WHERE-level early dedup.
Two shapes that returned silent wrong answers, errors or debug panics now follow SPARQL 1.1 on both surfaces, through two shared IR helpers: - In a grouping level, a HAVING or ORDER BY read of a variable that the WHERE binds and that is not a group key means SAMPLE(?v) (§18.2.4.1). The read is rewritten to a hoisted SAMPLE aggregate, reusing one the level already has. Keys, aggregate outputs, SELECT Extend outputs and variables nothing binds are left alone. It used to return no rows on the streaming lane, panic a debug build on the traditional lane, and reject ORDER BY with "Sort variable ... not found". - HAVING on a level that does not group is a Filter over its solutions (§18.2.4.2). It was dropped. The Filter reads the level's SELECT aliases as fresh, unbound variables, since HAVING cannot see SELECT expressions. SPARQL lowering of a SELECT level (top level and sub-SELECT) is now one shared function in §18.2.4 order: modifiers and aggregates, HAVING, then SELECT expressions. JSON-LD applies both rules to having and orderBy in queries and subqueries. The test that pinned an error for ORDER BY over a non-key variable now pins the sampled result.
A query level groups when it has a GROUP BY or an aggregate anywhere in its SELECT, HAVING or ORDER BY (SPARQL 1.1 §18.2.4.1). The projection check (V4) only counted SELECT aggregates, while lowering counted all three, so a query whose only aggregate was in HAVING or ORDER BY was lowered as grouped but never checked: a projected expression expanded into one row per solution, and a projected variable silently became an implicit GROUP BY key. Both now get the existing V4 error. The definition is one AST helper (SolutionModifiers::level_groups), and lowering asserts in debug builds that it agrees. An aggregate in SELECT, HAVING or ORDER BY that reads an alias assigned by the same SELECT clause is now a named error (V009). The alias is bound by the SELECT's Extend, after aggregation, so per the spec the aggregate sees it unbound (COUNT = 0 for every group). The error points at aggregating the expression itself.
After grouping, a variable that the WHERE binds but the grouping neither keys, aggregates nor binds is a per-group list (Binding::Grouped). Only a JSON-LD top-level projection may carry one, as documented. Every other reader used to invent its own meaning: FILTER and HAVING panicked a debug build and read unbound in release, a join matched anything, a GROUP BY merged every list into one group, and a JSON-LD subquery returned the list to its enclosing query, where a join on it silently matched nothing. - QueryOutput::Select carries UngroupedProjection: Reject by default, PerGroupList set only by the JSON-LD user-query entry (parse_query). A sub-query's projection is a plain variable list, so it cannot carry a per-group list at all. - Grouping::first_ungrouped_read is the one rule, enforced in apply_solution_modifiers, which the top-level and sub-query pipelines share: HAVING, the grouping's binds, ORDER BY binds and keys, and a Reject projection may not read such a variable. It fails the plan with a 4xx in every build and replaces the two ad-hoc ORDER BY checks. SPARQL and JSON-LD never reach it on HAVING or ORDER BY: lowering reads those as SAMPLE. It is the backstop for Cypher, hand-built IR and any future lowerer. - Without dependency sets (the sub-query pipeline), the traditional grouping lane now carries only keys and aggregate inputs, instead of every WHERE column as a list. - Grouping::assemble returns an error instead of dropping a HAVING or binds that have neither a key nor an aggregate. - A Grouped binding reaching scalar evaluation, a join or an OPTIONAL substitution is an Internal error in every build; the sort comparator and the group-key conversion, which cannot fail, keep a debug assertion.
The SPARQL JSON, NDJSON and XML writers expanded a Binding::Grouped cell into one row per list element: a cartesian product across list columns, and a dropped row for an empty list. That was the visible symptom of #1978, and it still reached JSON-LD queries rendered as SPARQL results (97 streamed rows from 2 groups on a two-list query). SPARQL queries can no longer produce such a cell, and SPARQL results have no list type, so the writers now refuse it with a FormatError, as they already refused a Cypher path or list. Deleted: disaggregate_row and the per-row list scans in format_string and stream_ndjson_rows (every row now streams cell by cell), the XML write_grouped_rows, and the CLI table's disaggregation fallbacks (the table renders a stray list ';'-joined, like CSV). The streaming endpoint refuses a JSON-LD query that projects a per-group list with a 4xx before the stream starts, naming the column, and plans the rest with per-group lists disallowed, so a wildcard projection streams keys and aggregates.
JSON-LD placed a select expression before grouping unless it read an aggregate output, so a key-only expression under groupBy -- the #1978 shape, e.g. (as (if (= ?a "Net") "network" "other") ?seg) -- came back as a per-group list of one value repeated, where SPARQL and Cypher give one value per group. JSON-LD now places its select expressions with the SPARQL rule (SelectExprPlacer), in one shared function for queries and subqueries. In a grouping query, an expression over keys, aggregate outputs or earlier per-group aliases, or a constant, runs once per group after having. Unchanged: an expression over a non-key variable keeps the documented per-group list, as does a chain over it and an alias an aggregate reads, and an alias that is itself a groupBy key is computed before grouping. A per-group alias may not reuse a variable the where clause binds: it would silently replace that variable's value after grouping, the groupBy key included (SPARQL rejects the same alias). The implicit output name of a bare aggregate ((count ?e) -> ?count), which downstream code reads by name, is now pinned by a test. Docs: the select-expression, groupBy and having sections of jsonld-query.md describe the grouping rules; the having example used a form that does not parse.
Cypher makes a non-aggregate RETURN expression a grouping key, so RETURN CASE ... AS seg, count(e) groups by the label, which is SPARQL's GROUP BY (IF(...) AS ?seg). Grouping by the key first and mapping after (WITH e.area AS a, count(e) AS n RETURN CASE ... AS seg, n) is what a SPARQL SELECT expression over GROUP BY ?a now means. The parity test pins both pairs on a non-injective mapping (three areas, two labels), where the two readings differ. Must-not-change guards for the grouped shapes fluree/solo runs today: a nested grouped sub-SELECT, SUM(IF(...)) buckets per key, aggregate- only sub-SELECTs under SELECT *, a compound expression of SAMPLEs beside a distinct count, and the grouped paging subquery ordered by its aggregate. SHACL sh:select constraints are lowered like SPARQL queries but skip the SPARQL validator, so the grouped semantics reach stored shapes directly: HAVING without grouping filters, HAVING over a non-key variable reads a sample, and a projected non-key variable fails the transaction with the plan error instead of reporting a violation at the focus node. Docs: sparql.md and compatibility.md describe grouped SELECT expressions (one per group, after HAVING, in SELECT order), SELECT * under grouping, implicit SAMPLE in HAVING and ORDER BY, HAVING without grouping, and the two shapes rejected instead of evaluated literally.
…ed lowering The planner's produced_vars_of gives the variables some pattern may bind (must_bind_vars gives the ones bound on every row). The grouped lowering computed the same set in four more places: the SPARQL level (pre_group_vars) and its SAMPLE rewrite, and the JSON-LD select computations and SAMPLE rewrite. It now lives in ir::pattern beside the other pattern-set helpers, and all five callers use it. No behavior change.
The plan-time grouped-read check (Grouping::first_ungrouped_read) runs in the planner, which has no variable registry, so its 4xx named the variable by internal id: "projected variable VarId(3) is neither a GROUP BY key nor an aggregate result". A JSON-LD subquery projecting a non-key variable reaches it, as does a JSON-LD select expression that mixes an aggregate with a non-key variable. The planner now returns a typed QueryError::UngroupedRead carrying the ids, and QueryError::name_variables turns it into InvalidQuery naming them wherever the query's registry is at hand: the runner's execute and execute_prepared (which also covers a sub-query, planned while the query runs), the API's view, dataset and stream prepare sites, and EXPLAIN. An id the registry does not hold prints as before. Where the typed error surfaces unnamed it keeps InvalidQuery's 400 and stream error code.
…ader J3 evaluates a JSON-LD select expression over group keys once per group. When a later expression reads that alias together with a non-key variable, as in `(as (strlen ?a) ?len)` then `(as (+ ?len (strlen (str ?e))) ?x)` under `groupBy ?a`, the reader ran per group as well, and the plan-time check rejected its non-key read. Before J3 the query returned per-group lists. An alias that a before-grouping expression reads now moves before grouping with it: the alias is constant within its group, so the reader stays per solution and both come back as per-group lists again. SelectExprPlacer places a level's expressions together (place_all) so it can see the readers. The SPARQL twin stays a V4 error, because a SELECT expression of a grouped level cannot read a non-key variable.
…gates `WITH p, count(f) AS c ORDER BY c + 1` and `RETURN …, count(f) AS c ORDER BY -c` failed: "an ORDER BY key is neither a GROUP BY key nor an aggregate result". The sort key lowers as a bind, and one level placed it in the aggregation's body, before grouping, where `c` does not exist. The clause now takes the two-level form whenever its WHERE or ORDER BY lowered anything the second stage can see, not only an output node's property: the aggregation, then the binds, the filter, the sort and the slice, as an explicit second WITH does. A clause that lowers nothing there keeps one level, and a read the second stage cannot see stays a plan error.
A property read joins every value, as `MATCH (p) WHERE p.age > 30` does. Reading it after the aggregation keeps the clause's own aggregates exact, but the rest of that model stands, and the reference now says so: in WHERE the row comes once per value that passes; in ORDER BY once per value, and DISTINCT keeps the copies (the sort key is projected with them); an aggregate in a later clause counts every copy. A read inside an aggregate's argument is still joined before grouping, so it repeats the group's rows for every aggregate of the clause. The test pins each in row order, so a change to the model flips them on purpose: an unsliced ORDER BY in a WITH and in a RETURN, a WITH DISTINCT … ORDER BY, count(*) and sum(c) after the WHERE, and collect(p.age) beside count(f).
A table test over raised_whatever_the_severity, one representative per QueryError variant (R2rml excepted: its error type lives in a crate this one does not depend on). The request's budgets, cancellation, memory, storage, catalog and policy access, data state and internal faults must fail the write whatever the severity; the constraint's own failures follow severity and mode. An exhaustive variant-name match makes a new variant fail to compile until the table places it. Three arms can't be reached today, and each now says why: the caller names variables before classifying, which turns an UngroupedRead into an InvalidQuery; NoSuitableIndex and UnsupportedMode are constructed nowhere.
A MustFire that failed because its site declined at open read "must proceed [proceeded: [..., group_by_object_count_topk]]". It now reads "must serve it" and lists the sites that declined at open beside those that proceeded.
As for select, a malformed value is now a 400; ask ignored both before.
Each entry of the table names a QueryError variant's pattern, the class it must get and its representatives. A macro builds the rows and a match over the same patterns with no wildcard, so a new variant fails to compile until it has an entry, which brings a class and a representative with it; each representative must match its own pattern. The earlier version forced only the variant's name, and its count was a manual step. R2rml has a row now (R2rmlError::Parse), through a dev-dependency on fluree-db-r2rml, which fluree-db-query already depends on; normal, no-default-features and wasm dependency trees are unchanged.
`WITH p, collect(f) AS fs ORDER BY size(fs)` was refused as an ORDER BY on a collect() list: the check refused any key that read the list. The reference says only a key on the list itself is refused, and a key that reads the list and yields one value now sorts after the aggregation, as other expression keys do: size(fs), a list predicate such as any(x IN fs WHERE …), a comparison. A key whose value is the list, or a list built from one (fs, tail(fs), fs + ['x']), is still refused, since sorting a list value is deferred. The reference also now says that `WITH DISTINCT` keeps the copies a multi-valued sort key makes and `RETURN DISTINCT` removes them.
`"orderBy": "(desc (count ?e))"` failed with "orderBy must be an array of
objects with 'var' field", which is not what orderBy takes. Sort keys
are variables; the error now lists the forms and says to select an
expression or aggregate under an alias and order by the alias. An
expression in a term's variable position (`["desc", "(count ?e)"]`,
`{"var": "(count ?e)"}`) gets the same error instead of "Invalid
variable syntax".
…y alias An aggregate can't appear in orderBy, so a grouped query is one with a groupBy or an aggregate in select or having. The orderBy section shows sorting on an aggregate through its select alias.
An sh:sparql constraint that cannot run was recorded with the severity of the shape that owns it. Checked as a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, or sh:node and sh:qualifiedValueShape on a property shape), that shape's own results surface only through the outer shape's, at its severity, but the failure kept the nested shape's: Violation by default, so it rejected the write under a Warning outer shape, and a Warning nested shape's failure was only logged under a Violation outer one. The failure now takes the severity of the outermost shape that reports the result: node-level combinators report under the node shape, the nested checks of a property shape under the property shape, and an enclosing nesting keeps its own. The nested shape's results keep their severity, which the conformance check reads. Errors that fail the write whatever the severity (budgets, cancellation, storage, catalogs, policy, internal faults) are raised before any failure is recorded, on the nested route as on any other. Tests cover every nesting, including one through a Violation middle shape, for Warning, Info and Violation outer shapes, a Warning inner shape, and a warn-mode graph, with shapes written in JSON-LD and with SPARQL UPDATE.
A sub-SELECT projecting an expression over its group keys, with an outer ORDER BY on the alias, ran the expression per solution before this branch: the alias was a per-group list, and the outer sort compared two lists (a debug_assert in the sort comparator, and no order in release builds). Evaluated once per group, it sorts. Pinned for SPARQL (STRLEN of a key, CONCAT over two keys, and COUNT * 2 as a control that already worked) and for a JSON-LD subquery.
openCypher's WITH applies its WHERE to the clause's results
(`WITH … [ORDER BY] [SKIP] [LIMIT] [WHERE]`), so
`WITH x ORDER BY x LIMIT 5 WHERE x > 3` keeps 4 and 5. Every WITH
lowering filtered before the slice and kept 4 to 8.
A WITH with SKIP or LIMIT and a WHERE now lowers without its WHERE,
and a stage over the sliced rows joins the WHERE's property accessors
and filters. The rows keep the clause's order, and the stage repeats
DISTINCT. This covers the plain, one-level aggregating and two-level
aggregating forms.
After a plain WITH, the sliced rows carry the earlier variables the
WHERE reads, including those of an exists { … } body, and the stage
drops them. After an aggregating or DISTINCT WITH, a WHERE read of
anything the clause does not project is an error rather than a free
variable.
A WITH without a slice lowers exactly as before.
The shape compiler scans SHACL predicates in a fixed order, and the sh:severity, sh:message, sh:name and sh:description arms wrote only to a subject a shape map already held. A node shape registered later lost them: one whose only constraint is sh:sparql is registered by its rdf:type sh:NodeShape after the scan, so it compiled as a Violation shape with no message, and a Warning or Info shape rejected the write. So did a shape whose registering predicate is in a later graph. The arms now record metadata for an unregistered subject, and claim_metadata gives it to the shape once the scan is done. It routes the metadata as the arms do (property shape first) and keeps any value the shape was given after it was registered. The nested-severity test no longer needs its inner shape to target an unused class.
…ction A JSON-LD @list or a Turtle collection stores one statement per member, and the sh:and / sh:or / sh:xone arms took each object for a shape. A SPARQL UPDATE stores `( … )` as an rdf:first / rdf:rest collection, so the arm got the collection's head, and validation failed with "Referenced shape … could not be resolved". List processing now replaces a member that heads a list in the graph with the list's members, as it already does for sh:ignoredProperties; sh:in, sh:languageIn and sh:path already read both forms. The and_list / or_list / xone_list fields, which nothing ever set, and their expansion loops are gone. The nested-severity test now writes every nesting with SPARQL UPDATE too.
…very query form Parameter substitution refuses a parameter that the query itself assigns (BIND, VALUES, AS), and for SELECT that includes the trailing VALUES clause. ASK, CONSTRUCT and DESCRIBE now take a trailing VALUES too, and their arms didn't walk it, so a parameter named there was substituted in the WHERE while the VALUES still bound the variable. All four forms now walk their trailing VALUES through one helper.
…riable it carries The item-31 test's exists cases gave the same rows whether the body's f was carried or free. Two cases tell them apart: exists over a carried f keeps Alice's one row whose friend she likes, and a body variable of its own (g) is free, so both of her rows pass.
`CALL (p) { MATCH (p)-[:knows]->(f) WITH f RETURN f.name }` crossed every
outer p with every knows edge. The WITH lowers to a subquery that
selects f alone, and a subquery correlates with its parent only on what
it selects, so the CALL's seeded p never reached it. A slice in the
body ran once for all of them.
Inside a CALL body, a subquery the Cypher lowering builds (a WITH, its
two-level and after-the-slice stages, an aggregating RETURN's first
level) now selects the imports it reads, groups by them when it groups,
and pins them. The imports it carries without projecting stay out of
scope after the WITH: a nested CALL can't import them, and WITH * doesn't
bring them back.
The planner treated a CALL whose body binds its import as uncorrelated,
because the nested WITH now selects it, and could place the CALL before
the import's producer. A pinned import is now always a correlation input
for placement. The executor still evaluates once and hash-joins when
the body allows it, and runs per row when a slice needs it.
…r their severity Whether a node conforms to a nested shape (sh:node, sh:not, sh:and, sh:or, sh:xone, sh:qualifiedValueShape, at node level and per value) counted only Violation results. SHACL §3.5 counts every result: a node conforms to a shape when validating it there reports none, and §2.1.4 has severity only categorize results. So a Warning nested shape that reported a result still conformed: sh:node passed and sh:not failed. The six sites now share one `conforms` check: the node-level combinators and check_value_against_nested_shape, which serves every per-value form and the literal-focus checks. The severity that decides a write is the outermost reporting shape's, on the result it reports. A constraint that cannot run is not a result, and is decided at that same outermost severity, as before.
…atever their severity
… no matches, as RETURN does
c67dc35 to
87576fd
Compare
|
The deviations from the design, each written up in full: moved out of the description to keep it under GitHub's size limit. They describe the branch at
|
|
Moved out of the description to keep it under GitHub's size limit: the full test list (current at Full test list
Added since the review, by binary:
Perf detail, when this PR was first openedNew bench. Local run: bench profile (fat LTO), Watched benches, which the design expected not to change: Tiny (CI's scale, 10% budget), four reps, two in each order. All 29 scenarios have a mean change within ±5.3%: Q9 +1.3%,
Small (5% budget). I ran the whole suite twice in opposite orders, the overlay matrix twice more, and then the outliers in isolation using criterion's filter, three alternations each. One overlay rep ran under a load spike of 44 and is excluded.
Quiet box, session 2, at
Part C,
Formatter micro-bench Gates, when this PR was first openedThe full gates ran at
Across the four interleaved A/B pairs, the head is faster on two (8%, 15%) and slower on two (6%, 21%), and the 21% — the last compaction run — ran at load 13–15 against 10–12 for its base run. Both compaction tests run no code this PR changes, and the same slowness shows on main, so it's the box, not this PR. Not run: the
Run locally (
|
|
This continues the non-vacuity log in the first comment, which covers the PR as first opened. It has every fix since then, from the review onward. Same method throughout: each fix was committed first. Then only the fix was reverted, using a needle taken from the post-fmt source and required to match an exact number of times. The tree was rebuilt, the named tests were run and watched fail on an assertion (never a compile error, except the one row below that is meant not to compile), and the files were restored with Grouping: EXISTS, HAVING and sort keys
Cypher
What the earlier SAMPLE reading did with a two-valued property. The
Alice appears twice in the fixed column because both of her ages pass the filter: the second stage reads every value. That matches the existing Cypher behavior for a multi-valued property, checked on the same fixture. Trailing VALUES, and ASK / CONSTRUCT / DESCRIBE
Fast paths and routing canariesThe canary pairs check that each fast path that gates on the grouping's binds serves the key-only query (MustFire) and declines the same query with a SELECT expression (MustNotFire).
SHACL
The outer-severity test without its workaround. Error names, formatting and the stream
The sub-SELECT sort panic, checked rather than mutated. A sub-SELECT's grouped SELECT expression, sorted in the outer ORDER BY, is fixed by this PR's per-group evaluation itself, not by one commit. So I ran the same queries on debug CLIs built at main ( Not mutated, and why
|
|
Thanks for these, @bplatz. All four ended up in this PR rather than as follow-ups, and the branch is rebased onto
A few more things turned up while tracing these. They're all pre-existing, and I folded them in rather than file them, so each has a row in the description's behavior table:
The last ones change the answers, or the write verdicts, of queries and shapes that run today without an error, so I'd especially like your eyes on these. Happy to talk through any of them:
|
This is more than #1978 needs on its own, and that's deliberate. What an expression means under GROUP BY, HAVING or implicit aggregation was decided in several places, so each shape broke its own way: #1362 with the expression as the group key, then #1978 with it in the SELECT. So rather than patch the formatter, grouping becomes one concept that owns its expressions: each is evaluated once per group, HAVING is just a filter, and a value the grouping can't supply is a plan-time error, not silently wrong rows. That leaves one place to reason about and tune. Fwiw, it's one of five PRs taking this approach, with #2007, #2008, #2009 and #2010.
Fixes #1978
Since the review
@bplatz approved
c67dc3516. This push is that branch rebased onto745736a04, with every review item addressed, plus a few pre-existing bugs in the same class that I decided to fold in rather than file. Briefly:WITH/RETURNreads a projected node's property after the aggregation; ash:sparqlconstraint that can't run follows its shape's severity and the graph's mode; at the top level, the generated binds now run after a trailing VALUES, as they do in a sub-SELECT; update WHERE errors name the variable; the stream refuses a bind-stage grouped read before it starts; and the secondhavingdoc example is fixed.ORDER BY ?nosuchorders nothing instead of failing with a 500; every fast path that gates on the grouping's binds has a routing-stamp canary pair; and HAVING reads a SELECT expression's alias (my call, over a diagnostic; it's an extension, documented in the compatibility reference).WITH'sWHEREfiltered before itsSKIP/LIMIT; aWITHinsideCALL (p) { … }lost its correlation withp; and three SHACL bugs: ansh:sparql-only node shape lost itssh:severity,sh:and/sh:or/sh:xonelists written by SPARQL UPDATE didn't resolve, and a nested shape's Warning results didn't count toward conformance. There are also a handful of smaller ones (error texts, JSON-LDaskoptions, per-group lists on indexed ledgers). Each of these changes what main returns today, so folding them in was a deliberate call rather than a review fix.Every one of these that changes behavior has a row in the table below.
What was wrong, and the fix
#1978 shows up in the SPARQL-results formatter, but the defect is upstream of it, in how we lower grouped SELECT expressions. A SELECT expression at a grouped query level is an
Extendover the group rows, evaluated once per group after HAVING, in SELECT order (SPARQL 1.1 §18.2.4.4). Our SPARQL lowering instead made every aggregate-free SELECT expression a per-solutionBINDbefore grouping, so the grouping stage carried the alias as a per-group list (Binding::Grouped) — and each reader of that list invented its own meaning for it:$thisrows below). A GROUP BY over it merged every group into one, and an UPDATE template refused it.So the class is: after the group step, something reads a WHERE variable that is neither a GROUP BY key nor an aggregate output. Rather than teach the formatter to cope, this PR closes that class structurally:
fluree-db-sparql). A level's SELECT expressions are now placed after its modifiers are lowered, from one shared rule (ir::SelectExprPlacer). In a grouping level (GROUP BY, or an aggregate anywhere in SELECT, HAVING or ORDER BY) every expression is a per-group Extend, in one list in SELECT order — except an expression whose alias is itself a group key (theGROUP BY (LCASE(?a))shortcut), which stays a WHERE bind.SELECT *of a grouping level projects its GROUP BY keys (nothing, under implicit grouping). The top level and sub-SELECTs share one function,lower_select_level.fluree-db-query). Per-group binds moved fromAggregationtoGrouping, so a dedup-onlyGROUP BYcan carry them (and removing the field turned every reader into a compile error).QueryOutput::Selectnow carries anUngroupedProjection:Rejectby default,PerGroupListonly for the JSON-LD user-query entry. A sub-query's projection is a plain variable list, so it can't carry a per-group list at all.Grouping::first_ungrouped_read, enforced inapply_solution_modifiers(which the top-level and sub-query pipelines share), rejects any HAVING, bind, ORDER BY orReject-projection read of such a variable with a 4xx, in every build, and replaces two ad-hoc ORDER BY checks. The lowerers rewrite HAVING and ORDER BY reads of such a variable toSAMPLE(?v)first (sample_ungrouped_reads, item 5), so for those the check is a backstop: it catches SELECT-expression and projection reads underReject, and IR from a lowerer that doesn't apply the rewrite (Cypher: a property of a node the clause doesn't project,WITH e.area AS a, count(*) AS c ORDER BY e.name). AGroupedvalue reaching scalar evaluation, a join or an OPTIONAL substitution is now an internal error instead of a debug panic / silent unbound.FormatError, as a Cypher path or list already was, anddisaggregate_row, the XMLwrite_grouped_rowsand the CLI table fallbacks are deleted. The streaming endpoint refuses a JSON-LD per-group list before the 200.5 and 6 below are calls I made on spec edges and on JSON-LD, and 7 is the rest of the class, which an adversarial review pass turned up — one piece of it, trailing VALUES, is also a call of mine. So that's probably where review time is best spent; happy to talk through any of them.
SAMPLE(?v)(§18.2.4.1);AggregateOverSelectAlias); per the spec it would see the alias unbound.docs/reference/compatibility.md). HAVING still reads a VALUES variable as unbound, as it would in the spec's order, unless the WHERE binds it or it's a GROUP BY key or an aggregate output — at the top level and in a sub-SELECT (the implicit SAMPLE had kept every group forHAVING (?v = 1) VALUES ?v { 1 });HAVING (EXISTS …)works (it was never resolved: always false), and a Cypher metadata read in an aggregatingWITH … WHEREsees policy-filtered flakes (it read as empty under a view policy);askgroup to spec instead of dropping GROUP BY / HAVING (the first version of this PR refused them; the review asked for the spec semantics);$this,$PATH) groups any level that projects it, so the sub-SELECT family that must project$thisevaluates, and a shape query the planner rejects is ash:sparqlconstraint failure naming the constraint, its shape and the variable, decided by the shape's severity and the graph's mode;select *undergroupBybefore the response starts when/querywould return per-group lists for it (the grouping'sGroupByOperatorlane). Main streamed those lists as cartesian rows, one per list element;/queryand on the stream ("projected variable ?x is unbound"). Main failed it on both with "Variable not found: Selected variable VarId(n) not found in query schema" — a 500 on the tracked/querypath, and on the stream an error record after the 200;?x), never internal ids, and describe a synthetic variable (a Cypher property access) instead of printing its internal name.Behavior changes
Class key: RESULTS CHANGE = the query still runs and returns different rows or values; NOW ERRORS = a query that runs today fails (4xx); ERROR → ROWS = a query that fails today returns rows; ERROR → COMMIT / COMMIT → ERROR = a write SHACL rejects today now commits, or the reverse. The shapes that don't change (each pinned by a test) and the Rust API changes are folded under the table.
("x" AS ?c) (COUNT(*) AS ?n)x, 0)SAMPLE(?v)semanticsSAMPLE(?v)SELECT *with an aggregate only in HAVING / ORDER BY(IF(…) AS ?seg) (COUNT(?seg) AS ?c)HAVING (EXISTS …)/(NOT EXISTS …), and an EXISTS SELECT expressionWITH … WHERE(lowered to HAVING) that reads node metadata (keys,labels,properties, …) under a non-root view policyMATCH … WHEREASK { … } HAVING (?a = "Nope")was true; CONSTRUCT built every solution's triplesgroupByaskwithgroupBy/havinghavingthat rejects every group answeredtrue)truewhen some group passeshavinggroupByselect *undergroupBywhen/queryreturns lists, i.e. on theGroupByOperatorlane) on the NDJSON stream / SPARQL JSON / SPARQL XMLuse /query or project the keys) /FormatErrorgroupBy/query: "Variable not found: Selected variable VarId(n) not found in query schema" (500 on the tracked path); stream: the same, after the 200sh:select(lowered like SPARQL, but the SPARQL validator does not run) with HAVING and no grouping, or HAVING over a non-key variableSAMPLE(?v)sh:selectprojecting a non-key variable of a grouped level, e.g.SELECT $this ?value … GROUP BY $this(V4 would reject it in a SPARQL query)sh:sparqlconstraint failure naming the constraint, its shape and?value, decided by severity and mode (row below)$this(every sub-SELECT must), e.g.{ SELECT $this ?v (COUNT(?x) AS ?c) … GROUP BY ?v } FILTER(?c > 1)$thiscame back as a list, which joined nothing)$thisgroups the level (it is constant per evaluation)WITH … WHERE) reading a SELECT expression's alias, e.g.(COUNT(?e) + 0 AS ?n) … HAVING (?n > 1),WITH p, count(f) + 0 AS c WHERE c > 1WITH … WHERE/ORDER BYor aggregatingRETURN … ORDER BYreading a property of a node the clause projects, e.g.WITH p, count(f) AS c WHERE p.age > 30,ORDER BY p.age + 1WHERE:[];ORDER BY: 400WITH p, c WHERE p.age > 30reads it, so a read inWHEREorORDER BYdoesn't change the clause's aggregates. A property with several values joins every value, asMATCH … WHEREdoes: inWHEREthe row comes once per value that passes; inORDER BYonce per value, andWITH DISTINCTkeeps those copies (the sort key is projected with them) whileRETURN DISTINCTremoves them; an aggregate in a later clause counts each copy. A read inside an aggregate's argument (collect(p.age)) is still joined before grouping, so it repeats the group's rows for every aggregate of that clause (unchanged; follow-up)WITH … ORDER BY/RETURN … ORDER BYon an expression over its aggregates, e.g.ORDER BY c + 1,ORDER BY -cORDER BYa value read from acollect()list, e.g.WITH p, collect(f) AS fs ORDER BY size(fs),ORDER BY any(x IN fs WHERE …)fs,tail(fs)), is still a 400orderByholding an expression or an aggregate, e.g."(desc (count ?e))",["desc", "(count ?e)"]orderBytakes and says to sort on an alias,(as <expr> ?k)ORDER BY ?nosuch, grouped or not, on/queryand the streamSELECT (SUM(?n * ?v) AS ?s) … VALUES ?v { 2 }?vunbound,SUM= 0 (a sub-SELECT already saw?v)COUNT(*) … VALUES ?a { "Net" }gave 6, not 3f:reifies*IRIsh:sparqlconstraint whose own query fails (it does not parse, lower or plan,$PATHhas nothing to bind, or it fails on its own terms, e.g.SUM(?nosuch)or an unknown function) on a Warning / Info shape, or in a warn-mode graphsh:sparqlconstraint that cannot run, in a shape checked as a nested shape (sh:node,sh:not,sh:and,sh:or,sh:xone,sh:qualifiedValueShape)VarId(0)/query; on the stream, the error arrived after the 200sh:selectSUM(?nosuch)sh:selectWITH p, count(f) AS c, sum(c) AS sASK … OFFSET nandASK … LIMIT 0OFFSET 2over two solutions andLIMIT 0answeredtruetrueonly when a solution remains after themVALUESafter ASK, CONSTRUCT or DESCRIBE (the grammar allows one after every query form)DESCRIBE ?x VALUES ?x { ex:a }describesex:aaskwithoffset,limitorvalues;constructwithvaluesselectWITHwith both aWHEREand aSKIPorLIMIT, e.g.UNWIND range(1, 10) AS x WITH x ORDER BY x LIMIT 5 WHERE x > 3WHEREfiltered before the sort and the slice: 4 to 8WHEREfilters the clause's sliced rows, as openCypher places it: 4 and 5. Plain, aggregating andDISTINCTWITHs alike; aWITHwithout a slice is unchangedDISTINCTWITH, aWHEREreading a variable the clause doesn't project, e.g.WITH DISTINCT p.age AS a ORDER BY a LIMIT 2 WHERE p.name = 'x'DISTINCT: filtered before theDISTINCTp, which is out of scope: after an aggregating or DISTINCT WITH, its WHERE sees only the variables the WITH projects"DISTINCT) / error text (aggregating)CALL (p) { … }, aWITH(or an aggregatingRETURNthat sorts on an expression) whose body reads the import, e.g.MATCH (p:P) CALL (p) { MATCH (p)-[:knows]->(f) WITH f RETURN f.name AS fn } RETURN p.name, fnpcrossed with the body's whole result: 24 rows instead of 6 on a four-person fixture. A slice,DISTINCTor aggregate in the body ran once for all of them, soWITH f ORDER BY f.age LIMIT 1gave everypthe same friendp's own friends, and a slice,DISTINCTor aggregate applies perp. As for an aggregating CALLRETURN(unchanged), an import whose body matches nothing gets no row from an aggregate, even one with no grouping key;OPTIONAL MATCHkeeps it as0(follow-up)sh:sparql(nothing but itsrdf:type sh:NodeShaperegisters it), withsh:severity sh:Warningorsh:Infosh:message,sh:nameandsh:description, and for a shape whose registering statement is in a later shapes graphsh:node,sh:not,sh:and,sh:or,sh:xoneorsh:qualifiedValueShape, at node level or on a property's values) that reports a Warning or Info result for the nodesh:nodeand the logical combinators passed,sh:notfailed, and a qualified count counted the valuesh:notis satisfied. Thesh:sparql-only severity fix (row above) made this reachable for a nestedsh:sparql-only Warning shape, which used to compile as Violationsh:and/sh:or/sh:xonelist written by a SPARQL UPDATE (( ex:A ex:B ), stored asrdf:first/rdf:rest)GROUP BY ?ocount top-k or the per-predicate directory count serves, e.g."select": ["?a", "?e", "(as (count ?e) ?n)"]withorderBy (desc ?n)/limitnullfor the list; the directory count failed ("Projected variable not in child schema")Shapes that don't change (pinned), and the Rust API changes
… GROUP BY ?a VALUES ?e { ex:e1 }(VALUES as a parameter)(Net, 1));VALUES ?x { 1 2 }doubles COUNTGROUP BY ?a HAVING (?v = 1) VALUES ?v { 1 }GROUP BY ?v HAVING (?v = 1) VALUES ?v { 1 2 }gives(1, 6), as before(as (strlen ?a) ?len)+(as (+ ?len (strlen (str ?e))) ?x)havingwith an["exists", …]formhaving)(count ?t)→?count)?count?count(now pinned)Grouping::assembleOption<Grouping>Result<Option<Grouping>, GroupingError>(HAVING / binds with no key and no aggregate are an error, not dropped)Aggregation.bindsGroupingvariants;Grouping::binds()unchangedQueryOutput::Select{ projection, restriction }+ ungrouped: UngroupedProjectionQueryErrorUngroupedRead(ir::UngroupedRead);QueryError::name_variables(&VarRegistry)turns it intoInvalidQueryShaclError::SparqlConstraint { constraint, .. }constraint: Sidconstraint: String, the constraint's IRIHavingOperator::new(child, expr)(child, expr, planning)Query.post_valuesOption<Pattern>Option<PostValues>(vars,rows,then: the binds after the join)fluree_db_shaclConstraintFailure,ConstraintFailures;ShaclEngine::with_constraint_failuresir::ConstructTemplatesubstitute_varfluree_db_sparql::ast::{AskQuery, ConstructQuery, DescribeQuery}values: Option<Box<GraphPattern>>(the trailing VALUES clause, as onSelectQuery); a struct literal outside the crate must set itir::ReadStage(new in this PR)UnboundAggregateInput,BoundAggregateOutput,RepeatedAggregateOutput: the grouping-stage plan errors about one variable, named byQueryError::name_variablesWorth calling out: stored artifacts. Saved queries that use any shape in the table above behave differently after upgrade — the RESULTS CHANGE rows change silently, and the NOW ERRORS rows start failing. And since SHACL shapes are stored in the ledger, a shape whose
sh:selecthas one of the SHACL shapes above changes which transactions commit. What that means for solo is just below.Downstream: fluree/solo
The short version for solo is that its own queries are fine, but stored artifacts may not be, and we can't enumerate those from the repo.
Per a survey of fluree/solo's checked-in and generated queries (at solo
main0262cf8, and re-diffed since with no relevant change), no static solo query changes result: every grouped query in the repo projects only keys, aggregates or expressions of aggregates, and solo's reads of the implicit?countcolumn name keep working (pinned here). The must-not-change guards in this PR cover solo's static grouped shapes.Stored artifacts do change, though: saved dashboards, decks and publications. Publications in particular write SPARQL-JSON rows into customers' Iceberg tables on every run and have no update endpoint. So before we bump the db pin, we should scan the stored dashboard/deck specs and the
_systemTablePublication queries for the shapes in the table above, in particular:HAVING (EXISTS …)/(NOT EXISTS …): now evaluated per group;askwithoffset,limitorvalues, andconstructwithvalues: now applied (was: ignored). A malformedoffsetorlimitonask(a string, a negative or fractional number, one beyond u64) is now a 400, as forselect(was: ignored);null, or an error);WITHwith both aWHEREand aSKIPorLIMIT: theWHEREnow filters the sliced rows;WITHinside aCALL (p) { … }body now stays correlated withp.Solo generates neither of those two Cypher shapes. Its Cypher surfaces (the client's
cypherQuery/cypherWrite, the query and transact lambdas, the Cypher import) pass a caller's statement through, and a grep of its Rust, TypeScript, Markdown and JSON at solomainc1efea7c6bfinds noWITH … LIMIT, noWITH … SKIPand no CypherCALLin its code, prompts or docs. So those reach solo only through user-written statements.SHACL is the other exposure (the SHACL and Rust-API greps below are at solo
main3bf84bc, whose db pin is stillba984c7f5). Solo's query Lambda runsvalidateover user-authored shapes and boundssh:sparqlevaluation with a fuel ceiling (fluree-lambda-query/src/handler.rs:1358,:1511); solo itself authors nosh:selectshapes (grep of3bf84bc: the onlysh:sparqlmentions are those two comments). So it's user-authored constraints that move:$thisnow fires — it never fired before, so writes it should have refused start being refused;sh:severityand the graph's mode. On a Warning or Info shape, or in a warn-mode graph, it's logged and writes commit; at base, parse and lower failures rejected every write; in a shape checked as a nested shape (sh:nodeand the logical constraints), the outermost reporting shape's severity applies;sh:sparqlnow commits with a warning (was: rejected). Solo authors nosh:sparqlshape (its query lambda only bounds user-authored ones), and the node shapes it generates carrysh:targetClass, which already registered them;sh:and/sh:or/sh:xonelists written by SPARQL UPDATE now resolve. Solo writes shapes as JSON-LD and authors nosh:and,sh:ororsh:xone;web/src/lib/ontology/writes/shacl.ts,sh:nodeto…-scheme). Its property shape…-innever carriessh:severity, so its results are Violations, which already counted. Solo authors nosh:not, logical combinators or qualified shapes. A user-authored nested Warning shape follows the new rule.On the Rust API (grep of solo
3bf84bc), solo references none of the changed items (Grouping::assemble,Aggregation,QueryOutput::Select,HavingOperator,ShaclError::SparqlConstraint,ReadStage,UngroupedRead,VariableNotFound), and itsQueryErrormatches all have a catch-all arm;fluree-lambda-transactconstructsShaclError::{CoreError, QueryError, InvalidPattern}, which are unchanged. Solomainf5d7518f72(2026-10-05) references none ofpost_values,ConstructTemplate,ConstraintFailure,with_constraint_failuresorQueryOutput::Ask, and no solo Rust or TypeScript builds a grouped ASK or CONSTRUCT. Re-run at solomain177bf9ac6a(2026-10-06): no reference toAskQuery,DescribeQuery,ReadStageorUngroupedRead(theConstructQueryhits are an unrelatedsparqlConstructQueryfield), none of the replaced error strings ("Aggregate input variable", "already exists in schema", "Duplicate aggregate output", "format_binding called without", "Projected variable not in child schema"), no ASK with OFFSET, LIMIT or VALUES, and no JSON-LDaskwith options. That's all from grep, fwiw — I haven't compiled this against solo. Dedicated query pools take the change only when their image tag moves.Lockstep for solo:
duplicate_pref_labelsshould switch to the expression form,"having": "(> (count ?concept) 1)", a one-line change influree-lambda-model/src/jsonld/queries/reports.rs:75. The query sends"having": [["filter", "(> (count ?concept) 1)"]], the form the old db doc example taught (that example is fixed here). db refuses it at parse ("filter operator must be a string") on base and on this head alike, so that report has presumably failed on every run. Solo's unit test only checks that thehavingkey exists. The fix doesn't depend on this PR, so it can land before or after the db pin bump.Nothing persisted changes: no commit, index, nameservice or graph-registry format. Stored queries and shapes run without migration, and behave per the table above.
Performance
The #1978 shapes get a lot cheaper. New bench
query_hot_grouped_projection(first commit, so its base-tree run is the baseline from the same bench code): N entities, N/200 groups, indexed file ledger; each scenario runs the query and renders the response (SPARQL JSON; JSON-LD for the JSON-LD scenario).key_count(control)expr_count(#1978)expr_mindedup_exprimplicit_const_countsubselect_exprjsonld_keyonly_exprThe "before" response is the exploded result (one row per solution). After the fix, a grouped SELECT expression costs what the key-only query costs, because it takes the same streaming
GroupAggregateOperatorplan (the expression runs once per group above it) instead ofGroupByOperator+AggregateOperator, which stored every input row and built a list per column. A dedup-only GROUP BY with an expression now gets WHERE-level early dedup too (plan pins init_query_explain), which is whydedup_exprlands below thekey_countcontrol.Quiet box. EC2 c7i.4xlarge, the fat-LTO bench profile, base and head in interleaved rounds; a change counts as a win or a loss only when the base and head ranges don't overlap and the median delta exceeds the bench's budget (5% at
small). These ran at the head that was reviewed,c67dc3516, and before it:The target path wins clearly.
expr_countwent 48.9 → 1.93 ms (25×) andjsonld_keyonly_expr27.8 → 2.00 ms (14×) in session 1, and atc67dc3516expr_countis 47.1 → 1.77 ms (27×). Thekey_countcontrol is stable.The count queries, settled. Session 1 found
count_novelty+9.4% andcount_cached+5.1%, both losses by the rule. There were two causes:c67dc3516as reviewed,51c153337after the rebase) collects them only for a stage that reads them;perf stat, the old merge rancount_noveltyat +0.09% instructions but +5.44% cycles.Session 2, at
c67dc3516:count_novelty+2.5%,count_cached+1.8% andcount_base+2.2%, all stable by the rule. The new merge is at instruction parity oncount_novelty(−0.05% instructions, +1.57% cycles), and bringscount_baseandcount_cachedinstructions back to base level (−0.5% and +0.4%).What remains: a 2–3% timing residual on the count queries, with the head slower in every round. It's layout, not logic: the instruction counts match base.
Everything else measured on the box is stable,
negation_count,bsbmq9 and the other overlay rows included.The local runs, the per-commit analysis, and both sessions' quiet-box tables are in a comment on this PR, moved out to keep this description under GitHub's size limit.
Hot paths touched, and what they cost now:
format_string, NDJSON rows, XML): the per-row scan of every cell forBinding::Groupedis gone.apply_solution_modifiers): one pass over HAVING, binds, order binds, sort keys and projection, scanning a few variables with no allocation, replacing two ad-hoc checks.RowPredicate): the per-row code is unchanged — the samefilter_batchcall when there is no EXISTS or metadata read, plus one extra.awaitper batch. Equivalent by inspection; not re-measured at the final head.8cfebf749); each such fast path returned a wrong answer for them. I don't have a timing, since a wrong answer gives no baseline to compare against. To keep the scan narrow, put the VALUES block inside the WHERE clause: it binds the variable before the scan, so the scan reads only the matching rows (and it is portable, per the compatibility note).bindable_sort_keys): borrows the ordering unchanged when every key is projected, grouped or an ORDER BY expression; otherwise one walk of the WHERE's produced variables, at plan time, once per sub-query operator (not per parent row).sh:sparqlfailures: nothing per row. Collecting one costs a mutex push, only when a constraint fails.Since the review, measured. All three use the same fixture: a release-build probe on an indexed 5,000-node ledger with 20,000
knowsedges.1a20ac6bc(the head after the first batch of review fixes, rebased on7ea640093) against56ea09988(the head after the second batch, on7ea640093), interleaved, three runs each of seven reps, medians of the per-run medians (µs per query, query plus formatting): count top-k SPARQL 27.6 → 27.4 and JSON-LD 21.6 → 21.4, directory count 14.8 → 14.7, 20,000 JSON-LD rows streamed 16,184 → 16,248 (+0.4%) and DOM 18,020 → 17,971, Cypher aggregatingWITH … WHERE c > 37,524 → 7,526 and aggregatingRETURN … ORDER BY16,868 → 16,629,ASK16.8 → 16.7. All within run-to-run spread on a shared machine (load 38–45 on 16 cores). The base binary can't format a JSON-LD per-group list of literals on this ledger (the bugc44d22acafixes), so that path has no baseline. The Cypher two-level form runs only for queries that used to fail or count wrong, and it is the plan an explicit secondWITHalready got. After the rebases the same patches ared6b2edfe6and4e30eaf0e. The fixes after those change no hot path (their one routing change, the two-level gate, takes only shapes that were a 400), except the two Cypher changes below.WITH'sWHEREafter its slice.7288a2f54againsta4aa18827on8632e33bc(the same patches after the rebase ared466ca9dfandb2ef72fba), which differ only in the Cypher lowering, interleaved, two runs of three rounds. AWITHwithout a slice, or with a slice and noWHERE, lowers to byte-identical IR (20 shapes dumped at both), and its timings are within noise (−4.7% to +5.7%, with overlapping ranges). AWITHwith a slice and aWHEREnow costs what the sameWITHcosts without itsWHERE(0–5%, 11% on one shape inside its spread). Against the old plan it costs more where theWHEREhad shrunk the input before the sort: 1.01–1.07× on three shapes; 1.2× forORDER BY p.age DESC LIMIT 100 WHERE p.age > 75; 1.6–1.9× for the two-level aggregate (ORDER BY p.age DESC LIMIT 50 WHERE c > 4) and forORDER BY p.name SKIP 100 WHERE p.age > 70; and 6× forORDER BY p.name LIMIT 10 WHERE p.area = 'A7', which the old plan answered from an index seek on 100 of the 5,000 nodes. That is the price of the openCypher answer: the old plan answered a different question, and the gap grows with the number of rows the slice ranks. Only Cypher, and only aWITHwith both.WITHinside aCALLbody.b6f867305against the same tree with07d0892ad's Cypher lowering and planner (before the fix), interleaved, three rounds.OPTIONAL MATCHwith a count, and a slicedRETURNran −3% to −12%. A CALL with no import ran −0.5%, and a non-CALL control −2.5%.pinned_vars, which only Cypher's CALL sets.Deviations from the design
Several of the items above are on this list too: trailing VALUES (an earlier commit here moved it after HAVING, per the spec, and
af22f8e49reverts that;docs/reference/compatibility.mdnow documents the deviation), HAVING on FILTER's evaluator, ASK / CONSTRUCT grouping, SHACLsh:select, and the stream's lane-dependentselect *refusal. The review fixes added a few more: HAVING reading SELECT aliases (an extension), EXISTS-body variables free over the group row, a CypherWITH'sWHEREafter its slice, aWITHinside aCALLbody, Cypher reads after the aggregation, nested SHACL conformance counting every result, andsh:sparqlfailures following severity and mode. Beyond those, the placement rule has three refinements over the design's literal version of the JSON-LD rule (JSON-LD keeps per-group lists for non-key variables), grouped-read errors get their names throughQueryError::name_variables,produced_vars_ofnow lives inir::patternwith all five callers on it, and the JSON-LD alias guard is standalone until #1867 lands — plus a handful of smaller structural ones. Each is written up in full in a comment on this PR, moved out to keep this description under GitHub's size limit.Follow-ups (not filed yet)
sh:selectskips SPARQL validation, so the SPARQL validator's checks (V4 and the rest) never run on a shape's query.SubqueryPattern, honored by every grouped-sub-query fast path and rewrite./query's JSON-LDselect *undergroupBydepends on the grouping lane: the keys and aggregates on the streaming lane, plus per-group lists on theGroupByOperatorlane (pre-existing). The stream follows it.QueryError::name_variables.WHEREgives the row once per value that passes,ORDER BYonce per value (WITH DISTINCTkeeps those copies, since the sort key is projected with them;RETURN DISTINCTremoves them), and an aggregate in a later clause counts each copy: at base,MATCH (p:P) WHERE p.age > 30 RETURN count(p)is 4 over three people when one of them has two ages above 30. A read inside an aggregate's argument (collect(p.age),avg(n.age)) is joined before grouping, so a property with several values repeats the group's rows for every aggregate of that clause (WITH p, count(f) AS c, collect(p.age) AS agescounts each friend once per age). openCypher has no multi-valued properties, so it never repeats a row; matching it needs a model-wide change (WHERE as a semi-join, a sort key of one value per row, aggregate arguments read per row), in the explicit-WITHandMATCHpaths at once, which is a lot more scope than this PR. The tests pin today's rows so that change flips them on purpose.+on strings returns null (RETURN p.name + 'x', at base too); openCypher concatenates strings (and lists) with+. SoORDER BY n + 'x'sorts on nulls.MATCH (p:P) WHERE p.name = 'Carol' CALL (p) { MATCH (p)-[:knows]->(f) RETURN count(f) AS c } RETURN p.name, cgives no row for a friendless Carol, where the standaloneMATCH … RETURN count(f)gives[[0]].collectdrops her[]the same way.lower_call_branchforRETURN, and in this PRcarry_importsforWITH), and an import with no body rows has no group.RETURN. This PR's CALL fix bringsWITHin a CALL body to the same model: at base every outer row got the whole body's wrong answer.cypher.md:OPTIONAL MATCH.count0,collect[],sum0.construct_and_describe_read_a_trailing_values's DESCRIBE checks (SPARQL, both lanes), asserting the described graph (see Sequencing).Sequencing
745736a04. Rebase history is folded below; every rebase was followed by a cleancargo check --workspace --all-targets --all-features.?__describeand a scan on it, and since fix(query): place BIND and FILTER after the triple that binds their variable #1995 the planner places that BIND after the scan, so on an indexed ledger the BIND drops every row (empty graph, both lanes, at base). This PR's DESCRIBE with a trailing VALUES inherits that. Applying fix(query): BIND and UNWIND do not wait for the patterns that bind their target #2013's planner change to this branch made single-target DESCRIBE, both DESCRIBE + VALUES forms and the equivalent SELECT and CONSTRUCT shapes correct in both lanes. So fix(query): BIND and UNWIND do not wait for the patterns that bind their target #2013 lands before this PR, and this PR's indexed-ledger DESCRIBE-with-trailing-VALUES test is added after it, asserting the correct graph (until then the trailing-VALUES DESCRIBE test runs on a memory ledger only).parse/lower.rslower_query/lower_subquery), plus Cypherstmt.rsnear theGrouping::assemblecalls. I kept the JSON-LD placement change as its own commit so that rebase stays mechanical — expect conflicts in both select-expression loops (Reject Cypher alias and JSON-LD bind collisions that silently returned zero rows #1867 addsselect_alias_binds+check_bind_targets; this PR replaceslower_select_expr_bindwithSelectExprPlacer). After that rebase, this PR's post-group alias guard should fold into Reject Cypher alias and JSON-LD bind collisions that silently returned zero rows #1867'scheck_bind_targets, and the comment there ("post-placed binds target aggregate outputs") needs updating, since it's no longer true now that key-only JSON-LD select expressions are evaluated once per group too.join::substitute_bindingaBinding::Groupedarm that returns this PR's internal error in all three positions, and drops this PR's two hunks atjoin.rs/optional.rs.next: at themain→nextmerge after this lands, RDF 1.2 triple terms: rdf:reifies links, triple terms as values, faster annotation queries #2017'sbinding_to_comparabletakes this PR's fail-closedGroupedarm (eval.rs), orgrouped_binding_fails_scalar_evaluationgoes red.Rebase history
The first commits rebased onto
2941d470cwith three one-line conflicts (#2004's multi-default-graph planning around this PR'sname_variablesmaps);cargo check --workspace --all-targets --all-featureswas clean right after, before any fix. The branch then rebased ontoac26b463e(#2003, the storage compare-and-swap fix), onto7ea640093(#2012, overflow integers keep xsd:integer;6a03e83a4, the transaction WHERE decodes through its own graph; docs and CI) and onto6b9d5b619(#2011, async WAL recovery; it shares two files with this PR, the CLI's query command and the benches doc, in disjoint hunks) with no conflicts, and the check was clean after each. #2012 changes how encoded overflow integers are decoded, so the indexed lanes suite now runs them through this PR's grouping: a per-group DATATYPE of the key, HAVING, MIN / MAX / SAMPLE and the implicit SAMPLE, a JSON-LD per-group list, the count top-k, and a group a decoded copy joins. Reverting either of #2012's DATATYPE or group-key changes turns it red.Next, the branch rebased onto
8632e33bc(#2018, #2020 and #2022). #2022 touched three files of this PR:view/stream_query.rs(itsrow_compactorbeside this PR'sensure_streamable; both kept),view/dataset_query.rs(its restructure, with this PR's twoname_variablesmaps applied to it), andit_iceberg_local_fs.rs, which ends as main's file, since this PR's canary moved to its own binary.git range-diffshows seven commits changed in context only and the rest patch-identical. The check was clean right after.Last, it rebased onto
745736a04(#2023, #2024 and #2025). The one conflict was inshacl_tests.rs, where #2024'sviolation_carries_resolved_resultssits beside this PR'ssh:sparqlhelpers; both are kept. Every other commit is patch-identical. #2024's SPARQL parameters walk each query form's AST, and they did not walk the trailing VALUES this PR adds after ASK, CONSTRUCT and DESCRIBE;bd3797dcccloses that seam. The 21 commits that touch a file #2023–#2025 also changed each passcargo check(the api lib in test mode andgrp_query_sparql, plus the sparql crate where they touch it), andcargo check --workspace --all-targets --all-featureswas clean after the rebase.Tests
New coverage is the SPARQL matrix (
it_query_sparql_grouped_projection.rs, each case asserting SPARQL-JSON rows and JSON-LD values on a non-injective fixture), its JSON-LD twin, UPDATE twins, stream refusals, explain pins, a Cypher-under-policy case, SHACL lib tests, unit tests, and a newit_grouped_projection_lanesbinary with MustFire / MustNotFire routing stamps — and, since the review, a test for each item above: grouped ASK / CONSTRUCT, EXISTS free over the group row, Cypher reads after the aggregation and the multi-value model pinned in row order,WITH … WHEREafter the slice on every lowering path,WITHinsideCALL (p), the trailing-VALUES cases at both levels and after every query form, the SHACL severity × mode cells, nested conformance through seven nestings on both write surfaces, and canary pairs on every binds-gated fast path. Every row of the non-vacuity table was committed first, then only the fix reverted, rebuilt, watched fail, restored, and the needle grepped.The full test list is in a comment on this PR. The non-vacuity logs, what was reverted and what went red, are in the first comment (the PR as first opened) and a follow-up comment (every fix since). Two things the follow-up log calls out: the mutations that remove only a fast path's binds gate stay green where the path also declines on its projection check (removing both turns it red; both runs are in the log), and the EXISTS seed mapping, which no shape reaches, is now a
debug_assert!(d6b2edfe6) instead of an untested branch. Release builds keep the spec reading (the list seeds as unbound).Gates
The gates ran at
31e238a7eon745736a04. The one commit after it,87576fdb1, changes onlydocs/query/cypher.md; at87576fdb1fmt is clean (the workspace andtestsuite-sparql) andmdbook buildsucceeds. At31e238a7e, every one exited 0:cargo fmt --all -- --check, the workspace andtestsuite-sparql;cargo clippy --workspace --all-targets --locked -- -D warnings, with--all-featuresand at default features, and clippy intestsuite-sparql: 0 warnings each;cargo nextest run --all-featuresonfluree-db-query,fluree-db-sparql,fluree-db-cypher,fluree-db-shaclandfluree-db-transact: 3,002 passed, 1 skipped;fluree-db-apilib plus 25 test binaries: the fivegrp_*groups this PR touches (grp_query,grp_query_sparql,grp_misc,grp_transact,grp_policy), six Cypher binaries, and 14 others on the paths it changes (explain, the indexed SPARQL suite, the SQL and Iceberg lanes, the lanes and fused-aggregate canaries, and the fast-path and count regression binaries): 3,436 passed, 13 skipped;testsuite-sparql(W3C): 36/36;testsuite-shaclis local only — no CI workflow runs it — and it doesn't compile on main today;d3df023d7adds its two missingValidateOptionsfields as its own commit. SHACL Core 81/98 and SHACL-SPARQL 17/22, with the same tests passing with and without the nested-conformance fix, and with and without thesh:sparql-only severity and list fixes; at the PR's first head it matched base test for test. The suite loads lists as indexed items and registers its shapes by targets or properties, so the list and severity fixes don't reach it, and none of its SHACL-SPARQL tests uses GROUP BY, HAVING or an aggregate. Theshacl_testspin those instead.Not run at this head: the workspace-wide nextest (CI's
testjob), and with it the otherfluree-db-apitest binaries and the server, CLI and python crate tests (all of which compile and pass clippy in the workspace runs above); the api doc tests; the wasm check and thewasm-smokebrowser suites; the SQL live bridge; thebench.ymlcompare against the committed baseline; and an openCypher reference engine for theWITH … WHEREchange (none was available, so that evidence is read, not executed). The workspace-wide run from when this PR was opened (13,765 run, 13,760 passed, at2a540f61e) is in the same comment as the perf detail, with its gate log.