Repository navigation
feat(query): union default graph, per ledger and per query - #2004
Conversation
A query that does not choose its own default graph can read the union of the ledger's default graph and all of its named graphs (the W3C service feature sd:UnionDefaultGraph): - f:queryDefaults / f:unionDefaultGraph on f:LedgerConfig, ledger-scoped - # PRAGMA union-default-graph and JSON-LD opts.unionDefaultGraph decide per query, either way; a non-boolean value is rejected - a ledger's SPARQL service description claims sd:UnionDefaultGraph while the setting is on The union runs as an implicit dataset of the ledger's own graphs: default graph patterns read every graph as one set, while GRAPH resolution, graph sources, SERVICE self-reference and as-of time behave as with no dataset. FROM naming the ledger itself reads the union; FROM naming a graph reads that graph. The reserved #txn-meta and #config graphs are never members. Writes are unaffected: transaction WHERE reads the default graph alone, as do the reads of a Cypher write statement. Property and shortest paths now traverse a default graph made of several graphs of one ledger instead of refusing it; a default graph spanning ledgers is still refused. Also fixes wrong results for joins over a default graph of several graphs on an indexed ledger: per-graph scans emitted encoded bindings into a scope with no graph view to decode them, so a join ran its right side unbound and returned a cross product. Multi-graph members now decode as they scan, and a GRAPH scope decodes what it hands to such a scope. execute() now plans a multi-member dataset as a set.
aaj3f
left a comment
There was a problem hiding this comment.
This is great, @bplatz. Two headline notes: (1) please do look at #2007 (I'll move it out of draft now). It doesn't stand in contradiction to this, but it does tidy up the design/abstraction of the same logics that this works with. The obvious call is to merge this and then rebase #2007 in such a way that it accommodates the wins of this PR (2) I do think it's worth some scrutiny into how this PR's widening rules happen (view/union_graph.rs:93-121) as I think they over-widen in unexpected ways (details below) and that widening might have both policy-bypass and performance consequences
Running the union as an implicit DataSet of the ledger's own graphs, with explicit_dataset() keeping GRAPH, SERVICE and graph-source addressing exactly where they were, reuses the dataset machinery instead of growing a new operator, and the two bugs it flushed out along the way (the cross product on indexed multi-graph joins, and property paths refused over one ledger's graphs) were worth fixing on their own. The mutation checks I ran agree with the body: admitting the reserved graphs breaks the COUNT tests, and dropping the member decode brings the cross product back.
This PR got no CI (stacked base, ci.yml only runs into main), so I ran what CI would have: clippy -D warnings on every touched crate, fmt --check, the new tests, and the full W3C testsuite-sparql on a trial merge with origin/main. All green, details below.
The two that need to land first:
fluree-db-api/src/view/union_graph.rs:100— with the union on,FROM <L> FROM <L#g1>reads every graph of L (I getAlice, Bob, Carolwhere the docs promiseAlice, Bob), in SPARQL and JSON-LD. Fix path: widen a ledger's default member only when no other default member names a graph of that ledger, plus a test.fluree-db-query/src/property_path.rs:371—read_edgescopies every scanned flake into a freshVecon the single-graph path, doubling peak memory for the closure / zero-length-universe scans of queries that never use the union. Fix path: return the lone graph'sVecdirectly, and letpath_snapshot()skip theVecthatpath_graphs()allocates per call.
The rest:
fluree-db-api/src/view/query.rs:1392— the off path resolves ledger config a second time per query on ledgers with a named graph.fluree-db-query/src/dataset_operator.rs:683— a question about how broad the eager decode needs to be.- The
query_hot_property_pathbench header asks PRs that touchproperty_path.rsto gate on it; numbers in the body would settle point 2 either way.
Heads-up on overlap with #2007 (mine, draft) — nothing for you to change here. #2007 rewrites the same addressing branches this PR moves onto explicit_dataset() (single_db_user_graph_* in context.rs, the SERVICE lookup at service.rs:160, the graph-source as_of_t at graph.rs:75, the R2RML precompute at runner.rs:893), so every textual conflict between the two is in fluree-db-query. I merged them locally with #2007's typed versions plus explicit_dataset() at each of those sites, and all 17 it_union_default_graph tests and an_implicit_dataset_names_graphs_as_a_single_ledger_does pass. Taking #2007's side verbatim instead fails three of them (service_on_the_same_ledger_reads_the_union comes back with only Alice, and graph_patterns_address_graphs_as_before and graph_scope_rows_join_the_union come back empty), so the split is doing real work, and I'll carry it through when I rebase #2007 onto this. Two more I'll handle on my side: union members built in member() (union_graph.rs:31) will need #2007's member kind, and user_graphs's alias filter (:27) needs reconciling with how #2007 enumerates a legacy alias-named graph.
Adherence to repo commitments:
- Patterns/abstractions: ✔ extends
DataSet/DatasetOperatorand the config groups rather than adding a union operator; the pragma andopts.unionDefaultGraphtwins meet inQuery::union_default_graph. - Performance (speed first, memory second): ✖ CRITICAL:
read_edgesdoubles peak memory on single-graph property-path scans (item 2); 🟠 a second config resolve per query on the off path; union-on cost is documented and the single-graph shortcuts are correctly disabled. - Deployment targets:
⚠️ shared engine and API code reaches the embedded host (solo's Lambdas and standalone binary), the server and wasm32; no background work or process-lifetime state added, but item 2's memory doubling lands hardest on the Lambda host's fixed memory ceiling. - Testing: ✔ 17 API + 4 server tests, unit tests for scope semantics and the service description, a fast-path battery from novelty and from an index, and the mutations above go red;
⚠️ no test forFROM <L> FROM <g>under the union (item 1), and no CI run on this head. - Conventions: ✔ thorough commit body; docs for the setting, the pragma, the
optskey and the service description; clippy and fmt clean locally.
Verified locally on this head and on a trial merge with origin/main (61b836e9a, clean): clippy -D warnings on the seven touched crates (the vector feature can't build on my machine, so vector/operator.rs is unlinted), fmt --check, it_union_default_graph (17), server union_default_graph (4), it_multi_graph_property_path (15), it_query_dataset (75), #1997's new tests, and testsuite-sparql (all 36 suites, including the stale-register check).
Approving now so you can merge without waiting on another pass from me — just be sure the first two are in before you do.
| ) -> Result<DataSet<'a>> { | ||
| let mut widened = Vec::with_capacity(dataset.default.len()); | ||
| for view in &dataset.default { | ||
| widened.push( |
There was a problem hiding this comment.
🔴 Must address before merge: with the union on, a FROM that names the ledger and one of its graphs reads every graph of the ledger.
runtime_dataset widens each default member that is a ledger's default graph without looking at its siblings, so FROM <L> FROM <L#g1> (or FROM <L> FROM <g1> in a ledger-scoped query, or JSON-LD "from": ["L", {"@id": "L", "graph": "g1"}]) gets L's default graph plus all of L's named graphs, g2 included. The new docs say the union applies when the query "does not choose a default graph of its own, which includes naming just the ledger itself", and that FROM <graph> reads just that graph; this query chose {default, g1}.
On the seed fixture from it_union_default_graph.rs with set_union(true), SELECT ?name FROM <L> FROM <g1> WHERE { ?s ex:name ?name } returns Alice, Bob, Carol; with the union off it returns Alice, Bob. The same happens with # PRAGMA union-default-graph: true on an unconfigured ledger, with the JSON-LD form above, and (once #2005 lands) with FROM <urn:default> FROM <g1>. Nothing errors, the extra rows are just there.
I think the rule wants to be: widen a ledger's default member only when no other default member names a graph of that same ledger, so multi-ledger FROM <L1> FROM <L2> still widens each one. A case next to from_the_ledger_reads_the_union_and_a_named_graph_narrows would pin it.
| for graph in graphs { | ||
| let flakes = | ||
| Self::read_graph_edges(ctx, graph, index, range_match(graph.snapshot)).await?; | ||
| out.extend(flakes); |
There was a problem hiding this comment.
🔴 Must address before merge: read_edges copies every scanned flake into a second Vec on the single-graph path, so property-path queries that never touch the union pay it too.
Before this PR each reader returned filter_edges' Vec as-is. Now read_edges starts an empty out and extends each graph's Vec into it, which for the ordinary one-graph case is a full memcpy of the scan with both buffers alive at once. The scans that go through here are the big ones: zero_length_universe (an unbounded SPOT scan for every * / ? closure), the wildcard closure (unbounded PSOT), and the per-predicate PSOT scans in composite_*. compute_closure never takes the ID lane, so ?s ex:p* ?o hits this on an indexed ledger too. A Flake is 176 bytes, so a full scan of a million-flake graph holds roughly another 176 MB at peak.
The fix is small — hand back the lone graph's Vec directly:
if let [graph] = graphs {
return Self::read_graph_edges(ctx, graph, index, range_match(graph.snapshot)).await;
}(or seed out from the first graph's Vec rather than an empty one).
Related and smaller: path_graphs() now allocates a Vec and bumps the policy Arc on every call, where require_single_graph() handed back borrowed refs, and it's called per BFS node in the wildcard forward_step / backward_step and in ShortestPathOperator::neighbors, and — through path_snapshot() — once per input row in process_correlated_row (:1461). path_snapshot() only needs the first graph's snapshot, so it can skip the Vec entirely. The query_hot_property_path bench (its header asks PRs touching this file to gate on it) exercises exactly these paths in scenarios 1, 3 and 4, so it'd be good to see its before/after in the PR body.
There was a problem hiding this comment.
Addressed in 949f1a2
Left the per-node path_graphs() in the wildcard BFS steps as is: a one-element Vec beside a range scan per node.
| // Settle the union default graph here, beside the other config | ||
| // defaults: execution reads only the settled switch. | ||
| executable.query.union_default_graph = Some( | ||
| self.reads_union_default_graph(db, parsed.union_default_graph) |
There was a problem hiding this comment.
🟠 Should address: with the union off, a ledger that has any named graph now resolves its config twice per query.
apply_reasoning_to_executable resolves the ledger config onto a clone of db (complete_config_defaults → resolve_and_attach_config) and drops it, and reads_union_default_graph then calls resolve_ledger_config_cached again on the original view. Views arrive without config on the ledger-scoped server routes (the comment in resolve_and_attach_config says that's every request), so every query on a ledger with a named graph pays a second resolve. At head that’s a cache hit (a loaded-handle lookup and a cache read); at a time-travel t the cache doesn't apply, so it's a second config-graph read per query.
It's per query, not per row, so I don't think it's a big number, but it's on the default path for a feature most queries don't use. Threading the config-completed view out of apply_reasoning_to_executable (or settling the union switch inside it) would make it free. Minor and non-blocking — but if you agree it's right, I'd rather see it folded in now than lost in the backlog.
| None | ||
| }; | ||
|
|
||
| self.decode_members = !multi_ledger && graphs.len() >= 2; |
There was a problem hiding this comment.
❓ Question: does every member of a multi-graph default graph need to decode everything?
This is more of a question than a suggestion. The cross-product fix is real — I set decode_members = false and joins_over_several_graphs_of_an_indexed_ledger came back Alice, Alice, Bob, Bob, Carol, Carol. But it turns on eager materialization for every member of every ≥2-graph one-ledger default graph, so an explicit FROM <L> FROM <L#g> that only counts or groups (no join) now decodes every row where it used to stay late-materialized. The new ScopeExit comment in graph.rs says only NUM_BIG handles are per-graph and the other encoded kinds decode against store-global dictionaries; if that holds here, a decode view for the multi-graph scope (any member's) would let these members keep late materialization and eagerly decode only NUM_BIG. I haven't measured the difference, so I may be overweighting it — happy to talk it through.
There was a problem hiding this comment.
Good question, and the premise holds: only NUM_BIG handles are per-graph. Subject, string and predicate IDs decode against the ledger's dictionaries, so in principle the members could stay late-materialized and eagerly decode NUM_BIG alone.
The obstacle is that the multi-graph scope has no graph view, and that's deliberate. has_binary_store() is false for a ≥2-member default graph because it also gates the single-graph fast paths keyed on one binary_g_id, and graph_view() is gated on it. Everything that decodes mid-pipeline goes through ctx.graph_view(): about 38 call sites across join, hash join, OPTIONAL, sort, aggregates and VALUES. Join substitution is where the cross product came from: substitute_pattern_with_store leaves an EncodedSid as a variable when the view is None. Keeping late materialization means passing a separate decode-only view to each of those sites, without turning the fast paths back on.
The cost only applies to default graphs made of several graphs of one ledger. That's the opt-in union, or an explicit multi-FROM on one ledger, which returned a cross product for joins on an indexed ledger before this PR. So I'd like to measure before building it: #2015 tracks comparing eager and late decoding on an indexed multi-graph default graph, and adding the decode-only view if the gap justifies it.
…a path-scan copy - runtime_dataset widened a ledger's default-graph member even when the same FROM list also named one of that ledger's graphs, so with the union on `FROM <L> FROM <g1>` read every graph of L. A ledger whose graphs the default graph names is no longer widened; other ledgers in the same FROM still read their own union. Pinned by from_the_ledger_and_one_of_its_graphs_reads_just_those (SPARQL and JSON-LD, single- and cross-ledger). - read_edges copied every scanned flake into a fresh Vec, doubling peak memory for single-graph property-path scans (zero-length universe, wildcard closure, composite PSOT). It now grows the first graph's Vec, so the single-graph read returns its scan as is. path_snapshot skips building the path graphs when there is no dataset. - build_executable_for_view resolved the ledger config a second time to settle the union switch. apply_reasoning_to_executable now returns the config-completed view and the switch is read off it.
…owering The pragma guard from #2002 requires every pragma to be listed. union-default-graph lowers into Query::union_default_graph, where its opts.unionDefaultGraph twin also lands, so no header, governance, tracking or alias-opts translator carries it.
Stacked on #2002 (
# PRAGMArequest options): the per-query switch here is a pragma. This PR's base isfeat/sparql-pragma-opts; it retargets tomainonce #2002 merges.CI (
ci.yml) runs only on PRs intomain, so this PR shows no checks until it is retargeted. Locally, the workspace suite (cargo nextest run --workspace --all-features), clippy with-D warnings(workspace and wasm32) andfmt --checkpass. The exceptions are three tests that also fail on the base commit locally and have nothing to do with this change.What this adds
This adds an option to read the default graph as the union of the ledger's default graph and all of its named graphs. It is the W3C service feature
sd:UnionDefaultGraph, and it helps anyone bringing data and queries that assume that model.f:queryDefaults→f:unionDefaultGraph trueonf:LedgerConfigin the#configgraph.f:GraphConfigand not subject to override control, likef:servingDefaults.# PRAGMA union-default-graph: true|falsein SPARQL, or"opts": {"unionDefaultGraph": true|false}in JSON-LD.GET /query/{ledger}claimssd:feature sd:UnionDefaultGraphwhile the setting is on.Semantics
FROM, or aFROMnaming only the ledger itself. That includes a time pin, as inFROM <ledger@t:5>and/query/ledger@t:5.#txn-metaand#configare never members.GRAPH <iri>andGRAPH ?gstill address the named graphs.GRAPH <ledger-alias>still names the default graph alone.FROM <graph>reads just that graph, andFROM NAMEDis taken literally.GRAPHpattern reading that graph is today. No access is added.WHEREare untouched; they run through the transact path, not the view path.MERGE/DELETEprobes must see what the write stages against, the same reason those probes already switch reasoning off.FROM … TO …) read the default graph alone.How it works
The union executes as a dataset of the ledger's own graphs, the one
FROM <L> FROM <g1> …would build, so it reuses the dataset scan and dedup path.The dataset is marked implicit (
DataSet::implicit).ExecutionContext::explicit_dataset()returnsNonefor it, and the branches that concern how a query addressed its graphs now consult that instead ofctx.dataset. Those branches cover:GRAPHresolution andsingle_db_user_graph_*as_of_tidiomBranches about which graphs a pattern scans keep
ctx.dataset.Where the switch is decided:
build_executable_for_view, beside the other config defaults.Fluree::runtime_dataset, which widens a default member that is a ledger's default graph.A ledger with no named graphs never leaves the single-graph path.
Existing behaviour this changes
Joins over a default graph of several graphs on an indexed ledger returned a cross product.
SELECT ?name FROM <L> FROM <g1> WHERE { ?s ex:knows ?o . ?o ex:name ?name }paired everyknowsrow with every name.DatasetOperatormembers of a ≥2-graph single-ledger scope decode as they scan (inopenandnext_batch, which builds its own context).GraphOperatoralso decodes every encoded binding leaving aGRAPHscope when the surrounding scope has no graph view; before, it decoded only NUM_BIG handles.Follow-up: Late decoding for a default graph of several graphs of one ledger #2015, to keep these members late-materialized and decode only NUM_BIG handles eagerly.
Property paths and shortest paths over a multi-graph default graph of one ledger now run instead of erroring.
ExecutionContext::path_graphs()replacesrequire_single_graph. A default graph spanning ledgers is still refused, since their SIDs aren't comparable, with a message saying so;it_multi_graph_property_path::q2now asserts that wording.Partially addresses Cross-snapshot property-path BFS: lift the multi-graph union-path guard (follow-up to #1425) #1405: the guard is lifted for several graphs of one ledger. Traversal across ledgers, which is what that issue asks for, remains open.
fluree_db_query::execute()now plans a multi-member dataset as a set. Before, a multi-member dataset through that entry point (the BM25/vector connection path) tripped the set-dedup guard with an internal error.sparql_service_descriptiontakes aunion_default_graphflag.Tests
fluree-db-api/tests/it_union_default_graph.rs(ingrp_query_sparql) covers:GRAPHaddressing, reserved graphs kept out, and set semantics (COUNT(*)over the union)FROMFROM <ledger>versusFROM <graph>, the SPARQL and JSON-LD connection routes, and config read as oftMERGEin all three probe forms (which don't)MAX, top-k, joins, paths), run on novelty and on a built indexfluree-db-server/tests/union_default_graph.rscovers the plain, tracked, streaming and time-pinned ledger routes, the pragma andoptsover HTTP, connectionFROM <ledger>, and the service description.opts.unionDefaultGraphexplicit_dataset/path_graphsscope semanticsGRAPHscope-exit decoding → theGRAPH-into-union joinexplicit_dataset→ the SERVICE testMERGEtestexecute()planning → the BM25 testNot covered
explain/explain_sparqlplan against the single graph and don't show the union. They also don't modelFROMdatasets today.GRAPH <graph-source>pattern inside a union query takes the same naming path as without the union, but no test here exercises an R2RML/Iceberg source under the union.