Skip to content

fix: unresolvable graph references fail closed, and TriG directives apply in document order - #1997

Merged
aaj3f merged 7 commits into
mainfrom
fix/unresolved-graph-refs-fail-closed
Sep 30, 2026
Merged

aaj3f merged 7 commits into
mainfrom
fix/unresolved-graph-refs-fail-closed

Conversation

@aaj3f

@aaj3f aaj3f commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

This bundles three small, silent data-integrity fixes, each in its own commit with its own tests. They're probably worth a v4.2.3:

  1. FROM NAMED <ledger#graph> answers again on POST /query (4314ba599). This is a v4.2.2 regression: the documented named-graph addressing form now returns a 400.
  2. An update's USING/WITH of an unknown graph reads an empty graph (d058cdb7a). Today it reads the ledger's real default graph, so a typo in USING makes a DELETE wipe default-graph data. This dates to the v4 baseline.
  3. TriG phase 1 applies @prefix/@base in document order (2447f8067). Today a later redefinition silently rewrites the IRIs written before it. That hits plain Turtle upserts whenever a literal happens to contain the word "graph", as well as TriG insert and bulk .trig import. Also v4 baseline.

The first two share a shape: a graph reference that doesn't resolve silently became something else. The third is the same kind of "means something other than what was written" bug on the parse side. Claude did most of the tracing and probing here, fwiw, including before/after runs of each fix against the previous commit. No issue is filed for any of the three, so there's no closing keyword.

1. FROM NAMED <ledger#graph> on POST /query (v4.2.2 regression)

On the connection-scoped POST /v1/fluree/query endpoint, the named-graph addressing form we document in docs/query/datasets.md ("Mixed Patterns", :468-487) and docs/query/sparql.md (:541-543) fails in v4.2.2:

SELECT ?e ?t
FROM NAMED <b2x:main#urn:ex:doc:1>
WHERE { GRAPH <b2x:main#urn:ex:doc:1> { ?e <http://example.org/title> ?t } }
Query error: Internal error: Nameservice error: Invalid ID format: Invalid ledger id 'b2x:main#urn:ex:doc:1': branch cannot contain '#' ('#' starts a graph fragment)

v4.2.1 returns the row. These were run on three fluree server run binaries (the v4.2.1 release, a ba984c7f5 build, and this branch) against the same TriG data. Every GRAPH block has a constant predicate, which is what the SQL pushdown lane admits:

Shape Route v4.2.1 (file) ba984c7f5 this branch (memory and file)
FROM NAMED <L#urn:ex:doc:1> + GRAPH <L#urn:ex:doc:1> SPARQL /query 200, 1 row 400 (memory and file) 200, 1 row
FROM <L> FROM NAMED <L#txn-meta> + GRAPH <L#txn-meta> { ?c f:t ?t } (the documented Mixed Patterns shape) SPARQL /query 200, 1 row 400 (memory) 200, 1 row
FROM NAMED <urn:fluree:L#txn-meta> + GRAPH SPARQL /query 200, 1 row 400 (memory) 200, 1 row
FROM NAMED <L#http://example.org/vocab#products> + GRAPH SPARQL /query 200, 1 row 400 (memory and file) 200, 1 row
COUNT(*) over GRAPH <L#urn:ex:doc:1> SPARQL /query 200, n=1 400 (memory) 200, n=1
JSON-LD "fromNamed": ["L#urn:ex:doc:1"] + ["graph", …] JSON-LD /query 400 (already broken) 400 (memory) 200, 1 row
row 1's SPARQL /stream/query in-stream error (already broken) in-stream error (memory) 1 row
GRAPH ?g, or all-variable { ?s ?p ?o } blocks SPARQL /query 200 200 200
row 1's SPARQL SPARQL /query/{ledger} 200, 1 row 200, 1 row 200, 1 row

Blocks the lane doesn't admit (GRAPH ?g, all-variable patterns) never reach the probe, which is probably why this looks narrower than it is at first.

Root cause. It seems to be two commits meeting, and both shipped in v4.2.2.

  • The SQL pushdown lane wraps every constant GRAPH <iri> { … } block it admits in a SqlBlockOperator (fluree-db-query/src/execute/where_plan.rs:3097-3101); sql_lane::admits (fluree-db-query/src/r2rml/sql_lane/mod.rs:180-185) only checks the block's shape.
  • When the operator opens, resolve_block (sql_lane/mod.rs:265-301) asks the provider's capability probe about any IRI that names a dataset member, and a ledger's own named graph is a dataset member.
  • FlureeR2rmlProvider::pushdown_capabilities (fluree-db-api/src/graph_source/r2rml.rs:2561-2571) calls sql_source (:1367-1389). That looked the IRI up in the nameservice and turned any lookup error into a query error. delta_source (:1395-1417) has the same shape.

What changed underneath it:

JSON-LD from/fromNamed and /stream/query already used the real providers, which is why their rows were already failing on v4.2.1's file backend. cc @bplatz, since this sits right where #1899 and #1958 meet. I don't think either PR is wrong on its own; it's the combination.

The fix. sql_source and delta_source now share a small helper, dispatch_record, that classifies the IRI before looking anything up:

  • If the IRI doesn't parse with LedgerId::parse, the grammar every backend applies to a lookup id, it can't name a graph source. That's Ok(None) with no lookup.
  • Otherwise, the parsed canonical id is looked up and every error propagates, exactly as before.

I think the fail-closed part matters: a failed lookup of a well-formed id still fails the query, so a nameservice or storage outage can't quietly plan a real SQL or Delta source as a native graph.

We could have matched the error instead (NameServiceError::InvalidId → "not a source"). But that variant also comes back when a stored record's identity fails validation (fluree-db-nameservice/src/file.rs:509, storage_ns.rs:295) and when a mounted record is localized (mount.rs:84), so matching it would turn a corrupt record into a silent change of plan. has_r2rml_mapping (r2rml.rs:2303-2314) maps every error to false, so this deliberately doesn't copy it; more on that at the bottom.

The probe runs once per distinct constant GRAPH IRI per query and is memoized in the catalog session, so there's nothing per row.

Class audit. Claude went through every lookup_graph_source( / .lookup_any( / LedgerId::parse( / LedgerRef::parse( / split_ledger_id( in the query, api and server crates, plus every query-side caller of the provider traits. Lines are at ba984c7f5:

Site Reachable with ledger#graph or a plain graph IRI? Same defect? Action
graph_source/r2rml.rs:1367 sql_source (via pushdown_capabilities :2561, sql_lane/mod.rs:299) Yes, by running (SPARQL and JSON-LD /query, /stream/query, query_from(), graph(..).query().with_r2rml()) Yes Fixed
graph_source/r2rml.rs:1395 delta_source Not today (callers only run for sources whose mapping resolved) Identical pattern Fixed via the shared helper
r2rml.rs:2303 has_r2rml_mapping (callers graph.rs:294,759,802,849, runner.rs:896,902,914); graph.rs:79, fused_aggregate.rs:2322 compiled_mapping(..).ok() Yes No 400, but they fail open on any nameservice error (a real R2RML source would read as an empty native graph during an outage) Not changed: needs an R2rmlProvider trait change
view/fluree_ext.rs:736 resolve_graph_source_at A plain graph IRI in a connection-route FROM NAMED: yes (<urn:ex:doc:1> → 400, <http://…> → 404, on ba984c7f5 and here alike) Different: what a plain graph IRI means there is a semantics decision Dataset-addressing work
server looks_like_ledger_ref; CLI base_ledger_id / query_targets_foreign_source Yes Different mechanism (routing classifiers) Dataset-addressing work (#1972, #1982)
service.rs:126; pin_key, same_ledger, snapshot.rs:318; bm25/vector/provider resolutions; graph_query_builder.rs:198 Mostly No: tolerant fallbacks, or resolutions where an error is correct None

2. USING/WITH of an unknown graph reads an empty graph

In a SPARQL UPDATE, a single USING <g> or WITH <g> naming a graph the ledger doesn't have made the WHERE read the ledger's real default graph. The JSON-LD update equivalents (from, or a top-level graph with no from) behaved the same. SPARQL 1.1 Update §3.1.3, via Query §13.2, makes the WHERE's default graph the graph named; a graph that doesn't exist is empty, so nothing should bind.

Run over HTTP on a fresh ledger per row. The seed is a default graph ex:a ex:v "d1" . ex:b ex:v "d2" plus GRAPH <http://example.org/g1> { ex:c ex:v "g1" }. The columns are the triples left in (default, g1, newg):

Case Update Before After
typo DELETE { ?s ex:v ?o } USING <http://example.org/typo> WHERE { ?s ex:v ?o } (0, 1, 0): default graph deleted (2, 1, 0)
WITH new graph WITH <http://example.org/newg> INSERT { ?s ex:copy ?o } WHERE { ?s ex:v ?o } (2, 1, 2): default graph copied into newg (2, 1, 0)
JSON-LD {"from": "http://example.org/typo2", "where": {"@id": "?s", "ex:v": "?o"}, "delete": {"@id": "?s", "ex:v": "?o"}} (0, 1, 0) (2, 1, 0)
another ledger typo case with USING <other:main> (0, 1, 0): this ledger's default graph deleted (2, 1, 0)
control: own ledger typo case with USING <u3:main> (0, 1, 0) (0, 1, 0)
control: two unknown USING <…/t1> USING <…/t2> (2, 1, 0) (2, 1, 0)
control: registered DELETE { GRAPH <g1> { ?s ex:v ?o } } USING <g1> WHERE { ?s ex:v ?o } (2, 0, 0) (2, 0, 0)
intended change USING <u8:main> USING <g1> (2, 1, 0): the ledger id was dropped (0, 1, 0)

Mechanism. fluree-db-transact/src/stage.rs:2690-2707 collects the WHERE default-graph IRIs (USING, else WITH, else JSON-LD from/graph). For exactly one IRI, :2726-2731 did resolve_graph_id(iri).unwrap_or(0), turning an IRI the registry doesn't know into g_id 0, the default graph. For two or more (:2791-2799), unknown IRIs were already skipped, and where_default_is_empty (:2783-2785) already modeled an empty default graph for USING NAMED alone. So the single-IRI branch was the odd one out. The server's SPARQL Protocol using-graph-uri parameters are rewritten into USING clauses (fluree-db-server/src/routes/transact.rs:2419-2423), so they took the same path. The same code is at the v4 baseline (5985d0f01); that's read in code, not bisected.

The ledger's own address. The one wrinkle: USING <u3:main> only worked because of that fallback. The graph registry never maps a ledger id (GraphRegistry::new_for_ledger, fluree-db-core/src/graph_registry.rs:162-171, seeds only #txn-meta and #config), while D-3 (docs/audit/burn-down/named-graph-dataset.md:227-229) says the ledger alias names its default graph. So the fix maps it explicitly. An IRI names this ledger's default graph when it parses with LedgerRef::parse, has no @ pin and no # fragment, and its id equals this ledger's canonical id. That accepts u3, u3:main and urn:fluree:u3:main. The query side only matches the exact string today; the dataset-addressing work lines the two up on this same table. Everything else goes to the registry, and an IRI the registry doesn't know contributes nothing.

The fix (stage.rs, +35/−18, about half comments). A new resolve_where_default_graph resolves each WHERE default-graph IRI once. With no IRIs, nothing changes. Otherwise the resolved graphs form the default-graph union, so a lone unknown IRI gives an empty one: the same ActiveGraphs::Many([]) mechanism where_default_is_empty already relies on. It costs one LedgerRef::parse per IRI per update, nothing per row.

Class audit, transact side. :2726-2731 was the only site reading the wrong graph. The multi-USING branch, where_default_is_empty, the USING NAMED/fromNamed allowlist (:2851-2864), ambient named graphs (:2865-2884), sync, graph management (CLEAR/DROP/COPY/MOVE/ADD), per-graph upsert retraction, and the API config-source selectors all already fail closed. Writes (WITH template graphs) aren't reads and are unchanged.

One pre-existing thing worth knowing: a WITH <g> INSERT … WHERE … that binds nothing still commits and registers <g> as an empty named graph, on both builds. So the WITH case above goes from "newg holds 2 wrongly copied triples" to "newg registered empty". That's through the existing path (see #1943); nothing new.

3. TriG phase 1 applies directives in document order

TriG phase 1 (parse_trig_phase1 → extract_phase1, fluree-db-transact/src/parse/trig_meta.rs) rebuilt the default graph as Turtle with every @prefix/@base first and every triple after them, and it expanded every graph block, and the <#txn-meta> block, with the document's final prefix map. So a later redefinition silently rewrote the IRIs written before it.

Upsert runs phase 1 whenever might_contain_graph_block (trig_meta.rs:446-453) sees a { or the bytes graph, in any case, anywhere in the text, including inside a string literal. That's how plain Turtle ends up affected. TriG insert and bulk .trig import run it for any document with blocks.

With @prefix ex: <http://a.org/> . ex:x ex:p "mentions the graph word" . @prefix ex: <http://b.org/> . ex:y ex:p "second" ., plus a TriG twin with GRAPH <g1> { ex:x ex:p "1" . } before the redefinition and GRAPH <g2> { ex:y ex:p "2" . } after it:

Case Before After
fluree upsert --format turtle, stored ex:x / ex:p http://b.org/x, http://b.org/p http://a.org/x, http://a.org/p (ex:y → http://b.org/y)
fluree insert --format turtle (streaming parser, never phase 1) http://a.org/x unchanged
fluree upsert --format trig, block before the redefinition g1: http://b.org/x g1: http://a.org/x; g2: http://b.org/y
fluree insert --format trig (TriG insert fallback) g1: http://b.org/x g1: http://a.org/x; g2: http://b.org/y

@base had the same hoisting problem in the default graph (a later @base re-based earlier relative IRIs). Relative IRIs inside blocks were already right, because resolve_iri (:809-816) runs while parsing.

The fix (trig_meta.rs, about +40/−36 in code). GraphBlock gains a prefixes snapshot taken when the block is parsed; directives can't occur inside a block, so that's exactly the map in effect for it. Every later expansion (named-graph blocks and the txn-meta subject, predicate and objects, on both extract and extract_phase1) uses it. One helper, default_graph_turtle, rebuilds the default graph for both extraction paths with directives and triples merged by byte offset in document order. The rebuilt Turtle still goes through the full parser. As first pushed, this actually cost two prefix-map clones per block (a snapshot at parse time and then a clone of that snapshot at extraction), not the same count as before; the review follow-up below makes the blocks share one map instead, so it now comes in cheaper than main (see the table there). Plus one sort of the spans per document.

Everything that consumes phase 1 gets the corrected order: upsert and the builder's Turtle operations (tx_builder.rs:461-472), the TriG insert fallback (tx.rs:3977-3995), bulk .trig import (import.rs:477, :639-700, :755), and .nq import (which goes through the same path; N-Quads has no directives). Sync and GSP unwrap blocks in place and never rebuilt the Turtle, so only their txn-meta block's prefix map changes. CLI is_trig_body only classifies. The chunked plain-Turtle import tracks its own prefixes and refuses any directive after data (PrefixAfterData, splitter.rs:323), so it never had this bug.

Tests

Every new test was proven non-vacuous: commit, revert only that fix, rebuild, watch the new tests fail, restore, then grep for the fix to confirm it's back.

  • Fix 1. fluree-db-api/tests/it_sql_pushdown_lane.rs (required-features = ["sql", "native"], so it can't pass without the feature that compiles the bug):

    • a_ledgers_own_graphs_are_not_sources_to_the_lane: five shapes through query_from(), including the JSON-LD fromNamed twin. Each also asserts, via the sql_block_pushdown routing stamp, that the block reached the lane and was declined. Two controls: a well-formed non-source id is still looked up, and an unknown ledger still fails as not-found.
    • a_failed_source_lookup_still_fails_the_query: the fail-closed check, using a nameservice whose graph-source lookups fail.
    • fluree-db-server/tests/sparql_dataset_semantics.rs: connection_route_reads_a_named_graph_addressed_by_ledger_fragment over POST /v1/fluree/query (reverted: left: 400, right: 200).
  • Fix 2. fluree-db-api/tests/it_named_graphs.rs (grp_graphsource), each case on a fresh ledger asserting all three graph counts:

    • test_using_or_with_an_unknown_graph_reads_an_empty_default_graph (SPARQL, indexed ledgers);
    • test_jsonld_update_from_an_unknown_graph_reads_an_empty_default_graph (the JSON-LD twin).

    Reverted, exactly the regression cases fail and all six controls pass. A second mutation disabling only the ledger-address branch fails exactly the two own-ledger controls, which pins the D-3 mapping.

  • Fix 3.

    • Unit tests in trig_meta.rs (prefix, @base, txn-meta snapshot).
    • it_trig_insert.rs (grp_transact): a_redefined_prefix_or_base_applies_only_after_it, through both upsert and insert. It asserts every stored quad with full IRIs; the cases are the repro, blocks before and after, @base, and a no-redefinition control.
    • it_import.rs (grp_import): import_trig_applies_a_redefined_prefix_only_after_it.

    Reverted, all of them fail except the control and the streaming-insert case, which is expected since it never used phase 1.

Gates

Run after the last commit, with edited files touched before clippy. These cover all three fixes:

  • cargo fmt --all -- --check: clean.
  • cargo clippy --all --all-features --all-targets --locked -- -D warnings: exit 0.
  • cargo nextest run -p fluree-db-transact --all-features: 380/380.
  • cargo nextest run -p fluree-db-api --all-features --no-fail-fast: 4541/4541 (27 #[ignore]d), including every new test above.
  • cargo test -p fluree-db-server --test grp_query: 153/153.
  • cargo nextest run -p fluree-db-cli: 457/457.
  • W3C testsuite-sparql (cargo test --test w3c_sparql): 36/36 manifests, including sparql11_update_tests, with both-way registers.

Not run: workspace-wide nextest, the rest of the server tests and the server under --all-features (its swagger-ui feature downloads assets at build time), doc tests, and the live SQL bridge job. testsuite-sparql itself isn't touched.

Left for the dataset-addressing work

A few things these fixes deliberately leave alone, because in each case the right fix is a design change rather than something to wedge into a hotfix:

  • has_r2rml_mapping and the compiled_mapping(..).ok() sites fail open on nameservice errors; fixing them changes the R2rmlProvider trait.
  • The typed dataset-reference work should make "is this dataset member a graph source" a property of the parsed reference, so the lane never probes a native member at all.
  • FROM NAMED <L@t:1> + GRAPH <L@t:1> returns 0 rows with no error, on v4.2.1 and here: the member is keyed without its time pin.
  • The query side's D-3 match should accept the same spellings of the ledger's own address that updates now do.

Follow-up: #1975, #1972

Review follow-ups (since 2447f8067)

Four commits on top of what @bplatz approved, addressing his three inline comments:

  • 0a67988da: WITH and JSON-LD graph naming the ledger's own address now read and write its default graph;
  • 90f5dd9b4: TriG graph blocks share their prefix map instead of copying it;
  • 27fea1105: regression tests on a ledger that has a graph registered under its own address;
  • 71392a9f2: docs.

WITH / JSON-LD graph and the ledger's own address

The second fix above made an update's WHERE read this ledger's own address as its default graph when USING, WITH or JSON-LD from/graph names it (the within-ledger FROM <ledger> convention), but the write side of WITH didn't follow. So WITH <it/x:main> DELETE { ?s ex:v ?o } INSERT { ?s ex:w ?o } WHERE { ?s ex:v ?o } read the default graph but wrote the ex:w triples into a new named graph called it/x:main, which is easy to miss because a query's GRAPH ?g doesn't list that graph (the query side reserves the canonical ledger id for the default graph).

This is a deliberately partial fix of the class. WITH and a JSON-LD top-level graph now use the default graph for the ledger's own address on both the read and the write side: lowering records the update's template default graph (Txn::template_default_graph), marks the templates that took that default (TripleTemplate::graph_from_template_default), and staging (stage_with_graph_delta, the only place that knows the ledger) maps just those templates to the default graph and doesn't register the address as a named graph. A template that names its own graph is never touched, even when it names the same IRI as the WITH. "The address" means any spelling LedgerRef::parse accepts (it/x, it/x:main, urn:fluree:it/x:main) with no #fragment and no time pin, and one predicate, ir::names_ledger, now decides it for both the WHERE side and the template side.

Every other position resolves the address through the graph registry exactly as it did before this PR: GRAPH <iri> in a template or in the WHERE, GRAPH ?g, USING NAMED and JSON-LD fromNamed, JSON-LD @graph and ["graph", …], INSERT DATA/DELETE DATA quads, TriG blocks (staged lanes and bulk import), and the graph-management verbs (CLEAR/DROP/COPY/MOVE/ADD, CREATE GRAPH). Taking the mapping into those positions too turned up a real hazard (Claude caught this one while wiring the write side): on a ledger that already has a graph registered under its own address (a TriG block can create one, and sparql_single_db_graph_alias_wins_over_colliding_named_graph builds one on purpose), the WHERE's GRAPH ?g/GRAPH <address> would still have read that graph while the templates wrote the default graph, so DELETE { GRAPH ?g {…} } WHERE { GRAPH ?g {…} } would have deleted default-graph triples that happened to match the other graph. Making the address mean the default graph in every position really needs a migration story for graphs like that, so the rest of the class (explicit GRAPH <address> in templates and the WHERE's named positions, bulk import, one spelling predicate shared with the query side, and migrating graphs already registered under the address) continues on our dataset-reference branch, whose commits and PR body will say explicitly where they pick up from bplatz's review comment.

One behavior note for a ledger that does have a graph registered under its address: WITH <address> / JSON-LD graph no longer touches that graph (it reads and writes the default graph; before this PR it read and wrote that graph), while GRAPH ?g and GRAPH <address> still read and delete it, in the WHERE and in the templates, as before. Folding such a graph into the default graph is ADD GRAPH <address> TO DEFAULT and then DROP GRAPH <address>.

Also not changed here, but worth flagging for the dataset-reference work: the query side and names_ledger recognise different spellings of the address. The query side (ExecutionContext::single_db_user_graph_id/single_db_user_graph_iris, resolve_within_ledger_graph) only matches the exact canonical id, while names_ledger also accepts the urn:fluree: and bare-name spellings, and I think one shared predicate belongs with that work.

Tests:

  • test_using_or_with_an_unknown_graph_reads_an_empty_default_graph (0a67988da) gains WITH <LEDGER> DELETE {…} INSERT {…} WHERE {…} with the address spelled name:branch and urn:fluree:name:branch, plus a WITH <urn:fluree:LEDGER#config> control that writes the config graph; every case asserts that no named graph gets registered under the address.
  • test_jsonld_update_from_an_unknown_graph_reads_an_empty_default_graph (0a67988da) covers the JSON-LD top-level graph twin, plus a graph #config control.
  • test_updates_on_a_graph_registered_under_the_ledger_address (27fea1105) runs on a ledger whose fixture registers the graph by a staged TriG block (the same way sparql_single_db_graph_alias_wins_over_colliding_named_graph builds its colliding graph), in both spellings, with JSON-LD twins where JSON-LD has the form: DELETE { GRAPH ?g {…} } WHERE { GRAPH ?g {…} } empties that graph and leaves the default graph intact, the same with GRAPH <address> on both sides does the same, and WITH <address> DELETE {…} WHERE {…} empties the default graph and leaves that graph untouched.
  • With the write-side mapping disabled, every WITH/graph case fails, including the WITH <address> rows on the ledger with a registered graph (verified by running), while the GRAPH ?g/GRAPH <address> rows and the controls pass either way. sparql_single_db_graph_alias_wins_over_colliding_named_graph passes unchanged.

TriG blocks share their prefix map instead of copying it

The third fix's perf note (corrected above) undercounted: it was two prefix-map clones per graph block, a snapshot per block at parse time plus a clone of that snapshot at extraction, where main cloned the document's final map once per block at extraction. Now the parser keeps its live prefix map in an Arc<FxHashMap<String, String>>, a directive updates it with Arc::make_mut (so it only copies while some block still holds the previous map), each block takes an Arc::clone, and extraction moves each block's Arc into NamedGraphBlock/RawTrigMeta instead of cloning it. Blocks with no directive between them share one map, so a document whose directives are all at the top copies the map zero times. NamedGraphBlock::prefixes and RawTrigMeta::prefixes become Arc<FxHashMap<String, String>>; readers still take &block.prefixes.

Measured on 10,000 blocks of 2 prefixed-name triples each, with 20 @prefix directives at the top, over 16 graph labels: release builds that differ only in trig_meta.rs (90f5dd9b4 carries the same trig_meta.rs as the measured build), 5 interleaved rounds (30 runs per round for phase 1, 15 for the full upsert into a fresh memory ledger). Time is the median of the per-round medians with their range in brackets; peak heap and allocations come from a counting allocator and were identical in every round.

Build Phase 1 time Phase 1 peak Phase 1 allocations Upsert time Upsert peak Upsert allocations
main e4793617f 23.7 ms [23.2–24.0] 42.3 MiB 890,136 76.1 ms [74.0–87.5] 58.6 MiB 1,242,540
2447f8067 31.4 ms [31.1–32.4] 62.8 MiB 1,300,137 83.2 ms [82.4–84.9] 62.8 MiB 1,652,541
90f5dd9b4 18.4 ms [18.2–18.7] 15.9 MiB 470,138 64.9 ms [63.9–66.8] 38.2 MiB 822,542

The box was shared with other builds at the time (load average 21–33 on 16 cores), but the fix's ranges still don't overlap main's. The steps between builds are about 41 allocations per block, i.e. one 20-entry map, which lines up with one map copy per block on main, two on 2447f8067, and none with the fix.

Docs

docs/query/sparql.md: the WITH bullet now says that with no USING/USING NAMED it also sets the WHERE's default graph, and that a graph that doesn't exist reads as empty and never falls back to the ledger's default graph; the USING bullets say the same for their cases (an unknown USING graph contributes nothing; USING NAMED without USING leaves the WHERE default graph empty); and a short paragraph says the ledger's own address names the default graph in USING and WITH only, while GRAPH blocks, USING NAMED and data quads resolve it like any other graph IRI. docs/transactions/update-where-delete-insert.md has the JSON-LD equivalents (top-level graph and from; per-node @graph and ["graph", …] resolve the address like any other graph IRI).

Gates at 71392a9f2

cargo fmt --all -- --check clean after the last edit; clippy on fluree-db-transact + fluree-db-api clean with default features and with --all-features; transact 381/381; server grp_query 153/153; CLI 457/457; W3C 36/36. fluree-db-api with --all-features was 4547 passed, 0 failed, and 6 killed at nextest's 360s limit on a heavily loaded box (load average 40–55): the two it_fwd_pack_compaction tests, which are unrelated and ran in about 2s in CI at 2447f8067, and four indexing tests that passed when re-run alone (127–314s). So those six weren't run to completion locally; CI runs them.

…e probe

On POST /v1/fluree/query, `FROM NAMED <ledger#graph>` with
`GRAPH <ledger#graph>`, the form docs/query/datasets.md documents, returned
400 "Invalid ID format: Invalid ledger id '...': branch cannot contain '#'"
in v4.2.2. v4.2.1 returned the rows.

The SQL pushdown lane asks the provider's capability probe about every
constant GRAPH IRI that names a dataset member. `sql_source` looked each
IRI up in the nameservice and failed the query on any lookup error. Since
d8d776d every nameservice backend rejects an id that is not
`name[:branch]`, which a ledger's own graph never is, and since 323f5bf
POST /query SPARQL runs with the real providers instead of the no-op one,
so the probe runs there too. JSON-LD `fromNamed` and /stream/query already
used the real providers.

The dispatch probes now classify the IRI first: one that does not parse as
a graph-source id names no source and is declined without a lookup. A
failed lookup of a well-formed id still fails the query, so a nameservice
outage cannot send a real SQL or Delta source down the native plan.
`delta_source` shares the helper.
…pty default graph

A single `USING <g>` or `WITH <g>`, or a JSON-LD update `from`/`graph`,
naming a graph the ledger does not have resolved to g_id 0 through
`.unwrap_or(0)`, so the WHERE read the ledger's real default graph.
`DELETE { ?s ex:v ?o } USING <typo> WHERE { ?s ex:v ?o }` deleted it, and
`WITH <newg> INSERT { ?s ex:copy ?o } WHERE { ?s ex:v ?o }` copied it into
newg. Under SPARQL 1.1 Update §3.1.3 and Query §13.2 the WHERE's default
graph is then the graph named, which is empty, so nothing binds. Two or
more USING graphs already skipped unknown IRIs.

The ledger's own address (`name[:branch]`, optionally `urn:fluree:`
prefixed, with no time pin or fragment) still names its default graph, per
the within-ledger FROM convention (D-3). It had worked only through the
fallback. Combined with other USING graphs it now also contributes the
default graph; before, it was dropped.
Phase 1 rebuilt a TriG document's default graph with every @prefix/@base
directive first and every default-graph triple after them, and expanded
each graph block, the txn-meta block included, with the document's final
prefix map. A later redefinition therefore rewrote the IRIs before it. With
`@prefix ex: <http://a.org/> . ex:x ex:p "mentions the graph word" .
@Prefix ex: <http://b.org/> . ex:y ex:p "second" .`, upsert stored
http://b.org/x. Upsert takes phase 1 whenever the text contains `{` or the
bytes "graph", even inside a literal; TriG insert and bulk .trig import take
it for any document with graph blocks.

The default graph now keeps directives and triples in document order, and
each block keeps the prefix map in effect where it appears. Relative IRIs
inside blocks were already resolved against the base in effect while
parsing.
@aaj3f aaj3f added the bug Something isn't working as expected label Sep 30, 2026

@bplatz bplatz left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The three fixes look correct and well-tested. I left one inline comment on how the ledger-id mapping interacts with WITH. Please read it before merging, and either address it here or file a follow-up.

// this ledger's own address names its default graph (the within-ledger
// `FROM` convention, D-3), a registered IRI names that graph, and anything
// else names a graph that does not exist here, so `None`.
let resolve_where_default_graph = |iri: &str| -> Option<GraphId> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This maps the ledger's own address to the default graph for the WHERE read. Under WITH, though, the same IRI also sets the template graph, and that side (lower_sparql_update.rs:1429) registers it verbatim. So the two halves of one update now disagree:

PREFIX ex: <http://example.org/>
WITH <it/x:main>
DELETE { ?s ex:v ?o } INSERT { ?s ex:w ?o } WHERE { ?s ex:v ?o }

On the unknown_graph_seed() data:

  • The WHERE binds the 2 default-graph rows, through this mapping.
  • DELETE/INSERT target a new named graph registered as it/x:main (g_id 4).
  • The default graph is unchanged, the ex:w triples land in a graph whose IRI is the ledger id, and GRAPH ?g doesn't list it.

A JSON-LD top-level graph naming the ledger takes the same path (parse_update_template_default_graph feeds both the write graph and the WHERE default).

This isn't a regression: the write side did the same before, via the old unwrap_or(0) fallback. But this PR now deliberately makes the ledger address mean the default graph for reads only. Options, roughly in order of preference:

  • apply the same mapping to the WITH/graph template graph;
  • reject a write graph that names this ledger with a 400;
  • limit the mapping to USING/from.

At minimum, add a WITH <LEDGER> case to test_using_or_with_an_unknown_graph_reads_an_empty_default_graph (and a graph case to the JSON-LD twin), plus a follow-up issue if you defer the fix.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for flagging this one, @bplatz, you're right that the two halves disagreed. I went with your first option, but only partway, and I want to be upfront about why.

In 0a67988da, WITH and a JSON-LD top-level graph that name this ledger's own address (any spelling LedgerRef::parse accepts: it/x, it/x:main, urn:fluree:it/x:main, with no #fragment and no time pin) now use the default graph on both the read and the write side, so your example reads and writes g_id 0 and registers nothing. Lowering records the update's template default (Txn::template_default_graph) and marks the templates that took it (TripleTemplate::graph_from_template_default); staging maps only those templates, and a template that names its own graph is never touched, even when it names the same IRI as the WITH. names_ledger is now the one predicate both halves use.

Where I stopped short: an explicit GRAPH <address> in a template or in the WHERE still resolves through the graph registry exactly as before. Pushing the mapping into every position turned up a real hazard (Claude caught this one while wiring the write side) on a ledger that already has a graph registered under its own address. A TriG block can create one, and sparql_single_db_graph_alias_wins_over_colliding_named_graph builds one on purpose. The WHERE's GRAPH ?g/GRAPH <address> would still have read that graph while the templates wrote the default graph, so DELETE { GRAPH ?g {…} } WHERE { GRAPH ?g {…} } would have deleted default-graph triples that just happened to match the other graph, which felt like the wrong trade to make in a hotfix.

Making the address mean the default graph in every position really needs a migration story for graphs like that, so the rest of the class (explicit GRAPH <address> in templates and the WHERE's named positions, bulk import, one spelling predicate shared with the query side, and migrating graphs already registered under the address) is continuing on our dataset-reference branch, and its commits and PR body will call out explicitly where they pick up from this comment. That's also why I didn't file a separate issue for it.

Tests:

  • test_using_or_with_an_unknown_graph_reads_an_empty_default_graph gains WITH <LEDGER> DELETE {…} INSERT {…} WHERE {…} in both spellings, plus a WITH <…#config> control;
  • test_jsonld_update_from_an_unknown_graph_reads_an_empty_default_graph has the top-level graph twin and a #config control;
  • test_updates_on_a_graph_registered_under_the_ledger_address (27fea1105) runs on a ledger with a graph registered under its address: DELETE {GRAPH ?g{…}} WHERE {GRAPH ?g{…}} and the GRAPH <address> form both empty that graph and leave the default graph intact (as before), and WITH <address> DELETE {…} WHERE {…} empties the default graph and leaves that graph untouched, in both spellings, with JSON-LD twins.

With the write-side mapping disabled, the WITH/graph cases fail and the rest pass. If you'd rather we take (a) all the way in this PR instead, happy to talk it through; I think the migration question is really the crux of it.

iri: graph_iri,
triples,
reified,
prefixes: self.prefixes.clone(),

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This snapshot is copied again at extraction (block.prefixes.clone() at :1560, :1848, :1859), so each block now costs two prefix-map copies, not the same count as before. Both extraction functions take self, so the maps can be moved out instead. Sharing an Arc until the next directive would also work. This matters for TriG files with many small blocks.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch, and it was actually a bit worse than the body said: the snapshot was taken per block at parse time and then cloned again at extraction, so two copies per block where main had one (I've corrected that line in the body too).

Fixed in 90f5dd9b4 by sharing rather than copying: the parser's live prefix map is now an Arc<FxHashMap<String, String>>, a directive updates it through Arc::make_mut (so it only copies while some block still holds the previous map), each graph block takes an Arc::clone, and extract/extract_phase1 now iterate std::mem::take(&mut self.graph_blocks) and move each block's Arc into NamedGraphBlock/RawTrigMeta instead of cloning it. Blocks with no directive between them share one map, so a document with its directives up top copies the map zero times.

Since you mentioned TriG files with lots of small blocks, I measured exactly that: 10,000 blocks of 2 prefixed-name triples each, 20 @prefix directives at the top, 16 graph labels, release builds differing only in trig_meta.rs. The numbers are the median over 5 interleaved rounds of the per-round medians (30 runs per round for phase 1, 15 for the full upsert into a fresh memory ledger), with heap from a counting allocator:

Build Phase 1 time Phase 1 peak heap Phase 1 allocations Upsert time Upsert peak heap Upsert allocations
main e4793617f 23.7 ms 42.3 MiB 890k 76.1 ms 58.6 MiB 1.24M
2447f8067 31.4 ms 62.8 MiB 1.30M 83.2 ms 62.8 MiB 1.65M
90f5dd9b4 18.4 ms 15.9 MiB 470k 64.9 ms 38.2 MiB 0.82M

So it ends up below main on every measure, and the per-round ranges (18.2–18.7 ms and 63.9–66.8 ms) don't overlap main's (23.2–24.0 ms and 74.0–87.5 ms) even though the box was fairly loaded at the time. The allocation steps work out to about 41 per block, i.e. one 20-entry map, which lines up with main copying once per block, 2447f8067 twice, and the fix never. Thanks for pushing on this one.

w.using_default_graph_iris.is_empty() && !w.using_named_graph_iris.is_empty()
});

// With no `USING`/`WITH`/`from`, the WHERE reads the ledger's default graph.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The WITH bullet in docs/query/sparql.md:1201 only mentions templates. It should also say that, with no USING, WITH sets the WHERE's default graph, and that a graph that doesn't exist reads as empty.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 71392a9f2, thanks. The WITH bullet in docs/query/sparql.md now says that with no USING/USING NAMED it also sets the WHERE's default graph, and that a graph that doesn't exist reads as empty and never falls back to the ledger's default graph. While I was there I did the same for the USING and USING NAMED bullets (an unknown USING graph contributes nothing; USING NAMED without USING leaves the WHERE default graph empty), and added a short paragraph saying the ledger's own address names the default graph in USING and WITH only, while GRAPH blocks, USING NAMED and data quads resolve it like any other graph IRI. docs/transactions/update-where-delete-insert.md has the JSON-LD equivalents (top-level graph and from; per-node @graph and ["graph", …] resolve the address like any other graph IRI).

… write its default graph

d058cdb made an update's WHERE read this ledger's own address (`name`,
`name:branch`, `urn:fluree:…`; no fragment, no time pin) as its default
graph when USING, WITH or JSON-LD from/graph names it. The write half of
WITH and of a JSON-LD update's top-level `graph` did not follow, so
`WITH <ledger> DELETE {…} INSERT {…} WHERE {…}` read the default graph and
wrote a new named graph called by the ledger id.

Those two positions, the update's template default graph, now write the
ledger's default graph when they name its address, and the address is not
registered as a named graph. Lowering records the IRI
(`Txn::template_default_graph`) and marks the templates that took it
(`TripleTemplate::graph_from_template_default`); staging, which knows the
ledger, maps them. `ir::names_ledger` decides "names this ledger" for both
halves.

Every other graph position still resolves the address through the graph
registry, as before: `GRAPH <iri>` in a template or in the WHERE,
`GRAPH ?g`, USING NAMED, JSON-LD fromNamed, `@graph` and `["graph", …]`,
INSERT/DELETE DATA quads and TriG blocks. On a ledger with a graph already
registered under its address, those positions read and write that graph as
before, and `WITH <address>` no longer touches it.
…opying it

The phase-1 fix gave each graph block a snapshot of the prefix map and then
cloned that snapshot again at extraction, so a block cost two copies of the
map where it had cost one.

The parser now keeps its live prefix map in an `Arc`. Each block takes a
reference, and a directive copies the map (`Arc::make_mut`) only while a
block still holds it. Extraction moves each block's map out instead of
cloning it, so blocks with no directive between them share one map.
`NamedGraphBlock::prefixes` and `RawTrigMeta::prefixes` become
`Arc<FxHashMap<String, String>>`; readers take `&block.prefixes` as before.
A TriG block can still register a named graph under the ledger's own
address; `sparql_single_db_graph_alias_wins_over_colliding_named_graph`
builds its colliding graph that way. On such a ledger:

- `GRAPH ?g` and `GRAPH <address>` in the WHERE and the templates read and
  delete that graph and leave the default graph alone, as they did before
  this branch.
- `WITH <address>` and a JSON-LD update's top-level `graph` read and delete
  the default graph and leave that graph alone.

The cases cover both spellings of the address (`name:branch`,
`urn:fluree:name:branch`), with JSON-LD twins for the constant-graph and
default-graph rows. Each case first checks that the fixture really
registered the graph.
…is the default graph

docs/query/sparql.md:
- With no USING or USING NAMED, WITH also sets the WHERE default graph.
- A graph named by WITH or USING that does not exist reads as empty and
  never falls back to the ledger's default graph.
- With USING NAMED but no USING, the WHERE default graph is empty.
- The ledger's own address names its default graph in USING and WITH only.
  GRAPH blocks, USING NAMED and data quads resolve it like any other graph
  IRI.
- `#config` is not the address.

docs/transactions/update-where-delete-insert.md: the same for a JSON-LD
update's top-level `graph` and `from`. A per-node `@graph` or a
`["graph", …]` template resolves the address like any other graph IRI.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working as expected

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants