Skip to content

fix(transact): retractions name stored facts; RDF text is parsed once - #2008

Open
aaj3f wants to merge 35 commits into
mainfrom
fix/write-fidelity-upsert-rdf-text
Open

aaj3f wants to merge 35 commits into
mainfrom
fix/write-fidelity-upsert-rdf-text

Conversation

@aaj3f

@aaj3f aaj3f commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

This is more than #1976 needs on its own, deliberately. Our write bugs have had two recurring shapes: deletes built from the request instead of from what's stored, which silently retracted nothing for some lanes and datatypes (#1976), and Turtle or TriG text parsed more than once, by paths that disagreed (#1977, #1511, #1930). So rather than patch each lane, a retraction can now only name a fact read from storage, through one resolver, and every transaction lane parses Turtle and TriG once, with one parser, so there's one place to reason about and tune. Fwiw, it's one of five PRs taking this approach, with #2006, #2007, #2009 and #2010.

Fixes #1976
Fixes #1977
Fixes #1511
Partially addresses #1930. What remains: bulk import (fluree create --from x.trig, the server's source import) still reads block contents with the phase-1 parser. That moves to the same driver in a follow-up (listed below).
Partially addresses #1849. What remains: the TriG insert fallback now parses once into templates, but it is not yet a flake-level path like plain Turtle insert.
Partially addresses #1979, item 3 only, on the transaction lanes. Blank nodes inside GRAPH blocks are now accepted rather than refused with a token-level error, and a nested block is refused as nested graph block: …. Bulk import keeps the old messages until it moves to the same driver. Items 1–2 are not in this PR.

This PR is two write-fidelity fixes, in two commit groups. A1 makes every retraction name a fact that is actually stored — read from storage through one resolver, instead of rebuilt from query bindings — and A2 makes every transaction lane that takes Turtle or TriG parse it once, with the conformant Turtle parser, instead of round-tripping it through JSON-LD. Both change behavior (the table below is worth a careful read), and there are three calls in here I'd like a reviewer to explicitly ratify before this merges.

Three calls to ratify

1. Modify policy on deletes and on restated values — a product-behavior decision. Modify policy is evaluated for every delete intent, whether or not it matches a stored fact, so a delete the identity may not perform is refused with the same answer regardless of what is stored (unchanged from before this PR). The alternative would be resolving deletes only against facts the identity can view, which would stop a modify-but-not-view identity from deleting values it can't see. As maintainer I agree with this posture, but it asserts a particular perspective on product behavior.

The same rule now covers a write that restates a stored value as it is: an upsert of an unchanged value, a DELETE and INSERT of the same value, or a graph sync that keeps a value. Such a write stages nothing, because the retraction of the stored value and the assertion of the same value cancel.

Behavior change from main. Under a non-root policy, when the identity cannot view the restated value, both sides are now checked against modify policy, and the write is refused if the identity may not modify the value.

  • On main, a write restating a stored value was not checked against modify policy: plain and list restatements committed as no-ops.
  • Restating a value the identity can view is unchanged: it commits nothing, with no modify check. A language-tagged restatement now behaves the same way. main refused it, because the retraction it built carried no language tag and named nothing stored.
  • Root, and writes that restate nothing, are unaffected.
  • Tests cover SPARQL and JSON-LD updates, JSON-LD and Turtle upserts (plain, language-tagged and list values), and graph sync, over novelty and over an index.

Cost. Only under a non-root policy, and only for restated values. Each restated value gets one view check. A value the identity cannot view also goes through the modify check, both sides.

Measured with dev-fast builds on 2,000 subjects of five values each. The base is this branch just before the check (11b5872a6). Each table is 6 interleaved rounds of 15 runs; median of the round medians, range in brackets.

Upserts that restate all 10,000 values:

policy before the check with the check ratio
identity can view the values (view and modify allowed except on one unrelated property) 23.65 ms [22.59–24.40] 25.80 ms [24.98–25.99] 1.09
identity cannot view the values but may modify them 23.70 ms [22.78–24.01] 27.03 ms [26.12–27.60] 1.14
root 23.61 ms [22.55–24.30] 23.75 ms [22.58–23.90] 1.01

Writes that restate nothing, measured in an earlier run with the first policy:

write before the check with the check ratio
upsert changing all 10,000 values, under the policy 59.65 ms [56.03–62.36] 59.69 ms [55.87–61.92] 1.00
the same, as root 58.89 ms [55.14–61.65] 59.09 ms [56.27–61.09] 1.00
insert of 2,000 new subjects, under the policy 29.94 ms [27.80–30.53] 29.46 ms [28.51–31.66] 0.98
the same, as root 29.34 ms [27.40–30.17] 28.73 ms [27.90–30.56] 0.98

That is about 0.2 µs per restated value the identity can view (the view check), and about 0.33 µs per value it cannot view (the view check and the modify check). A policy whose view or modify rules run queries or match on classes costs more per value, since each restated value runs them. Writes that restate nothing pay nothing measurable, and neither does root.

Docs. The policy docs now say how restatements are checked. They also say that granting modify without view does not keep a value confidential from that identity, since it can tell whether a given value is stored by writing it, and that such policies should be configured with care.

Please ratify or push back on both parts.

2. One of the tests from #1987 (merged) changes: DELETE DATA of an ill-typed literal that isn't stored now commits nothing instead of erroring. it_typed_literal_index::sparql_update_rejects_ill_typed_literal asserted that DELETE DATA refuses an ill-typed xsd:date the way INSERT DATA does. A DELETE now names such a literal as stored (a Turtle write keeps it as its lexical form with the declared datatype), so the test is now sparql_insert_data_rejects_an_ill_typed_literal: INSERT DATA still refuses the literal and names the datatype, and DELETE DATA of one that is not stored commits nothing. Please confirm this is the intended reading of #1987.

3. Upsert replaces every value of a (graph, subject, property), across languages. That's the documented upsert semantics, and it's what the branch implements: the replacement unit is every value of (g,s,p), all languages and list positions, so a partial-language upsert (only prefLabel@fr) now removes the other languages' values, where it used to accumulate them. The alternative is per-(s, p, lang) replacement, which would leave the other languages alone. The replacement unit lives in one function, upsert_replacement_slots. Please ratify or push back.

What changed

A1: a retraction names a stored fact. Upsert, DELETE, graph sync, CLEAR/COPY/MOVE, the annotation cascade, merge/rebase conflict retractions and revert now all read the facts they retract from storage, through one resolver (CurrentFacts), and retract those facts as stored — graph, datatype, language tag and list position included. A retraction can only be built from a stored fact (StoredFact::retract → Retraction), so an intent that names nothing stored stages nothing. The old behaviors this replaces:

  • Upsert used to rebuild retractions from query bindings, which carry no language tag or list position, so "a"@en and list entries were never removed and an identical refresh committed phantom retractions (upsert doesn't retract language-tagged (or @list) values: the retraction flakes are built with m: None #1976).
  • DELETE DATA of an absent triple committed a phantom retraction.
  • DELETE { x } INSERT { x } with x absent lost the insert.
  • Index-resident named-graph values survived take-source merges and rebases.
  • The annotation cascade retracted body triples outside their graph.

A fact read from the persisted index is also named by its index key — which keeps no XSD subtype for an integer beyond 64 bits, and keeps a dateTime or time to the microsecond — so a DELETE naming such a value still deletes it. A DELETE row whose value an OPTIONAL or a single UNION branch bound is the decode of a stored fact, as a plain WHERE row is, and retracts without a read; that covers Cypher SET and DETACH DELETE. Modify policy judges every delete a transaction asks for, whether or not it names a stored fact, and judges a stored value a write restates when the identity cannot view it (call 1 above).

A2: RDF text is parsed once. Every transaction lane that takes Turtle or TriG — upsert, the TriG insert fallback, graph sync and the Graph Store routes, and fluree sync — now reads it once with the conformant Turtle parser, into one TemplateSink. It used to go through a JSON-LD round trip plus a second, smaller TriG block parser, and that round trip:

A token-only locator finds graph blocks without interpreting a term, and one parser then reads the segments from the locator's tokens in document order — so the document is lexed once, prefix and base declarations carry forward, and one sink per document keeps blank-node labels document-scoped. Nothing converts RDF text to JSON-LD on a transaction lane any more: parse_to_json is documented as lossy and disallowed in the fluree-db-transact, fluree-db-api and fluree-db-cli clippy.toml.

Behavior changes

Area Before After
DELETE DATA / JSON-LD delete / Cypher REMOVE of an absent fact commits a phantom retraction at t+1 stages nothing; no commit, t unchanged, and a mixed transaction drops only the phantom
DELETE {x} INSERT {x} with x absent insert lost x inserted
DELETE term datatype compatible datatypes matched exact datatype (a delete naming xsd:int does not remove a stored xsd:integer). Once a fact is indexed, the match is against what the index stores: an integer beyond 64 bits keeps no XSD subtype and a dateTime/time keeps microseconds, so such a fact is named by the value under any integer datatype, or to the microsecond, as on main. A DELETE naming a big integer under a different subtype is therefore a no-op while the fact is in novelty, and deletes it once indexed: the same outcomes as on main
Upsert replacement unit values of (s,p) as the binding rebuilt them: tags and list positions missed, stale values kept every value of (g,s,p): all languages and list positions. A partial-language upsert (only prefLabel@fr) now removes the other languages' values, where it used to accumulate them. (Call 3 above.)
Identical re-upsert of lang/list/blank-node data commits phantom retractions and re-assertions no commit, including payload-scoped blank-node subjects. Not for a value stored with less detail than it was written (upgrade notes)
Under a non-root policy, a write that restates a stored value the identity cannot view (an upsert of an unchanged value, a DELETE and INSERT of one value, a graph sync that keeps it) plain and list values: committed nothing, with no modify check checked against modify policy on both sides: refused unless the identity may modify the value. Restating a value the identity can view still commits nothing, with no check
Revert of a commit holding a legacy upsert phantom could assert a tagless rdf:langString skips retractions no stored fact could match, with a warning count
Turtle/TriG upsert, TriG insert, sync, Graph Store JSON-LD round trip (losses above) parsed once: stores what a Turtle insert stores
Ill-typed literal ("abc"^^xsd:integer) on upsert/sync error kept with its declared datatype, as insert always did
DELETE naming a stored ill-typed literal refused when the DELETE is lowered, so no DELETE could remove one deleted: a DELETE template keeps the lexical form as stored. An INSERT still refuses one. The trade-off: a mistyped literal in DELETE DATA now commits nothing instead of erroring, since it names nothing stored. (Call 2 above.)
Relative IRI reference in a GRAPH block (its label or its contents), no @base kept as written on insert/upsert; sync and Graph Store refused one in block contents kept as written on every lane. Relative references in GRAPH blocks keep today's verbatim behavior: documented cross-ledger config relies on it (f:ledger <org/governance:main>). Sync and Graph Store block bodies now match insert/upsert. The default graph is unchanged (refused without a base)
Anonymous default-graph block { … } error accepted (TriG default graph)
TriG structural errors (nested, unclosed, directive in a block, blank-node label) Parse error: … Turtle parse error: … (400, TURTLE_PARSE) at the document's byte offset. Undefined prefix in a block: Turtle parse error instead of a transaction error
Upsert blank-node identity for RDF text hash of the JSON-LD conversion order-insensitive hash of the parsed statements (below)
Raw transaction stored for a Turtle upsert the JSON-LD conversion the text, as for a Turtle insert
Fluree::upsert_turtle* plain Turtle only also TriG, as the builders did
fluree sync Turtle converted to JSON-LD client-side RDF text sent as written; a 415 from a pre-RDF server gets the old conversion as a fallback (Turtle only)

The upgrade notes (for the release notes) and the API changes are folded under this section.

Fuel. A Turtle upsert reports tally.fuel on every builder lane, with and without a policy, over novelty and over an index. The charge is the same as for any upsert: 10 per transaction plus 0.001 per committed flake. For the test payload, main committed 9 flakes (10.009), 4 of its 5 retractions phantoms that removed nothing; this branch commits 9 real flakes (10.009). Payloads where main staged phantoms now stage fewer flakes, so reported fuel is lower by 0.001 per phantom (an identical refresh reports the 10.000 baseline and no commit). That is a billing change for callers who bill from the tally.

#1988 (merged). SPARQL DELETE DATA of a typed literal ("1"^^xsd:int, xsd:long, xsd:dateTime) used to lower the value as a string, which committed phantom retractions that removed nothing. With this PR alone it would commit nothing, and the values would still stay; with #1988's coercion, the intent names the stored term exactly and the resolver deletes it. it_delete_stored_facts::delete_data_of_typed_literals_deletes_them pins this over novelty and over an index (its own commit).

On top of #1997. This branch is rebased on 61b836e9a, and #1997's semantics are unchanged. How the DELETE witness rule reads #1997's graph mapping, and what happened to #1997's own tests, is folded below.

Named graphs and anchors. Bundles built from RDF text get their f:reifiesGraph anchor from one emit rule, graph_scope_emit — a local stand-in with the contract of GraphScope::emit from the JSON-LD graph-scoping PR (#2009), which replaces it when this rebases over that PR. bundle_templates never states an anchor, and a block label reaches staging as written, so the rules staging applies to write-graph names apply to TriG blocks unchanged.

Upgrade notes (for the release notes)
  • The first re-upsert of an unchanged Turtle/TriG document that holds blank nodes mints them once more: the identity is now computed from the parsed statements. The previous nodes stay, as they do after any edit. docs/transactions/turtle.md has a query that lists blank-node subjects nothing references. JSON-LD upsert and graph-sync identities are unchanged; a test pins that a graph synced through the old conversion re-syncs from the same text without a commit.
  • The first re-upsert or re-sync of a document with collections or language-tagged values commits a one-time correction. Collection entries are restated with their positions, and values the old path failed to retract are removed.
  • Relative references in TriG GRAPH blocks, as the label or in the contents, keep today's verbatim behavior; documented cross-ledger config relies on it. Sync and Graph Store now accept them in block contents too.
  • A DELETE of an absent fact no longer commits, so fewer commits advance t.
  • An identical re-upsert or re-sync is not a no-op for a value stored with less detail than it was written: an integer beyond 64 bits written with an XSD subtype, or a dateTime/time with digits past the microsecond. Once the value is indexed (for temporals, once it is reloaded from its commit), re-writing it as written commits a retraction and an assertion of one stored fact and can remove it. main behaves the same way. upsert.md now says so and how to write such values, every other idempotence claim in the transaction docs names the exception, and two tests document it (ignored until the fix lands).
  • A DELETE can name a stored ill-typed literal ("abc"^^xsd:integer, which Turtle writes store as its lexical form); lowering used to refuse it. In exchange, a mistyped literal in DELETE DATA ("1990-13-01"^^xsd:date) commits nothing instead of erroring.
API changes
  • Removed: Fluree::stage_transaction_with_trig_meta, stage_transaction_with_named_graphs, stage_transaction_with_named_graphs_tracked (the JSON-LD form is now stage_transaction_tracked), transact_with_trig_meta, transact_with_named_graphs. From fluree-db-transact: extract_trig_txn_meta, unwrap_trig_graph_blocks, UnwrappedTrig, TrigMetaResult, apply_cancellation and dedup_retractions (the generate::cancellation module), and FlakeGenerator::generate_retractions (now the crate-private generate_retract_intents).
  • Changed: FlakeAccumulator::push_retractions takes Retractions.
  • New: parse_rdf_text, parse_rdf_text_txn, parse_trig_txn, has_graph_blocks, Placement, TemplateSink, CurrentFacts/StoredFact/Retraction, FlakeAccumulator::finalize_with_restated, DELETE_WITNESSED_SITE, BinaryRangeProvider::persisted_object_key, TransactError::{Turtle, PayloadGraphMismatch}, fluree_graph_turtle::{RelativeIris, ParserOptions::relative_iris} (default Resolve), fluree_graph_turtle::SegmentParser and parser::TokenStream (Parser gains a token-source type parameter, defaulted to StreamingLexer), and in the CLI RemoteLedgerClient::sync_rdf and RemoteLedgerError::UnsupportedMediaType.
  • parse_trig_phase1/resolve_trig_meta stay, for bulk import only.
#1997: the DELETE witness rule over its graph mapping, and its tests

WITH and a JSON-LD update's top-level graph naming the ledger's own address read and write the ledger's default graph. An explicit GRAPH <address> resolves through the graph registry. The DELETE witness rule compares the graph ids that mapping produces: the WHERE's default graph when it is one ledger graph, each GRAPH <name>'s graph, and the graph the ledger has under a template's IRI. A default graph that reads as empty (an unknown USING/WITH graph) witnesses nothing. So under USING <address>, a GRAPH <address> template that writes a graph registered under that name is not treated as witnessed (it_delete_stored_facts::a_graph_named_by_the_ledger_address_gets_what_its_templates_write). No transaction lane reads TriG with phase 1 any more; bulk import still does, with #1997's shared prefix maps. #1997's redefined-prefix test, it_trig_insert::a_redefined_prefix_or_base_applies_only_after_it, now runs on the parse-once lanes. Two of #1997's phase-1 unit tests read the default graph through parse_to_json, which fluree-db-transact's clippy.toml now disallows; they read it with the Turtle parser instead.

Solo lockstep

From an audit of fluree/solo's write paths, these want a lockstep change on the solo side:

  • MCP jsonld_upsert hints ("prefLabel, one per language") and the legacy /v1/fluree/tm PATCH send one language per upsert. Under the new replacement unit (call 3 above) that deletes the other languages, so we'll want the hints and tool text to send the full multilingual set, or to use update with a language filter.
  • The "Update model from file" doc (update-a-model-from-a-file.md:51-53) should describe the one-time re-mint, then stable ids.
  • More no-op receipts: phantom-only deletes and identical refreshes. Map the empty-content commit id to commit: null.
  • Turtle upsert keeps populating result.tally (pinned by a test). Fuel is lower where main billed phantom flakes.
  • TriG structural errors are now Turtle parse error: …, which solo's router maps to 502, so the prefix should go in CLIENT_ERROR_PREFIXES.
  • Since that audit, solo's new MCP jsonld_update tool text says "Retracting a triple that is not there does nothing", which is true only with this PR — and hints.rs:85 still says skos:prefLabel "(one per language)", so the language-hint change above is still needed.

Performance

The short version: Turtle upsert is 22–30% faster (one parse instead of the JSON-LD round trip), TriG upsert and insert are flat to faster, and on #1997's many-block TriG shape upsert is 10% faster and insert at parity, with 36–37% fewer allocations and a 25% lower peak, which Claude's review pass reproduced against 61b836e9a on its own (upsert 72.0 → 65.4 ms, insert 62.5 → 61.8 ms, allocations −36/−37%, peak 38.4 → 28.9 MiB). A1's upsert staging reads 7–12% faster in isolation (new_subjects +0.5%). JSON-LD and Turtle insert, whose paths A2 changes only in dispatch, read flat by their fastest reps, and bulk import is untouched by A2 and reads flat.

Quiet box. EC2 c7i.4xlarge, the fat-LTO bench profile, base and head in interleaved rounds; a change counts as a win or a loss only when the base and head ranges don't overlap and the median delta exceeds the bench's budget (5% at small). Four wins and no loss: named_novelty −82.5% (120 → 21 ms), upsert_turtle −36%, trig −21% and trig_blocks −17% (the new many-blocks case, in its first A/B). Every other measured row is stable, transact_filtered_delete included (+1.2% and +0.8%). Of the three end-to-end A1 upsert rows the local runs flagged, named_indexed (−3.8%) and novelty_heavy (−5.1%) are stable, and a second session resolved new_subjects at −0.9%, with the head faster in all 7 rounds, so the first session's +14% doesn't reproduce.

The quiet-box tables

Session 1 (2026-10-01): three rounds, in AB, BA, AB order. The base is main at 61b836e9a plus this PR's two bench-only commits (8024ea1d2, 41579486a), so both sides run the same bench source; the head is GitHub's merge of this PR (refs/pull/2008/merge, 449b3dc1c).

bench scenario scale base median [range] head median [range] Δ budget verdict
insert_formats jsonld/10txn_100nodes small 20.676 ms [20.027 ms–20.807 ms] 20.512 ms [19.813 ms–20.717 ms] -0.8% 5% stable
insert_formats trig/10txn_100nodes small 25.342 ms [24.755 ms–25.485 ms] 19.958 ms [19.479 ms–19.987 ms] -21.2% 5% win
insert_formats trig_blocks/10txn_100nodes small 22.134 ms [21.456 ms–22.172 ms] 18.276 ms [17.649 ms–18.295 ms] -17.4% 5% win
insert_formats turtle/10txn_100nodes small 14.169 ms [13.748 ms–14.174 ms] 14.066 ms [13.621 ms–14.129 ms] -0.7% 5% stable
insert_formats upsert_turtle/10txn_100nodes small 29.499 ms [28.745 ms–29.686 ms] 18.954 ms [18.727 ms–19.050 ms] -35.7% 5% win
transact_filtered_delete no_lists/small small 18.825 ms [17.715 ms–18.996 ms] 19.053 ms [17.689 ms–19.169 ms] +1.2% 5% stable
transact_filtered_delete with_lists/small small 30.336 ms [28.517 ms–30.575 ms] 30.588 ms [28.051 ms–30.747 ms] +0.8% 5% stable
transact_upsert_replace named_indexed/small small 29.147 ms [28.986 ms–29.187 ms] 28.025 ms [27.949 ms–30.354 ms] -3.8% 5% stable
transact_upsert_replace named_novelty/small small 120.359 ms [119.490 ms–120.887 ms] 21.106 ms [20.942 ms–23.315 ms] -82.5% 5% win
transact_upsert_replace new_subjects/small small 17.310 ms [16.955 ms–19.578 ms] 19.761 ms [17.516 ms–19.783 ms] +14.2% 5% stable
transact_upsert_replace novelty_heavy/small small 29.399 ms [28.826 ms–30.314 ms] 27.907 ms [26.892 ms–29.479 ms] -5.1% 5% stable
transact_upsert_replace wide_predicates/small small 99.471 ms [96.793 ms–99.796 ms] 93.480 ms [91.864 ms–102.432 ms] -6.0% 5% stable

Session 2: the one row session 1 left inconclusive, over 7 rounds in alternating AB/BA order. head>base counts the paired rounds in which the head was slower.

bench scenario scale base median [range] head median [range] Δ head>base (paired rounds) budget verdict
transact_upsert_replace new_subjects/small small 16.726 ms [15.442 ms–16.920 ms] 16.583 ms [15.330 ms–16.852 ms] -0.9% 0/7 5% stable

Where it reads slower, fwiw: on an indexed ledger, the constant-object DELETE costs 1.09–1.13× current main (61b836e9a; 1.13× in this branch's re-run, 1.09× in Claude's review pass) — the cost of reading every slot instead of committing 100k phantom retractions, since the DELETE matches nothing. One POST read per (predicate, object) pair would serve all of them, but that stays a follow-up because it's a real scope addition rather than a tweak: it needs a novelty read by pair beside today's per-slot one, and a bound for popular pairs, where one POST read would scan every subject holding the value while the DELETE names a few. Three end-to-end A1 upsert rows also read 1.13–1.20× locally, but neither the staging probe, which covers the only code A1 changes, nor the quiet box (above) reproduces them.

Two caveats I want to be upfront about. The A1 numbers are pre-rebase: they were measured before the rebases (onto bf523e24e, then onto 61b836e9a with #1997) and before the late A1 commits. The #1997-shape runs and the many-small-blocks table (re-measured against 61b836e9a) are on the current base; the A2 criterion rows (including the 22–30% Turtle-upsert figure) predate #1997 and the late commits too. And the tables ran on one shared 16-core machine with load from other work (1-minute load average 5–85), where end-to-end rows vary up to ±30% between reps of the same binary, so the regression-budget gate on a quiet runner is the authority for the 5% budget, and the quiet box above applied the same budgets.

Every perf table, with methodology

dev-fast profile unless noted, on one shared 16-core machine with load from other work during the runs (1-minute load average 5–85). Each table compares binaries run interleaved.

A1 numbers are pre-rebase; the quiet box (above) has since measured the PR head. The two A1 tables were measured before the rebases (onto bf523e24e, then onto 61b836e9a with #1997) and before the late A1 commits (index-key matching, the policy check, OPTIONAL/UNION witnesses).

A1, upsert staging in isolation (a scratch probe staging each upsert 40× at N=2000 or 20× at N=8000 on clones of one fixture; p50 ms, median of 3 interleaved rounds at N=2000, rounds 1/2 at N=8000). Base is main plus this PR's first commit (the new bench).

scenario base N=2k A1 N=2k Δ base N=8k A1 N=8k
default_indexed 11.46 10.23 −11% 48.3 / 51.2 41.8 / 46.3
novelty_heavy 12.04 10.61 −12% 52.0 / 59.0 44.4 / 49.9
named_indexed 12.54 11.18 −11% 52.1 / 55.5 48.4 / 52.2
wide_predicates 36.36 32.32 −11% 155.3 / 179.3 147.9 / 161.8
lang_list 27.29 25.36 −7% 117.4 / 131.1 118.5 / 124.0
new_subjects 6.47 6.50 +0.5% 26.3 / 29.1 26.1 / 29.7

A1, criterion end to end (stage + file-backed commit), small scale, 3 reps, median ms unless noted; same base, before the rebase:

bench base A1 ratio
transact_upsert_replace/default_novelty 23.0 19.5 0.85
…/default_indexed 26.9 23.8 0.89
…/named_novelty 98.9 25.9 0.26
…/named_indexed 28.5 32.9 1.16
…/lang_list 55.4 53.5 0.97
…/wide_predicates 97.8 71.7 0.73
…/novelty_heavy 39.2 47.2 1.20
…/new_subjects 29.4 33.2 1.13
transact_filtered_delete/no_lists 31.6 30.6 0.97
transact_filtered_delete/with_lists 45.6 46.2 1.01
transact_commit/fresh_ledger (µs) 96.5 101 1.04
transact_commit/populated_ledger (µs) 283 284 1.01
novelty_replay/replay_chain 14.2 12.5 0.88
query_overlay_only_range/narrow_subject (µs) 16.3 16.3 1.00
query_overlay_only_range/wide_subject (µs) 122 123 1.01

The end-to-end rows vary up to ±30% between reps of the same binary on this machine. The three upsert rows above 1.1 are not reproduced by the staging probe, which covers the only code A1 changes: commit is unchanged. The regression-budget gate on a quiet runner is the authority for the 5% budget, and on the quiet box (above) none of the three is a loss.

A2, criterion end to end on the branch at bf523e24e (before #1997), before the late commits (the segment parser, verbatim block labels, the late A1 commits): the A1 tip plus the A2 bench commit against the A2 tip of that time. insert_formats runs 10 transactions per iteration, each iteration on a fresh ledger (ms). Medians of 3 interleaved reps. The first run read 1.09 on trig, so the TriG and 100-node insert rows were run again, 5 interleaved reps; for those rows the last column is the fastest of all 8 reps of each binary.

bench nodes/txn A1 A2 ratio 5-rep rerun fastest rep A1 / A2
jsonld (insert) 10 1.42 1.40 0.99
turtle (insert) 10 1.01 1.02 1.01
trig (TriG upsert, one block) 10 1.78 1.93 1.09 0.85 1.76 / 1.77
upsert_turtle 10 1.91 1.37 0.72
upsert_turtle_replace 10 1.93 1.41 0.73
trig_mixed (TriG insert, default graph + block) 10 1.91 1.76 0.92 0.83 1.84 / 1.69
jsonld 100 14.0 14.7 1.05 0.86 13.2 / 13.2
turtle 100 9.04 9.49 1.05 0.83 8.89 / 8.87
trig 100 16.7 17.0 1.02 0.91 16.4 / 15.7
upsert_turtle 100 18.6 13.1 0.70
upsert_turtle_replace 100 19.3 15.0 0.78
trig_mixed 100 17.2 16.6 0.96 0.86 16.7 / 15.1
import_bulk single-threaded (5 reps) small 237 228 0.96 224 / 217
import_bulk default threads (5 reps) small 216 222 1.02 214 / 211

TriG on #1997's shape. 10,000 blocks of 2 prefixed-name triples each, 20 @prefix directives at the top, 16 graph labels, through upsert_turtle and insert_turtle into a fresh memory ledger. dev-fast builds of main (61b836e9a), of the A1 commits plus the bench commit, and of this branch; 5 interleaved rounds of 15 runs; median of the per-round medians, with the range in brackets. Peak heap and allocations come from a counting allocator and were identical in every round.

build upsert (ms) insert (ms) peak allocations, upsert / insert
main 72.2 [71.0–77.8] 62.5 [62.0–63.6] 38.4 MiB 852,594 / 832,732
A1 commits 72.5 [71.7–73.4] 62.6 [62.4–64.3] 38.4 MiB 852,593 / 832,734
this branch 64.7 [64.5–73.2] 62.2 [61.4–67.5] 28.9 MiB 543,779 / 523,915

Upsert is 10% faster and insert is at parity, with 36–37% fewer allocations and a 25% lower peak. Claude's review pass reproduced it against 61b836e9a on its own: upsert 72.0 → 65.4 ms, insert 62.5 → 61.8 ms, allocations −36/−37%, peak 38.4 → 28.9 MiB. Before the parser read the locator's tokens, this branch lexed every block twice and measured 78.2 ms and 75.0 ms on the same shape. The parser is generic over its token source, and the default lexer path is unchanged:

  • Parsing 21.3 MB of plain Turtle (840k triples) through fluree_graph_turtle::parse alone, release builds, 6 interleaved rounds of 7 runs: 360.6 ms [358.5–365.6] (59.1 MB/s) before the change, 350.3 ms [342.7–362.5] (60.9 MB/s) after.
  • A plain Turtle insert of 180k flakes, dev-fast, 6 interleaved rounds of 5 runs: 495.7 ms [479.1–505.0] before, 492.9 ms [476.0–508.6] after.

Many small blocks, and DELETE rows the matched lane used to read (dev-fast CLI end to end, including about 0.2 s of start-up; this branch's head against main at 61b836e9a, with #1997; 3 interleaved reps; seconds, median). The TriG documents are 100k one-triple blocks under 20 prefixes over 50 graphs (blocks100k) and 100k default-graph statements plus a block (mixed100k); the DELETEs run over 100k subjects:

case main this branch ratio
TriG insert, blocks100k 0.42 0.41 0.98
TriG upsert, blocks100k 0.44 0.39 0.89
TriG insert, mixed100k 0.65 0.38 0.58
TriG upsert, mixed100k 0.62 0.31 0.50
DELETE WHERE, witnessed, novelty 0.40 0.38 0.95
DELETE, OPTIONAL-bound, novelty 0.78 0.78 1.00
DELETE, constant object, novelty 0.37 0.32 0.86
DELETE WHERE, witnessed, indexed 0.30 0.31 1.03
DELETE, OPTIONAL-bound, indexed 0.35 0.36 1.03
DELETE, constant object, indexed 0.24 0.27 1.13

Claude's review pass measured the TriG rows independently against 61b836e9a at 0.97, 0.90, 0.59 and 0.53, and the indexed constant-object DELETE at 1.09.

Before the segment parser and the OPTIONAL/UNION witnesses, an adversarial review pass (Claude) measured this branch at 1.6–1.9× main on the block documents and 1.5× on the indexed OPTIONAL-bound DELETE.

insert_formats/trig_blocks (every statement in its own block, sixteen extra prefixes) now covers the many-blocks shape in the criterion suite. It hasn't been run A/B on this machine, but the quiet box has it at −17.4% (above). The constant-object DELETE matches nothing: main commits 100k phantom retractions, and this branch reads each slot and commits nothing. On the indexed ledger that costs 1.09–1.13× main: reading every slot instead of committing 100k phantoms (see Follow-ups).

Turtle upsert is 22–30% faster: one parse instead of the JSON-LD round trip. TriG upsert and insert are flat to faster. JSON-LD and Turtle insert, whose paths A2 changes only in dispatch, read flat by their fastest reps. Bulk import is untouched by A2 and reads flat. transact_upsert_replace (JSON-LD upsert over existing values) ran in the same session and shows no consistent direction, with ±50% between reps of one binary at that load.

Deviations from the design

  • Witnesses: the design never let an OPTIONAL or UNION triple witness a DELETE row. A triple that is the only pattern of an OPTIONAL or of one UNION branch now does, when its object variable appears nowhere else, and graphs compare as the ledger graph ids the WHERE reads (rather than per graph name).
  • Relative references in GRAPH blocks with no @base are kept as written (today's behavior), where the design made them an error. That takes one parser option, ParserOptions::relative_iris (default Resolve, so every other caller stays strict). The design expected no change to fluree-graph-turtle; it also gains SegmentParser, one parser over the pieces of a TriG document, fed the tokens the locator lexed.

The smaller, structural ones, and the rebase note for #1555/#1569 (the whole fluree-graph-turtle diff, +201/−7), are folded here:

Structural deviations, and the fluree-graph-turtle diff for #1555/#1569

Follow-ups

  • Bulk import on the same driver (the import half of TriG GRAPH blocks reject anonymous blank nodes [ … ] on /upsert and import #1930), and then deleting the TriG phase-1 reader.
  • The pre-existing transact_filtered_delete and novelty_replay benches drop their file-backed TempDir inside the timed routine. This PR's new bench doesn't; those two should follow.
  • Values stored with less detail than written: an integer beyond 64 bits keeps no XSD subtype in the index, and a dateTime/time is stored to the microsecond. An identical re-upsert or re-sync of such a value, once indexed (temporals: once reloaded from the commit), retracts the stored decode and asserts the written term under one key, and can remove the value. main does the same. The fix belongs where values are written (normalize to storage precision) or in the index (keep the subtype). Claude's review pass found more values that read back with less detail than written, on main too: after indexing, a big plain xsd:integer reads back as xsd:decimal; -0.0 loses its sign; and "1.50" reads back as 1.5.
  • A constant-object DELETE over many subjects reads one slot per subject: 1.09–1.13× main on an indexed ledger, where main commits phantom retractions without reading. One POST read per (predicate, object) pair would serve all of them. It needs a novelty read by pair beside today's per-slot one, and a bound for popular pairs, where one POST read would scan every subject holding the value while the DELETE names a few.
  • Per-(s, p, lang) upsert replacement is open for a decision (call 3 above); the replacement unit lives in one function, upsert_replacement_slots.
  • Query-side datatype fidelity for big integers, on main too: a SELECT over an indexed big integer reads xsd:decimal, and BIND gives one xsd:integer whatever its subtype. So on a novelty-only ledger a DELETE row whose value BIND bound does not name a big integer stored under a subtype.
  • Tracing span-capture tests are racy in bundled runs: a test that installs a thread-local capture (span_capture::init_test_tracing) can miss its event while other tests in the same process install theirs (it_annotation_filter_pushdown, it_limit_stops_work). Unfiled.
  • f:ledger is read only as an IRI reference, and a slashed ledger id has no absolute spelling. Accepting a string literal there (and documenting it) would let config avoid relative references.

Tests

Both groups have regression suites — A1's stored-retraction, delete, policy (deletes and restated values), witness-routing (paired MustFire/MustNotFire stamps on one ledger), annotation, conflict-retraction and legacy-revert tests, and A2's RDF-text parity matrix over 11 upsert/insert entry points and 4 graph-payload lanes — and each regression test was run with its fix reverted and watched fail. testsuite-sparql (W3C) is 36/36, and the five triple-term entries that were blocked on TriG GRAPH-block parsing now load their data (none passes yet; each is re-attributed to the failure it now reaches). The full list and the non-vacuity runs are folded here:

Full test list, and the non-vacuity runs
  • A1:
    • it_upsert_stored_retractions: lang, list, identical refresh, named graph, blank-node refresh, annotated cascade. Each case runs on three lanes (novelty, indexed, novelty over index), before and after reindex.
    • it_delete_stored_facts: absent DELETE twins (SPARQL, JSON-LD, Cypher), DELETE/INSERT of an absent fact, exact terms, language replacement twins, the big-number witnessed delete with and without lists, and typed-literal DELETE DATA (with fix: keep typed literal values through the index (reindex affected ledgers) #1988). Also:
      • indexed facts named by index key, each on a novelty-only and an indexed ledger: big integers under four subtypes (DELETE DATA, JSON-LD delete, constant template), rows bound by OPTIONAL/UNION/BIND (SPARQL, JSON-LD OPTIONAL), Cypher SET and DETACH DELETE, a named graph, sub-microsecond dateTime/time, and a different big value that is not deleted;
      • a delete refused by policy with one answer whether or not the value is stored (DELETE DATA of the stored and two absent values, a WHERE-driven guess, JSON-LD twins), and a permitted delete of an absent value that commits nothing;
      • an upsert (JSON-LD plain, language-tagged and list; Turtle), a DELETE and INSERT of one value (SPARQL, JSON-LD), and a graph sync, each refused by policy with one answer whether or not the value is stored; restating a value the identity can view commits nothing without modify permission, and as root;
      • a ledger holding a graph registered under its own address, updated with USING and GRAPH naming that address;
      • a stored ill-typed literal deleted through SPARQL and JSON-LD, and still refused by INSERT DATA.
    • it_delete_witness_routing: paired MustFire/MustNotFire stamps on one ledger; OPTIONAL and UNION rows (JSON-LD and SPARQL), Cypher SET and DETACH DELETE take the witnessed lane, and a UNION binding the object twice does not (JSON-LD and SPARQL).
    • shacl_tests::a_refresh_upsert_satisfies_unique_lang, and two ignored tests for values stored with less detail than written (big-integer subtype, sub-microsecond dateTime).
    • it_annotation_delete_named_graph, it_named_graph_conflict_retractions (merge, rebase, revert), it_revert_legacy_phantom.
    • it_typed_literal_index::sparql_insert_data_rejects_an_ill_typed_literal, Index build silently destroys temporal and xsd:long literals (xsd:date → 1970-01-01), and the resulting triples cannot be retracted by literal match #1987's test updated (see Interaction with Index build silently destroys temporal and xsd:long literals (xsd:date → 1970-01-01), and the resulting triples cannot be retracted by literal match #1987).
    • Unit tests for current_facts and delete_witness (graph ids, OPTIONAL and UNION witnesses).
  • A2:
    • it_rdf_text_parity:
      • IRI matrix over 11 upsert/insert entry points and 4 graph-payload lanes;
      • upsert stores what insert stores;
      • Turtle list upsert;
      • idempotence, order-insensitivity and edits;
      • fuel over 4 lanes;
      • the raw transaction is the text;
      • errors at document offsets, and refusals inside a block;
      • synced collection order;
      • sync identity against the old conversion;
      • the config recipe (trig upsert with anonymous blank nodes causes parse error #1511);
      • a ledger id in a GRAPH block stored as written on upsert, TriG insert and sync, and under a relative block label on upsert and TriG insert (a sync target must be an absolute IRI); the default-graph twin refused; a base resolving label and contents;
      • the documented orphan queries, in the default graph and graph by graph, and a node described in one graph and referenced from another, which they do not list (a check within the node's own graph would).
    • it_trig_insert::a_redefined_prefix_or_base_applies_only_after_it (fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997's): its redefined-prefix and redefined-base cases, now through the parse-once upsert and insert lanes.
    • Unit tests for rdf_text and locate: spellings, refusals, offsets, directive order, block labels under redefined declarations, block parity, blank-node scope, literals, txn-meta, content identity, placement, stable ids, reserved predicates, relative references. SegmentParser carries declarations across pieces.
    • CLI: RDF sent as written; 415 fallback for Turtle only; locator detection.
    • The existing cross-ledger config tests (29, which name model ledgers as <org/governance:main> in TriG blocks) pass unchanged.
  • Non-vacuity. Each regression test was run with the fix reverted, and failed as expected:
    • A1: old upsert wave; forced witnessed; no witnesses; no enrichment; old cascade; old rebase read; old undo fold. Also:
      • no index-key match → the five indexed big-integer and sub-microsecond tests fail (the different-value control passes either way);
      • unmatched intents kept from the policy check → the policy test fails at the absent value;
      • restated values kept from the policy check → every stored-value case of the upsert, DELETE-and-INSERT and sync policy tests commits nothing instead of being refused (13 cases), and the viewable-value test still passes;
      • graph ids not compared → graph_contexts_must_agree fails;
      • OPTIONAL/UNION candidates off → the witness unit test fails, and the routing test fails on all four JSON-LD and SPARQL OPTIONAL/UNION cases;
      • another UNION branch's binding ignored → both routing cases that bind the object twice fail, each retracting a phantom;
      • ill-typed lexical forms refused again → the ill-typed DELETE test fails;
      • retractions without their language tag → the sh:uniqueLang refresh test fails.
    • A2:
      • upsert back through phase 1 + parse_to_json → IRI, parity, list, idempotence and error tests fail;
      • raw txn as JSON-LD → raw-txn test fails;
      • tracker dropped → fuel test fails (fuel 0.0 on all 4 lanes);
      • TriG insert back through phase 1 → the redefined-prefix/base insert cases and 5 insert IRI lanes failed (run before fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997, which has since fixed those cases in phase 1);
      • graph payload back through the conversion → graph IRI and synced collection tests fail;
      • anonymous blank labels changed → sync identity test fails;
      • firewall off → reserved-predicate test fails;
      • stable ids ignored → stable-id test fails;
      • no 415 fallback → fallback test fails;
      • phase-1 TriG detection → [ … ]-in-block detection fails;
      • block contents resolved like the default graph → both ledger-id tests (unit and integration) fail;
      • block labels resolved like the default graph → the relative-label cases of both fail;
      • the label cache keyed by the label alone → the block-label test fails (a label after a redefined prefix or base keeps its first IRI).
    • trig upsert with anonymous blank nodes causes parse error #1511 reproduced on main: expected object, found '['.
  • testsuite-sparql (W3C): 36/36 groups pass. The five eval-triple-terms entries registered as "blocked on TriG GRAPH-block parsing" now load their data. Each is re-attributed to the failure it reaches (non-isomorphic reifier results, unlowered triple-term functions, SPARQL UPDATE quoted triples, GRAPH ?g-valued annotation); none passes yet.
  • Special float values are excluded from the parity fixture; they are tracked separately.

Gates

This is 35 commits on 61b836e9a, with the head at 99d5776c9. The gates ran on the head commit as first written, before an amend that only moved one helper within stage.rs; fmt, clippy and the policy tests were re-run at the head. There, fmt (after the last edit) and clippy -D warnings are clean, every commit type-checks on its own, and the workspace --all-features check is clean. cargo test -p fluree-db-api is 3,950 passed / 1 failed / 121 ignored: the failure is a tracing-capture test that fails on main too in a bundled run, and it passes under cargo nextest, one process per test, as CI runs them. The SHACL, server, consensus/memory and #1997 suites pass, and testsuite-sparql is 36/36 SPARQL and 3/3 RDF groups. A few results are from 987a9379c, in crates and workspaces nothing has changed since: fluree-db-query, fluree-graph-turtle, fluree-graph-format/fluree-db-r2rml, testsuite-sparql's fmt/clippy, and it_sql_pushdown_lane. In an earlier run there, it_limit_stops_work was intermittent, failing in 1 of 7 bundled runs on this branch and 0 of 18 on main; this branch touches neither tracing nor the property join, and it passes under nextest. Not run: testsuite-shacl (not a CI job), the workspace-wide --all-features nextest that CI runs, and the fluree-db-api tests gated on the graph-source and vector features, except it_sql_pushdown_lane.

Full gate log

At the head commit as first written, on 61b836e9a, before an amend that only moved one helper within stage.rs. The head, 99d5776c9, differs only in where that helper sits; fmt, clippy and the policy tests were re-run there.

  • cargo fmt --all -- --check, after the last edit: clean.
  • cargo clippy --all --all-features --all-targets --locked -- -D warnings: clean.
  • Every commit type-checks on its own (cargo check --all-targets over the query, transact, api, Turtle, CLI and server crates). The workspace --all-features check is clean.
  • cargo test -p fluree-db-api (the lib and every integration group, default features): 3,950 passed, 1 failed, 121 ignored.
    • The failure is it_annotation_filter_pushdown::annotation_body_threshold_reduces_scan_work_on_both_surfaces. It fails on main too in a bundled run: its tracing capture shares the process with the rest of grp_misc.
    • Under cargo nextest, one process per test as CI runs them, it passes.
    • In an earlier run at 987a9379c, it_limit_stops_work::a_selective_star_sizes_later_chunks_from_its_yield missed its tracing event the same way (expected one property join, got []). It is intermittent: it failed in 1 of 7 bundled runs of its group at this branch and in none of 18 on main. This branch touches neither tracing nor the property join, and it passes under nextest.
  • cargo test -p fluree-db-api --features shacl: the lib, 922 passed; the targets with SHACL-gated tests (grp_ledger, grp_graphsource, it_optimistic_rebase, it_staged_view_dict), 501 passed.
  • fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997's tests pass: it_named_graphs (including test_updates_on_a_graph_registered_under_the_ledger_address), it_trig_insert (including the redefined-prefix test), it_import, it_sql_pushdown_lane (--features sql, run at 987a9379c), and the server's sparql_dataset_semantics in the server run below.
  • cargo nextest run: fluree-db-server (every target) 662 passed; fluree-db-consensus and fluree-db-memory 153 passed.
  • cargo test: fluree-db-transact --lib 329 and fluree-db-cli 460 passed. Doc tests of fluree-graph-turtle, fluree-db-transact and fluree-db-api: 1 run and passed, 104 marked ignore.
  • At 987a9379c, with no change to these crates since: fluree-db-query --lib 1,648, fluree-graph-turtle 197, and fluree-graph-format and fluree-db-r2rml 188 passed.
  • testsuite-sparql: 36/36 SPARQL groups and 3/3 RDF groups pass at the tip; fmt and clippy -D warnings clean at 987a9379c, with no change to that workspace since.
  • Not run: testsuite-shacl (not a CI job); the workspace-wide --all-features nextest, which CI runs; and the fluree-db-api tests gated on the graph-source and vector features, except it_sql_pushdown_lane.

Nothing in the suite measured the upsert wave that retracts the current
values of every (graph, subject, predicate) an upsert names: insert_formats
upserts only new subjects and transact_commit measures inserts. Eight
scenarios: default and named graph, each unindexed and indexed; language
tags plus @list; 16 predicates per subject; an indexed base under 4x
unrelated novelty; and new subjects as the #1549 control. Each scenario
asserts its lane in setup and its retract/assert counts after every
measured upsert, so a drifting fixture fails instead of timing the wrong
path.
…tion LRU alone

binary_range_eq_v3 sent every probe through the cross-call translation
LRU when the overlay reported a content version, and NoOverlay reports
one. Staging reads the persisted index through NoOverlay once per
retraction group, and the LRU key includes to_t, so each commit's first
such probe per (graph, index order) inserted an empty entry into an
eight-slot cache and could evict a live query's novelty translation.
The translation step moves into probe_overlay, which returns an empty
translation for an effectively empty overlay before touching the cache.
A retraction staged for commit must name a fact that is currently
stored. Today staging rebuilds retractions from query bindings, templates
and edge keys, which drops language tags and list positions and invents
retractions of facts that do not exist. This adds the read those sites
will share: per (graph, subject, predicate) slot, one seek into the
persisted index through NoOverlay plus a Sid-space novelty seek bounded
by the slot's sentinels, lifecycle-resolved, graph stamped. A graph whose
novelty is small next to the slots asked of it is walked once instead.

StoredFact (private fields, built only here) and Retraction (private
field, built from StoredFact::retract) carry the invariant in the types;
the next commits route staging through them.
The upsert wave queried each (graph, subject, predicate) with
execute_pattern and rebuilt every retraction from the ?o binding with
m: None. A language-tagged value was retracted without its tag and a list
entry without its position, so neither retraction matched anything: the
old values stayed, and an identical refresh of such values committed
two flakes every time (#1976). A named graph without a range provider
also walked the graph's whole novelty once per slot.

The wave now asks CurrentFacts for each slot's current facts and retracts
them as stored. The #1549 absence pre-check and its counters are
unchanged; pattern_queries now counts slot reads. binding_to_flake_object,
query_novelty_for_graph and the per-group materializer are gone.

The replacement unit is one function, upsert_replacement_slots: every
current value of each (graph, subject, predicate) the payload names, all
languages and list positions together. Blank subjects join the wave when
the skolem scope is deterministic (the upsert payload scope), so an
identical re-upsert of a document with blank nodes is now a no-op too.
…rrentFacts

The whole-graph scan moves from scan_graph_flakes to
CurrentFacts::whole_graph, which returns StoredFacts with the graph
stamped. Graph sync, CLEAR/DROP and the COPY/MOVE destination and source
clears now build their retractions with StoredFact::retract instead of
flipping op and t on scanned flakes. Behavior is unchanged: the scan,
its limit, the stamping and the policy model are the same.
A DELETE template instantiated into a retraction of a fact the ledger does
not hold used to be committed anyway. DELETE DATA of an absent triple
committed a new t with one retraction, which docs/transactions/retractions.md
says is a no-op. Worse, DELETE {x} INSERT {x} WHERE {} with x absent
committed nothing and never wrote x: the phantom retraction cancelled the
insert in the accumulator, against SPARQL 1.1 Update's (DS - D) ∪ I. Both
twins (SPARQL and JSON-LD) behaved the same.

Every DELETE instantiation is now a RetractIntent, and the accumulator
accepts only Retraction, which CurrentFacts produces:
- An intent whose template a WHERE triple witnesses (delete_witness: same
  subject, predicate and object variable, the object bound nowhere else,
  the same graph) is the decode of a fact the WHERE just matched. On a
  ledger that holds no @list position (ListFree) it retracts as decoded with
  no read, which keeps bulk filtered deletes at zero lookups. On a
  list-bearing ledger it reads its slot for the list position, under the
  rule hydration had: the first stored list entry with the same term.
- Every other intent (template constants, DELETE DATA, values bound by
  another pattern) is matched against the stored facts of its slot: exact
  datatype (#1990's rule, one function), case-insensitive language tag, the
  named list position or else the first list entry, else the plain value.
  An intent that names no stored fact stages nothing.

hydrate_list_index_meta_for_retractions is gone, and so is the unused
generate::cancellation module, which pushed arbitrary flakes as
retractions. The delete_witnessed routing stamp reports which lane ran.

The it_fast_stats_1391 fixture relied on a DELETE writing a no-op
retraction; it now hand-builds that commit, the shape legacy and replayed
commits still carry.
…r graph

The cascade that follows a retracted edge to its annotations rebuilt each
f:reifies* bundle from the decoded EdgeKey, and read annotation subjects
with range_with_overlay(&novelty). Its body-cleanup pass (a by-id
annotation delete in LPG mode) did not stamp the graph on index-resident
rows, so the body's retraction landed in the default graph and the body
survived in its named graph. And nothing retracted the bundle's
f:reifiesGraph anchor, which the by-id delete does not name, so it stayed
behind describing nothing.

All three passes now read annotation subjects with CurrentFacts::of_subject,
graph stamped, and retract bundles and bodies as stored. A complete by-id
bundle retract also retracts every other stored slot of that bundle,
unless the transaction re-points the reifier.
A take-source merge or revert, and a take-branch rebase, retract the
losing side's current values on each conflicting (subject, predicate,
graph). current_asserted_for_key read them with a range read and kept the
rows whose graph equalled the key's, but index-resident rows come back
without their graph, so in a named graph every indexed value was dropped:
the take-source merge left the target's value next to the source's, and
merge_preview showed no target value. Novelty-resident values and the
default graph were unaffected.

current_asserted_for_key now reads the slot through CurrentFacts, which
stamps the graph, and the retractions come from StoredFact::retract.
Upserts written before the retraction resolver retracted "a"@en as a
tagless "a"^^rdf:langString. That retraction matched nothing and is inert
for reads and indexing, but a revert inverts every flake of the reverted
commits, and inverting it asserted "a" as an rdf:langString with no language
tag (LANG ""), a literal that cannot exist.

The undo fold now skips any retraction whose shape no stored fact has: an
rdf:langString without a tag, or a tag on another datatype. It logs how
many it skipped. Legacy retractions of list entries without their position
are well-formed plain literals, so no shape check can catch them; telling
those apart needs a historical stored-fact read and is left as a follow-up.
…only stored facts

upsert.md: the replacement unit is the whole (graph, subject, predicate), all
languages and list positions together, and an identical refresh commits
nothing, blank nodes included. The "Empty Replacement" and comparison-table
text said an upsert removes predicates the payload does not name; it never
did, and the page's own summary said so. retractions.md: a DELETE of a triple
that is not stored commits nothing, names terms exactly (datatype exact,
language tag case-insensitive), and a DELETE/INSERT of an absent triple
inserts it. update-where-delete-insert.md: the same, and the comparison
table's upsert granularity.
With #1988, SPARQL UPDATE lowers `"1"^^xsd:int`, `xsd:long` and
`xsd:dateTime` literals as JSON-LD does, so a DELETE DATA intent names the
stored term exactly and the retraction resolver deletes it, over novelty
and over an index. Before #1988 the intent named a string, and the delete
removed nothing.
The persisted index keeps less than some terms carry. An integer too large
for i64 is keyed by value alone (NUM_BIG), with no XSD subtype, and a
dateTime or time is keyed to the microsecond. So an indexed fact's decode
can differ from the term that names it. The index decode says xsd:integer
where a big xsd:nonNegativeInteger was written, a query row says
xsd:decimal, and a nanosecond literal decodes to its microsecond.

Matching by term alone therefore missed these facts on an indexed ledger.
DELETE DATA, a JSON-LD delete or a constant DELETE template naming a big
integer under its subtype committed nothing. So did a DELETE row whose
value an OPTIONAL, UNION or BIND bound. Cypher SET kept the old big value
next to the new one, DETACH DELETE left it behind, and a sub-microsecond
dateTime or time survived DELETE DATA. main deleted each of these, because
its retraction carried the same index key.

A fact read from the index is now also named by its index key: the
intent's object encoded as the index keys it (o_type + o_key), compared
with the fact's. Terms with one key are one fact to the index, which
merges every retraction by that key. The resolver still retracts the fact
as read, so every retraction names a stored fact. Term identity stays
first, and novelty facts keep it alone. The key comparison runs only
when no stored term matches, so a DELETE of a present term costs what it
did. The witnessed-row list lookup uses the same rule.

BinaryRangeProvider::persisted_object_key exposes the encoding the
overlay merge already applies to retractions.
A DELETE intent that names no stored fact stages nothing. Under a policy,
that also kept it from the modify-policy check, which only saw the staged
flakes. A delete the identity may not perform was then refused when its
target was stored and committed nothing when it was not.

Intents that name no stored fact now go through the modify-policy check
with the staged flakes, and are still dropped from the commit. A delete
the identity may not perform is refused the same way whether or not its
target is stored, as it was before intents were resolved against storage.
A permitted delete of an absent value still commits nothing. Only a
non-root policy context collects the unmatched intents.
…L and UNION rows included

Two changes to the rule that lets a DELETE row retract without a read.

Graphs compare as ledger graph ids. The rule used to compare the WHERE
default's IRI with the template's IRI, and a dataset alias by name. It now
takes the graph ids the WHERE dataset actually reads: its default graph
when that is one ledger graph (the ledger's own with no USING, WITH or
from, else the one graph those name resolves to, where the ledger's own
address is its default graph), and each GRAPH <name>'s graph. It compares
them with the graph the ledger has under the template's IRI. A name the
WHERE resolves to one graph and a template writes to another is never a
witness. USING or WITH naming the ledger's own address reads the default
graph, while a GRAPH <address> template writes a graph registered under
that name. A test covers a ledger holding such a graph.

OPTIONAL and UNION rows can witness. A triple that is the only pattern of
a top-level OPTIONAL, or of one branch of a top-level UNION (or of a UNION
an OPTIONAL holds alone), witnesses a template whose object variable
nothing else binds. A row that binds that variable is the decode of the
fact the triple matched, and a row that does not emits no intent. These
rows, which include Cypher SET and the outbound half of DETACH DELETE,
now retract without reading every slot. A UNION whose other branch binds
the object from another predicate still goes through the resolver.
…h less detail

An identical re-upsert commits nothing, except for a value stored with less
detail than it was written: an integer beyond 64 bits written with an XSD
subtype, or a dateTime or time with digits past the microsecond. Re-upserting
such a value as written commits a retraction and an assertion of the same
stored fact, and the value can be lost. This is not new in this branch. The
upsert page now says so and says how to write such values. A second ignored
U5 test pins the temporal case next to the big-integer one.

Also:
- A test pins that a refresh upsert of every language's label satisfies
  sh:uniqueLang, and that a second label in a language is still refused.
- The of_slots doc no longer claims duplicate slots are read once.
- A stale comment about SPARQL typed literals is removed.
Adds three columns over the existing people data: upsert_turtle (the
Turtle bodies upserted into a fresh ledger), upsert_turtle_replace (the
same upserts over a ledger that already holds them, so every subject's
values are read and restated), and trig_mixed (a TriG insert of each body
with its second half in a trailing graph block, the documents the
streaming Turtle parse stops on). These are the RDF-text lanes the next
commits move off the JSON-LD round trip; this records them first.
The locator finds each graph block's label, as written, and its extent by
reading tokens only. It never interprets a term, so it cannot disagree
with the parser about what the text means: a brace or the word "graph"
inside a literal is not a block. It blanks the block syntax in place, so
every segment it returns is plain Turtle at the document's own byte
offsets, and a parse error inside a block can report where it is in the
document.

It names the construct it refuses (a nested block, a directive inside a
block, a blank-node label, an unclosed block, a stray `}`) as a Turtle
parse error at the offending token. `has_graph_blocks` exposes it for
callers that must tell TriG from Turtle before choosing a lane.
`parse_rdf_text` reads Turtle and TriG with the conformant parser into one
`TemplateSink`, for every verb that writes RDF text. Plain Turtle takes a
single parse. A TriG document is parsed segment by segment in document
order from the locator's segments, each seeded with the prefix and base
declarations made before it: a redefinition applies to what follows it and
to nothing before it. A block label resolves through the parser, under the
declarations in force where the block appears; parse errors report the
document's byte offset.

The sink converts literals exactly as `FlakeSink` does (lenient: an
ill-typed lexical form is kept with its declared datatype), keeps list
positions, takes blank-node and literal `rdf:type` objects like any other
predicate's, and applies the reserved-predicate firewall. Blank-node labels
are document-scoped across default statements and blocks; anonymous nodes
get fresh `-b{N}` labels. RDF 1.2 annotations become the `f:reifies*`
bundle through `bundle_templates`, which states no `f:reifiesGraph`: the
graph scope that places a template in a named graph adds the anchor, here
through a local stand-in with the contract of the graph-scope emitter.

`<#txn-meta>` (with no base in force) becomes commit metadata with the
refusals the TriG txn-meta parser applies. `content_id` is a commutative
sum of per-statement 128-bit hashes, so it ignores statement order;
`parse_rdf_text_txn` derives the upsert blank-node skolem scope from it.
`Placement::Into` homes every statement in one graph for graph-scoped
writes and refuses a block for another graph or a mix of blocks and
default-graph statements, naming the graph.
Turtle and TriG upserts were parsed into a graph, re-encoded as JSON-LD and
expanded again, and TriG blocks went through a second parser. The round
trip refused IRIs whose scheme the compact-IRI guard took for an undefined
prefix (`tag:`, `kb:`), stored collections as unordered values, refused an
annotated `rdf:type` edge, and parsed literals more strictly than insert.

Every upsert entry point (`Fluree::upsert_turtle*`, the owned and cached
builders with or without a policy, the graph builder) now stages from
`parse_rdf_text_txn`, parsed against the state it is staged on:
`OpPlan::Rdf` carries the text through the cached-handle retry loop. The
upsert stores what a Turtle insert of the same document stores, blank
nodes are scoped to the document's content, the raw transaction is the
text, and a parse error is a Turtle parse error at the document's offset.
`Fluree::upsert_turtle` now reads TriG as the builders do.

Fuel is charged as for any upsert: the baseline plus one unit per staged
flake, reported in the tally on every builder lane.
A document the streaming Turtle parser stops on is read as TriG when the
locator finds a graph block in it (`<#txn-meta>` included), and is then
parsed once into templates, as a TriG upsert is, and staged with insert
semantics. It used to go through the phase-1 rewrite and the JSON-LD round
trip, so a TriG insert refused `tag:`/`kb:` IRIs and applied a prefix or
base redefinition to the statements before it as well as after. Plain
Turtle still streams to flakes, and a document with no graph block keeps
its Turtle error; a malformed block is reported as one, by the locator.

The routing stamp is unchanged: `turtle_insert` proceeds for plain Turtle
and falls back for TriG. The directive-order cases are the ones the
phase-1 fix pins, for both lanes.
A graph-scoped RDF body (graph sync, the graph store routes) had its TriG
blocks unwrapped and the whole text converted to JSON-LD, so it refused
`tag:`/`kb:` IRIs and stored a collection as unordered values. It is now
parsed once into templates homed on the target graph
(`Placement::Into`): a block must name the target, a body holds the
target's statements either in blocks or as default-graph statements, and
the refusals keep their messages and their 400. An empty body is refused
as before unless a sync confirms it.

Sync keeps its graph-scoped blank-node scope, and the parse labels blank
nodes as the JSON-LD conversion did, so a graph last synced through the
conversion re-syncs from the same text without a commit.
With every RDF-text lane staging from the one parse, nothing reads a TriG
document as JSON-LD any more. `convert_named_graphs_to_templates`, the
trig-meta and named-graph staging variants (`stage_transaction_with_*`,
`transact_with_*`, now `stage_transaction_tracked` for JSON-LD), the block
hashing in `upsert_payload_id`, `extract_trig_txn_meta` and
`unwrap_trig_graph_blocks` go. The TriG reader stays for bulk import only
(phase 1 plus `resolve_trig_meta`); its tests now read documents that way,
so they keep covering the import lane.

What the deleted converter pinned is pinned on the parse: a hand-written
`f:reifies*` statement in a block is refused as a transaction error, an
undefined prefix in a block is a Turtle parse error, and a stable `_:fdb-`
id in a block addresses the stored node.

`parse_to_json` is documented as lossy, and the transact and api crates
disallow it in their clippy.toml, so no transaction lane can go back to it
without saying why.
`fluree sync` converted Turtle to JSON-LD client-side, so a Turtle export
lost collection order and refused `tag:`/`kb:` IRIs before it reached the
server. RDF text is now sent as written: locally as a `GraphPayload::Rdf`,
remotely as `text/turtle` (or `application/trig` for a body with graph
blocks), and parsed once where it is staged. Only a server from before
`/sync` read RDF bodies answers 415, before staging anything; for that one
a Turtle body is converted to JSON-LD, as older CLIs did, and sent again.
TriG has no JSON-LD form for its blocks and is not converted.

TriG detection (`is_trig_body`, used by sync, `validate` and `--shacl`) is
the locator's: a token pass, so a block holding `[ … ]` or a collection is
TriG, not "not TriG" as the old content parse made it.
The harness loads named-graph data as TriG through `upsert_turtle`, and
the five eval-triple-terms tests registered as "blocked on TriG
GRAPH-block parsing" now load their data. None of them passes yet; each is
re-attributed to the failure it now reaches: graphs-1/2 to the non-
isomorphic results of the reifier model, expr-1 to the unlowered
triple-term functions, update-1 to quoted triples in SPARQL UPDATE
lowering, and update-2 to an annotation valued by a `GRAPH ?g` binding.
The harness notes no longer describe the phase-1 reader.
turtle.md describes how every transaction endpoint now reads Turtle and
TriG: directives in document order, relative IRIs against the base in
force (an error without one), literals kept as written, any IRI scheme,
collections in order, document-scoped blank-node labels, errors at the
document's offset. It narrows the #1930 limitation to bulk import, and
documents the upsert blank-node scope with a query that lists blank nodes
nothing references, pinned by a test.

The ledger-config recipe upserts as written again (tested); `fluree sync`
docs describe RDF bodies and the 415 fallback; edge-annotations no longer
lists upsert and sync among the paths that refuse an annotated
`rdf:type` edge; the compact-IRI troubleshooting entry notes that Turtle
and TriG never raise it.
The TriG insert fallback asked the locator whether the document has a
graph block and then parsed it, which located it again: two token passes
over every TriG insert. `parse_trig_txn` locates once and returns `None`
for a document with no block, so the streaming parser's Turtle error
stands as before. The routing stamp now fires once the document parses as
TriG.
Cross-ledger configuration names its model ledger by id, as an IRI
reference in the config graph's block (`f:ledger <org/governance:main>`,
the form docs/security/cross-ledger-policy.md uses), and an id with a `/`
before its `:` is a relative reference. The TriG block reader kept such a
reference as written when no `@base` was in force; the one parse refused
it, which broke every TriG cross-ledger config.

A new `ParserOptions::relative_iris` (default `Resolve`, so every other
caller stays strict) lets the RDF-text driver keep a relative reference
as written (`Verbatim`) in a labeled block's contents when no base is in
force: upsert and TriG insert as before, and sync bodies now alike. The
default graph and block labels still resolve against the base and are
refused without one, and a base in force still resolves a reference in a
block.

The parity test also syncs its fixture: sync stores what insert stores.
With no `@base` in force, a block's label keeps a relative reference
exactly as written, as the block's contents do and as the TriG reader
always did: `GRAPH <cfg/local> { … }` names the graph `cfg/local`. A
ledger id is a relative reference, so a label that names a ledger reaches
staging as written, and what it denotes is decided there with every other
write graph name. The default graph still refuses a relative reference
without a base, and a base in force still resolves label and contents.
…ne parser

The driver used to start a fresh parser for every segment of a TriG
document. Each one was seeded from a copy of every prefix declared so far,
and each block label was resolved by its own probe parse. Each segment was
also lexed a second time, after the locator had lexed the whole document to
find the blocks. A document of many small blocks paid all of it per block:
100k one-triple blocks under 20 prefixes ingested 1.6 to 1.9 times slower
than the phase-1 reader.

fluree-graph-turtle gains SegmentParser, one parser over the pieces of a
document, fed the tokens the locator lexed. The parser is generic over its
token source, a streaming lexer by default, so the lexer path compiles as
before. Declarations, the base, the term caches and the sink carry from
piece to piece as they do between statements, the relative-IRI mode is set
per piece, and error positions are the document's. The locator hands over
each segment's tokens instead of an edited copy of the text, so nothing is
lexed twice or copied per block. The driver resolves each distinct label
once per declaration epoch, which a prefix or base declaration advances,
so a label under a redefined prefix still names the new graph.
A TriG upsert leaves orphaned blank nodes in the graphs its blocks write,
where the documented default-graph query cannot see them. turtle.md adds
the same check over GRAPH ?g, and the documented-query test pins it on an
edited block.
A Turtle write keeps an ill-typed literal ("abc"^^xsd:integer) as its
lexical form with the declared datatype, and upsert and sync now store one
the way insert always did. SPARQL and JSON-LD DELETE templates refused the
same literal at lowering, so no DELETE could name that stored term.

DELETE templates now keep an ill-typed literal as that same stored form:
DELETE DATA, DELETE and DELETE WHERE templates in SPARQL, and the JSON-LD
delete. A DELETE names only what is stored, so a term that is not stored
still retracts nothing. INSERT templates, and JSON-LD VALUES, still refuse
an ill-typed literal.
upsert.md's Idempotency section says that a value stored with less detail
than it was written is the one exception to a retry-safe upsert. The
page's opening line, its behavior list and its comparison table, the
transactions README, and the WHERE/DELETE/INSERT comparison table still
called replace mode idempotent without it. They now name the exception and
point to that section, and the README's retry checklist says how to write
such values.
…SPARQL too

The routing test's doc named the SPARQL shapes but ran only JSON-LD. It now
runs SPARQL OPTIONAL and UNION twins, and a SPARQL UNION whose other branch
binds the object from another predicate, which stays on the resolver. It
collects every mismatch before asserting, so one run shows each case.
A blank-node label names one node across a TriG document, so a node
described in one graph can be referenced from another graph or from the
default graph. The per-graph orphan query checked references only within
the node's own graph and listed such a node, and the default-graph query
missed references from named graphs. Both now look for references in every
graph. A test pins a node referenced from another graph and one referenced
from a named graph, and shows the own-graph checks listing them.
…not view

A write that restates a stored value as it is stages nothing: the
retraction of the stored fact and the assertion of the same fact cancel
before the modify check, so such a write was not checked against modify
policy.

Under a non-root policy, the accumulator now reports the pairs that
cancel, and both sides of each pair whose stored fact the identity cannot
view go to the modify check. Restating a value the identity can view
still commits nothing without a check, root is unaffected, and a write
with no restated value builds no view check. This covers upsert, a DELETE
and INSERT of one value on any surface, and graph sync.

The policy docs say how restatements are checked, and that granting
modify without view does not keep a value confidential from that
identity.
@aaj3f aaj3f added bug Something isn't working as expected area:transact Staging, commit, import/bulk-import, novelty, retraction semantics labels Oct 1, 2026
@aaj3f
aaj3f marked this pull request as ready for review October 2, 2026 02:07
@aaj3f
aaj3f requested review from bplatz and zonotope October 2, 2026 02:07

@bplatz bplatz left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. Please look at the review comments before merging, in particular the two in current_facts.rs (the ghost fact and the cost on wide slots), the fluree sync fallback, and my notes on the three calls.


/// Resolve base rows plus novelty ops to the current facts, stamping
/// the graph and the origin.
fn resolve(&self, flakes: Vec<Flake>, g_sid: Option<&Sid>) -> Vec<StoredFact> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

resolve merges index rows with novelty ops by flake identity, but storage and queries merge by index key (o_type, o_key). When a novelty retraction has the same key but a different identity, the value disappears from queries, yet CurrentFacts still reports it and then cancels a later assertion of it.

Repro on an indexed ledger:

  1. Store ex:s ex:n "123456789012345678901234567890"^^xsd:integer.
  2. Run DELETE WHERE { ex:s ex:n ?o }.
  3. Upsert the same value again.

The upsert commits nothing (flake count 0, t unchanged), and the value stays gone until the next reindex. Base commits it. Reverting an indexed insert of a BIG^^xsd:nonNegativeInteger value has the same effect.

A second cause is in stage.rs: materialize_one_binding takes the witnessed retraction's datatype from dt_sids()[dt_id]. For a big integer that is the DECIMAL placeholder, so it retracts (BigInt, xsd:decimal); the query materializer uses overflow_numeric_datatype_sid() here.

Two fixes would close both: make this merge key-aware (drop an index row when a newer novelty op has the same persisted key; IndexKeys already computes the keys), and use the overflow datatype in materialize_one_binding.

/// holding the term, in index order, else the plain value: the contract
/// hydration had, under which a value asserted at N list positions loses
/// exactly one entry per distinct DELETE row.
fn match_stored<'f>(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deletes over a wide slot are quadratic. match_stored scans the slot twice per intent, the witnessed path scans it once per row, and IndexKeys::hits allocates a Vec<bool> per intent.

One subject with n values of one predicate, n = 5k / 10k / 20k:

Case Base Head
Indexed, JSON-LD delete of all n values 71 / 145 / 299 ms 0.65 / 2.41 / 9.26 s
Witnessed DELETE WHERE, ledger with any @list, indexed 78 / 149 / 309 ms 0.34 / 1.20 / 4.49 s

Witnessed deletes on a ledger with no lists stay linear. The benches use one value per slot, so they don't show this. A per-slot hash keyed by (o, dt, lowercased lang), plus a per-slot flag for whether the slot has list entries, would fix it.

let mut flake = intent.flake;
if list_index_of(&flake).is_none() {
let is_entry = |f: &StoredFact| list_index_of(&f.flake).is_some();
let entry = slot_facts

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Every identical witnessed row takes the first stored entry. So DELETE WHERE { ex:s ex:list ?o } over {"@list": ["v", "v", "w"]} leaves one "v" behind, and the same happens with ["v", {"@list": ["v"]}]. This is pre-existing on main. Since the slot is read here anyway, the k-th identical row could take the k-th stored entry.

/// [`finalize`](Self::finalize), plus one `(retraction, assertion)` pair
/// for each fact the transaction retracted and asserted again: a stored
/// fact restated as it is, which stages nothing.
pub fn finalize_with_restated(self) -> (Vec<Flake>, Vec<(Flake, Flake)>) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Call 1: ratify.

One gap remains: ADD/COPY/MOVE skip re-homed assertions that already exist in the destination (stage.rs ~1733), and they skip them without the restatement check. Under a non-root policy, whether ADD <mine> TO <hidden> succeeds then depends on whether the destination already holds the source's facts, which is the same existence signal this change closes for upsert and sync. Either apply the same rule there or document the exception.

/// so those are left out. Upsert has no WHERE, so its templates are
/// instantiated for exactly one solution, and `upsert_blank_subject` mints
/// the Sid that solution's assertions carry.
fn upsert_replacement_slots(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Call 3: ratify the (g,s,p) unit; it's the documented contract.

It does break clients in practice. On main a language-tagged upsert never retracted anything, so clients that upsert one language at a time have been accumulating languages, and after this change each upsert removes the others.

.sync_rdf(ledger, graph, text, content_type, dry_run, allow_empty)
.await
{
Err(RemoteLedgerError::UnsupportedMediaType(_)) if !*trig => {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

v4.2.0 and v4.2.1 don't answer a Turtle body on /sync with a 415. They return 400 sync accepts application/json (JSON-LD); convert Turtle payloads client-side (fluree-db-server/src/routes/transact.rs in those tags). So against the servers this fallback targets, fluree sync x.ttl now fails where it used to convert and succeed.

Falling back on that 400 as well is safe, because the server refuses before staging. The test stub answers 415, and docs/cli/sync.md and server-integration.md both describe the 415.

The fallback is also silent. A warning that the conversion is lossy (collection order, tag:/kb: IRIs) would help.

TokenKind::LBrace | TokenKind::KwGraph => {
return Err(locate_error(
tok.start as usize,
"nested graph block: TriG graph blocks cannot contain graph blocks",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

scan_block ignores a } while bracket depth is above 0, so an unclosed [ inside a block is reported at the next block. GRAPH ex:g { ex:a ex:p [ ex:q 1 } GRAPH ex:h { … } gives nested graph block, and an open [ at the end of the document gives unclosed graph block. Both are still 400s; only the message points at the wrong thing.

document mints new blank nodes; the old ones stay (see
[Stable blank-node ids](update-where-delete-insert.md#editing-blank-node-structures-stable-_fdb--ids)
for editing a stored blank node in place). For Turtle and TriG the
identity comes from the parsed statements, not their order or spelling;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The identity does depend on spelling. Each of these gives a different id, so a re-upsert mints new blank nodes:

  • 1 vs "1"^^xsd:integer
  • true vs "true"^^xsd:boolean
  • @en-US vs @en-us
  • a repeated statement
  • renaming _:b to _:c

Prefixed vs full IRIs, and 'x' vs "x", do give the same id.

t=2: DELETE { ex:alice schema:age 30 }
Result: No change (triple didn't exist)
t=2: DELETE DATA { ex:alice schema:age 30 }
Result: nothing is committed; the ledger stays at t=1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This doesn't hold for a graph the ledger doesn't have yet. DELETE DATA { GRAPH <unregistered> { … } } makes a commit with flake count 0 that advances t and registers the graph.

/// `f:reifies*` statement a transaction error. Neither reads as an engine
/// fault.
#[tokio::test]
async fn refusals_inside_a_block_are_the_user_s_errors() {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

After rebasing onto main: main's urn:default write refusals hold on these lanes (checked on a scratch merge), but main's tests only cover TriG upsert and JSON-LD sync. Twin tests for TriG insert and RDF sync/insert with GRAPH <urn:default> would pin them.

@aaj3f
aaj3f removed the request for review from zonotope October 8, 2026 13:41

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:transact Staging, commit, import/bulk-import, novelty, retraction semantics bug Something isn't working as expected

Projects

None yet

2 participants