Repository navigation
Conversation
Nothing in the suite measured the upsert wave that retracts the current values of every (graph, subject, predicate) an upsert names: insert_formats upserts only new subjects and transact_commit measures inserts. Eight scenarios: default and named graph, each unindexed and indexed; language tags plus @list; 16 predicates per subject; an indexed base under 4x unrelated novelty; and new subjects as the #1549 control. Each scenario asserts its lane in setup and its retract/assert counts after every measured upsert, so a drifting fixture fails instead of timing the wrong path.
…tion LRU alone binary_range_eq_v3 sent every probe through the cross-call translation LRU when the overlay reported a content version, and NoOverlay reports one. Staging reads the persisted index through NoOverlay once per retraction group, and the LRU key includes to_t, so each commit's first such probe per (graph, index order) inserted an empty entry into an eight-slot cache and could evict a live query's novelty translation. The translation step moves into probe_overlay, which returns an empty translation for an effectively empty overlay before touching the cache.
A retraction staged for commit must name a fact that is currently stored. Today staging rebuilds retractions from query bindings, templates and edge keys, which drops language tags and list positions and invents retractions of facts that do not exist. This adds the read those sites will share: per (graph, subject, predicate) slot, one seek into the persisted index through NoOverlay plus a Sid-space novelty seek bounded by the slot's sentinels, lifecycle-resolved, graph stamped. A graph whose novelty is small next to the slots asked of it is walked once instead. StoredFact (private fields, built only here) and Retraction (private field, built from StoredFact::retract) carry the invariant in the types; the next commits route staging through them.
The upsert wave queried each (graph, subject, predicate) with execute_pattern and rebuilt every retraction from the ?o binding with m: None. A language-tagged value was retracted without its tag and a list entry without its position, so neither retraction matched anything: the old values stayed, and an identical refresh of such values committed two flakes every time (#1976). A named graph without a range provider also walked the graph's whole novelty once per slot. The wave now asks CurrentFacts for each slot's current facts and retracts them as stored. The #1549 absence pre-check and its counters are unchanged; pattern_queries now counts slot reads. binding_to_flake_object, query_novelty_for_graph and the per-group materializer are gone. The replacement unit is one function, upsert_replacement_slots: every current value of each (graph, subject, predicate) the payload names, all languages and list positions together. Blank subjects join the wave when the skolem scope is deterministic (the upsert payload scope), so an identical re-upsert of a document with blank nodes is now a no-op too.
…rrentFacts The whole-graph scan moves from scan_graph_flakes to CurrentFacts::whole_graph, which returns StoredFacts with the graph stamped. Graph sync, CLEAR/DROP and the COPY/MOVE destination and source clears now build their retractions with StoredFact::retract instead of flipping op and t on scanned flakes. Behavior is unchanged: the scan, its limit, the stamping and the policy model are the same.
A DELETE template instantiated into a retraction of a fact the ledger does
not hold used to be committed anyway. DELETE DATA of an absent triple
committed a new t with one retraction, which docs/transactions/retractions.md
says is a no-op. Worse, DELETE {x} INSERT {x} WHERE {} with x absent
committed nothing and never wrote x: the phantom retraction cancelled the
insert in the accumulator, against SPARQL 1.1 Update's (DS - D) ∪ I. Both
twins (SPARQL and JSON-LD) behaved the same.
Every DELETE instantiation is now a RetractIntent, and the accumulator
accepts only Retraction, which CurrentFacts produces:
- An intent whose template a WHERE triple witnesses (delete_witness: same
subject, predicate and object variable, the object bound nowhere else,
the same graph) is the decode of a fact the WHERE just matched. On a
ledger that holds no @list position (ListFree) it retracts as decoded with
no read, which keeps bulk filtered deletes at zero lookups. On a
list-bearing ledger it reads its slot for the list position, under the
rule hydration had: the first stored list entry with the same term.
- Every other intent (template constants, DELETE DATA, values bound by
another pattern) is matched against the stored facts of its slot: exact
datatype (#1990's rule, one function), case-insensitive language tag, the
named list position or else the first list entry, else the plain value.
An intent that names no stored fact stages nothing.
hydrate_list_index_meta_for_retractions is gone, and so is the unused
generate::cancellation module, which pushed arbitrary flakes as
retractions. The delete_witnessed routing stamp reports which lane ran.
The it_fast_stats_1391 fixture relied on a DELETE writing a no-op
retraction; it now hand-builds that commit, the shape legacy and replayed
commits still carry.
…r graph The cascade that follows a retracted edge to its annotations rebuilt each f:reifies* bundle from the decoded EdgeKey, and read annotation subjects with range_with_overlay(&novelty). Its body-cleanup pass (a by-id annotation delete in LPG mode) did not stamp the graph on index-resident rows, so the body's retraction landed in the default graph and the body survived in its named graph. And nothing retracted the bundle's f:reifiesGraph anchor, which the by-id delete does not name, so it stayed behind describing nothing. All three passes now read annotation subjects with CurrentFacts::of_subject, graph stamped, and retract bundles and bodies as stored. A complete by-id bundle retract also retracts every other stored slot of that bundle, unless the transaction re-points the reifier.
A take-source merge or revert, and a take-branch rebase, retract the losing side's current values on each conflicting (subject, predicate, graph). current_asserted_for_key read them with a range read and kept the rows whose graph equalled the key's, but index-resident rows come back without their graph, so in a named graph every indexed value was dropped: the take-source merge left the target's value next to the source's, and merge_preview showed no target value. Novelty-resident values and the default graph were unaffected. current_asserted_for_key now reads the slot through CurrentFacts, which stamps the graph, and the retractions come from StoredFact::retract.
Upserts written before the retraction resolver retracted "a"@en as a tagless "a"^^rdf:langString. That retraction matched nothing and is inert for reads and indexing, but a revert inverts every flake of the reverted commits, and inverting it asserted "a" as an rdf:langString with no language tag (LANG ""), a literal that cannot exist. The undo fold now skips any retraction whose shape no stored fact has: an rdf:langString without a tag, or a tag on another datatype. It logs how many it skipped. Legacy retractions of list entries without their position are well-formed plain literals, so no shape check can catch them; telling those apart needs a historical stored-fact read and is left as a follow-up.
…only stored facts upsert.md: the replacement unit is the whole (graph, subject, predicate), all languages and list positions together, and an identical refresh commits nothing, blank nodes included. The "Empty Replacement" and comparison-table text said an upsert removes predicates the payload does not name; it never did, and the page's own summary said so. retractions.md: a DELETE of a triple that is not stored commits nothing, names terms exactly (datatype exact, language tag case-insensitive), and a DELETE/INSERT of an absent triple inserts it. update-where-delete-insert.md: the same, and the comparison table's upsert granularity.
With #1988, SPARQL UPDATE lowers `"1"^^xsd:int`, `xsd:long` and `xsd:dateTime` literals as JSON-LD does, so a DELETE DATA intent names the stored term exactly and the retraction resolver deletes it, over novelty and over an index. Before #1988 the intent named a string, and the delete removed nothing.
The persisted index keeps less than some terms carry. An integer too large for i64 is keyed by value alone (NUM_BIG), with no XSD subtype, and a dateTime or time is keyed to the microsecond. So an indexed fact's decode can differ from the term that names it. The index decode says xsd:integer where a big xsd:nonNegativeInteger was written, a query row says xsd:decimal, and a nanosecond literal decodes to its microsecond. Matching by term alone therefore missed these facts on an indexed ledger. DELETE DATA, a JSON-LD delete or a constant DELETE template naming a big integer under its subtype committed nothing. So did a DELETE row whose value an OPTIONAL, UNION or BIND bound. Cypher SET kept the old big value next to the new one, DETACH DELETE left it behind, and a sub-microsecond dateTime or time survived DELETE DATA. main deleted each of these, because its retraction carried the same index key. A fact read from the index is now also named by its index key: the intent's object encoded as the index keys it (o_type + o_key), compared with the fact's. Terms with one key are one fact to the index, which merges every retraction by that key. The resolver still retracts the fact as read, so every retraction names a stored fact. Term identity stays first, and novelty facts keep it alone. The key comparison runs only when no stored term matches, so a DELETE of a present term costs what it did. The witnessed-row list lookup uses the same rule. BinaryRangeProvider::persisted_object_key exposes the encoding the overlay merge already applies to retractions.
A DELETE intent that names no stored fact stages nothing. Under a policy, that also kept it from the modify-policy check, which only saw the staged flakes. A delete the identity may not perform was then refused when its target was stored and committed nothing when it was not. Intents that name no stored fact now go through the modify-policy check with the staged flakes, and are still dropped from the commit. A delete the identity may not perform is refused the same way whether or not its target is stored, as it was before intents were resolved against storage. A permitted delete of an absent value still commits nothing. Only a non-root policy context collects the unmatched intents.
…L and UNION rows included Two changes to the rule that lets a DELETE row retract without a read. Graphs compare as ledger graph ids. The rule used to compare the WHERE default's IRI with the template's IRI, and a dataset alias by name. It now takes the graph ids the WHERE dataset actually reads: its default graph when that is one ledger graph (the ledger's own with no USING, WITH or from, else the one graph those name resolves to, where the ledger's own address is its default graph), and each GRAPH <name>'s graph. It compares them with the graph the ledger has under the template's IRI. A name the WHERE resolves to one graph and a template writes to another is never a witness. USING or WITH naming the ledger's own address reads the default graph, while a GRAPH <address> template writes a graph registered under that name. A test covers a ledger holding such a graph. OPTIONAL and UNION rows can witness. A triple that is the only pattern of a top-level OPTIONAL, or of one branch of a top-level UNION (or of a UNION an OPTIONAL holds alone), witnesses a template whose object variable nothing else binds. A row that binds that variable is the decode of the fact the triple matched, and a row that does not emits no intent. These rows, which include Cypher SET and the outbound half of DETACH DELETE, now retract without reading every slot. A UNION whose other branch binds the object from another predicate still goes through the resolver.
…h less detail An identical re-upsert commits nothing, except for a value stored with less detail than it was written: an integer beyond 64 bits written with an XSD subtype, or a dateTime or time with digits past the microsecond. Re-upserting such a value as written commits a retraction and an assertion of the same stored fact, and the value can be lost. This is not new in this branch. The upsert page now says so and says how to write such values. A second ignored U5 test pins the temporal case next to the big-integer one. Also: - A test pins that a refresh upsert of every language's label satisfies sh:uniqueLang, and that a second label in a language is still refused. - The of_slots doc no longer claims duplicate slots are read once. - A stale comment about SPARQL typed literals is removed.
Adds three columns over the existing people data: upsert_turtle (the Turtle bodies upserted into a fresh ledger), upsert_turtle_replace (the same upserts over a ledger that already holds them, so every subject's values are read and restated), and trig_mixed (a TriG insert of each body with its second half in a trailing graph block, the documents the streaming Turtle parse stops on). These are the RDF-text lanes the next commits move off the JSON-LD round trip; this records them first.
The locator finds each graph block's label, as written, and its extent by reading tokens only. It never interprets a term, so it cannot disagree with the parser about what the text means: a brace or the word "graph" inside a literal is not a block. It blanks the block syntax in place, so every segment it returns is plain Turtle at the document's own byte offsets, and a parse error inside a block can report where it is in the document. It names the construct it refuses (a nested block, a directive inside a block, a blank-node label, an unclosed block, a stray `}`) as a Turtle parse error at the offending token. `has_graph_blocks` exposes it for callers that must tell TriG from Turtle before choosing a lane.
`parse_rdf_text` reads Turtle and TriG with the conformant parser into one
`TemplateSink`, for every verb that writes RDF text. Plain Turtle takes a
single parse. A TriG document is parsed segment by segment in document
order from the locator's segments, each seeded with the prefix and base
declarations made before it: a redefinition applies to what follows it and
to nothing before it. A block label resolves through the parser, under the
declarations in force where the block appears; parse errors report the
document's byte offset.
The sink converts literals exactly as `FlakeSink` does (lenient: an
ill-typed lexical form is kept with its declared datatype), keeps list
positions, takes blank-node and literal `rdf:type` objects like any other
predicate's, and applies the reserved-predicate firewall. Blank-node labels
are document-scoped across default statements and blocks; anonymous nodes
get fresh `-b{N}` labels. RDF 1.2 annotations become the `f:reifies*`
bundle through `bundle_templates`, which states no `f:reifiesGraph`: the
graph scope that places a template in a named graph adds the anchor, here
through a local stand-in with the contract of the graph-scope emitter.
`<#txn-meta>` (with no base in force) becomes commit metadata with the
refusals the TriG txn-meta parser applies. `content_id` is a commutative
sum of per-statement 128-bit hashes, so it ignores statement order;
`parse_rdf_text_txn` derives the upsert blank-node skolem scope from it.
`Placement::Into` homes every statement in one graph for graph-scoped
writes and refuses a block for another graph or a mix of blocks and
default-graph statements, naming the graph.
Turtle and TriG upserts were parsed into a graph, re-encoded as JSON-LD and expanded again, and TriG blocks went through a second parser. The round trip refused IRIs whose scheme the compact-IRI guard took for an undefined prefix (`tag:`, `kb:`), stored collections as unordered values, refused an annotated `rdf:type` edge, and parsed literals more strictly than insert. Every upsert entry point (`Fluree::upsert_turtle*`, the owned and cached builders with or without a policy, the graph builder) now stages from `parse_rdf_text_txn`, parsed against the state it is staged on: `OpPlan::Rdf` carries the text through the cached-handle retry loop. The upsert stores what a Turtle insert of the same document stores, blank nodes are scoped to the document's content, the raw transaction is the text, and a parse error is a Turtle parse error at the document's offset. `Fluree::upsert_turtle` now reads TriG as the builders do. Fuel is charged as for any upsert: the baseline plus one unit per staged flake, reported in the tally on every builder lane.
A document the streaming Turtle parser stops on is read as TriG when the locator finds a graph block in it (`<#txn-meta>` included), and is then parsed once into templates, as a TriG upsert is, and staged with insert semantics. It used to go through the phase-1 rewrite and the JSON-LD round trip, so a TriG insert refused `tag:`/`kb:` IRIs and applied a prefix or base redefinition to the statements before it as well as after. Plain Turtle still streams to flakes, and a document with no graph block keeps its Turtle error; a malformed block is reported as one, by the locator. The routing stamp is unchanged: `turtle_insert` proceeds for plain Turtle and falls back for TriG. The directive-order cases are the ones the phase-1 fix pins, for both lanes.
A graph-scoped RDF body (graph sync, the graph store routes) had its TriG blocks unwrapped and the whole text converted to JSON-LD, so it refused `tag:`/`kb:` IRIs and stored a collection as unordered values. It is now parsed once into templates homed on the target graph (`Placement::Into`): a block must name the target, a body holds the target's statements either in blocks or as default-graph statements, and the refusals keep their messages and their 400. An empty body is refused as before unless a sync confirms it. Sync keeps its graph-scoped blank-node scope, and the parse labels blank nodes as the JSON-LD conversion did, so a graph last synced through the conversion re-syncs from the same text without a commit.
With every RDF-text lane staging from the one parse, nothing reads a TriG document as JSON-LD any more. `convert_named_graphs_to_templates`, the trig-meta and named-graph staging variants (`stage_transaction_with_*`, `transact_with_*`, now `stage_transaction_tracked` for JSON-LD), the block hashing in `upsert_payload_id`, `extract_trig_txn_meta` and `unwrap_trig_graph_blocks` go. The TriG reader stays for bulk import only (phase 1 plus `resolve_trig_meta`); its tests now read documents that way, so they keep covering the import lane. What the deleted converter pinned is pinned on the parse: a hand-written `f:reifies*` statement in a block is refused as a transaction error, an undefined prefix in a block is a Turtle parse error, and a stable `_:fdb-` id in a block addresses the stored node. `parse_to_json` is documented as lossy, and the transact and api crates disallow it in their clippy.toml, so no transaction lane can go back to it without saying why.
`fluree sync` converted Turtle to JSON-LD client-side, so a Turtle export lost collection order and refused `tag:`/`kb:` IRIs before it reached the server. RDF text is now sent as written: locally as a `GraphPayload::Rdf`, remotely as `text/turtle` (or `application/trig` for a body with graph blocks), and parsed once where it is staged. Only a server from before `/sync` read RDF bodies answers 415, before staging anything; for that one a Turtle body is converted to JSON-LD, as older CLIs did, and sent again. TriG has no JSON-LD form for its blocks and is not converted. TriG detection (`is_trig_body`, used by sync, `validate` and `--shacl`) is the locator's: a token pass, so a block holding `[ … ]` or a collection is TriG, not "not TriG" as the old content parse made it.
The harness loads named-graph data as TriG through `upsert_turtle`, and the five eval-triple-terms tests registered as "blocked on TriG GRAPH-block parsing" now load their data. None of them passes yet; each is re-attributed to the failure it now reaches: graphs-1/2 to the non- isomorphic results of the reifier model, expr-1 to the unlowered triple-term functions, update-1 to quoted triples in SPARQL UPDATE lowering, and update-2 to an annotation valued by a `GRAPH ?g` binding. The harness notes no longer describe the phase-1 reader.
turtle.md describes how every transaction endpoint now reads Turtle and TriG: directives in document order, relative IRIs against the base in force (an error without one), literals kept as written, any IRI scheme, collections in order, document-scoped blank-node labels, errors at the document's offset. It narrows the #1930 limitation to bulk import, and documents the upsert blank-node scope with a query that lists blank nodes nothing references, pinned by a test. The ledger-config recipe upserts as written again (tested); `fluree sync` docs describe RDF bodies and the 415 fallback; edge-annotations no longer lists upsert and sync among the paths that refuse an annotated `rdf:type` edge; the compact-IRI troubleshooting entry notes that Turtle and TriG never raise it.
The TriG insert fallback asked the locator whether the document has a graph block and then parsed it, which located it again: two token passes over every TriG insert. `parse_trig_txn` locates once and returns `None` for a document with no block, so the streaming parser's Turtle error stands as before. The routing stamp now fires once the document parses as TriG.
Cross-ledger configuration names its model ledger by id, as an IRI reference in the config graph's block (`f:ledger <org/governance:main>`, the form docs/security/cross-ledger-policy.md uses), and an id with a `/` before its `:` is a relative reference. The TriG block reader kept such a reference as written when no `@base` was in force; the one parse refused it, which broke every TriG cross-ledger config. A new `ParserOptions::relative_iris` (default `Resolve`, so every other caller stays strict) lets the RDF-text driver keep a relative reference as written (`Verbatim`) in a labeled block's contents when no base is in force: upsert and TriG insert as before, and sync bodies now alike. The default graph and block labels still resolve against the base and are refused without one, and a base in force still resolves a reference in a block. The parity test also syncs its fixture: sync stores what insert stores.
With no `@base` in force, a block's label keeps a relative reference
exactly as written, as the block's contents do and as the TriG reader
always did: `GRAPH <cfg/local> { … }` names the graph `cfg/local`. A
ledger id is a relative reference, so a label that names a ledger reaches
staging as written, and what it denotes is decided there with every other
write graph name. The default graph still refuses a relative reference
without a base, and a base in force still resolves label and contents.
…ne parser The driver used to start a fresh parser for every segment of a TriG document. Each one was seeded from a copy of every prefix declared so far, and each block label was resolved by its own probe parse. Each segment was also lexed a second time, after the locator had lexed the whole document to find the blocks. A document of many small blocks paid all of it per block: 100k one-triple blocks under 20 prefixes ingested 1.6 to 1.9 times slower than the phase-1 reader. fluree-graph-turtle gains SegmentParser, one parser over the pieces of a document, fed the tokens the locator lexed. The parser is generic over its token source, a streaming lexer by default, so the lexer path compiles as before. Declarations, the base, the term caches and the sink carry from piece to piece as they do between statements, the relative-IRI mode is set per piece, and error positions are the document's. The locator hands over each segment's tokens instead of an edited copy of the text, so nothing is lexed twice or copied per block. The driver resolves each distinct label once per declaration epoch, which a prefix or base declaration advances, so a label under a redefined prefix still names the new graph.
A TriG upsert leaves orphaned blank nodes in the graphs its blocks write, where the documented default-graph query cannot see them. turtle.md adds the same check over GRAPH ?g, and the documented-query test pins it on an edited block.
A Turtle write keeps an ill-typed literal ("abc"^^xsd:integer) as its
lexical form with the declared datatype, and upsert and sync now store one
the way insert always did. SPARQL and JSON-LD DELETE templates refused the
same literal at lowering, so no DELETE could name that stored term.
DELETE templates now keep an ill-typed literal as that same stored form:
DELETE DATA, DELETE and DELETE WHERE templates in SPARQL, and the JSON-LD
delete. A DELETE names only what is stored, so a term that is not stored
still retracts nothing. INSERT templates, and JSON-LD VALUES, still refuse
an ill-typed literal.
upsert.md's Idempotency section says that a value stored with less detail than it was written is the one exception to a retry-safe upsert. The page's opening line, its behavior list and its comparison table, the transactions README, and the WHERE/DELETE/INSERT comparison table still called replace mode idempotent without it. They now name the exception and point to that section, and the README's retry checklist says how to write such values.
…SPARQL too The routing test's doc named the SPARQL shapes but ran only JSON-LD. It now runs SPARQL OPTIONAL and UNION twins, and a SPARQL UNION whose other branch binds the object from another predicate, which stays on the resolver. It collects every mismatch before asserting, so one run shows each case.
A blank-node label names one node across a TriG document, so a node described in one graph can be referenced from another graph or from the default graph. The per-graph orphan query checked references only within the node's own graph and listed such a node, and the default-graph query missed references from named graphs. Both now look for references in every graph. A test pins a node referenced from another graph and one referenced from a named graph, and shows the own-graph checks listing them.
…not view A write that restates a stored value as it is stages nothing: the retraction of the stored fact and the assertion of the same fact cancel before the modify check, so such a write was not checked against modify policy. Under a non-root policy, the accumulator now reports the pairs that cancel, and both sides of each pair whose stored fact the identity cannot view go to the modify check. Restating a value the identity can view still commits nothing without a check, root is unaffected, and a write with no restated value builds no view check. This covers upsert, a DELETE and INSERT of one value on any surface, and graph sync. The policy docs say how restatements are checked, and that granting modify without view does not keep a value confidential from that identity.
bplatz
left a comment
There was a problem hiding this comment.
Approving. Please look at the review comments before merging, in particular the two in current_facts.rs (the ghost fact and the cost on wide slots), the fluree sync fallback, and my notes on the three calls.
|
|
||
| /// Resolve base rows plus novelty ops to the current facts, stamping | ||
| /// the graph and the origin. | ||
| fn resolve(&self, flakes: Vec<Flake>, g_sid: Option<&Sid>) -> Vec<StoredFact> { |
There was a problem hiding this comment.
resolve merges index rows with novelty ops by flake identity, but storage and queries merge by index key (o_type, o_key). When a novelty retraction has the same key but a different identity, the value disappears from queries, yet CurrentFacts still reports it and then cancels a later assertion of it.
Repro on an indexed ledger:
- Store
ex:s ex:n "123456789012345678901234567890"^^xsd:integer. - Run
DELETE WHERE { ex:s ex:n ?o }. - Upsert the same value again.
The upsert commits nothing (flake count 0, t unchanged), and the value stays gone until the next reindex. Base commits it. Reverting an indexed insert of a BIG^^xsd:nonNegativeInteger value has the same effect.
A second cause is in stage.rs: materialize_one_binding takes the witnessed retraction's datatype from dt_sids()[dt_id]. For a big integer that is the DECIMAL placeholder, so it retracts (BigInt, xsd:decimal); the query materializer uses overflow_numeric_datatype_sid() here.
Two fixes would close both: make this merge key-aware (drop an index row when a newer novelty op has the same persisted key; IndexKeys already computes the keys), and use the overflow datatype in materialize_one_binding.
| /// holding the term, in index order, else the plain value: the contract | ||
| /// hydration had, under which a value asserted at N list positions loses | ||
| /// exactly one entry per distinct DELETE row. | ||
| fn match_stored<'f>( |
There was a problem hiding this comment.
Deletes over a wide slot are quadratic. match_stored scans the slot twice per intent, the witnessed path scans it once per row, and IndexKeys::hits allocates a Vec<bool> per intent.
One subject with n values of one predicate, n = 5k / 10k / 20k:
| Case | Base | Head |
|---|---|---|
| Indexed, JSON-LD delete of all n values | 71 / 145 / 299 ms | 0.65 / 2.41 / 9.26 s |
Witnessed DELETE WHERE, ledger with any @list, indexed |
78 / 149 / 309 ms | 0.34 / 1.20 / 4.49 s |
Witnessed deletes on a ledger with no lists stay linear. The benches use one value per slot, so they don't show this. A per-slot hash keyed by (o, dt, lowercased lang), plus a per-slot flag for whether the slot has list entries, would fix it.
| let mut flake = intent.flake; | ||
| if list_index_of(&flake).is_none() { | ||
| let is_entry = |f: &StoredFact| list_index_of(&f.flake).is_some(); | ||
| let entry = slot_facts |
There was a problem hiding this comment.
Every identical witnessed row takes the first stored entry. So DELETE WHERE { ex:s ex:list ?o } over {"@list": ["v", "v", "w"]} leaves one "v" behind, and the same happens with ["v", {"@list": ["v"]}]. This is pre-existing on main. Since the slot is read here anyway, the k-th identical row could take the k-th stored entry.
| /// [`finalize`](Self::finalize), plus one `(retraction, assertion)` pair | ||
| /// for each fact the transaction retracted and asserted again: a stored | ||
| /// fact restated as it is, which stages nothing. | ||
| pub fn finalize_with_restated(self) -> (Vec<Flake>, Vec<(Flake, Flake)>) { |
There was a problem hiding this comment.
Call 1: ratify.
One gap remains: ADD/COPY/MOVE skip re-homed assertions that already exist in the destination (stage.rs ~1733), and they skip them without the restatement check. Under a non-root policy, whether ADD <mine> TO <hidden> succeeds then depends on whether the destination already holds the source's facts, which is the same existence signal this change closes for upsert and sync. Either apply the same rule there or document the exception.
| /// so those are left out. Upsert has no WHERE, so its templates are | ||
| /// instantiated for exactly one solution, and `upsert_blank_subject` mints | ||
| /// the Sid that solution's assertions carry. | ||
| fn upsert_replacement_slots( |
There was a problem hiding this comment.
Call 3: ratify the (g,s,p) unit; it's the documented contract.
It does break clients in practice. On main a language-tagged upsert never retracted anything, so clients that upsert one language at a time have been accumulating languages, and after this change each upsert removes the others.
| .sync_rdf(ledger, graph, text, content_type, dry_run, allow_empty) | ||
| .await | ||
| { | ||
| Err(RemoteLedgerError::UnsupportedMediaType(_)) if !*trig => { |
There was a problem hiding this comment.
v4.2.0 and v4.2.1 don't answer a Turtle body on /sync with a 415. They return 400 sync accepts application/json (JSON-LD); convert Turtle payloads client-side (fluree-db-server/src/routes/transact.rs in those tags). So against the servers this fallback targets, fluree sync x.ttl now fails where it used to convert and succeed.
Falling back on that 400 as well is safe, because the server refuses before staging. The test stub answers 415, and docs/cli/sync.md and server-integration.md both describe the 415.
The fallback is also silent. A warning that the conversion is lossy (collection order, tag:/kb: IRIs) would help.
| TokenKind::LBrace | TokenKind::KwGraph => { | ||
| return Err(locate_error( | ||
| tok.start as usize, | ||
| "nested graph block: TriG graph blocks cannot contain graph blocks", |
There was a problem hiding this comment.
scan_block ignores a } while bracket depth is above 0, so an unclosed [ inside a block is reported at the next block. GRAPH ex:g { ex:a ex:p [ ex:q 1 } GRAPH ex:h { … } gives nested graph block, and an open [ at the end of the document gives unclosed graph block. Both are still 400s; only the message points at the wrong thing.
| document mints new blank nodes; the old ones stay (see | ||
| [Stable blank-node ids](update-where-delete-insert.md#editing-blank-node-structures-stable-_fdb--ids) | ||
| for editing a stored blank node in place). For Turtle and TriG the | ||
| identity comes from the parsed statements, not their order or spelling; |
There was a problem hiding this comment.
The identity does depend on spelling. Each of these gives a different id, so a re-upsert mints new blank nodes:
1vs"1"^^xsd:integertruevs"true"^^xsd:boolean@en-USvs@en-us- a repeated statement
- renaming
_:bto_:c
Prefixed vs full IRIs, and 'x' vs "x", do give the same id.
| t=2: DELETE { ex:alice schema:age 30 } | ||
| Result: No change (triple didn't exist) | ||
| t=2: DELETE DATA { ex:alice schema:age 30 } | ||
| Result: nothing is committed; the ledger stays at t=1 |
There was a problem hiding this comment.
This doesn't hold for a graph the ledger doesn't have yet. DELETE DATA { GRAPH <unregistered> { … } } makes a commit with flake count 0 that advances t and registers the graph.
| /// `f:reifies*` statement a transaction error. Neither reads as an engine | ||
| /// fault. | ||
| #[tokio::test] | ||
| async fn refusals_inside_a_block_are_the_user_s_errors() { |
There was a problem hiding this comment.
After rebasing onto main: main's urn:default write refusals hold on these lanes (checked on a scratch merge), but main's tests only cover TriG upsert and JSON-LD sync. Twin tests for TriG insert and RDF sync/insert with GRAPH <urn:default> would pin them.
This is more than #1976 needs on its own, deliberately. Our write bugs have had two recurring shapes: deletes built from the request instead of from what's stored, which silently retracted nothing for some lanes and datatypes (#1976), and Turtle or TriG text parsed more than once, by paths that disagreed (#1977, #1511, #1930). So rather than patch each lane, a retraction can now only name a fact read from storage, through one resolver, and every transaction lane parses Turtle and TriG once, with one parser, so there's one place to reason about and tune. Fwiw, it's one of five PRs taking this approach, with #2006, #2007, #2009 and #2010.
Fixes #1976
Fixes #1977
Fixes #1511
Partially addresses #1930. What remains: bulk import (
fluree create --from x.trig, the server's source import) still reads block contents with the phase-1 parser. That moves to the same driver in a follow-up (listed below).Partially addresses #1849. What remains: the TriG insert fallback now parses once into templates, but it is not yet a flake-level path like plain Turtle insert.
Partially addresses #1979, item 3 only, on the transaction lanes. Blank nodes inside
GRAPHblocks are now accepted rather than refused with a token-level error, and a nested block is refused asnested graph block: …. Bulk import keeps the old messages until it moves to the same driver. Items 1–2 are not in this PR.This PR is two write-fidelity fixes, in two commit groups. A1 makes every retraction name a fact that is actually stored — read from storage through one resolver, instead of rebuilt from query bindings — and A2 makes every transaction lane that takes Turtle or TriG parse it once, with the conformant Turtle parser, instead of round-tripping it through JSON-LD. Both change behavior (the table below is worth a careful read), and there are three calls in here I'd like a reviewer to explicitly ratify before this merges.
Three calls to ratify
1. Modify policy on deletes and on restated values — a product-behavior decision. Modify policy is evaluated for every delete intent, whether or not it matches a stored fact, so a delete the identity may not perform is refused with the same answer regardless of what is stored (unchanged from before this PR). The alternative would be resolving deletes only against facts the identity can view, which would stop a modify-but-not-view identity from deleting values it can't see. As maintainer I agree with this posture, but it asserts a particular perspective on product behavior.
The same rule now covers a write that restates a stored value as it is: an upsert of an unchanged value, a DELETE and INSERT of the same value, or a graph sync that keeps a value. Such a write stages nothing, because the retraction of the stored value and the assertion of the same value cancel.
Behavior change from
main. Under a non-root policy, when the identity cannot view the restated value, both sides are now checked against modify policy, and the write is refused if the identity may not modify the value.main, a write restating a stored value was not checked against modify policy: plain and list restatements committed as no-ops.mainrefused it, because the retraction it built carried no language tag and named nothing stored.Cost. Only under a non-root policy, and only for restated values. Each restated value gets one view check. A value the identity cannot view also goes through the modify check, both sides.
Measured with dev-fast builds on 2,000 subjects of five values each. The base is this branch just before the check (
11b5872a6). Each table is 6 interleaved rounds of 15 runs; median of the round medians, range in brackets.Upserts that restate all 10,000 values:
Writes that restate nothing, measured in an earlier run with the first policy:
That is about 0.2 µs per restated value the identity can view (the view check), and about 0.33 µs per value it cannot view (the view check and the modify check). A policy whose view or modify rules run queries or match on classes costs more per value, since each restated value runs them. Writes that restate nothing pay nothing measurable, and neither does root.
Docs. The policy docs now say how restatements are checked. They also say that granting modify without view does not keep a value confidential from that identity, since it can tell whether a given value is stored by writing it, and that such policies should be configured with care.
Please ratify or push back on both parts.
2. One of the tests from #1987 (merged) changes:
DELETE DATAof an ill-typed literal that isn't stored now commits nothing instead of erroring.it_typed_literal_index::sparql_update_rejects_ill_typed_literalasserted thatDELETE DATArefuses an ill-typedxsd:datethe wayINSERT DATAdoes. A DELETE now names such a literal as stored (a Turtle write keeps it as its lexical form with the declared datatype), so the test is nowsparql_insert_data_rejects_an_ill_typed_literal:INSERT DATAstill refuses the literal and names the datatype, andDELETE DATAof one that is not stored commits nothing. Please confirm this is the intended reading of #1987.3. Upsert replaces every value of a (graph, subject, property), across languages. That's the documented upsert semantics, and it's what the branch implements: the replacement unit is every value of
(g,s,p), all languages and list positions, so a partial-language upsert (onlyprefLabel@fr) now removes the other languages' values, where it used to accumulate them. The alternative is per-(s, p, lang) replacement, which would leave the other languages alone. The replacement unit lives in one function,upsert_replacement_slots. Please ratify or push back.What changed
A1: a retraction names a stored fact. Upsert, DELETE, graph sync, CLEAR/COPY/MOVE, the annotation cascade, merge/rebase conflict retractions and revert now all read the facts they retract from storage, through one resolver (
CurrentFacts), and retract those facts as stored — graph, datatype, language tag and list position included. A retraction can only be built from a stored fact (StoredFact::retract→Retraction), so an intent that names nothing stored stages nothing. The old behaviors this replaces:"a"@enand list entries were never removed and an identical refresh committed phantom retractions (upsertdoesn't retract language-tagged (or@list) values: the retraction flakes are built withm: None#1976).DELETE DATAof an absent triple committed a phantom retraction.DELETE { x } INSERT { x }withxabsent lost the insert.A fact read from the persisted index is also named by its index key — which keeps no XSD subtype for an integer beyond 64 bits, and keeps a
dateTimeortimeto the microsecond — so a DELETE naming such a value still deletes it. A DELETE row whose value an OPTIONAL or a single UNION branch bound is the decode of a stored fact, as a plain WHERE row is, and retracts without a read; that covers CypherSETandDETACH DELETE. Modify policy judges every delete a transaction asks for, whether or not it names a stored fact, and judges a stored value a write restates when the identity cannot view it (call 1 above).A2: RDF text is parsed once. Every transaction lane that takes Turtle or TriG — upsert, the TriG insert fallback, graph sync and the Graph Store routes, and
fluree sync— now reads it once with the conformant Turtle parser, into oneTemplateSink. It used to go through a JSON-LD round trip plus a second, smaller TriG block parser, and that round trip:tag:,kb:) (Turtleupsert(andsync) rejects absolute IRIs whose scheme isn't on the JSON-LD allowlist:<kb:n#x>→ "Unresolved compact IRI" #1977);[ … ]and( … )insideGRAPHblocks (TriG GRAPH blocks reject anonymous blank nodes [ … ] on /upsert and import #1930, trig upsert with anonymous blank nodes causes parse error #1511);rdf:typeobject as an IRI named by its label;A token-only locator finds graph blocks without interpreting a term, and one parser then reads the segments from the locator's tokens in document order — so the document is lexed once, prefix and base declarations carry forward, and one sink per document keeps blank-node labels document-scoped. Nothing converts RDF text to JSON-LD on a transaction lane any more:
parse_to_jsonis documented as lossy and disallowed in thefluree-db-transact,fluree-db-apiandfluree-db-cliclippy.toml.Behavior changes
DELETE DATA/ JSON-LD delete / CypherREMOVEof an absent factt+1tunchanged, and a mixed transaction drops only the phantomDELETE {x} INSERT {x}withxabsentxinsertedxsd:intdoes not remove a storedxsd:integer). Once a fact is indexed, the match is against what the index stores: an integer beyond 64 bits keeps no XSD subtype and adateTime/timekeeps microseconds, so such a fact is named by the value under any integer datatype, or to the microsecond, as onmain. A DELETE naming a big integer under a different subtype is therefore a no-op while the fact is in novelty, and deletes it once indexed: the same outcomes as onmain(s,p)as the binding rebuilt them: tags and list positions missed, stale values kept(g,s,p): all languages and list positions. A partial-language upsert (onlyprefLabel@fr) now removes the other languages' values, where it used to accumulate them. (Call 3 above.)rdf:langString"abc"^^xsd:integer) on upsert/syncDELETE DATAnow commits nothing instead of erroring, since it names nothing stored. (Call 2 above.)GRAPHblock (its label or its contents), no@basef:ledger <org/governance:main>). Sync and Graph Store block bodies now match insert/upsert. The default graph is unchanged (refused without a base){ … }Parse error: …Turtle parse error: …(400,TURTLE_PARSE) at the document's byte offset. Undefined prefix in a block: Turtle parse error instead of a transaction errorFluree::upsert_turtle*fluree syncThe upgrade notes (for the release notes) and the API changes are folded under this section.
Fuel. A Turtle upsert reports
tally.fuelon every builder lane, with and without a policy, over novelty and over an index. The charge is the same as for any upsert: 10 per transaction plus 0.001 per committed flake. For the test payload,maincommitted 9 flakes (10.009), 4 of its 5 retractions phantoms that removed nothing; this branch commits 9 real flakes (10.009). Payloads wheremainstaged phantoms now stage fewer flakes, so reported fuel is lower by 0.001 per phantom (an identical refresh reports the 10.000 baseline and no commit). That is a billing change for callers who bill from the tally.#1988 (merged). SPARQL
DELETE DATAof a typed literal ("1"^^xsd:int,xsd:long,xsd:dateTime) used to lower the value as a string, which committed phantom retractions that removed nothing. With this PR alone it would commit nothing, and the values would still stay; with #1988's coercion, the intent names the stored term exactly and the resolver deletes it.it_delete_stored_facts::delete_data_of_typed_literals_deletes_thempins this over novelty and over an index (its own commit).On top of #1997. This branch is rebased on
61b836e9a, and #1997's semantics are unchanged. How the DELETE witness rule reads #1997's graph mapping, and what happened to #1997's own tests, is folded below.Named graphs and anchors. Bundles built from RDF text get their
f:reifiesGraphanchor from one emit rule,graph_scope_emit— a local stand-in with the contract ofGraphScope::emitfrom the JSON-LD graph-scoping PR (#2009), which replaces it when this rebases over that PR.bundle_templatesnever states an anchor, and a block label reaches staging as written, so the rules staging applies to write-graph names apply to TriG blocks unchanged.Upgrade notes (for the release notes)
docs/transactions/turtle.mdhas a query that lists blank-node subjects nothing references. JSON-LD upsert and graph-sync identities are unchanged; a test pins that a graph synced through the old conversion re-syncs from the same text without a commit.GRAPHblocks, as the label or in the contents, keep today's verbatim behavior; documented cross-ledger config relies on it. Sync and Graph Store now accept them in block contents too.t.dateTime/timewith digits past the microsecond. Once the value is indexed (for temporals, once it is reloaded from its commit), re-writing it as written commits a retraction and an assertion of one stored fact and can remove it.mainbehaves the same way.upsert.mdnow says so and how to write such values, every other idempotence claim in the transaction docs names the exception, and two tests document it (ignored until the fix lands)."abc"^^xsd:integer, which Turtle writes store as its lexical form); lowering used to refuse it. In exchange, a mistyped literal inDELETE DATA("1990-13-01"^^xsd:date) commits nothing instead of erroring.API changes
Fluree::stage_transaction_with_trig_meta,stage_transaction_with_named_graphs,stage_transaction_with_named_graphs_tracked(the JSON-LD form is nowstage_transaction_tracked),transact_with_trig_meta,transact_with_named_graphs. Fromfluree-db-transact:extract_trig_txn_meta,unwrap_trig_graph_blocks,UnwrappedTrig,TrigMetaResult,apply_cancellationanddedup_retractions(thegenerate::cancellationmodule), andFlakeGenerator::generate_retractions(now the crate-privategenerate_retract_intents).FlakeAccumulator::push_retractionstakesRetractions.parse_rdf_text,parse_rdf_text_txn,parse_trig_txn,has_graph_blocks,Placement,TemplateSink,CurrentFacts/StoredFact/Retraction,FlakeAccumulator::finalize_with_restated,DELETE_WITNESSED_SITE,BinaryRangeProvider::persisted_object_key,TransactError::{Turtle, PayloadGraphMismatch},fluree_graph_turtle::{RelativeIris, ParserOptions::relative_iris}(defaultResolve),fluree_graph_turtle::SegmentParserandparser::TokenStream(Parsergains a token-source type parameter, defaulted toStreamingLexer), and in the CLIRemoteLedgerClient::sync_rdfandRemoteLedgerError::UnsupportedMediaType.parse_trig_phase1/resolve_trig_metastay, for bulk import only.#1997: the DELETE witness rule over its graph mapping, and its tests
WITHand a JSON-LD update's top-levelgraphnaming the ledger's own address read and write the ledger's default graph. An explicitGRAPH <address>resolves through the graph registry. The DELETE witness rule compares the graph ids that mapping produces: the WHERE's default graph when it is one ledger graph, eachGRAPH <name>'s graph, and the graph the ledger has under a template's IRI. A default graph that reads as empty (an unknownUSING/WITHgraph) witnesses nothing. So underUSING <address>, aGRAPH <address>template that writes a graph registered under that name is not treated as witnessed (it_delete_stored_facts::a_graph_named_by_the_ledger_address_gets_what_its_templates_write). No transaction lane reads TriG with phase 1 any more; bulk import still does, with #1997's shared prefix maps. #1997's redefined-prefix test,it_trig_insert::a_redefined_prefix_or_base_applies_only_after_it, now runs on the parse-once lanes. Two of #1997's phase-1 unit tests read the default graph throughparse_to_json, whichfluree-db-transact'sclippy.tomlnow disallows; they read it with the Turtle parser instead.Solo lockstep
From an audit of fluree/solo's write paths, these want a lockstep change on the solo side:
jsonld_upserthints ("prefLabel, one per language") and the legacy/v1/fluree/tmPATCH send one language per upsert. Under the new replacement unit (call 3 above) that deletes the other languages, so we'll want the hints and tool text to send the full multilingual set, or to use update with a language filter.update-a-model-from-a-file.md:51-53) should describe the one-time re-mint, then stable ids.commit: null.result.tally(pinned by a test). Fuel is lower wheremainbilled phantom flakes.Turtle parse error: …, which solo's router maps to 502, so the prefix should go inCLIENT_ERROR_PREFIXES.jsonld_updatetool text says "Retracting a triple that is not there does nothing", which is true only with this PR — andhints.rs:85still saysskos:prefLabel"(one per language)", so the language-hint change above is still needed.Performance
The short version: Turtle upsert is 22–30% faster (one parse instead of the JSON-LD round trip), TriG upsert and insert are flat to faster, and on #1997's many-block TriG shape upsert is 10% faster and insert at parity, with 36–37% fewer allocations and a 25% lower peak, which Claude's review pass reproduced against
61b836e9aon its own (upsert 72.0 → 65.4 ms, insert 62.5 → 61.8 ms, allocations −36/−37%, peak 38.4 → 28.9 MiB). A1's upsert staging reads 7–12% faster in isolation (new_subjects+0.5%). JSON-LD and Turtle insert, whose paths A2 changes only in dispatch, read flat by their fastest reps, and bulk import is untouched by A2 and reads flat.Quiet box. EC2 c7i.4xlarge, the fat-LTO bench profile, base and head in interleaved rounds; a change counts as a win or a loss only when the base and head ranges don't overlap and the median delta exceeds the bench's budget (5% at
small). Four wins and no loss:named_novelty−82.5% (120 → 21 ms),upsert_turtle−36%,trig−21% andtrig_blocks−17% (the new many-blocks case, in its first A/B). Every other measured row is stable,transact_filtered_deleteincluded (+1.2% and +0.8%). Of the three end-to-end A1 upsert rows the local runs flagged,named_indexed(−3.8%) andnovelty_heavy(−5.1%) are stable, and a second session resolvednew_subjectsat −0.9%, with the head faster in all 7 rounds, so the first session's +14% doesn't reproduce.The quiet-box tables
Session 1 (2026-10-01): three rounds, in AB, BA, AB order. The base is
mainat61b836e9aplus this PR's two bench-only commits (8024ea1d2,41579486a), so both sides run the same bench source; the head is GitHub's merge of this PR (refs/pull/2008/merge,449b3dc1c).Session 2: the one row session 1 left inconclusive, over 7 rounds in alternating AB/BA order.
head>basecounts the paired rounds in which the head was slower.Where it reads slower, fwiw: on an indexed ledger, the constant-object DELETE costs 1.09–1.13× current
main(61b836e9a; 1.13× in this branch's re-run, 1.09× in Claude's review pass) — the cost of reading every slot instead of committing 100k phantom retractions, since the DELETE matches nothing. One POST read per (predicate, object) pair would serve all of them, but that stays a follow-up because it's a real scope addition rather than a tweak: it needs a novelty read by pair beside today's per-slot one, and a bound for popular pairs, where one POST read would scan every subject holding the value while the DELETE names a few. Three end-to-end A1 upsert rows also read 1.13–1.20× locally, but neither the staging probe, which covers the only code A1 changes, nor the quiet box (above) reproduces them.Two caveats I want to be upfront about. The A1 numbers are pre-rebase: they were measured before the rebases (onto
bf523e24e, then onto61b836e9awith #1997) and before the late A1 commits. The #1997-shape runs and the many-small-blocks table (re-measured against61b836e9a) are on the current base; the A2 criterion rows (including the 22–30% Turtle-upsert figure) predate #1997 and the late commits too. And the tables ran on one shared 16-core machine with load from other work (1-minute load average 5–85), where end-to-end rows vary up to ±30% between reps of the same binary, so the regression-budget gate on a quiet runner is the authority for the 5% budget, and the quiet box above applied the same budgets.Every perf table, with methodology
dev-fast profile unless noted, on one shared 16-core machine with load from other work during the runs (1-minute load average 5–85). Each table compares binaries run interleaved.
A1 numbers are pre-rebase; the quiet box (above) has since measured the PR head. The two A1 tables were measured before the rebases (onto
bf523e24e, then onto61b836e9awith #1997) and before the late A1 commits (index-key matching, the policy check, OPTIONAL/UNION witnesses).A1, upsert staging in isolation (a scratch probe staging each upsert 40× at N=2000 or 20× at N=8000 on clones of one fixture; p50 ms, median of 3 interleaved rounds at N=2000, rounds 1/2 at N=8000). Base is
mainplus this PR's first commit (the new bench).A1, criterion end to end (stage + file-backed commit), small scale, 3 reps, median ms unless noted; same base, before the rebase:
The end-to-end rows vary up to ±30% between reps of the same binary on this machine. The three upsert rows above 1.1 are not reproduced by the staging probe, which covers the only code A1 changes: commit is unchanged. The regression-budget gate on a quiet runner is the authority for the 5% budget, and on the quiet box (above) none of the three is a loss.
A2, criterion end to end on the branch at
bf523e24e(before #1997), before the late commits (the segment parser, verbatim block labels, the late A1 commits): the A1 tip plus the A2 bench commit against the A2 tip of that time.insert_formatsruns 10 transactions per iteration, each iteration on a fresh ledger (ms). Medians of 3 interleaved reps. The first run read 1.09 ontrig, so the TriG and 100-node insert rows were run again, 5 interleaved reps; for those rows the last column is the fastest of all 8 reps of each binary.jsonld(insert)turtle(insert)trig(TriG upsert, one block)upsert_turtleupsert_turtle_replacetrig_mixed(TriG insert, default graph + block)jsonldturtletrigupsert_turtleupsert_turtle_replacetrig_mixedimport_bulksingle-threaded (5 reps)import_bulkdefault threads (5 reps)TriG on #1997's shape. 10,000 blocks of 2 prefixed-name triples each, 20
@prefixdirectives at the top, 16 graph labels, throughupsert_turtleandinsert_turtleinto a fresh memory ledger. dev-fast builds ofmain(61b836e9a), of the A1 commits plus the bench commit, and of this branch; 5 interleaved rounds of 15 runs; median of the per-round medians, with the range in brackets. Peak heap and allocations come from a counting allocator and were identical in every round.mainUpsert is 10% faster and insert is at parity, with 36–37% fewer allocations and a 25% lower peak. Claude's review pass reproduced it against
61b836e9aon its own: upsert 72.0 → 65.4 ms, insert 62.5 → 61.8 ms, allocations −36/−37%, peak 38.4 → 28.9 MiB. Before the parser read the locator's tokens, this branch lexed every block twice and measured 78.2 ms and 75.0 ms on the same shape. The parser is generic over its token source, and the default lexer path is unchanged:fluree_graph_turtle::parsealone, release builds, 6 interleaved rounds of 7 runs: 360.6 ms [358.5–365.6] (59.1 MB/s) before the change, 350.3 ms [342.7–362.5] (60.9 MB/s) after.Many small blocks, and DELETE rows the matched lane used to read (dev-fast CLI end to end, including about 0.2 s of start-up; this branch's head against
mainat61b836e9a, with #1997; 3 interleaved reps; seconds, median). The TriG documents are 100k one-triple blocks under 20 prefixes over 50 graphs (blocks100k) and 100k default-graph statements plus a block (mixed100k); the DELETEs run over 100k subjects:mainblocks100kblocks100kmixed100kmixed100kClaude's review pass measured the TriG rows independently against
61b836e9aat 0.97, 0.90, 0.59 and 0.53, and the indexed constant-object DELETE at 1.09.Before the segment parser and the OPTIONAL/UNION witnesses, an adversarial review pass (Claude) measured this branch at 1.6–1.9×
mainon the block documents and 1.5× on the indexed OPTIONAL-bound DELETE.insert_formats/trig_blocks(every statement in its own block, sixteen extra prefixes) now covers the many-blocks shape in the criterion suite. It hasn't been run A/B on this machine, but the quiet box has it at −17.4% (above). The constant-object DELETE matches nothing:maincommits 100k phantom retractions, and this branch reads each slot and commits nothing. On the indexed ledger that costs 1.09–1.13×main: reading every slot instead of committing 100k phantoms (see Follow-ups).Turtle upsert is 22–30% faster: one parse instead of the JSON-LD round trip. TriG upsert and insert are flat to faster. JSON-LD and Turtle insert, whose paths A2 changes only in dispatch, read flat by their fastest reps. Bulk import is untouched by A2 and reads flat.
transact_upsert_replace(JSON-LD upsert over existing values) ran in the same session and shows no consistent direction, with ±50% between reps of one binary at that load.Deviations from the design
GRAPHblocks with no@baseare kept as written (today's behavior), where the design made them an error. That takes one parser option,ParserOptions::relative_iris(defaultResolve, so every other caller stays strict). The design expected no change tofluree-graph-turtle; it also gainsSegmentParser, one parser over the pieces of a TriG document, fed the tokens the locator lexed.The smaller, structural ones, and the rebase note for #1555/#1569 (the whole
fluree-graph-turtlediff, +201/−7), are folded here:Structural deviations, and the
fluree-graph-turtlediff for #1555/#1569Retractiontyping landed with the DELETE commit rather than on its own, and the locator is its own module (parse/rdf_text/locate.rs).resolve_trig_meta), so they still cover that lane.unwrap_trig_graph_blockswas deleted too, since it has no caller after the sync lane moved.TransactError::PayloadGraphMismatch.fluree-graph-turtlediff (+201/−7):options.rs:RelativeIris, andParserOptions::relative_iriswithwith_relative_iris.parser.rs: aTokenStreamtrait, implemented byStreamingLexer;Parser<'a, 'input, S, L = StreamingLexer<'input>>, with the constructor's body moved into a privatewith_lexer;parsesplit intoparse_statements; aRelativeIris::Verbatimarm where a relative reference meets no base; andSegmentParser.lib.rs: exportsRelativeIrisandSegmentParser, and documentsparse_to_jsonas lossy.Follow-ups
transact_filtered_deleteandnovelty_replaybenches drop their file-backedTempDirinside the timed routine. This PR's new bench doesn't; those two should follow.dateTime/timeis stored to the microsecond. An identical re-upsert or re-sync of such a value, once indexed (temporals: once reloaded from the commit), retracts the stored decode and asserts the written term under one key, and can remove the value.maindoes the same. The fix belongs where values are written (normalize to storage precision) or in the index (keep the subtype). Claude's review pass found more values that read back with less detail than written, onmaintoo: after indexing, a big plainxsd:integerreads back asxsd:decimal;-0.0loses its sign; and"1.50"reads back as1.5.mainon an indexed ledger, wheremaincommits phantom retractions without reading. One POST read per (predicate, object) pair would serve all of them. It needs a novelty read by pair beside today's per-slot one, and a bound for popular pairs, where one POST read would scan every subject holding the value while the DELETE names a few.upsert_replacement_slots.maintoo: a SELECT over an indexed big integer readsxsd:decimal, and BIND gives onexsd:integerwhatever its subtype. So on a novelty-only ledger a DELETE row whose value BIND bound does not name a big integer stored under a subtype.span_capture::init_test_tracing) can miss its event while other tests in the same process install theirs (it_annotation_filter_pushdown,it_limit_stops_work). Unfiled.f:ledgeris read only as an IRI reference, and a slashed ledger id has no absolute spelling. Accepting a string literal there (and documenting it) would let config avoid relative references.Tests
Both groups have regression suites — A1's stored-retraction, delete, policy (deletes and restated values), witness-routing (paired
MustFire/MustNotFirestamps on one ledger), annotation, conflict-retraction and legacy-revert tests, and A2's RDF-text parity matrix over 11 upsert/insert entry points and 4 graph-payload lanes — and each regression test was run with its fix reverted and watched fail.testsuite-sparql(W3C) is 36/36, and the five triple-term entries that were blocked on TriGGRAPH-block parsing now load their data (none passes yet; each is re-attributed to the failure it now reaches). The full list and the non-vacuity runs are folded here:Full test list, and the non-vacuity runs
it_upsert_stored_retractions: lang, list, identical refresh, named graph, blank-node refresh, annotated cascade. Each case runs on three lanes (novelty, indexed, novelty over index), before and after reindex.it_delete_stored_facts: absent DELETE twins (SPARQL, JSON-LD, Cypher), DELETE/INSERT of an absent fact, exact terms, language replacement twins, the big-number witnessed delete with and without lists, and typed-literalDELETE DATA(with fix: keep typed literal values through the index (reindex affected ledgers) #1988). Also:DELETE DATA, JSON-LD delete, constant template), rows bound by OPTIONAL/UNION/BIND (SPARQL, JSON-LD OPTIONAL), CypherSETandDETACH DELETE, a named graph, sub-microseconddateTime/time, and a different big value that is not deleted;DELETE DATAof the stored and two absent values, a WHERE-driven guess, JSON-LD twins), and a permitted delete of an absent value that commits nothing;USINGandGRAPHnaming that address;it_delete_witness_routing: pairedMustFire/MustNotFirestamps on one ledger; OPTIONAL and UNION rows (JSON-LD and SPARQL), CypherSETandDETACH DELETEtake the witnessed lane, and a UNION binding the object twice does not (JSON-LD and SPARQL).shacl_tests::a_refresh_upsert_satisfies_unique_lang, and two ignored tests for values stored with less detail than written (big-integer subtype, sub-microseconddateTime).it_annotation_delete_named_graph,it_named_graph_conflict_retractions(merge, rebase, revert),it_revert_legacy_phantom.it_typed_literal_index::sparql_insert_data_rejects_an_ill_typed_literal, Index build silently destroys temporal and xsd:long literals (xsd:date → 1970-01-01), and the resulting triples cannot be retracted by literal match #1987's test updated (see Interaction with Index build silently destroys temporal and xsd:long literals (xsd:date → 1970-01-01), and the resulting triples cannot be retracted by literal match #1987).current_factsanddelete_witness(graph ids, OPTIONAL and UNION witnesses).it_rdf_text_parity:GRAPHblock stored as written on upsert, TriG insert and sync, and under a relative block label on upsert and TriG insert (a sync target must be an absolute IRI); the default-graph twin refused; a base resolving label and contents;it_trig_insert::a_redefined_prefix_or_base_applies_only_after_it(fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997's): its redefined-prefix and redefined-base cases, now through the parse-once upsert and insert lanes.rdf_textandlocate: spellings, refusals, offsets, directive order, block labels under redefined declarations, block parity, blank-node scope, literals, txn-meta, content identity, placement, stable ids, reserved predicates, relative references.SegmentParsercarries declarations across pieces.<org/governance:main>in TriG blocks) pass unchanged.witnessed; no witnesses; no enrichment; old cascade; old rebase read; old undo fold. Also:graph_contexts_must_agreefails;sh:uniqueLangrefresh test fails.parse_to_json→ IRI, parity, list, idempotence and error tests fail;[ … ]-in-block detection fails;main:expected object, found '['.testsuite-sparql(W3C): 36/36 groups pass. The five eval-triple-terms entries registered as "blocked on TriG GRAPH-block parsing" now load their data. Each is re-attributed to the failure it reaches (non-isomorphic reifier results, unlowered triple-term functions, SPARQL UPDATE quoted triples,GRAPH ?g-valued annotation); none passes yet.Gates
This is 35 commits on
61b836e9a, with the head at99d5776c9. The gates ran on the head commit as first written, before an amend that only moved one helper withinstage.rs; fmt, clippy and the policy tests were re-run at the head. There, fmt (after the last edit) and clippy-D warningsare clean, every commit type-checks on its own, and the workspace--all-featurescheck is clean.cargo test -p fluree-db-apiis 3,950 passed / 1 failed / 121 ignored: the failure is a tracing-capture test that fails onmaintoo in a bundled run, and it passes undercargo nextest, one process per test, as CI runs them. The SHACL, server, consensus/memory and #1997 suites pass, andtestsuite-sparqlis 36/36 SPARQL and 3/3 RDF groups. A few results are from987a9379c, in crates and workspaces nothing has changed since:fluree-db-query,fluree-graph-turtle,fluree-graph-format/fluree-db-r2rml,testsuite-sparql's fmt/clippy, andit_sql_pushdown_lane. In an earlier run there,it_limit_stops_workwas intermittent, failing in 1 of 7 bundled runs on this branch and 0 of 18 onmain; this branch touches neither tracing nor the property join, and it passes under nextest. Not run:testsuite-shacl(not a CI job), the workspace-wide--all-featuresnextest that CI runs, and thefluree-db-apitests gated on the graph-source and vector features, exceptit_sql_pushdown_lane.Full gate log
At the head commit as first written, on
61b836e9a, before an amend that only moved one helper withinstage.rs. The head,99d5776c9, differs only in where that helper sits; fmt, clippy and the policy tests were re-run there.cargo fmt --all -- --check, after the last edit: clean.cargo clippy --all --all-features --all-targets --locked -- -D warnings: clean.cargo check --all-targetsover the query, transact, api, Turtle, CLI and server crates). The workspace--all-featurescheck is clean.cargo test -p fluree-db-api(the lib and every integration group, default features): 3,950 passed, 1 failed, 121 ignored.it_annotation_filter_pushdown::annotation_body_threshold_reduces_scan_work_on_both_surfaces. It fails onmaintoo in a bundled run: its tracing capture shares the process with the rest ofgrp_misc.cargo nextest, one process per test as CI runs them, it passes.987a9379c,it_limit_stops_work::a_selective_star_sizes_later_chunks_from_its_yieldmissed its tracing event the same way (expected one property join, got []). It is intermittent: it failed in 1 of 7 bundled runs of its group at this branch and in none of 18 onmain. This branch touches neither tracing nor the property join, and it passes under nextest.cargo test -p fluree-db-api --features shacl: the lib, 922 passed; the targets with SHACL-gated tests (grp_ledger,grp_graphsource,it_optimistic_rebase,it_staged_view_dict), 501 passed.it_named_graphs(includingtest_updates_on_a_graph_registered_under_the_ledger_address),it_trig_insert(including the redefined-prefix test),it_import,it_sql_pushdown_lane(--features sql, run at987a9379c), and the server'ssparql_dataset_semanticsin the server run below.cargo nextest run:fluree-db-server(every target) 662 passed;fluree-db-consensusandfluree-db-memory153 passed.cargo test:fluree-db-transact --lib329 andfluree-db-cli460 passed. Doc tests offluree-graph-turtle,fluree-db-transactandfluree-db-api: 1 run and passed, 104 markedignore.987a9379c, with no change to these crates since:fluree-db-query --lib1,648,fluree-graph-turtle197, andfluree-graph-formatandfluree-db-r2rml188 passed.testsuite-sparql: 36/36 SPARQL groups and 3/3 RDF groups pass at the tip; fmt and clippy-D warningsclean at987a9379c, with no change to that workspace since.testsuite-shacl(not a CI job); the workspace-wide--all-featuresnextest, which CI runs; and thefluree-db-apitests gated on the graph-source and vector features, exceptit_sql_pushdown_lane.