Repository navigation
Conversation
Two benches for the request shapes whose dataset references are parsed,
classified and resolved per request, so the typed-reference work that
follows has before numbers to compare against.
fluree-db-server dataset_http: warm in-process HTTP requests with SPARQL
FROM / FROM NAMED, JSON-LD from, and GRAPH scopes on /query and
/query/{ledger}, a bearer-token request, a query with a 2 KB VALUES block,
and three streaming requests on /stream/query/{ledger} (10k default-graph
rows with and without a policy input, and 16 FROM NAMED members under a
policy input). Each case checks its row count once before timing.
fluree-db-api dataset_build: build_dataset_view and FROM NAMED x 1/4/16
through query_from() on a file-backed connection, for members that are
graphs of one ledger and for members that are separate ledgers, plus the
same query with the graph-source providers attached (iceberg feature).
A ledger, a time, a graph of a ledger and a bare graph IRI are all written
as text, and each surface split that text itself: at HEAD there are seven
parsers of the address grammar, and they disagree about '#', '@' and
'urn:fluree:'. This adds the one typed form they can all share, in a new
fluree-db-core module, dataset_ref:
- GraphIri: an absolute IRI (validate_absolute_graph_iri), the only way to
name a graph by IRI.
- GraphSel { Default, TxnMeta, Config, Named(GraphIri) }: which graph of one
ledger. The one keyword table for 'default' / 'txn-meta' / 'config'. A
relative name is refused instead of being looked up literally.
- LedgerRef { id, at: Option<TimeSpec>, graph: GraphSel }, private fields:
the parsed address for ledger positions. Replaces the raw-string struct
of the same name (its pin was a String and its fragment unvalidated).
- DatasetRef { Address | GraphIri | Ambiguous }: what a dataset-position
string (FROM, FROM NAMED, a JSON-LD from string) names before resolution.
Ambiguous carries both readings of text that is lexically both, such as
'mydb:main', for a resolver that holds a graph registry to choose.
- TimeSpec moves here from fluree-db-api (re-exported at both of its old
paths, with From<LedgerIdTimeSpec>), together with its grammar tests.
Two grammar rules keep the address grammar from swallowing graph IRIs:
- A branch may not begin with '/': LedgerId::parse refuses one, so
'http://ex.org/g' is no longer read as ledger 'http', branch
'//ex.org/g'. Legacy branches that merely contain '/' still parse, and
stored records keep parse_persisted.
- '@' starts a pin only before a known tag. In a ledger position any
other '@' is a bad pin; in a dataset position it means the text is not an
address ('http://ex.org/@alice/g', 'mailto:a@b').
With the first rule a hierarchical graph IRI no longer parses as an id, so the SQL and
Delta dispatch probes, which receive the GRAPH IRIs a query writes, now
classify the IRI first: one that does not parse as a graph-source id names
no source and is declined without a lookup, while a failed lookup of a
well-formed id still fails the query: the same guard #1997 added
(dispatch_record).
The existing LedgerRef callers (server path parsing, scope_id, the MCP
ledger argument, the graph-source fallback, same_ledger) move to the typed
accessors.
A dataset member was a GraphSource with an `identifier: String` that every
consumer re-read: a SPARQL FROM IRI went through its own stringifier
(prefixed names unexpanded), then a third address parser that split '@' and
'#' textually, and a graph IRI on a connection route was looked up in the
nameservice as a ledger. Members are now parsed once into typed values and
resolved by one table.
- GraphSource keeps the text as written, what it names (MemberRef: a graph
keyword, or a DatasetRef), its time and its per-source options, behind
private fields. It is built by parsing (GraphSource::parse, TryFrom,
FromStr) or from a typed address (GraphSource::ledger); `identifier` is
gone. dataset::GraphSelector is now an alias of core's GraphSel.
- One JSON-LD dataset parser: DatasetSpec::from_json reads the `opts` and
`ledger` spellings that only from_query_json used to. An object's `@id`
is a ledger address and its `graph` a graph of that ledger; naming a
graph in both is refused.
- SPARQL dataset clauses are read from resolve_dataset_clause, so prefixed
and BASE-relative FROM / FROM NAMED IRIs expand against the prologue on
every route (DatasetSpec::from_sparql / from_sparql_ast). A dataset IRI
written in full is taken as written: FROM <mydb:main> is not a misused
prefix even when the query declares `mydb:`.
- core gains MemberRef and TargetLedger, the resolution table for a
ledger-scoped surface: the ledger's own address in any spelling is its
default graph; the written text as an exact registry IRI (reserved slots
included) is that graph; a keyword or `L#graph` selects within; another
ledger is CrossLedger; an unknown graph IRI is GraphNotFound. The view
path's within-ledger FROM and every graph selection (`db("L#g")`, a
member's graph) resolve through it, so the short alias `FROM <L>`,
`FROM <L#iri>`, `FROM <L#config>`, a keyword, and the main branch's
`urn:fluree:L:main#config` read from a branch all work there (#1512).
- On a connection surface a bare graph IRI or keyword names no ledger: a
400 that names the fix, instead of a nameservice lookup of the IRI.
- A named member is known by its alias, else its text as written, pin
included: `FROM NAMED <L@t:2>` answers `GRAPH <L@t:2>` and two pins of
one ledger are two members. The time-stripped name it had before stays a
non-enumerated alias while exactly one member claims it, so
`GRAPH <L>` over one pinned member still matches (DataSetDb gains the
alias map).
- Caller mistakes in this area are no longer 500s: an unknown graph is a
404 (new ApiError::GraphNotFound, which is not is_not_found(), so it
never falls back to a graph-source lookup); a missing dataset, a
cross-ledger or pinned within-ledger member, and the within-ledger
history-range refusal are 400s. An unknown pin tag in a dataset position
reports the pin, not "not an IRI".
- sparql_dataset_ledger_ids returns the canonical ids a clause names by
address (graph IRIs are not ledgers), and the CLI's "already has a
dataset clause" guards use the new sparql_has_dataset_clause.
- Fluree::is_single_ledger_fast_path (test-only callers) is removed.
The ledger route keeps its own FROM classifier for now and only moves to
the typed constructors; routing it through TargetLedger follows.
…robed `R2rmlProvider::has_r2rml_mapping` returned `bool`, and the api's provider turned every nameservice error into "no mapping"; the GRAPH operator and the fused R2RML aggregate also `.ok()`-ed `compiled_mapping`. During a nameservice outage a GRAPH-addressed R2RML, Iceberg or SQL source therefore read as an empty native graph, with a 200. The probe now returns `QueryResult<bool>`. An IRI that is no graph-source id is still `Ok(false)` without a lookup; a lookup that fails fails the query, and so does a failed mapping load for a graph found to be a source. Failing closed alone would have made every query that carries the R2RML provider (every `query_from()` and `graph().query()` on an iceberg build) depend on the graph-source registry, because the executor probed the primary snapshot and each dataset member on every query. Views now say what they read: a `MemberKind` (`Native`, `GraphSource`, or `Unclassified` for callers that do not say) on each runtime `GraphRef`, and on `ContextConfig` for the no-dataset primary. The executor's precomputed set, the GRAPH operator, the fused aggregate and the SQL lane never ask about a graph loaded as a ledger graph. Graph-source members are probed as before, and a GRAPH IRI that is not a member is the one remaining run-time probe. A probe failure surfaces as the provider's existing "Nameservice error" (`QueryError::Internal`, which the api maps to 400 like the provider's other nameservice errors). Test: it_sql_pushdown_lane `graph_source_probes_fail_closed_and_skip_ledger_graphs` runs against real SQL sources with every graph-source lookup failing: `GRAPH <source>` errors on and off the lane, in SPARQL and JSON-LD; the native twin answers exactly as with a healthy nameservice through its own view, `FROM`, and `FROM NAMED` + `GRAPH`; a GRAPH IRI that is no id answers without a lookup.
… ledger's from
`graph()` stored the raw string and handed it to the ledger loader, which
takes only `name[:branch]`, so `graph("L#txn-meta")`, `graph("L#<graph IRI>")`
and `graph("urn:fluree:L")` failed although `db()` accepted them (#1961).
The handle now parses a `LedgerRef` once. Queries and `load()` read the
address's graph through the loader `db()` uses. A pin in the address
(`L@t:5`) is honored, and one that disagrees with `graph_at`'s is a 400;
`db()` honors a pin the same way. A transaction applies to the whole ledger
at HEAD, so `.transact()` through a graph-qualified or pinned handle is
refused with a 400. The commit builder reads the address's ledger.
A JSON-LD query on a view reads the view's ledger, but a `from` or
`fromNamed` naming another ledger, or one that does not exist, was ignored
and the query silently answered from the view. All four view entry points
refuse it now with a 400 that points at `query_from()`. A `from` naming the
view's own ledger, in any spelling, still answers. The other within-ledger
JSON-LD `from` shapes keep their current reading; honoring them on the view
path belongs to lane consolidation (part 2).
The tracked view paths report the within-ledger FROM resolver's own status
(404 for a graph the ledger lacks), as the untracked paths already do.
Tests: it_named_graphs `graph_handles_take_the_addresses_db_takes` and
`a_missing_ledger_reads_as_not_found_at_every_entry_point`; it_query_dataset
`a_view_refuses_a_jsonld_from_naming_another_ledger`.
…edger `SERVICE <fluree:ledger:X>` looked X up among the dataset's members and, when X was not one, ran the block against the current dataset. An endpoint naming another ledger, or no ledger at all, therefore answered with this dataset's rows. A non-member endpoint is now a 400 that points at FROM NAMED, and it contributes no rows under SILENT. It is never loaded on the spot, because a SERVICE endpoint is not authorized the way a dataset member is. With no dataset, the endpoint must still name the query's own ledger, as before. The endpoint's address is parsed with the shared ledger grammar, so the `urn:fluree:` spelling works. A pin or a graph in it is malformed: the member's time and graphs come from the dataset. Tests: service unit tests for the endpoint grammar; it_service_cross_ledger_iri `service_to_a_non_member_is_refused_not_answered_by_the_dataset` and `a_service_endpoint_is_never_loaded` (a recording nameservice sees no lookup of the endpoint's ledger).
A single `USING <g>` or `WITH <g>`, or a JSON-LD update `from` / `graph`,
that the graph registry did not resolve fell back to g_id 0, so the WHERE
read the ledger's real default graph. `DELETE { ?s ex:v ?o } USING <typo>
WHERE { ?s ex:v ?o }` deleted it, `WITH <new> INSERT … WHERE …` copied it,
and `USING <L#config>` read the default graph instead of the config graph.
Two or more `USING` graphs skipped an unresolved IRI, which dropped
the ledger's own address. JSON-LD refused `"from": "L:main"` as an
undefined prefix.
Every WHERE-dataset reference (`USING`, `USING NAMED`, `WITH`, and the
JSON-LD `from`, `fromNamed` and `graph` keys) now resolves once, in the
transaction's ledger, through the table a query's FROM uses:
- the ledger's own address, in any spelling, names its default graph;
- `L#config` or its URN names the config graph;
- a registered IRI names its graph.
A graph the ledger does not have contributes nothing, so a WHERE over it
binds nothing (SPARQL 1.1 Update §3.1.3, Query §13.2). Another ledger's
address, or a pinned address, is a 400, since the WHERE reads only this
ledger as it stands. The single and multi branches share the resolver. The
JSON-LD dataset keys are dataset references: a prefix the context defines
still expands, and any other text is kept as written for the resolver.
#1997 resolved an update's `USING` / `WITH` default graph by a narrower
rule (own address, else registry, else empty). This is its generalization;
the rebase onto #1997 keeps this side.
Test: it_named_graphs
`update_where_dataset_references_resolve_in_the_transactions_ledger`, each
shape above in SPARQL and JSON-LD, with the cross-ledger and pinned
refusals.
…dger
The ledger routes (`/query/{L}`, `/stream/query/{L}`, `/explain/{L}`) loaded,
authorized and compared the path as spelled, so `/query/b2x` with
`FROM <b2x:main>` or `"from": "b2x:main"` was a ledger mismatch, and
`/query/urn:fluree:b2x:main` could not load at all (#1982). `PathLedger` now
carries the canonical id (plus any graph the path selects), and every
downstream use reads it.
Each SPARQL `FROM` / `FROM NAMED` IRI, and each JSON-LD `from` / `fromNamed`
source, was classified by string shape: an IRI with `#` or `@` was taken for
ledger syntax ("Ledger mismatch" for `<http://ex.org/vocab#products>`,
"Invalid time travel" for `<http://ex.org/@alice/g>`), the short
alias or the URN of the path's own ledger read as a graph name (500 Unknown
named graph), and a graph IRI in `fromNamed` was loaded as a
ledger (404). The clause is now expanded against the prologue first
(prefixed and BASE-relative IRIs), and each reference resolves once, in
the path's ledger, through the table the view path and updates use, with the
ledger's graph registry:
- the ledger's own address in any spelling names it, or the graph the
address names;
- a keyword or a registered IRI names that graph;
- an IRI the registry does not hold is left to loading, which answers 404
`err:db/GraphNotFound`;
- another ledger is a 400.
A single JSON-LD `from` must name the path's ledger. A `from` array and
`fromNamed` may also name other ledgers, as before. The dataset lane and its
per-member loading are unchanged.
`ApiError::GraphNotFound` now maps to 404 with the new `@type`
`err:db/GraphNotFound` on the server (it fell through to a 500).
Tests: sparql_dataset_semantics
`ledger_route_dataset_references_resolve_in_the_paths_ledger` (graph IRIs
with `#` and `@`, the config graph by its URN, each spelling of the path and
of the ledger in `FROM` and `from`, a prefixed `FROM NAMED`, a graph IRI in
`fromNamed`, an unknown graph, another ledger in each language); unit tests for the path id and the JSON-LD
normalization.
`fluree query` chose between the target's own view and the connection path by comparing base names (`mydb:feature-x` and `mydb` share `mydb`) and by reading only a string or string-array JSON-LD `from` (#1972). A branch other than the target's ran against the target's branch; a JSON-LD `from` naming one of the ledger's graphs, in string or object form, or naming a ledger that does not exist, ran on the view path, which ignored it and answered from the default graph; SPARQL `FROM` of another branch failed there. The router now resolves each reference in the target ledger through the table every surface shares, using the ledger's graph registry when it is local. Locally: - the view path keeps what it reads: its own ledger in any spelling, and SPARQL graphs of that ledger; - another ledger or branch, a time pin, and every JSON-LD source other than the whole ledger (a graph, `fromNamed`, a history `to`) take the connection path, which reads them. For a remote target, the ledger route resolves references against its own registry, so only a reference that names another ledger leaves it. The base name comparison and the `from` scraper are gone. Tests: unit routing decisions for the local and remote rules; CLI integration `local_query_reads_the_graph_or_branch_its_dataset_names` (graph by address and object form, config graph, another branch in JSON-LD and SPARQL, a missing ledger).
Callers outside db read which ledgers a query names before running it,
through `DatasetSpec::from_query_json` (each member's `identifier` and
`policy_override`) and `sparql_dataset_ledger_ids`, so these two must report
every ledger the query can read. The typed members had made `identifier` and
`policy_override` private and changed what `sparql_dataset_ledger_ids`
reported.
Both fields are public again. `identifier` is the reference as the query
named it, with an address's `@` pin removed and a `#graph` kept, as before.
Both are derived from the typed parse, and resolution reads the typed
reference, never them. `sparql_dataset_ledger_ids` reports every member's
identifier (minus `#graph`) and the history range's ledger, as before. A
reference whose reading depends on a ledger's graph registry is reported
too.
The typed ids the connection path loads are the new `sparql_dataset_ledgers`,
which the server's own authorization, refresh and comparisons now use.
A dataset clause's prefixed name whose prefix is not declared
(`FROM ledger:main`) is again the ledger address as written; a declared
prefix still expands.
Tests: dataset unit tests `from_query_json_reports_every_named_ledger` and
`sparql_dataset_ledger_ids_reports_every_named_ledger` (string, array and
object `from`, `{"@id", "t"}`, `fromNamed` object, legacy `from-named`,
`opts.from`, `ledger`, `FROM<x>` with no space, a bare `FROM ledger:main`,
several clauses, a commented-out clause, a history range, a per-source
policy); it_query_dataset
`reported_dataset_ledgers_cover_every_ledger_the_engine_reads`, which runs a
corpus through a recording nameservice and checks that every ledger whose
record the engine reads was reported. The recording nameservice moves to
the shared test support.
A dataset that names only named graphs has an empty default graph (SPARQL 1.1 §13.2), so patterns outside `GRAPH` match nothing. HTTP says so in an `x-fdb-warning` header computed by two route-level walkers, one per language, which do not look inside `EXISTS`. Embedders and the CLI were told nothing. `QueryAdvisory` (for now `EmptyDefaultGraph`) is computed once in the API from the resolved dataset and the lowered query. One IR walker serves both languages and reaches `OPTIONAL`, `UNION`, `MINUS`, sub-queries and `EXISTS` / `NOT EXISTS`. It is carried on `QueryResult::advisories` (an additive field), and `query_from()` can return it alongside a formatted result (`execute_formatted_with_advisories`). The CLI prints each advisory to stderr on the view and connection paths; rows are unchanged. The HTTP header keeps its route-level walkers for now. Tests: it_query_dataset `a_named_only_dataset_advises_when_the_query_reads_its_empty_default_graph` (SPARQL and JSON-LD, a pattern only inside `FILTER EXISTS`, `GRAPH`-only and default-graph controls); CLI integration `a_named_only_dataset_warns_on_stderr`.
`fluree query --remote … --at` carried the time inside the query. For SPARQL it injected `FROM <ledger@t:N>`, which left a `GRAPH` block nothing to match (the dataset then has no named graphs) and refused a query that had its own FROM / FROM NAMED. For JSON-LD it overwrote `from` with `<ledger>@t:N`, so a `from` naming a graph or another member was lost. The server's ledger routes have taken a pinned path (`/query/<ledger>@t:N`, `/explain/<ledger>@t:N`) since v4.2.2. The CLI now posts the query as written to the pinned path, so a query with no dataset clause reads the ledger at that time, named graphs included, and a query's own dataset keeps what it names and is read at that time. A server that predates path pins parses the whole path as a ledger id and refuses it: v4.1.6 and v4.2.1 both answer a 500, "Invalid ledger ID format '<ledger>@t:N': expected 'name' or 'name:branch'" (run against both binaries). On an "invalid ledger id" error that names the pinned path, and only then, the CLI falls back to the older rewrite. So does an alias that names a graph, because a pin cannot ride with a `#graph` in the path. NDJSON and TSV/CSV output keep the older form until lane consolidation (part 2). Tests: CLI unit `remote_at_pins_the_ledger_path` and `only_a_server_without_path_pins_falls_back` (v4.2.1's reply verbatim, and the errors that must not fall back); server_memory `remote_at_pins_the_path_and_keeps_the_querys_dataset` (no dataset, a `GRAPH` with no dataset clause, a JSON-LD `from` naming a graph, and SPARQL FROM NAMED, each at t=1 of two commits) and `remote_at_falls_back_on_a_server_without_path_pins` (a stand-in that replies as v4.2.1 does: the pinned request, then the older rewrite on the unpinned path).
…dataset position) `LedgerRef` had a private parse mode, `UnknownPin::NotAnAddress`, that answered "not an address" for an `@` whose tag is not a known one, and `DatasetRef::parse` used it. `DatasetRef::parse` also falls back to the IRI reading for any text that fails the address grammar, and an unknown tag always fails it (the pin parser accepts only the known tags), so the mode never changed an answer. It is gone. `LedgerRef::parse` is the one address parser (an unknown tag is a malformed pin there), and `DatasetRef::parse` states the rule where it applies it. No behavior change; the unit tests for the rule are unchanged.
…r surface - CLI `issue_1972_reproduction_reads_the_graph_each_from_names`: #1972's commands as filed (a JSON-LD `from` naming a graph and `#txn-meta`, and SPARQL `FROM <mydb:main#txn-meta>`, with the ledger active and with another ledger active). - CLI `issue_1512_sparql_from_reads_the_config_graph`: SPARQL `FROM <ledger#config>` in every spelling of the address, and the keyword. - api `a_branch_reads_config_by_the_main_urn_and_urn_named_graphs_stay_graphs`: a branch reads its config graph by the main branch's URN, which it inherited, by its own URN and by keyword, and a user graph named `urn:fluree:…#…` stays a graph; SPARQL `FROM` and `FROM NAMED` + `GRAPH` on the branch's view, and the JSON-LD `from` object through `query_from()`. The resolver's unit tests held this with a hand-built registry; this runs it against a real branch. - api `solo_jsonld_dataset_shapes_still_answer`: the JSON-LD dataset shapes solo sends (a pinned object reading `txn-meta` by `t` and by `at`, an object naming a user graph, the bare and branch-qualified ledger, and a history range) answer through `query_from()`. - api `the_embedder_surface_on_the_hard_keep_list_still_compiles`: the multi-query struct literals, the agent-JSON context fields, the paths and streaming entry points embedders use, the SPARQL dataset clause, the time-travel split and conversion, and the dataset parsers an authorizer reads, named the way an embedder names them.
The pages that described behavior this branch changes, or documented a form that did not work: - query/sparql.md: a SERVICE endpoint must be a dataset member (or the query's own ledger), takes no pin or graph, is never loaded on its own, and contributes no rows under SILENT; the `urn:fluree:` endpoint spelling; SPARQL comments in the examples. `USING` / `WITH` resolve in the ledger being updated. - transactions/update-where-delete-insert.md: the JSON-LD `from` / `fromNamed` twin of that. - guides/cookbook-branching.md: the branch-diff example names both branches in its dataset, as SERVICE now requires. - ledger-config/README.md: `FROM <config>` names the config graph where a query targets one ledger (the page said it was refused). - concepts/datasets-and-named-graphs.md: on the connection endpoint a graph is named with its ledger (a bare graph IRI is a 400); the ledger-scoped keywords and address spellings; three examples put FROM NAMED before SELECT. - reference/graph-identities.md: full `https://` IRIs and a BASE do not name a ledger (never implemented); a BASE-relative dataset IRI expands to a graph IRI, and the `urn:fluree:` form is how to name a ledger under a BASE. - api/errors.md: `err:db/GraphNotFound` (404) and the 400s for another ledger or a malformed reference. - cli/server-integration.md, guides/cookbook-time-travel.md: `--remote --at` uses the pinned ledger path, with the fallback; where a SPARQL `FROM` pin works. - indexing-and-search/geospatial.md: `@t:100`, not `?t=100`. - graph-sources/overview.md: SPARQL comments. The "SPARQL Execution Modes" section of query/datasets.md (#1975) waits for lane consolidation (part 2).
One function now reads a JSON-LD query's dataset in the ledger a surface
addresses, and the server and the CLI both route by it:
- `resolve_jsonld_dataset_in_target` rewrites the body so each source names
what the target ledger reads for it: a graph of the ledger (by keyword,
registered IRI or `L#<g>`) becomes `{"@id": <ledger>, "graph": <g>}`, the
ledger's own address in any spelling stays as written, and another ledger
is refused in a lone `from` on a ledger route or left for the dataset lane.
It reports the lane the rewritten query runs on and whether it names
another ledger.
- `jsonld_lane`, `jsonld_view_ledger`, `jsonld_dataset_ledger` and
`jsonld_names_dataset` read the dataset keys with the precedence of
`DatasetSpec::from_json` (`opts` before the top level, `from` before
`ledger`, `fromNamed` before `from-named`). A lone `from` runs on a view
only when it names a whole ledger at head.
- Graph IRIs resolve through a borrowed lookup into the ledger's cached head
(`LedgerView::graph_id_for_iri`). The server's `ScopeRegistry` and the
CLI's endpoint registry no longer copy the registry per request.
They read the head with `LedgerHandle::peek`, which skips the read-side
compaction check `snapshot` runs. That check visits every graph in
novelty, and the query's own load runs it anyway: a second one per
request cost 6-8% CPU on a ledger with 20,000 named graphs (server CPU
per ledger-route request with `FROM` + `FROM NAMED`, debug build: 2.82 ms
on the base, 3.20 ms with the second check, 2.83 ms peeking).
Server: the ledger routes' `normalize_ledger_scoped_from`, the view/lane
choice (`requires_dataset_features`) and `get_ledger_id` delegate to these.
A body whose dataset sits in `opts` is now read where the parser reads it:
`{"opts": {"from": "a:main"}}` on `/query` runs on `a:main` instead of
failing with a missing ledger, and an `opts.from` naming another ledger on a
ledger route is a 400 however the top level reads.
The server test `dataset_options_and_envelope_defaults_cannot_replace_authority`
sends `{"from": L, "opts": {"from": {"@id": other}}}` with a token for L
only. The connection route still answers 404 (`other` is outside the token,
as if it did not exist). The ledger route now refuses it as the lone `from`
naming another ledger, a 400 that says no more than the request did, where
it used to read the top-level `from` and leave `other` to the token check.
CLI: the router resolves the body with the same function. A local query
whose dataset leaves the view is sent to the connection path with the body
its ledger resolved, so `"from": "config"` or a registered graph IRI reaches
the connection path named as that ledger's graph.
A JSON-LD query on a view read its dataset keys only to refuse another ledger. Anything else a dataset can say was ignored and answered from the view: a time pin, a graph of the view's own ledger (by address, keyword or object), named graphs, a history range, or a dataset that does not parse. The view now reads the dataset as `DatasetSpec::from_json` does (`opts` first) and accepts only the ledger's own address, in any spelling, as a default graph with nothing of its own beside the address. Everything else is a 400 that points at `query_from()`. Every view entry point applies it: buffered and tracked, with and without the graph-source providers, and the streaming planner. The connection path's one-ledger shortcut and a single-ledger dataset hand a view that already holds the query's dataset (loaded at the member's time and graph) to the same entry points; they mark the execution options so the view does not read the dataset keys a second time.
… 500 A graph-source probe or lookup whose nameservice call failed was reported as `QueryError::Internal`, which the API and the server map to 400 with every other query error. A backend outage then read as a caller mistake: clients did not retry it and monitoring counted it as bad requests. `QueryError::Nameservice` now carries these failures from every graph-source site (the R2RML provider's probe, mapping, table and SQL lookups). The API maps it to 500, and the server types it as the nameservice error it is (`errors::NAMESERVICE`, 500), both ahead of the generic query arm. Tests: the fail-closed lane tests assert the 500. `a_failed_source_lookup_ still_fails_the_query` now asks about a `GRAPH` IRI on the ledger's view, where the probe still runs (a dataset member's `GRAPH` is never probed), and the graph-source lane test no longer pins what an undeclared graph source answers through `GRAPH` on a plain view; it keeps the fail-closed cases.
…ling `GraphSource::identifier` (what `DatasetSpec::from_query_json` reports) and `sparql_dataset_ledger_ids` reported an address as written minus its pin, so `urn:fluree:g:main` came back with its `urn:fluree:` wrapper. A consumer that authorizes the reported ledgers compares `name[:branch]` strings, and a wrapped one matched nothing. An address is now reported as `name[:branch][#graph]` regardless of spelling: the `urn:fluree:` wrapper and the `@` pin are removed, and a bare name stays bare (it is not rewritten to `name:main`). Anything that is not an address is still reported as written. `reported_dataset_ledgers_cover_every_ledger_the_engine_reads` no longer canonicalizes the reported side. It reads each reported string as a consumer does (dropping `#graph` and `@pin`, keeping the branch, a bare name standing for its `main` branch) and checks that it covers every ledger the engine read. Its corpus now also covers `opts` over the top level, named objects, a sub-query's own dataset keys, and SERVICE, nested or not.
…ne time core: `DatasetRef::parse` read `urn:fluree:L:main@t:abc` as a graph IRI, so the malformed pin surfaced as a graph that does not exist, while the same address without `urn:fluree:` reported the pin. Both spellings now report it. An address with nothing after its `#` (`dsx:main#`, `urn:fluree:dsx:main#`) is likewise the address's error, a 400, instead of a graph IRI with an empty fragment; an IRI that is no address keeps its empty fragment (`http://ex.org/ns#`). query: a SERVICE naming a ledger read whichever of that ledger's members it found first, and named members sit in a map with no order, so a dataset holding the ledger at two times (two pins are two members) answered from either. `DataSet::service_member` picks the first default graph, else the named member first by name, and refuses a ledger the dataset holds at more than one time (400, or no rows under SILENT).
…one way This picks up the rest of the review comment on #1997 (#1997 (comment)). #1997 made the ledger's own address name its default graph in an update's default-graph positions (SPARQL `WITH`, a JSON-LD top-level `graph`); every other graph position still read it as a graph IRI, so a template, a data quad or a TriG block named by the address created a graph under it, and a `GRAPH ?g` bound to it through `USING NAMED` read the default graph but wrote such a graph. One table now answers what an IRI names in a graph position of a ledger (`TargetLedger::graph_position`), for reads and writes alike: - the ledger's own address, in any spelling: its default graph; - the address with an IRI graph (`L#<g>`): the graph `<g>`, created by a write when the ledger does not have it; - any other IRI: the graph registered under it exactly, or none (a write creates it). No write creates a graph under an address of this ledger: `L#<L>` with no such graph and the address with a time are refused as write targets. A reserved graph keeps the IRIs it is registered under. Only an IRI that starts with the ledger's name is parsed; any other costs one registry lookup. Where it applies: - transact: every template graph, `Txn::write_graphs` (`CREATE GRAPH` included) and a sync target resolve through the table before staging, and each `GRAPH ?g` binding through the same table as the WHERE streams, so a `?g` template writes the graph its WHERE read. The template-default special case #1997 added (`Txn::template_default_graph`, `TripleTemplate::graph_from_template_default`) is removed: it is a case of the table. Graph management keeps registry semantics for the graphs it names but does not create one under the address. - import: a bulk-import TriG block named by the address loads into the default graph, and `L#<g>` into `<g>`. - The update's WHERE with no `USING` reads `GRAPH <iri>` through the table (a `DataSet` name resolver, one lookup per name the dataset holds no key for), and a single-ledger query's `GRAPH <iri>` and `GRAPH ?g` do the same. A graph registered under the address before this change is listed by `GRAPH ?g` as `L#<address>` on both sides, which reads it back. - JSON-LD: a top-level `graph`, a `["graph", …]` template or a node `@graph` spelled as a ledger address that strict compact-IRI expansion refuses (`mydb:main` with no `mydb` prefix) is kept as written, as `from` is, and staging accepts it only as an address of the ledger it writes.
One contract for what a graph position reads and writes for the ledger's address, in concepts/datasets-and-named-graphs.md, which the SPARQL and JSON-LD update pages now point to instead of stating their own. It picks up the rest of the review comment on #1997 (#1997 (comment)): the address in any spelling is the default graph in every graph position, `L#<g>` is the graph `<g>`, no write creates a graph under the address, a `GRAPH ?g` template writes the graph its WHERE read, and a graph an earlier version registered under the address is reached, and listed, as `L#L`, with the graph-management recipe that moves it into the default graph. sparql.md drops the paragraph that limited the address to `USING` and `WITH` (it contradicted the one after it), and says the same of `GRAPH <iri>`, templates and data quads. The connection route's refusal is worded as it behaves: a keyword, or an IRI that cannot be a ledger address, names no ledger; an IRI that could be one is looked up as a ledger.
…d graph A graph position read the ledger's own address with a reserved keyword (`L#config`, `L#txn-meta`) as a plain IRI, so a write named that way created a user graph called by that text: a data quad, a template `GRAPH`, `WITH`, `CREATE GRAPH`, and the JSON-LD forms. The same graph's `urn:fluree:` name wrote the config graph, and a write to `urn:fluree:L:main#txn-meta` was refused. The address with a reserved keyword, in any spelling and with no time, now names that reserved graph in every position, reads and writes alike, as its `urn:fluree:` form does (`TargetLedger::reserved_graph_iri`): - a graph position writes the config graph, and a `#txn-meta` write meets the existing reserved-graph refusal; - a dataset position reads it first, as it reads the address alone, before the exact-registry step; a graph an earlier version registered under the literal text is reached, and listed by `GRAPH ?g`, as `L#<text>`; - CLEAR, DROP, ADD, COPY and MOVE refuse it as they refuse the `urn:fluree:` form. Docs: the graph-position table and the SPARQL update page. Tests: core `an_address_with_a_reserved_keyword_names_the_reserved_graph`; it_named_graphs `test_an_address_with_a_reserved_keyword_names_the_reserved_graph` (each write position in both languages, the txn-meta and graph-management refusals, WHERE and FROM reads) and `test_a_graph_registered_under_a_reserved_keyword_address_is_reached_through_it`; server `ledger_route_reserved_keyword_addresses_name_the_reserved_graphs`; CLI `local_reserved_keyword_addresses_name_the_reserved_graphs`.
CLEAR, DROP, ADD, COPY and MOVE named graphs by their exact registry IRI, so a graph named through the ledger's address was out of their reach. With a graph an earlier version registered under `L#config`, `ADD`, `COPY` or `MOVE GRAPH <L#L#config>` reported that the source does not exist, and `DROP GRAPH <L#L#config>` succeeded without dropping anything. They now read `L#<g>` as the graph `<g>`, as every other update position does, and the address with a reserved keyword as that reserved graph, which they refuse (`TargetLedger::graph_management_iri`). The address itself (`L`, `L#default`) keeps registry semantics: `DROP GRAPH <L>` is the graph registered under `L`, never the default graph. Docs: the graph-position page and the SPARQL page. Test: it_named_graphs `test_graph_management_reads_a_graph_named_through_the_address`: for a graph registered under `L#config` and for one under `L#txn-meta`, ADD, COPY, MOVE (from it, and back into it as the destination) and DROP through `L#L#<keyword>` act on that graph; a plain graph is reached through `L#<g>`; `DROP GRAPH <L>` acts on the graph registered under the address and leaves the default graph untouched. Core unit assertions for `graph_management_iri`.
|
Full mutation log for this PR's non-vacuity runs, moved out of the description to keep it under GitHub's size limit. Each new regression test was run against its fix reverted (the file text saved, mutated, the named test run, the text restored, and the tree checked clean after each). Every mutation below turned its test red. Earlier rounds ran on the commits as they were then, before the rebase onto api: typed dataset members (ac7394c)
query/api: graph-source probes fail closed (3591cb4)Test: it_sql_pushdown_lane::graph_source_probes_fail_closed_and_skip_ledger_graphs.
api: graph() and db() take any ledger address; view refuses another ledger's from (9dde79a)
query: SERVICE endpoint names a dataset member or the query's own ledger (6597d79)
transact: an update's WHERE dataset resolves in the transaction's ledger (1e6c466)
server: the ledger routes resolve dataset references in the path's ledger (9e486b6)
cli: route a query after resolving its dataset in the target ledger (6c72522)
api: keep the dataset parsers' public shape (6b1a876)
api/cli: typed query advisories (cb9cd66)Test: it_query_dataset::a_named_only_dataset_advises_when_the_query_reads_its_empty_default_graph.
cli:
|
The ledger route parses its path into an id once, then parsed that id again twice per request that names a dataset: to find the ledger's cached handle for the graph registry (`ledger_cached(id.as_str())`), and to read a `FROM <L>` written exactly as the id. `Fluree::ledger_handle(&LedgerId)` finds the cached handle by an id already parsed; `ledger_cached` parses, then does the same. It does not await `ledger_handle`: that nests one more future in every caller's, and a test future already at the type-layout depth limit (`it_absent_subject_scan_narrowing`, default features) overflowed it. `resolve_in_target` recognizes the target's id exactly as it stands, before parsing anything, as its default graph: what resolution's first step reads for that spelling, registry or not. An id stored before the current grammar, with `@`, `#` or `://` in it, still takes the full parse, and debug builds check the shortcut against it. Per request, in-process: 5 fewer allocations and 2 fewer reallocations on the dataset_http w1 and c shapes, 2 and 1 fewer on w2. Test: target_dataset `the_targets_own_id_reads_as_its_other_spellings` (the id as it stands, the short name, the `urn:fluree:` form and `L#default`, with a graph registered under the id's text, with and without a registry; a stored id with `@` in its name is not read as the address it spells).
…ames To resolve a request's dataset references, the ledger route and the CLI router built a whole `LedgerView` of the cached head, copying its nameservice record's strings, and boxed it, for one lookup per reference. `LedgerHandle::graph_names` returns only what that lookup reads, the snapshot's graph registry and the binary index store's, as two shared handles (`GraphNames`). `LedgerView::graph_id_for_iri` and `GraphNames::graph_id_for_iri` share one lookup, in the order every read path uses. `LedgerHandle::peek`, which this branch added for the route and the CLI, is gone, and `LedgerHandle::snapshot` is as it was on main. Per request, in-process: 3 fewer allocations on the dataset_http w1, w2 and c shapes.
A `GraphSource` kept its text twice, as written and as its identifier, though the two are the same text unless the member is spelled `urn:fluree:…` or with a pin, or given as an object's `@id`. It now keeps the identifier, and the text as written only where the two differ; `written()` and `name()` read the same text as before. `GraphSource::ledger` builds `name:branch[#graph]` once, at its final length: it was formatted, regrown twice and copied into a new allocation. `GraphSel::as_str` gives a graph's text for both it and `Display`. The public `identifier` field is unchanged; it is derived from the parse, to be read, not set. Per request, in-process: 3 fewer allocations and 2 fewer reallocations on the dataset_http w1 and c shapes, 1 fewer allocation on w2. Test: dataset `a_member_keeps_its_text_once_unless_it_differs_from_its_identifier` (the written text is the identifier's own storage for an address `GraphSource::ledger` builds, a canonical address, a graph IRI and a keyword, and separate for a `urn:fluree:` or pinned spelling).
… copying IRIs `TargetLedger::graph_position` copied every IRI it read into a new `Arc<str>`, and a single-ledger `GRAPH <iri>` reads one up to seven times per request (`single_db_user_graph_id` and its neighbours), dropping each copy. `GraphPosition` now borrows the text it was asked about, the whole text or the `<g>` of `L#<g>`; only a reserved graph's `urn:fluree:` IRI, which the text does not contain, is owned. A write takes ownership where it records the graph (`WriteGraph::Named`), as it did before. Resolving a member copied its IRI again when the registry held it (`TargetLedger::resolve` and `TargetLedger::graph`); it now shares the member's own `GraphIri`. Per request, in-process: 7 fewer allocations on the dataset_http d shape, 2 fewer on w1, w2 and c. Test: dataset_ref `graph_positions_and_resolved_graphs_copy_no_iri` (a position's IRI is the caller's own text, the `<g>` of `L#<g>` included, and a resolved or selected graph shares the member's `GraphIri`).
This is more than any one issue below needs, on purpose. Dataset and graph references were strings that each site interpreted on its own, and the sites disagreed: SPARQL couldn't read the config graph JSON-LD could (#1512),
graph()refused namesdb()took (#1961), and the CLI silently ignored afromnaming the ledger's own graph (#1972). So rather than patch each site, this parses a reference once, at the edge, into a typed value, with one grammar and one resolver for reads and writes, so every surface reads it the same way and there's one place to reason about and tune. Fwiw, it's one of five PRs taking this approach, with #2006, #2008, #2009 and #2010.A dataset or graph reference —
FROM,FROM NAMED, JSON-LDfrom/fromNamed,USING/WITH, aSERVICEendpoint,db()/graph(), the ledger path of a route — used to be text that every surface split and classified on its own. Before this PR there were seven parsers of the address grammar and nine string-shape classifiers, and they disagreed about#,@andurn:fluree:. This PR parses a reference once, at the edge, into typed values influree-db-core, and resolves it through one table. Most of the fixes below fall out of that: each was a surface where a classifier guessed wrong.This is part 1 of two. It changes how references are parsed and resolved, not which execution lane the server or the embedded API picks for a shape: each shape runs where it ran before, with two exceptions that follow from reading a query the way the engine's parser reads it. A JSON-LD dataset written under
optsnow takes the lane the same dataset takes at the top level (the parser has always readoptsfirst), and the CLI's router sends a query whose dataset names a graph or another branch to the connection path, which reads it (#1972). Moving shapes between lanes is lane consolidation, which comes in a follow-up PR (part 2); see the end.Where #1997's review left off. This also picks up the rest of bplatz's review comment on #1997: #1997 (comment). #1997 made the ledger's own address name its default graph in an update's default-graph positions (
WITH, a JSON-LD top-levelgraph). This PR makes it do so in every graph position of a query and an update, including templates, data quads, TriG blocks and bulk import — see "Graph positions" below.Fixes #1972
Fixes #1961
Fixes #1982
Fixes #1512
Follow-up: #1975
Follow-up: #1996
What changes for users
Grammar
A ledger address is
[urn:fluree:]name[:branch][@<tag>:<value>][#<graph>], parsed by one parser (LedgerRef::parse). Two rules keep it from swallowing graph IRIs:/.http://ex.org/gused to parse as ledgerhttp, branch//ex.org/g; it is now never a ledger id. The rule lives inLedgerId::parse, so it holds for every caller of that function, not only at the input edges. Stored records deserialize throughLedgerId::parse_persistedand are unaffected, and branches that merely contain/(feature/x) still parse.@starts a time pin only before a known tag (t:,time:/iso:,commit:,recorded:,snapshot:). In a ledger position any other@is a malformed pin (400); in a dataset position it means the text is not an address, so<http://ex.org/@alice/g>andmailto:a@bare graph IRIs.A dataset-position string is classified as an address, a graph IRI, or ambiguous (lexically both, like
mydb:main):scheme://…or two or more:before any#/@is a graph IRI (aurn:fluree:address excepted); a pinned address is an address;name:branchis ambiguous; a barenameis an address. An address with a known pin tag and a bad value (mydb:main@t:abc) or with nothing after its#(mydb:main#) is a malformed address (400) in either spelling,urn:fluree:included; an IRI that is no address keeps an empty fragment (http://ex.org/ns#). SPARQL dataset IRIs are expanded against the prologue (prefixed names,BASE) on every route before classification. An IRI written in full is taken as written, and a prefixed name whose prefix the query does not declare (FROM ledger:main) is the address as written.One resolution table
Where a surface reads one ledger (a ledger route, a view, an update's
WHERE, the CLI with a target), each reference in a dataset position resolves in that ledger, in order:mydb,mydb:main,urn:fluree:mydb:main), with no pin and no graph: the default graph.urn:fluree:mydb:main#configresolves on a branch, and a user graph may itself be namedurn:fluree:….default,txn-meta,config) or an address of this ledger with a graph (mydb:main#config,mydb#http://ex.org/g): that graph.err:db/GraphNotFound.The connection route (
POST /query,query_from()) has no target ledger, so there an address names a ledger (and its graph). A keyword, or an IRI that cannot be a ledger address (http://ex.org/g), is a 400 that says to name the ledger, instead of a nameservice lookup of the IRI as a ledger id. An IRI that can also be read as an address (urn:g1,ex:g) is still looked up as a ledger there, and is a 404 when there is none.A
FROM NAMEDmember is named by its IRI as written, pin included:FROM NAMED <L@t:2>answersGRAPH <L@t:2>,GRAPH ?gbindsL@t:2, and two pins of one ledger are two members. The time-stripped name (GRAPH <L>) still matches while exactly one member claims it; it is not enumerated byGRAPH ?g.Graph positions
Every graph position of a ledger reads the ledger's own address one way, for reads and writes alike:
GRAPH <iri>in a query or in an update'sWHERE, an update template'sGRAPH <iri>, anINSERT DATA/DELETE DATAquad, a TriG block in a transaction or a bulk import,WITH, and the JSON-LD forms (top-levelgraph, a node's@graph,["graph", …]). One table influree-db-core(TargetLedger::graph_position) answers:mydb:main)mydb:main#config,mydb#txn-metaurn:fluree:name reads it#txn-metawrite is refusedmydb:main#http://ex.org/ghttp://ex.org/ghttp://ex.org/gNo write registers a graph under the ledger's own address. The address with a time names no graph in these positions, and the reserved graphs keep the IRIs they are registered under. A
GRAPH ?gtemplate writes the graph its binding names in this table, which is the graph theWHEREread. Every resolution is one borrowed registry lookup; only an IRI that starts with the ledger's name is parsed as an address.What changes, on ledger
mydb:main:GRAPH <mydb:main>(any spelling)GRAPH ?gdid not listGRAPH <mydb:main#http://ex.org/g>(and the quad, TriG and import forms)mydb:main#http://ex.org/ghttp://ex.org/g, the graphFROM <mydb:main#http://ex.org/g>readsINSERT { GRAPH ?g {…} } USING NAMED <mydb:main> WHERE { GRAPH ?g {…} }GRAPH <mydb:main>in an update'sWHEREwith noUSINGmydb:main, if anyGRAPH <urn:fluree:mydb:main>orGRAPH <mydb>in a query (GRAPH <mydb:main>already read the default graph)GRAPH ?gover a graph registered under the address by an earlier versionmydb:mainin an update'sWHEREmydb:main#mydb:mainby both, a name that reads it backGRAPH <mydb:main@t:1>, or toGRAPH <mydb:main#mydb:main>when no such graph existsCOPY/MOVE/ADD … TO <mydb:main>with no graph registered thereDEFAULTCLEAR/DROP/ADD/COPY/MOVEnamingGRAPH <mydb:main#<g>><g>;DROP GRAPH <mydb:main>still names the graph registered under the address, never the default graphCREATE GRAPH <mydb:main>mydb:mainGRAPH <mydb:main#config>or<mydb#txn-meta>, in any form#txn-meta, as for theurn:fluree:names"graph": "mydb:main"(top-level,["graph", …]or a node's@graph) with nomydbprefix defined"from": "mydb:main"already read; a name that is no address of the ledger being written is still that 400A graph an earlier version registered under the address keeps its data. It is reached as
<mydb:main#mydb:main>(one undermydb:main#configas<mydb:main#mydb:main#config>) in every position above and by the graph-management verbs, which also reach it by the address alone. To move it into the default graph:ADD GRAPH <mydb:main> TO DEFAULT ; DROP GRAPH <mydb:main>. The contract is written once, indocs/concepts/datasets-and-named-graphs.md("The ledger's own address in a graph position"), and the SPARQL and JSON-LD update pages point to it.JSON-LD update templates take no graph variable, so the
GRAPH ?gshapes have no JSON-LD twin; every fixed-graph shape has one.Status codes
FROMerr:db/GraphNotFound(new@type)#), in either spellingFROMmember on a view (one view is one snapshot)query_from()SERVICEto a ledger that is not a dataset memberSILENTSERVICEto a ledger the dataset holds at more than one timeSILENTerr:system/NameServiceErrorL#Lwith no such graph, or a graph-management destination at the addressPer surface
/query/{L},/stream/query/{L},/explain/{L}, /query/{ledger} parses the path ledger once but authorizes, loads and compares its raw spelling #1982): the path is parsed once into the canonical id, and auth, loading and every comparison read that id./query/urn:fluree:L:mainloads, and/query/LwithFROM <L:main>(or the reverse) is no longer a mismatch. EachFROM/FROM NAMEDand JSON-LDfrom/fromNamedresolves through the table with the ledger's graph registry: an IRI containing#or@is a graph, not ledger syntax; the short alias and the URN of the path's own ledger name it; a graph IRI infromNamedis a graph. A single JSON-LDfrommust name the path's ledger; afromarray andfromNamedmay still name other ledgers, as before. The JSON-LD dataset is read where the engine's parser reads it (optsbefore the top level,frombeforeledger,fromNamedbeforefrom-named), so anopts.fromnaming another ledger is a 400 however the top level reads, and anoptsdataset takes the lane its top-level twin takes (it used to be ignored:{"opts": {"from": "L@t:1"}}answered at head)./explain/{L}accepts the within-ledger graph IRIs/query/{L}accepts (partly: see Deviations).POST /query): the ledger a JSON-LD body's view reads, and the ledger its request is labeled and authorized by, are read with the same precedence, so{"from": "a:main", "opts": {"from": "b:main"}}runs onb:main, which the parser reads, and{"opts": {"from": "a:main"}}runs ona:main(it failed with a missing ledger).graph()anddb()(graph() does not accept #graph fragments or urn:fluree: ids that db() accepts #1961):graph()takes every addressdb()takes (L#txn-meta,L#<graph IRI>,urn:fluree:L,L@t:5). A pin in the address is honored by both; one that disagrees withgraph_at's is a 400..transact()through a graph-qualified or pinned handle is a 400, since a transaction applies to the whole ledger at head.query_from(), which reads them. The connection path's one-ledger shortcut and a one-ledger dataset hand the view a dataset it already holds, and are not refused. SPARQLFROM <L#config>,FROM <L#iri>,FROM <L>and keywords resolve on the view (Config Graph query doesn't work with SPARQL #1512).SERVICE <fluree:ledger:X>: X must be a dataset member, or, with no dataset, the query's own ledger. Anything else is a 400 that points atFROM NAMED, and contributes no rows underSILENT. An endpoint is never loaded on the spot.urn:fluree:spellings work; a pin or a graph in an endpoint is malformed. X is read at one time: a dataset holding X at two times (two pins are two members) is refused the same way, and with one time the member read is fixed (the default graph, else the first named member by name) instead of whichever a hash map yielded first.USING,USING NAMED,WITHand the JSON-LD update keysfrom,fromNamedandgraphresolve in the transaction's ledger through the same table. The ledger's own address names its default graph, in the single and the multi-USINGbranch alike;L#confignames the config graph; a registered IRI names its graph; a graph the ledger does not have contributes nothing, soDELETE … USING <typo> WHERE …deletes nothing (it used to read the real default graph). Another ledger or a pinned address is a 400. JSON-LD"from": "L:main"is accepted (it was an undefined-prefix error), and so is"graph": "L:main"for the ledger being written. fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997's check for the ledger's own address and resolution step 0 are one predicate now (LedgerRef::is_own_address), soL#defaultreads the same in both. The graph positions are above.L#<g>into<g>, through the same table.R2rmlProvider::has_r2rml_mappingreturns aResult, and a failed lookup fails the query instead of reading as an empty native graph. Graphs loaded as ledger graphs (the primary and native dataset members) are never probed, so a query over native data does not depend on the graph-source registry being reachable. A failed nameservice lookup at any graph-source site (the probe, and the mapping, table and SQL-source lookups) is a newQueryError::Nameservice, which the API and the server answer with 500err:system/NameServiceError, the typeApiError::NameServicealready has: it is the backend's fault, so clients may retry it and monitoring does not count it as a bad request.fluree query: afromnaming the target ledger's own named graph or#txn-metais silently ignored #1972):fluree querydecides between the target's view and the connection path by resolving the dataset in the target ledger, not by comparing base names. JSON-LD bodies go through the same API function the server uses (see Public Rust API), with the parser's key precedence. A JSON-LDfromnaming one of the ledger's graphs (by keyword, registered IRI or address, in string or object form), its#txn-meta, or another branch now reads what it names, and one naming a missing ledger reports it missing: the body reaches the connection path with each graph of the ledger rewritten as{"@id": <ledger>, "graph": <g>}.fluree query --remote … --atposts the query as written to the pinned path (/query/{L}@t:N, available since v4.2.2): aGRAPHwith no dataset clause reads the ledger's named graph at that time, a JSON-LDfromnaming a graph is kept (it was overwritten), and a SPARQL query with its ownFROMis accepted (it was refused). A server older than v4.2.2 refuses the pinned path; on that reply only, the CLI falls back to the old rewrite (see Deviations). Locally,--atwith a JSON-LDfromnaming a graph of the ledger is now a usage error ("--atis not supported on the connection/federated query path"); it used to read the wrong graph without a word.QueryResult::advisoriesnow carries a typedQueryAdvisory::EmptyDefaultGraphwhen the query reads that default graph, computed once from the lowered query for both languages, including patterns insideEXISTS/NOT EXISTS. The CLI prints it to stderr; rows are unchanged. The HTTPx-fdb-warningheader keeps its route-level walkers for now.Public Rust API
The typed references change a few public shapes, so embedders will notice:
LedgerRef's fields become private typed accessors,GraphSource's parsed fields become accessors (GraphSource::parse/TryFrom<&str>/FromStrreplacenew/from_str),HistoryTimeRangetakes aLedgerId,R2rmlProvider::has_r2rml_mappingreturns aResult,ApiError::GraphNotFoundandQueryError::Nameserviceare new variants on enums that aren't#[non_exhaustive], andDataSetDbandTxngain pub fields that a struct literal needs. The embedder surface on the hard keep list is pinned by a compile test (the_embedder_surface_on_the_hard_keep_list_still_compiles). The full list:Every public API change, and what's kept source-compatible
Changed:
fluree-db-core: new moduledataset_refwithGraphIri,GraphSel,LedgerRef,DatasetRef,MemberRef,TargetLedger,TargetGraph,TargetError,GraphPosition.LedgerRefreplaces the raw-string struct of the same name (pubid,at: Option<String>,fragment: Option<String>,time_spec()): its fields are private and typed (id(),at(),graph(),is_bare(),is_own_address(),with_at,with_graph,into_parts).TargetLedgerresolves dataset positions (resolve,graph) and graph positions (graph_position,enumeration_name,names_this_ledger).TimeSpecandACCEPTED_TIME_SPEC_SPELLINGSmove here fromfluree-db-api, re-exported at both old paths,From<LedgerIdTimeSpec>included.LedgerId::parserefuses a branch beginning with/.fluree-db-api::GraphSourcekeepspub identifier: Stringandpub policy_override. The pub fieldstime_spec,source_aliasandgraph_selectorbecome accessors (time_spec(),alias(),reference(),address(),written());GraphSource::new/from_str/with_graphbecomeGraphSource::parse/TryFrom<&str>/FromStr/GraphSource::ledger(LedgerRef).with_time,with_alias,with_policyare unchanged.dataset::GraphSelectoris an alias ofGraphSel(Iri(String)becomesNamed(GraphIri), plusConfig).ledger_info::GraphSelectoris a different type and is untouched.DatasetSpec::from_sparql_clause(&SparqlDatasetClause)becomesfrom_sparql(&ResolvedDatasetClause)/from_sparql_ast(&SparqlAst). New:DatasetSpec::sources(),ledgers().HistoryTimeRange { identifier: String, … }becomes{ ledger: LedgerId, … }, and so doesHistoryTimeRange::new.fluree-db-api: moduletarget_dataset, re-exported at the root, the one reader of a JSON-LD dataset in the ledger a surface targets:resolve_jsonld_dataset_in_target(rewrites the body so each source names what the target reads for it, and reports the lane),resolve_in_target,jsonld_lane,jsonld_view_ledger,jsonld_dataset_ledger,jsonld_names_dataset,InTarget,JsonLdLane,JsonLdInTarget,SingleFrom,GraphLookup;LedgerView::graph_id_for_iri(a borrowed lookup into the cached head),LedgerHandle::graph_namesandGraphNames(that lookup alone),Fluree::ledger_handle(a cached handle by a parsed id);sparql_dataset_ledgers(the typed ids the connection path loads),sparql_has_dataset_clause,QueryAdvisory,QueryResult::advisories(additive),execute_formatted_with_advisories,ApiError::GraphNotFound(ApiErroris not#[non_exhaustive], so an exhaustive match needs the arm),DataSetDb::named_aliases(a pub field, so a struct literal ofDataSetDbneeds it) andwith_named_alias. Re-exported:resolve_dataset_clause,ResolvedDatasetClause.Fluree::is_single_ledger_fast_path(test-only callers).fluree-db-query:R2rmlProvider::has_r2rml_mappingreturnsResult<bool>, so implementors must change (the three in-tree ones and the test doubles did). New variantQueryError::Nameservice(String):QueryErroris not#[non_exhaustive], so an exhaustive match needs the arm. NewMemberKindonGraphRef::kind(a pub field) andContextConfig::primary_kind/ExecutionContext::primary_kind, plusExecutionContext::graph_is_native,DataSet::service_memberandDataSet::with_name_resolver(GraphNameResolver).fluree-db-transact:Txn::address_graph_names(a pub field, so a struct literal ofTxnneeds it);WriteGraph,GraphTableandFlakeGenerator::set_graph_table. fix: unresolvable graph references fail closed, and TriG directives apply in document order #1997'sTxn::template_default_graph,TripleTemplate::graph_from_template_defaultandTripleTemplate::in_template_default_graphare removed: that case is a row of the table (no release carried them).fluree-vocab:errors::GRAPH_NOT_FOUND.Kept source-compatible, and pinned by a compile test (
the_embedder_surface_on_the_hard_keep_list_still_compiles):MultiQueryRequest { context, as_of, opts, queries }andMultiQuerySubquery { language, query, opts }as struct literals; theAgentJsonContextfields;fluree_db_api::dataset::GovernanceOptions;fluree_db_core::validate_branch_name;ledger_info::GraphSelector::ByIri;build_stream_dataset,build_stream_dataset_for_sparql,plan_stream_query_dataset,run_stream_query_dataset,OwnedStreamQuery,run_stream_query;fluree_db_sparql::resolve_dataset_clauseandResolvedDatasetClause { default_graphs, named_graphs, to_graph };fluree_db_core::split_time_travel_suffix;TimeSpec::from(LedgerIdTimeSpec);DatasetSpec::from_query_json,GraphSource::identifier/policy_override, andsparql_dataset_ledger_ids.Commits
28 commits on
61b836e9a:The commit list
7ecf7cd00bench: dataset_http and dataset_buildd07f13ca3core: typed dataset and graph referencesac7394cacapi: typed dataset members, one resolver for a target ledger3591cb400query/api: graph-source probes fail closed; ledger graphs are never probed9dde79a4aapi: graph() and db() take any ledger address; a view refuses another ledger's from6597d7936query: a SERVICE endpoint names a dataset member or the query's own ledger1e6c46685transact: an update's WHERE dataset resolves in the transaction's ledger9e486b663server: the ledger routes resolve dataset references in the path's ledger6c7252243cli: route a query after resolving its dataset in the target ledger6b1a876b7api: keep the dataset parsers' public shape for callers that authorizecb9cd6652api/cli: typed query advisories on the result0f512cdefcli:--remote --atpins the ledger path159487c41core: one home for the unknown-@-tag rule (an unknown@tag in a dataset position)92b8d2f52test: the filed reproductions, solo's dataset shapes, and the embedder surfaced7e4372d0docs: dataset references, SERVICE membership, update datasets, CLI --atacfbaeb89api: resolve a JSON-LD dataset in the ledger a surface targetsb318ab3eaapi: a view answers a JSON-LD query only for its own ledger, wholeeb1d92815query/api/server: a nameservice lookup that fails inside a query is a 500a17fbfd63api: report every dataset address as name[:branch], whatever its spelling8b4dd8403core/query: a malformed URN address is its own error; SERVICE reads one time23450ec00core/query/transact: every graph position reads the ledger's address one waye2e13e365docs: the ledger's own address in a graph positionc2adf42a2core/transact: the address with a reserved keyword names that reserved graph206071c3ftransact: graph management readsL#<g>as the graph<g>aa75eb960server/api: resolve a ledger route's dataset with its path id as parsed5c2deab2dapi/server/cli: resolve dataset references against the head's graph names5f679ef75api: keep a dataset member's text onceed58916cdcore/query/transact: read graph positions and resolve members without copying IRIsTests
New integration tests cover each surface — the API's dataset and graph-handle suites,
SERVICE, the update-dataset matrix, graph positions, graph management and bulk import, the address with a reserved keyword (L#config,L#txn-meta), the graph-source probes against real SQL sources, the server's ledger routes, and the CLI's routing (including #1972 and #1512 reproduced from the issues verbatim) — each with its SPARQL and JSON-LD twin where the surface has both, plus unit tests for the grammar, the classifier and both resolution tables. Seven existing tests changed their expectation on purpose, including two of #1997's and the server's policy regression test; each one is explained in the fold.Each new regression test was also run against its fix reverted (the file text saved, mutated, the named test run, the text restored, and the tree checked clean after each), and every mutation turned its test red. The full mutation log is in the first comment below: it would push this description past GitHub's size limit.
The new tests, and the existing tests whose expectation changed on purpose
New integration tests, each with its SPARQL and JSON-LD twin where the surface has both:
fluree-db-api(grp_query):sparql_within_ledger_from_accepts_every_spelling_of_the_ledger,sparql_within_ledger_from_names_a_graph_through_the_address,sparql_within_ledger_from_rejections_are_typed,pinned_from_named_members_are_named_as_written(and the compatible alias),a_connection_query_refuses_a_bare_graph_iri,connection_dataset_iris_expand_against_the_prologue,a_view_refuses_a_jsonld_dataset_beyond_its_own_ledger(afromnaming a ledger that does not exist, and every other shape a view does not read: all five entry points,optsincluded, and the own-ledger spellings that still answer),reported_dataset_ledgers_cover_every_ledger_the_engine_reads(a recording nameservice; see Solo),a_named_only_dataset_advises_when_the_query_reads_its_empty_default_graph,a_branch_reads_config_by_the_main_urn_and_urn_named_graphs_stay_graphs,solo_jsonld_dataset_shapes_still_answer,the_embedder_surface_on_the_hard_keep_list_still_compiles.fluree-db-api(grp_query,it_service_cross_ledger_iri):service_to_a_non_member_is_refused_not_answered_by_the_dataset,a_service_endpoint_is_never_loaded,a_service_to_a_ledger_held_at_two_times_is_refused(and, with one time, the same member on each of eight freshly built datasets).fluree-db-api(grp_graphsource,it_named_graphs):graph_handles_take_the_addresses_db_takes(graph() does not accept #graph fragments or urn:fluree: ids that db() accepts #1961),a_missing_ledger_reads_as_not_found_at_every_entry_point,update_where_dataset_references_resolve_in_the_transactions_ledger(each update dataset shape, cross-ledger and pinned refusals, both languages); for graph positions,test_writes_through_the_ledger_address_name_the_default_graph(the address andL#<g>in quads, aGRAPH ?gtemplate bound throughUSING NAMED(the INSERT reproduction from Claude's review), TriG, a JSON-LD node@graphand top-levelgraph; the refusals; no graph registered under the address),test_updates_on_a_graph_registered_under_the_ledger_address(the DELETE reproduction from Claude's review, and every read/write position over a graph registered under the address, both spellings),test_query_and_update_list_the_address_graph_under_one_name(the query and the update list it asL#L, that name reads it back, and a?gtemplate writes it) andtest_a_graph_registered_under_the_address_migrates_to_the_default_graph(the documented recipe).fluree-db-api(grp_import):import_trig_blocks_named_by_the_ledger_address.it_named_graphs, a server and a CLI test (SPARQL and JSON-LD, reads and writes), andtest_graph_management_reads_a_graph_named_through_the_address.fluree-db-api(it_sql_pushdown_lane, featuressql,native):graph_source_probes_fail_closed_and_skip_ledger_graphs(against real SQL sources with every graph-source lookup failing; the failure is a 500).fluree-db-server:ledger_route_dataset_references_resolve_in_the_paths_ledger(graph IRIs with#and@, the config graph by its URN, each spelling of the path and of the ledger, a prefixedFROM NAMED, a graph IRI infromNamed, an unknown graph, another ledger in each language).fluree-db-cli:local_query_reads_the_graph_or_branch_its_dataset_names,issue_1972_reproduction_reads_the_graph_each_from_names(the issue's commands verbatim),issue_1512_sparql_from_reads_the_config_graph,a_named_only_dataset_warns_on_stderr,remote_at_pins_the_path_and_keeps_the_querys_dataset(afluree server run --memorychild),remote_at_falls_back_on_a_server_without_path_pins(a stand-in that replies as v4.2.1 does), plus unit tests for the routing decisions (local_routing_hands_the_connection_the_resolved_bodyamong them) and the fallback predicate.fluree-db-core::dataset_ref(a malformed address in theurn:fluree:spelling and with nothing after#,a_graph_position_reads_the_address_as_the_default_graph,a_graph_is_listed_under_a_name_that_reads_it_back); the dataset parsers influree-db-api::dataset;fluree-db-api::target_dataset(key precedence, resolution in the target, a graph member that names a second graph); the status of a nameservice failure in the API and the server; the SERVICE endpoint grammar; the server's path id, JSON-LD normalization and key precedence.Existing tests whose expectation changed on purpose:
sparql_within_ledger_from_alias_spelling_mismatch_is_rejectedlockedFROM <wl>(the short alias) on a view as a rejection, a deferred limitation; the short alias now names the default graph, and the test is replaced bysparql_within_ledger_from_accepts_every_spelling_of_the_ledger.sparql_reserved_graphs_stay_unreachable_when_not_named_in_fullrefused the bare keywordsconfig/txn-metain a view'sFROM; a keyword is now an explicit naming of the ledger's reserved graph, as on the ledger route, so the test keeps refusing the relative#config/#txn-metaand another ledger's reserved IRI, andGRAPH ?gstill never enumerates a reserved graph.a_view_refuses_a_jsonld_from_naming_another_ledgerbecomesa_view_refuses_a_jsonld_dataset_beyond_its_own_ledger: the within-ledger shapes it let through (a pin, a graph, named graphs) are now refused too.test_updates_on_a_graph_registered_under_the_ledger_addresslockedGRAPH <address>in aWHEREand a template as that graph. It now reads and writes the default graph there, and the graph is reached asL#<address>. Since no write can register such a graph any more, its fixture, and that ofsparql_single_db_graph_alias_wins_over_colliding_named_graph, builds one the way an earlier version could leave it:mainwrites a graph named by a branch's address, and the branch inherits it.a_failed_source_lookup_still_fails_the_queryasked aboutGRAPH <maybe-source:main>through a one-ledger dataset. It keeps testing that a failed lookup fails the query, with aGRAPHIRI this PR still probes (on the ledger's view), and now also that the failure is a 500. The graph-source lane test no longer asserts what an undeclared graph source answers throughGRAPHon a plain view; it keeps the fail-closed cases.dataset_options_and_envelope_defaults_cannot_replace_authority(fluree-db-server):{"from": L, "opts": {"from": {"@id": other}}}with a token forLonly is now a 400 on/query/L(a lonefromnaming another ledger, read where the parser reads it, refused before loading) and still a 404 on/query; neither response carries the restricted document.reported_dataset_ledgers_cover_every_ledger_the_engine_readscompared both sides canonicalized, so it could not see a report in a form a consumer does not parse (aurn:fluree:wrapper stayed green). It now reads each reported string as a consumer does (dropping#graphand@pin, keeping the branch; only the id grammar's own equivalence, a bare name for itsmainbranch, is allowed), and its corpus coversoptsover the top level, named objects, a sub-query's own dataset keys, and SERVICE, nested or not.Gates
At the head,
ed58916cd(on61b836e9a), run locally: fmt is clean, and so is workspace clippy-D warnings, with--all-features, with default features and for CI's wasm32 engine-stack lint. fluree-db-api's lib and its dataset, query, transact, import, policy, ledger and misc groups (3,449), fluree-db-server (716), fluree-db-cli (465) and the W3C SPARQL suite (36 of 36 groups against its registers) pass, but for one flake. The full crate test run was on this tree with one difference,ledger_cachedstill awaitingledger_handle: 9,582 passed / 150 ignored / 1 failed over 131 test binaries, the one failure that grp_misc flake, which isn't this PR's (it failed the same way on #1997's head, and passes alone, 3 of 3 on this head).testsuite-sparql's own fmt/clippy pair wasn't run, since nothing there changed. Details:Full gate log
Run locally on
61b836e9a. On the head,ed58916cd:cargo fmt --all -- --check: clean.cargo clippy --all --all-targets --locked -- -D warnings, with--all-featuresand with default features, and CI's wasm32 engine-stack lint: clean.testsuite-sparql/is not touched, so its own fmt/clippy pair was not run.The full run, on this tree with
ledger_cachedstill awaitingledger_handle(the one difference):cargo test --all-features --locked --no-fail-fastover core, api, server, query, transact, cli and sparql, 131 test binaries: 9582 passed, 150 ignored, 1 failed,annotation_body_threshold_reduces_scan_work_on_both_surfaces, a grp_misc flake that failed the same way on #1997's head (bf523e24e) and passes alone (3 of 3 on this head).Performance
The connection path gets faster, and no case regresses in both passes.
dataset_build'squery/*cases are 3.4–9.3× faster, and the saving grows with the member count (query/within/10.108 → 0.031 ms,query/cross/161.315 → 0.157 ms), which matches what commit 4 (3591cb400) removed: the executor no longer asks the graph-source registry about native members. That attribution is inferred from the scaling, fwiw, not measured commit by commit. Over HTTP the connection-route cases are 18–50% faster in the first pass (one JSON-LDfromnaming a graph is flat there) and more in the second,build_dataset_viewalone is 0–17% faster, and the ledger-route and single-ledger stream cases sit within this box's noise (the JSON-LD wide-registry case is faster in both passes). The box was shared with other builds (1-minute load 9–22), so small differences don't mean much. (These criterion runs are from before the route stopped re-running the read-side compaction check, a change that only removes work from ledger-route and stream-route requests, and before the last four commits, which cut per-request allocations; both are below.)Quiet box. EC2 c7i.4xlarge, the fat-LTO bench profile with the server's default features, base and head in interleaved rounds; a change counts as a win or a loss only when the base and head ranges don't overlap and the median delta exceeds the bench's budget (5% at
small). Seven wins and no loss:dataset_build'squeryandquery_r2rmlcases at 16 members are 91–96% faster (2.2–5.3 ms → about 190 µs), and over HTTP the connection-route casesa_connandb_connare 45% and 80% faster, ands5, the 16-member stream, 43%. The otherdataset_httprows are stable, though that bench is noisy enough on the box that its stable rows are weak evidence either way. The 10,000-graph ledger route (w1) read +19.8% in the first session and +11.1% over 7 rounds in the second, with the head slower in 8 of the 10 rounds, which timing on that box can't resolve; by instruction count (206071c3fagainst61b836e9a, fat-LTO release builds) it's flat, with nothing that grows with the registry, so that was noise or placement. Its JSON-LD twin (w2) is 30% fewer instructions.Per-request cost. A per-request count of the w1 path (retired instructions and heap allocations, in-process, fat LTO, against base
61b836e9awith the bench commit) put the head about 1% over base on every ledger-route request. That's a constant cost, the same on a 16-graph ledger, which is why EC2's +11–20% on w1 reads as noise or placement rather than work that grows with the registry. The last four commits take that cost out: the route resolves its dataset with the path's id as parsed and reads only the head's graph names, a member keeps its text once, and graph positions and resolved members borrow their IRIs instead of copying them. Per request now: w1 652 allocations (base 657) and −0.6% instructions, c 652 (657) and −1.1%, d 554 (555) and +0.4%, and w2 432 (472) and −31%.One regression from an earlier round of this branch is worth calling out, because it's fixed here: that round copied the ledger's whole graph registry on every ledger-route or stream-route request that named a dataset, and Claude's review pass measured
/query/{L}withFROM+FROM NAMEDon a 20,000-graph ledger at 2.0 → 19.4 ms p50. Every reference now resolves with a borrowed lookup into the ledger's cached head (LedgerView::graph_id_for_iri). Re-measured with the review pass's fixture, as server CPU time per request (which load on a shared box skews far less than wall time) and mean wall time:61b836e9a/query/big:main,FROM+FROM NAMED/stream/query/big:main, the same body/query/big:main, no dataset (control)/query,FROM+FROM NAMED <big:main#…>(control)The two dataset requests now match base: 3.56 vs 3.62 ms CPU on the ledger route, and 3.58 vs 3.66 on the stream route, where the previous round took 32.49 and 34.15. The first version of this fix left a residue of 6–8% CPU on those two requests; a
sampleprofile put it inLedgerHandle::snapshot, which runs the read-side compaction check (Novelty::needs_tier_compaction, a visit to every graph in novelty), and the route ran it once more per request than the query's own load does. The route and the CLI now read only the head's graph names (LedgerHandle::graph_names), which skips it (same run: 2.82 ms CPU base, 3.20 with the extra check, 2.83 without it). The method, the wide-registry cases and the per-case numbers are folded:Full benchmark method and numbers
cargo bench --profile dev-fast --locked -p fluree-db-server -p fluree-db-api --features fluree-db-api/iceberg --bench dataset_http --bench dataset_build. BEFORE is the base61b836e9awith the bench commit (7ecf7cd00), and AFTER is this branch before the route stopped re-running the read-side compaction check (below), a change that only removes work from ledger-route and stream-route requests. Both sets of executables were built first and then run back to back in two passes: BEFORE then AFTER, then AFTER then BEFORE, so drift in load shows up as the passes disagreeing. The box was shared with other builds and test runs (1-minute load 9–22 during the runs), so small differences are not meaningful.POST /query,query_from(), and the dataset stream with 16FROM NAMEDmembers): faster in both passes.dataset_build'squery/*cases are 3.4–9.3× faster, and the saving grows with the member count (query/within/10.108 → 0.031 ms,query/cross/161.315 → 0.157 ms). That pattern matches what commit 4 removed: on every query the executor asked the graph-source registry about the primary and each member, and it no longer asks about native members. This attribution is inferred from the scaling; it was not measured commit by commit. Over HTTP the same cases are 18–50% faster in the first pass (f, onefromnaming a graph, is flat there), and more in the second.build_dataset_viewalone: 0–17% faster in both passes.No case regresses in both passes.
The wide-registry cases are new:
dsx-wide:mainregisters 10,000 graphs (250 per commit; a commit registers at most 256), and w2 names its graph asL#<g>, which the base also answers.The 20,000-graph re-measure. The previous round copied the whole graph registry per ledger-route or stream-route request that named a dataset (Claude's review pass: 2.0 → 19.4 ms p50 on 20,000 graphs); each reference is now one borrowed lookup (
LedgerView::graph_id_for_iri). Re-measured with the review pass's fixture (debug binaries, 20,000 graphs in 80 commits), the three binaries interleaved: server CPU per request (median of 3 rounds of 200), which load skews far less than wall time, and mean wall time:(The table is in the section above.)
The first version of the fix left a residue of 6–8% CPU on the two dataset requests. A
sampleprofile put it inLedgerHandle::snapshot, which runs the read-side compaction check (Novelty::needs_tier_compaction, a visit to every graph in novelty): the route ran it once more per request than the query's own load does. The route and the CLI now read only the head's graph names (LedgerHandle::graph_names), which skips it (same run: 2.82 ms CPU base, 3.20 with the extra check, 2.83 without it).Per-case numbers
dataset_http (fluree-db-server)
dataset_build (fluree-db-api,
--features iceberg)"after vs before, p2" comes from the second pass, where BEFORE ran against AFTER's saved baseline, converted to read the same way as p1.
Solo
Two changes reach solo's call sites; neither needs a lockstep code change, but its tests may need a test-only update. Solo itself was not built against this branch.
DatasetSpec::from_query_json(each member'sidentifier, andpolicy_override) andsparql_dataset_ledger_ids. Both keep their signatures and fields. Identifiers are reported in thename[:branch]form regardless of spelling: an address'surn:fluree:wrapper and@pin are removed and its#graphkept, a bare name stays bare (not rewritten toname:main), and anything that is not an address is reported as written. Before, aurn:fluree:spelling came back with its wrapper, which an authorizer comparingname[:branch]strings matched against nothing. Solo's extractor tests that pin a raw URN string need that update.sparql_dataset_ledger_idsreports every member's identifier (without#graph) and the history range's ledger, as before. A reference whose reading depends on a ledger's graph registry (mydb:mainis a ledger or a graph IRI) is reported too.from_query_json_reports_every_named_ledger/sparql_dataset_ledger_ids_reports_every_named_ledgerport solo's extractor expectations, each dataset key and spelling included, andreported_dataset_ledgers_cover_every_ledger_the_engine_readschecks, for each query in its corpus, the report the way a consumer reads it against every ledger whose record the engine reads (see Tests). The typed ids db itself loads are the newsparql_dataset_ledgers.QueryError::Nameservice: a new variant, so an exhaustivematchonQueryErrordownstream needs an arm. A nameservice failure inside a query is now a 500 where some were a 400. (Solo'sApiErrormatches have wildcard arms, and solo implements noR2rmlProvider.)urn:fluree:…#…: resolution step 2.a_branch_reads_config_by_the_main_urn_and_urn_named_graphs_stay_graphsruns both against a real branch: the main URN, the branch's own URN and theconfigkeyword, as SPARQLFROMandFROM NAMED+GRAPHon the branch's view and as the JSON-LDfromobject throughquery_from(). The HTTP ledger route (/query/L:feature) resolves through the same table but has no branch test of its own.FROM NAMED <L@t:N>+GRAPH <L>: still matches through the compatible alias (pinned_from_named_members_are_named_as_written).ledger_cached,ledger_infoandrefreshon a well-formed missing id still say "not found" (a_missing_ledger_reads_as_not_found_at_every_entry_point).a_service_endpoint_is_never_loadedcounts nameservice lookups).solo_jsonld_dataset_shapes_still_answerruns solo's shapes throughquery_from():{"@id": L, "t": 1, "graph": "txn-meta"}, the same with"at", an object naming a user graph,"L"and"L:main", and afrom/tohistory range. HistoryFROM … TOin SPARQL is covered by the existing history suites, unchanged and green.FROM <L>query answering the same rows with fast paths on and off, the injectedFROM <L>/FROM <L@t:N>/FROM <L:main@iso:…>spellings answering under the whole-ledgerFROMconvention, and the embedder stream path answering the same rows as buffered.resolve_dataset_clause,ResolvedDatasetClause,split_time_travel_suffixandTimeSpec::from(LedgerIdTimeSpec); all four are unchanged (pinned by the compile test).Deviations from the design
x-fdb-warningheader is still computed by the two route-level walkers;QueryResult::advisoriesis not yet what renders it, and the MCP envelope does not render advisories./explain/{L}accepts within-ledger graph IRIs through the typed ids and plans on the ledger view; a dataset's explain still plans the single view.ledger_info()andledger_cached()still parse their argument withLedgerId::parse, so aurn:fluree:id is refused there ("branch cannot contain ':'"). The design moves them ontoLedgerRef::parse; that waits for a follow-up PR (part 2). Their "not found" text is unchanged.@. Run against the v4.1.6 and v4.2.1 binaries, both answer a 500, "Invalid ledger ID format '@t:N': expected 'name' or 'name:branch'". The CLI falls back on an "invalid ledger id" reply that names the pinned path, at any of 400/404/500, and on nothing else.--remote --atkeep the old rewrite; they move to pinned paths in the follow-up PR (part 2)./queryrequest still parses its dataset more than once (auth, refresh and min-t collection each read it); the design's one parse per request waits for the walkers below.Pre-existing, not changed here
FROM NAMEDmember sometimes fails with a 400 FormatError: every time withAccept: application/json, and about one run in three in other formats, solo's pinnedFROM NAMEDshapes included. It reproduces onmainand on this branch alike, and is to be filed separately.Not in this PR: lane consolidation, in a follow-up PR (part 2)
FROMconvention (docs: the "Ledger-Bound Mode" example (FROM <mydb:main>+GRAPH <g>) returns no rows #1975): a clause that is exactly oneFROMnaming a whole ledger selects the ledger with its named graphs./stream/query/{L}@pin, the CLI's NDJSON--at(local and remote) and TSV/CSV--remote --aton pinned paths.from/fromNamedon the view path resolved as within-ledger datasets (this PR refuses them there).QueryResult::advisories.refreshable_ledger_id,refreshable_ledger_id_and_t,collect_refreshable_jsonld_ledgers,collect_jsonld_min_t_requirements, andcollect_sparql_min_t_requirementswithiri_to_string): they skip any text that starts withurn:or contains://, read JSON-LDfrom/fromNamedat the top level only (notopts, notledger), and do not expand a SPARQL prefixed name. A ledger named in any of those ways is not refreshed before the query, and its@t:pin sets no min-t requirement.needs_default_graph_injectionreads only the top-levelfrom/fromNamed/from-named,jsonld_dataset_semantics_warning_headerslikewise, and the SPARQL twin of the warning does not look insideEXISTS. So thex-fdb-warningheader can differ fromQueryResult::advisories, which reads the lowered query.pin_jsonld_dataset(a pinned ledger path) walks every dataset key, top level andoptsalike, where the parser reads one of each pair.query/multi/snapshot.rs:explicit_pin_at,bare_ledger_id, the SPARQL span splicer) keys ledgers by their text throughLedgerId::parse: aurn:fluree:spelling does not match the envelope's snapshot, and a prefixed name is skipped.load_addressrenders a parsed address back to text for the loader to parse again. It round-trips; no divergence is known.