Repository navigation
feat(graphql): GraphQL endpoint over a schema derived from the ledger - #1748
Conversation
Adds a GraphQL surface to every ledger with no registration step: no
.graphqls upload, no resolvers, no build step. The schema is derived from
what the ledger already contains and sharpens in three tiers, each selected
by the data rather than by config:
1. Inferred — HEAD statistics alone. Every class a type, every observed
property a nullable list field.
2. Shaped — a sh:NodeShape with sh:targetClass adds cardinality,
datatypes, enums, reverse fields, documentation and closed types.
3. Curated — a graphql:Schema instance decides what is published, names
it, marks abstract classes as interfaces, and — only if asked —
opens a write surface.
Because the customization layer is SHACL in the ledger plus the shared
datashapes.org/graphql# vocabulary, one artifact drives both validation and
the API contract; it is versioned with the data and time-travels with t.
Execution lowers each request to a JSON-LD query *document*, so the engine's
own parser validates and encodes it and policy, SHACL, time travel and
formatting all apply unchanged. Mutations are ordinary transactions: a
rejected write comes back as a GraphQL error having written nothing.
New crate fluree-db-graphql holds the schema model, naming, datatype
mapping, the three builders, lowering and reshaping; it takes plain IRIs, so
ledger access (SID resolution, policy pruning, stats assembly) stays in
fluree-db-api and the crate is testable without ledger fixtures.
Surfaces:
POST/GET /v1/fluree/graphql/<ledger>
GET /v1/fluree/graphql-schema/<ledger>
fluree graphql [--schema|--bootstrap|--explain]
Behind a `graphql` feature (default on in server and CLI, in `full` for the
api), which implies `shacl` since tier 2 reads shapes.
Two changes outside the GraphQL layer, each useful on its own:
* JSON-LD queries gain per-level ordering and paging inside a hydration. A
nested selection's value may now be an object of select/orderBy/limit/
offset instead of a bare array, bounding how many values *each subject*
shows — a different question from how many rows the query returns, and
one the WHERE clause cannot express. Modifiers are part of the hydration
cache key, or two levels differing only in `limit` would collide.
* SHACL compiles sh:description, sh:order and sh:defaultValue as annotation
properties. They constrain nothing, but are what a schema generator has
to work with. sh:defaultValue is deliberately never materialized: a
default is a statement about presentation, not about what the graph
holds, and asserting the triple would make sh:minCount 1 self-satisfying.
Also fixes an incorrect `values` arity in the JSON-LD query docs, and fills
in the ten workspace crates missing from the crate map.
Design decisions and every deviation from the original plan are recorded in
GRAPHQL.md, including the deliberate gaps: nested `where`, nested arguments
on a reverse field, multiple graphql:Schema instances per ledger, and ::n
intra-mutation references.
`docs/api/endpoints.md` pointed at `../design/cli-server-contract.md`, which
was never tracked in git — a planning-doc reference that shipped. The CLI
contract it wanted is `docs/cli/server-integration.md` ("Implementing Server
Support For Fluree CLI"), whose clone/pull section documents exactly the
remote-`t` preflight the note is about; repointed there.
Sweeping the rest of docs/ found the same class of problem elsewhere. All
verified against the actual target before repointing:
memory/README.md, memory/concepts/recall-and-ranking.md
../docs/... resolved to docs/docs/ — dropped the extra segment
cli/publish.md
design/server-implementation.md never existed; server-integration.md is
that guide
cli/validate.md
no cli/import.md; bulk import is documented under `create --from`
api/headers.md
no transactions/sparql.md; SPARQL UPDATE lives in query/sparql.md
concepts/edge-annotations.md
cypher.md moved under docs/query/
contributing/tracing-guide.md
dev-docs/ is not in this repo; bullet removed rather than left dangling
And eight dangling anchors, mostly headings that were renamed under links
that were not: rust-api "Bulk import (high throughput)", cypher "Names, and
opting into IRIs", policy-model "Combining algorithm", sharing-data's
numbered heading, and a `drop_ledger` link to a section this file never had
(now cli/drop.md).
`graph-sources/iceberg.md` had "Enabling local tables" as bold text with
three links already treating it as an anchor; promoted to a real heading so
they resolve.
Verified with a GitHub-compatible slug checker over all 280 files: zero
broken file links, zero broken anchors. Two subtleties that produced false
positives on the way, in case anyone rechecks: GitHub maps spaces to hyphens
1:1 rather than collapsing runs (so `--` anchors are legal), and it keeps
underscores in slugs.
…scan
`format_subject` asked the level twice per (subject, predicate): once via
`select_predicate` for the sub-spec, then again via `predicate_modifiers`.
An `Explicit` level answers both by scanning `forward`, so this walked that
list twice for every predicate of every subject — on the shared hydration
path, not just for GraphQL.
Measured over an Explicit level, one lookup per property:
properties one scan two scans
8 93.1 ns 243.8 ns (2.6x)
40 1.523 us 4.456 us (2.9x)
100 10.15 us 29.49 us (2.9x)
`select_predicate_with_modifiers` returns both from a single pass, putting
the loop back to one scan (10.19 us at 100 properties, within noise of the
one-scan baseline). Wildcard levels were never affected: they answer the
modifier question from the discriminant without scanning.
`select_predicate` and `predicate_modifiers` stay for callers that want only
one of the two.
`SystemTime::now()` panics on wasm32-unknown-unknown. It compiles there, so nothing would have caught it until a browser build called `create_<T>` without an explicit `id` — the one path that mints. Uniqueness now comes from a process-lifetime counter, hashed under a per-process random seed so minted IRIs do not read as a visible sequence. Uniqueness is the property actually relied on: knowing another subject's IRI grants nothing by itself, since reads are governed by policy. The previous comment claimed unpredictability was a security boundary; it is not, and now says so. The rest of the crate is already clock- and filesystem-free. The workspace's other `SystemTime` uses sit in file-storage code that is `cfg(not(target_arch = "wasm32"))` gated, so they are safe by exclusion; this one had no such gate. Not reachable from the wasm32 CI gate, which checks fluree-db-api with --no-default-features and so never compiles the graphql feature — the fix is for anyone who does enable it on that target.
aaj3f
left a comment
There was a problem hiding this comment.
@bplatz this is really neat! And if you'd asked me before this PR, I would not have realized how "easily" (in terms of runtime, not in terms of the work you did) we could derive and provide something like this.
The one thing you may want to consider landing before (or immediately with) merge: it seems GraphQL reads bypass the server's query timeout and client-disconnect cancellation that bound every other read surface. Fluree::query here runs with QueryExecutionOptions::default() outside any run_query_task, so a GraphQL query is unbounded by default where the identical JSON-LD query is cut, for example, at 15 minutes, and aliasing lets one request fan out into many concurrent unbounded queries. Threading QueryExecutionOptions through LedgerExecutor and wrapping the route in run_query_task would close this; adding limit_depth/limit_complexity (finding #2) is the companion hardening a public GraphQL endpoint wants. The minted-IRI width, a policy-pruning test, and the double-parse are all optional thought possibly worth considering now rather than later.
Adherence checklist
- Patterns / abstractions — ✔ Exemplary. Lowers to the shared JSON-LD read path and IR; no parallel evaluator (nested fields are pass-through resolvers, execution is
query_with_optionsthrough the ordinary engine); isolated crate. The 5 sharedfluree-db-queryedits EXTEND the IR (amodifiersfield on the existingForwardItem::Property+ newNestedModifierstypes + a JSON-LD-only parser object-form), not a GraphQL branch in the shared parse path; SHACL/vocab changes additive. JSON-LD and SPARQL unaffected — verified green at default features (fluree-db-query 1536/1536, grp_query 428/428, grp_query_sparql 369/369). - Performance —
⚠️ New work is isolated and the shared hydration/parser changes are perf-neutral (common path untouched, sort-key resolution hoisted, cache key updated). But the endpoint ships without a query timeout or cancellation (blocking#1) and without depth/complexity limits (#2) — a default-active resource-bounding gap on the product's headline axis. - Testing — ✔/
⚠️ 71 crate + 40 e2e + 12 HTTP + 6 nested-select tests, all wired and green; strong names. Verified locally the shared IR change does not regress the siblings. Gap: the policy-pruning/cache-bypass security path is untested (#4). - Conventions — ✔ New crate opts into
[lints] workspace=true, inherits BUSL/version, ownthiserrorerror,async-graphqlpinned at workspace, bench registered inregression-budget.json. fmt/clippy clean. Self-describing multi-line commits. Base ismain(real CI).
| let response = if fluree_db_api::graphql::is_mutation(&request) { | ||
| execute_mutation(&state, &ledger, &headers, &bearer, &credential, &request).await? | ||
| } else { | ||
| let view = policy_view(&state, &ledger, &headers, &bearer, &credential).await?; |
There was a problem hiding this comment.
Blocking — GraphQL reads bypass query_timeout_ms and client-disconnect cancellation.
Every other read surface bounds queries: routes/query.rs, stream_query.rs, validate.rs all run inside query_control::run_query_task(state.config.query_timeout_ms, …) / current_query_execution_options(…), which installs a QueryCancellation handle plus a timeout task (default DEFAULT_QUERY_TIMEOUT_MS = 15*60*1000, fluree-db-api/src/server_defaults.rs:17). This route calls state.fluree.graphql(&view, &request) directly — not inside run_query_task — and LedgerExecutor::resolve calls the 2-arg self.fluree.query(&db, &q) (fluree-db-api/src/graphql.rs:884), which is query_with_options(…, QueryExecutionOptions::default()) (view/query.rs:286-289): no cancellation, no timeout. The timeout lives in a tokio::task_local! in fluree-db-server that fluree-db-api cannot read, so nothing installs it here.
Fuel is not a backstop either: the lowered query carries no opts.maxFuel, tracker_for_limits (query/helpers.rs:610) returns Tracker::disabled() when none is present, and the GraphQL surface exposes no way to set one — so the timeout is the only ceiling that would apply, and this is the path that drops it. An expensive GraphQL query runs to completion regardless of the configured (or default 15-min) timeout, where the same query over /v1/fluree/query is cancelled; a client that disconnects mid-query does not cancel it. Amplified by aliasing (finding #2): async-graphql resolves root fields concurrently, so one small document with N aliased expensive root fields launches N concurrent, individually-unbounded queries.
Thread QueryExecutionOptions into LedgerExecutor (a field), have resolve/read_back/mutate's read-backs call query_with_options(&db, &q, options.clone()), and wrap the route body in crate::query_control::run_query_task(state.config.query_timeout_ms, …) so a shared cancellation handle covers all of a request's root-field queries. This is the same wiring routes/query.rs:573 already uses.
| /// | ||
| /// `mutations` are the write fields to expose; empty means a read-only schema, | ||
| /// which is every tier below 3 and any tier-3 schema that did not opt in. | ||
| pub fn build_schema(model: &SchemaModel, mutations: &[MutationField]) -> Result<Schema> { |
There was a problem hiding this comment.
Should-fix (fold in now) — no depth / complexity / aliasing limits on the schema.
Schema::build here is finished with no .limit_depth(_) / .limit_complexity(_); async-graphql defaults both to unlimited. Deep nesting recurses parse_query, selection::walk (selection.rs:123), build_level, and reshape, and the inferred schema has cyclic types (Person.knows: [Person]) so depth is attacker-chosen. Note the first two run in run_graphql (api/graphql.rs:1152,1157) before schema.execute() (:1169), so a schema-level limit_depth alone wouldn't cover them.
Either vector is cheap to exploit: { a: persons(where:…){…} b: persons … } × thousands of aliases (a few hundred KB, within the body limit) → thousands of concurrent queries (each unbounded per blocking #1); or one deeply-nested document → deep recursion before any budget applies.
Add .limit_depth(N).limit_complexity(M) on the builder in build_schema, plus a cheap explicit depth guard in selection::extract/walk (it already threads state for the cyclic-fragment guard, so a depth counter is a small addition). Fixing #1 bounds each query's wall-clock; this bounds the count and the pre-execution recursion. Conservative production defaults, ideally configurable alongside query_timeout_ms.
| /// | ||
| /// Deliberately no clock: `SystemTime::now()` *panics* on | ||
| /// `wasm32-unknown-unknown`, and it would compile fine on the way there. | ||
| fn new_id() -> String { |
There was a problem hiding this comment.
Optional — minted subject IRIs are 64-bit; consider widening to avoid a silent create-merge.
new_id mints a 16-hex (64-bit) local id from a per-process AtomicU64 counter hashed under a per-process RandomState. Across a clustered/replicated write deployment each process seeds independently, so uniqueness is birthday-bound (~2³² mints), and a collision makes create_<T> (a Verb::Insert, :178) silently merge the new subject's facts onto an existing one. Low probability, silent outcome.
A 128-bit id (UUIDv4 or ULID) drops the collision risk to negligible and keeps the no-clock property the last commit was after.
| if entry.count == 0 { | ||
| continue; | ||
| } | ||
| if policy.is_some_and(|p| crate::cypher_procedures::class_denied(p, &entry.class_sid)) { |
There was a problem hiding this comment.
Optional — add a test for policy-pruned introspection and the cache bypass.
The filters that keep a denied class/predicate out of the derived schema (:300-314, :594-643) and the cache_key → None that stops a policy view sharing a cached schema (:229-231) are a security boundary with no direct test (no it_graphql* references policy/Deny/default_allow).
An end-to-end test — a policy-restricted view's SDL/introspection omits the denied type, and re-deriving under a different identity does not serve a cached (root) schema — should go red if the pruning is removed. Worth folding in now while the mechanism is fresh.
| /// over `POST`, so the method says nothing about intent. An unparseable | ||
| /// document reads as a query, so the parse error surfaces from the read path | ||
| /// where it is the only thing wrong. | ||
| pub fn is_mutation(request: &GraphQlRequest) -> bool { |
There was a problem hiding this comment.
Optional — parse the request document once. is_mutation (:91-95, called from the route) parses the document, then run_graphql (:1152) parses it again and selection::extract walks it. Cheap vs. the query, but the parsed ExecutableDocument could be produced once and threaded, deciding read-vs-write from the same parse.
| } else { | ||
| ordering | ||
| }; | ||
| } |
There was a problem hiding this comment.
Optional, cosmetic — nested orderBy on large integers loses precision. compare_sort_values compares via a.as_f64().partial_cmp(&b.as_f64()), which misorders Long/Decimal past 2⁵³. It only orders one subject's already-materialized values, so stakes are low; if easy, compare serde_json::Number as i64/decimal before falling back to f64.
# Conflicts: # fluree-db-api/Cargo.toml # fluree-db-api/src/lib.rs
main's `autotests = false` landed after this branch forked, so the merge silently stopped auto-discovering tests/graphql_http_integration.rs — the crate's own harness_coverage guard is what caught it.
…nt limits GraphQL was the one read surface with no ceiling. `LedgerExecutor::resolve` ran `Fluree::query`, which is `query_with_options(.., default())` — no cancellation handle, no timeout — and the route called `state.fluree.graphql` outside `run_query_task`, so the handle the server installs for every other read could not reach it. Fuel is no backstop either: the lowered query carries no `opts.maxFuel`, so it runs untracked. Thread `QueryExecutionOptions` through `LedgerExecutor` and wrap the route in `run_query_task`, the same wiring `routes/query.rs` uses. One document's root fields resolve concurrently, so they share a single handle: a timeout or a client disconnect cancels the whole fan-out rather than whichever field checks next. `graphql_with_options` / `graphql_transact_with_options` are the new entry points; the existing pair delegates with defaults. A timeout bounds how long one query runs, not how many a document launches, and does nothing about recursion that happens before execution — so also add depth and complexity limits. A derived schema is cyclic wherever one class references another, so nesting depth is the caller's choice. Depth is checked twice: `selection::walk` counts levels the way async-graphql's `DepthCalculate` does, because `parse_query` and `selection::extract` run before `schema.execute()` and a schema-level limit alone would let a deep document recurse through extraction first. Complexity is left to the schema, where it bounds fields across every alias and fragment. Both are baked into the registered schema, which is cached per ledger version, so they join `RegisteredKey` — otherwise the first request through would fix the ceiling for every later one. Defaults are 15 and 1000, configurable as `graphql_max_depth` / `graphql_max_complexity`, with 0 disabling either the way `query_timeout_ms = 0` disables the timeout. Also covers the policy boundary, which had no direct test: a denied class or property is absent from the derived schema rather than present-and-empty, and a policy view declines the schema cache key so one identity is never served another's schema. Each new test was checked against a reverted fix.
`async_graphql_parser` counts recursion depth while *building the AST*
(`MAX_RECURSION_DEPTH = 64`), but pest has already descended the grammar
recursively with no limit of its own by the time that counter runs. A document
of a few hundred KB — `{ persons { persons { … } } }` nested ten thousand deep —
overflows the stack inside `parse_query` and aborts the process, which no caller
can catch. Measured: 1,000 levels is refused, 10,000 aborts.
The depth limit added alongside cannot cover this: it runs during
`selection::extract`, after the parse. So guard the raw document instead — an
O(n) scan bounding brace nesting, skipping strings and comments, at the same 64
the parser would itself have enforced. Both entry points check it, `is_mutation`
included: the route parses there first, so it would otherwise be the hole.
Also drop GRAPHQL.md. It was the working design record for this branch and the
design is settled; what outlives it — the gaps a user needs to know about —
moves into docs/query/graphql.md, which is where the two code comments that
referenced it now point.
Three review follow-ups, all cheaper to take before the feature lands than after. **Minted IRIs are 128 bits.** `new_id` hashed a process-lifetime counter under one random seed and emitted 64 bits. Each writer in a clustered deployment seeds independently, so uniqueness across the cluster was birthday-bound at ~2^32 mints, and the failure is silent: `create_<T>` lowers to an Insert, so a collision merges the new subject's facts onto whatever already holds that IRI. Two independent seeds put the bound at ~2^64, which is UUIDv4's. No clock and no new dependency — `SystemTime::now()` panics on wasm32, which rules out a ULID or UUIDv7, and the seeds come from the same `RandomState` the 64-bit form already used. This is the one item here that is genuinely irreversible: a minted IRI is permanent, so the width cannot be widened for ledgers already written. **Each document is parsed once.** It was three times: once to decide whether the operation writes, once to extract the selection tree, and once more inside async-graphql. `PreparedRequest` does the parse, answers `writes()`, and carries the document on to `selection::extract` and then to the executor through `Request::set_parsed_query` — which `execute` uses as-is, with validation, the depth and complexity limits included, still running on it. This also collapses the stack guard to one site; it previously had to be repeated wherever a parse happened, which is exactly the shape a future caller forgets. Replaces the `is_mutation` free function. **Nested `orderBy` compares integers exactly.** `compare_sort_values` went through `as_f64`, which collapses `xsd:long` values past 2^53 onto one float, so distinct values tied and a stable sort then ordered them by whatever the hydration happened to produce. Compare as i64/u64 first. Pinned by unit tests on the comparator rather than end to end: the arrival order masks the tie, and an integration test that passes with the bug reverted is worse than none.
|
Thanks — this was a genuinely useful review, and the blocking finding was correct in every particular, including the detail that fuel is no backstop. All five findings are addressed. I'd originally planned to defer the three optional ones and you'd have been right to push back on that: "don't churn an approved branch" is an argument about shipped code, and none of this has shipped. The minted-IRI one especially — that's a permanent on-disk identifier, so it's the one item here that genuinely cannot be fixed later. Branch is rebased on #1 — timeout and cancellation ·
|
Adds a GraphQL endpoint to every ledger, with the schema derived from what the ledger already contains. No
.graphqlsupload, no resolvers, no build step — point a client at an existing ledger and it introspects, filters, paginates and queries.Three tiers
The schema sharpens as the ledger says more about itself. Which tier applies is decided by what the ledger contains, not by config:
sh:NodeShapewithsh:targetClassgraphql:SchemainstanceThe customization layer is SHACL in the ledger, plus the
http://datashapes.org/graphql#vocabulary for the parts SHACL does not cover. So one artifact drives both validation and the API contract: it is versioned with the data and time-travels witht.{ persons(where: { name: { RE: "^A" }, knows: { name: { EQ: "Bob" } } }, limit: 20) { id name knows(orderBy: { name: ASC }, limit: 5) { id name } } }Surfaces
POST/GET /v1/fluree/graphql/<ledger>andGET /v1/fluree/graphql-schema/<ledger>fluree graphql [--schema | --bootstrap | --explain]Fluree::graphql(read) andFluree::graphql_transact(write)Behind a
graphqlfeature — default on in server and CLI, infullfor the api — which impliesshacl, since tier 2 reads shapes.How it works
A request is parsed, its selection tree extracted from the document, and each root field lowered to a JSON-LD query document. That document goes through the engine's ordinary read path, so its own parser validates and encodes it and policy, SHACL, time travel and formatting all apply unchanged. The result is reshaped back into the GraphQL response — aliases, fragments,
__typenameand list contracts restored.extensions.explainreturns that lowered query, so the JSON-LD a field compiled to is inspectable:Mutations lower to ordinary transactions through the same write path any other client uses, so a rejected write comes back as a GraphQL error having written nothing.
The new
fluree-db-graphqlcrate holds the schema model, naming, datatype mapping, the three builders, lowering and reshaping. It takes plain IRIs; SID resolution, policy pruning and stats assembly stay influree-db-api, which keeps the crate testable without ledger fixtures.What the schema claims
Tier 1 is deliberately weak, because statistics cannot support more. Every field is a nullable list — statistics can say a property has never been seen twice on one subject, not that it never will be — and
orderByaccepts onlyid, since ordering solutions by a multi-valued key would repeat subjects. No interfaces and no reverse fields: both need someone to say which classes are abstract and what the other direction is called.xsd:integermaps to a customLongrather than 32-bitInt, andxsd:decimaltoDecimalrather than rounding intoFloat.Shapes supply what statistics cannot:
sh:maxCount 1gives a single value,sh:minCount ≥ 1gives!,sh:ingives anenum,sh:inversePathgives a reverse field, andsh:closed truedrops observed-but-undeclared properties.A
graphql:Schemamakes the endpoint a deliberate contract: only listed shapes are published, so it does not grow a type the moment someone writes an instance.publicShapegets root fields,protectedShapeis reachable only through a reference, andprivateShapedegrades references toNode— the IRI stays visible without naming a type the caller cannot query.Policy applies by pruning: a class or property an identity cannot read is absent from introspection rather than present-but-empty.
Mutations
Off unless a
graphql:Schemasetsf:graphqlEnableMutations, and only in tier 3 — a schema derived from whatever a ledger happens to contain should never become a write surface by accident. Each published type then getscreate_,update_anddelete_.Writing needs a
LedgerState, which a read view does not carry, so mutation fields are registered only on the write path: the SDL a read endpoint serves matches what it can answer.f:graphqlIriBaseis required to mint an IRI, with no default, since a wrong guess writes identifiers that cannot be un-minted.update_lowers towhere/delete/insertrather than an upsert: a property set tonullmust be retracted with nothing put back, which an upsert cannot express, and this keeps every property and subject in one atomic transaction. Bothupdate_anddelete_anchor on@type, sodelete_Personon a Company's IRI is a no-op rather than a wipe.Also included
Two changes outside the GraphQL layer, both independently useful:
select/orderBy/limit/offsetinstead of a bare array, bounding how many values each subject shows — a different question from how many rows the query returns, and one the WHERE clause cannot express. Documented indocs/query/jsonld-query.md.sh:description,sh:orderandsh:defaultValueas annotation properties. They constrain nothing but are what a schema generator reads.sh:defaultValueis carried and never materialized: a default is a statement about presentation, not about what the graph holds, and asserting the triple would makesh:minCount 1self-satisfying.Performance
Both the derived model and the registered executable schema are cached per ledger version, keyed so that any write, shape edit or context change invalidates them. Steady state, a small query costs ~84 µs against ~39 µs for the equivalent hand-written JSON-LD query; the difference is document parsing, lowering, execution and reshaping. Tracked by
fluree-db-api/benches/graphql_schema.rs.Testing
71 tests in
fluree-db-graphql, 40 end-to-end influree-db-api, 12 over HTTP influree-db-server, and 6 for the new JSON-LD nested-selection syntax. Full workspace suite green at 11,724.Docs
docs/query/graphql.md(reference),docs/cli/graphql.md(CLI, with the full SHACL and curation tables), a GraphQL vocabulary section indocs/reference/vocabulary.md, and the endpoints indocs/api/endpoints.md.GRAPHQL.mdat the root records the design and the deliberate gaps: nestedwhere, nested arguments on a reverse field, severalgraphql:Schemainstances per ledger, and::nintra-mutation references.Also in this PR: the
valuesarity in the JSON-LD query docs was incorrect and is corrected; every broken internal link and anchor indocs/is repaired (0 of 280 files now broken, was 20); and the ten workspace crates missing from the crate map are filled in.