Skip to content

feat(graphql): GraphQL endpoint over a schema derived from the ledger - #1748

Merged
bplatz merged 9 commits into
mainfrom
feat/graphql-endpoint
Sep 3, 2026
Merged

bplatz merged 9 commits into
mainfrom
feat/graphql-endpoint

Conversation

@bplatz

@bplatz bplatz commented Aug 30, 2026 •

Copy link
Copy Markdown
Contributor

Adds a GraphQL endpoint to every ledger, with the schema derived from what the ledger already contains. No .graphqls upload, no resolvers, no build step — point a client at an existing ledger and it introspects, filters, paginates and queries.

Three tiers

The schema sharpens as the ledger says more about itself. Which tier applies is decided by what the ledger contains, not by config:

Tier Trigger What it adds
Inferred nothing Every class a type, every observed property a nullable list field
Shaped a sh:NodeShape with sh:targetClass Cardinality, datatypes, enums, reverse fields, docs, closed types
Curated a graphql:Schema instance What is published, its names, interfaces, and — opt-in — mutations

The customization layer is SHACL in the ledger, plus the http://datashapes.org/graphql# vocabulary for the parts SHACL does not cover. So one artifact drives both validation and the API contract: it is versioned with the data and time-travels with t.

{
  persons(where: { name: { RE: "^A" }, knows: { name: { EQ: "Bob" } } }, limit: 20) {
    id
    name
    knows(orderBy: { name: ASC }, limit: 5) { id name }
  }
}

Surfaces

  • POST/GET /v1/fluree/graphql/<ledger> and GET /v1/fluree/graphql-schema/<ledger>
  • fluree graphql [--schema | --bootstrap | --explain]
  • Fluree::graphql (read) and Fluree::graphql_transact (write)

Behind a graphql feature — default on in server and CLI, in full for the api — which implies shacl, since tier 2 reads shapes.

How it works

A request is parsed, its selection tree extracted from the document, and each root field lowered to a JSON-LD query document. That document goes through the engine's ordinary read path, so its own parser validates and encodes it and policy, SHACL, time travel and formatting all apply unchanged. The result is reshaped back into the GraphQL response — aliases, fragments, __typename and list contracts restored.

extensions.explain returns that lowered query, so the JSON-LD a field compiled to is inspectable:

?explain=true   |   extensions: {"explain": true}   |   fluree graphql --explain

Mutations lower to ordinary transactions through the same write path any other client uses, so a rejected write comes back as a GraphQL error having written nothing.

The new fluree-db-graphql crate holds the schema model, naming, datatype mapping, the three builders, lowering and reshaping. It takes plain IRIs; SID resolution, policy pruning and stats assembly stay in fluree-db-api, which keeps the crate testable without ledger fixtures.

What the schema claims

Tier 1 is deliberately weak, because statistics cannot support more. Every field is a nullable list — statistics can say a property has never been seen twice on one subject, not that it never will be — and orderBy accepts only id, since ordering solutions by a multi-valued key would repeat subjects. No interfaces and no reverse fields: both need someone to say which classes are abstract and what the other direction is called. xsd:integer maps to a custom Long rather than 32-bit Int, and xsd:decimal to Decimal rather than rounding into Float.

Shapes supply what statistics cannot: sh:maxCount 1 gives a single value, sh:minCount ≥ 1 gives !, sh:in gives an enum, sh:inversePath gives a reverse field, and sh:closed true drops observed-but-undeclared properties.

A graphql:Schema makes the endpoint a deliberate contract: only listed shapes are published, so it does not grow a type the moment someone writes an instance. publicShape gets root fields, protectedShape is reachable only through a reference, and privateShape degrades references to Node — the IRI stays visible without naming a type the caller cannot query.

Policy applies by pruning: a class or property an identity cannot read is absent from introspection rather than present-but-empty.

Mutations

Off unless a graphql:Schema sets f:graphqlEnableMutations, and only in tier 3 — a schema derived from whatever a ledger happens to contain should never become a write surface by accident. Each published type then gets create_, update_ and delete_.

Writing needs a LedgerState, which a read view does not carry, so mutation fields are registered only on the write path: the SDL a read endpoint serves matches what it can answer. f:graphqlIriBase is required to mint an IRI, with no default, since a wrong guess writes identifiers that cannot be un-minted.

update_ lowers to where/delete/insert rather than an upsert: a property set to null must be retracted with nothing put back, which an upsert cannot express, and this keeps every property and subject in one atomic transaction. Both update_ and delete_ anchor on @type, so delete_Person on a Company's IRI is a no-op rather than a wipe.

Also included

Two changes outside the GraphQL layer, both independently useful:

  • JSON-LD queries gain per-level ordering and paging inside a hydration. A nested selection's value may be an object of select/orderBy/limit/offset instead of a bare array, bounding how many values each subject shows — a different question from how many rows the query returns, and one the WHERE clause cannot express. Documented in docs/query/jsonld-query.md.
  • SHACL compiles sh:description, sh:order and sh:defaultValue as annotation properties. They constrain nothing but are what a schema generator reads. sh:defaultValue is carried and never materialized: a default is a statement about presentation, not about what the graph holds, and asserting the triple would make sh:minCount 1 self-satisfying.

Performance

Both the derived model and the registered executable schema are cached per ledger version, keyed so that any write, shape edit or context change invalidates them. Steady state, a small query costs ~84 µs against ~39 µs for the equivalent hand-written JSON-LD query; the difference is document parsing, lowering, execution and reshaping. Tracked by fluree-db-api/benches/graphql_schema.rs.

Testing

71 tests in fluree-db-graphql, 40 end-to-end in fluree-db-api, 12 over HTTP in fluree-db-server, and 6 for the new JSON-LD nested-selection syntax. Full workspace suite green at 11,724.

Docs

docs/query/graphql.md (reference), docs/cli/graphql.md (CLI, with the full SHACL and curation tables), a GraphQL vocabulary section in docs/reference/vocabulary.md, and the endpoints in docs/api/endpoints.md. GRAPHQL.md at the root records the design and the deliberate gaps: nested where, nested arguments on a reverse field, several graphql:Schema instances per ledger, and ::n intra-mutation references.

Also in this PR: the values arity in the JSON-LD query docs was incorrect and is corrected; every broken internal link and anchor in docs/ is repaired (0 of 280 files now broken, was 20); and the ten workspace crates missing from the crate map are filled in.

bplatz added 2 commits August 30, 2026 06:42
Adds a GraphQL surface to every ledger with no registration step: no
.graphqls upload, no resolvers, no build step. The schema is derived from
what the ledger already contains and sharpens in three tiers, each selected
by the data rather than by config:

  1. Inferred — HEAD statistics alone. Every class a type, every observed
     property a nullable list field.
  2. Shaped   — a sh:NodeShape with sh:targetClass adds cardinality,
     datatypes, enums, reverse fields, documentation and closed types.
  3. Curated  — a graphql:Schema instance decides what is published, names
     it, marks abstract classes as interfaces, and — only if asked —
     opens a write surface.

Because the customization layer is SHACL in the ledger plus the shared
datashapes.org/graphql# vocabulary, one artifact drives both validation and
the API contract; it is versioned with the data and time-travels with t.

Execution lowers each request to a JSON-LD query *document*, so the engine's
own parser validates and encodes it and policy, SHACL, time travel and
formatting all apply unchanged. Mutations are ordinary transactions: a
rejected write comes back as a GraphQL error having written nothing.

New crate fluree-db-graphql holds the schema model, naming, datatype
mapping, the three builders, lowering and reshaping; it takes plain IRIs, so
ledger access (SID resolution, policy pruning, stats assembly) stays in
fluree-db-api and the crate is testable without ledger fixtures.

Surfaces:
  POST/GET /v1/fluree/graphql/<ledger>
  GET      /v1/fluree/graphql-schema/<ledger>
  fluree graphql [--schema|--bootstrap|--explain]

Behind a `graphql` feature (default on in server and CLI, in `full` for the
api), which implies `shacl` since tier 2 reads shapes.

Two changes outside the GraphQL layer, each useful on its own:

* JSON-LD queries gain per-level ordering and paging inside a hydration. A
  nested selection's value may now be an object of select/orderBy/limit/
  offset instead of a bare array, bounding how many values *each subject*
  shows — a different question from how many rows the query returns, and
  one the WHERE clause cannot express. Modifiers are part of the hydration
  cache key, or two levels differing only in `limit` would collide.

* SHACL compiles sh:description, sh:order and sh:defaultValue as annotation
  properties. They constrain nothing, but are what a schema generator has
  to work with. sh:defaultValue is deliberately never materialized: a
  default is a statement about presentation, not about what the graph
  holds, and asserting the triple would make sh:minCount 1 self-satisfying.

Also fixes an incorrect `values` arity in the JSON-LD query docs, and fills
in the ten workspace crates missing from the crate map.

Design decisions and every deviation from the original plan are recorded in
GRAPHQL.md, including the deliberate gaps: nested `where`, nested arguments
on a reverse field, multiple graphql:Schema instances per ledger, and ::n
intra-mutation references.
`docs/api/endpoints.md` pointed at `../design/cli-server-contract.md`, which
was never tracked in git — a planning-doc reference that shipped. The CLI
contract it wanted is `docs/cli/server-integration.md` ("Implementing Server
Support For Fluree CLI"), whose clone/pull section documents exactly the
remote-`t` preflight the note is about; repointed there.

Sweeping the rest of docs/ found the same class of problem elsewhere. All
verified against the actual target before repointing:

  memory/README.md, memory/concepts/recall-and-ranking.md
    ../docs/... resolved to docs/docs/ — dropped the extra segment
  cli/publish.md
    design/server-implementation.md never existed; server-integration.md is
    that guide
  cli/validate.md
    no cli/import.md; bulk import is documented under `create --from`
  api/headers.md
    no transactions/sparql.md; SPARQL UPDATE lives in query/sparql.md
  concepts/edge-annotations.md
    cypher.md moved under docs/query/
  contributing/tracing-guide.md
    dev-docs/ is not in this repo; bullet removed rather than left dangling

And eight dangling anchors, mostly headings that were renamed under links
that were not: rust-api "Bulk import (high throughput)", cypher "Names, and
opting into IRIs", policy-model "Combining algorithm", sharing-data's
numbered heading, and a `drop_ledger` link to a section this file never had
(now cli/drop.md).

`graph-sources/iceberg.md` had "Enabling local tables" as bold text with
three links already treating it as an anchor; promoted to a real heading so
they resolve.

Verified with a GitHub-compatible slug checker over all 280 files: zero
broken file links, zero broken anchors. Two subtleties that produced false
positives on the way, in case anyone rechecks: GitHub maps spaces to hyphens
1:1 rather than collapsing runs (so `--` anchors are legal), and it keeps
underscores in slugs.
@bplatz bplatz added enhancement New feature or request area:query Query execution, planning, fast paths, overlay, result formatting labels Aug 30, 2026
@bplatz
bplatz requested review from aaj3f and zonotope August 30, 2026 11:05
bplatz added 2 commits August 30, 2026 07:45
…scan

`format_subject` asked the level twice per (subject, predicate): once via
`select_predicate` for the sub-spec, then again via `predicate_modifiers`.
An `Explicit` level answers both by scanning `forward`, so this walked that
list twice for every predicate of every subject — on the shared hydration
path, not just for GraphQL.

Measured over an Explicit level, one lookup per property:

    properties      one scan     two scans
             8      93.1 ns      243.8 ns   (2.6x)
            40      1.523 us     4.456 us   (2.9x)
           100      10.15 us     29.49 us   (2.9x)

`select_predicate_with_modifiers` returns both from a single pass, putting
the loop back to one scan (10.19 us at 100 properties, within noise of the
one-scan baseline). Wildcard levels were never affected: they answer the
modifier question from the discriminant without scanning.

`select_predicate` and `predicate_modifiers` stay for callers that want only
one of the two.
`SystemTime::now()` panics on wasm32-unknown-unknown. It compiles there, so
nothing would have caught it until a browser build called `create_<T>`
without an explicit `id` — the one path that mints.

Uniqueness now comes from a process-lifetime counter, hashed under a
per-process random seed so minted IRIs do not read as a visible sequence.
Uniqueness is the property actually relied on: knowing another subject's IRI
grants nothing by itself, since reads are governed by policy. The previous
comment claimed unpredictability was a security boundary; it is not, and now
says so.

The rest of the crate is already clock- and filesystem-free. The workspace's
other `SystemTime` uses sit in file-storage code that is
`cfg(not(target_arch = "wasm32"))` gated, so they are safe by exclusion;
this one had no such gate.

Not reachable from the wasm32 CI gate, which checks fluree-db-api with
--no-default-features and so never compiles the graphql feature — the fix is
for anyone who does enable it on that target.

@aaj3f aaj3f left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@bplatz this is really neat! And if you'd asked me before this PR, I would not have realized how "easily" (in terms of runtime, not in terms of the work you did) we could derive and provide something like this.

The one thing you may want to consider landing before (or immediately with) merge: it seems GraphQL reads bypass the server's query timeout and client-disconnect cancellation that bound every other read surface. Fluree::query here runs with QueryExecutionOptions::default() outside any run_query_task, so a GraphQL query is unbounded by default where the identical JSON-LD query is cut, for example, at 15 minutes, and aliasing lets one request fan out into many concurrent unbounded queries. Threading QueryExecutionOptions through LedgerExecutor and wrapping the route in run_query_task would close this; adding limit_depth/limit_complexity (finding #2) is the companion hardening a public GraphQL endpoint wants. The minted-IRI width, a policy-pruning test, and the double-parse are all optional thought possibly worth considering now rather than later.

Adherence checklist

  1. Patterns / abstractions — ✔ Exemplary. Lowers to the shared JSON-LD read path and IR; no parallel evaluator (nested fields are pass-through resolvers, execution is query_with_options through the ordinary engine); isolated crate. The 5 shared fluree-db-query edits EXTEND the IR (a modifiers field on the existing ForwardItem::Property + new NestedModifiers types + a JSON-LD-only parser object-form), not a GraphQL branch in the shared parse path; SHACL/vocab changes additive. JSON-LD and SPARQL unaffected — verified green at default features (fluree-db-query 1536/1536, grp_query 428/428, grp_query_sparql 369/369).
  2. Performance — ⚠️ New work is isolated and the shared hydration/parser changes are perf-neutral (common path untouched, sort-key resolution hoisted, cache key updated). But the endpoint ships without a query timeout or cancellation (blocking #1) and without depth/complexity limits (#2) — a default-active resource-bounding gap on the product's headline axis.
  3. Testing — ✔/⚠️ 71 crate + 40 e2e + 12 HTTP + 6 nested-select tests, all wired and green; strong names. Verified locally the shared IR change does not regress the siblings. Gap: the policy-pruning/cache-bypass security path is untested (#4).
  4. Conventions — ✔ New crate opts into [lints] workspace=true, inherits BUSL/version, own thiserror error, async-graphql pinned at workspace, bench registered in regression-budget.json. fmt/clippy clean. Self-describing multi-line commits. Base is main (real CI).

Comment thread fluree-db-server/src/routes/graphql.rs Outdated
let response = if fluree_db_api::graphql::is_mutation(&request) {
execute_mutation(&state, &ledger, &headers, &bearer, &credential, &request).await?
} else {
let view = policy_view(&state, &ledger, &headers, &bearer, &credential).await?;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocking — GraphQL reads bypass query_timeout_ms and client-disconnect cancellation.

Every other read surface bounds queries: routes/query.rs, stream_query.rs, validate.rs all run inside query_control::run_query_task(state.config.query_timeout_ms, …) / current_query_execution_options(…), which installs a QueryCancellation handle plus a timeout task (default DEFAULT_QUERY_TIMEOUT_MS = 15*60*1000, fluree-db-api/src/server_defaults.rs:17). This route calls state.fluree.graphql(&view, &request) directly — not inside run_query_task — and LedgerExecutor::resolve calls the 2-arg self.fluree.query(&db, &q) (fluree-db-api/src/graphql.rs:884), which is query_with_options(…, QueryExecutionOptions::default()) (view/query.rs:286-289): no cancellation, no timeout. The timeout lives in a tokio::task_local! in fluree-db-server that fluree-db-api cannot read, so nothing installs it here.

Fuel is not a backstop either: the lowered query carries no opts.maxFuel, tracker_for_limits (query/helpers.rs:610) returns Tracker::disabled() when none is present, and the GraphQL surface exposes no way to set one — so the timeout is the only ceiling that would apply, and this is the path that drops it. An expensive GraphQL query runs to completion regardless of the configured (or default 15-min) timeout, where the same query over /v1/fluree/query is cancelled; a client that disconnects mid-query does not cancel it. Amplified by aliasing (finding #2): async-graphql resolves root fields concurrently, so one small document with N aliased expensive root fields launches N concurrent, individually-unbounded queries.

Thread QueryExecutionOptions into LedgerExecutor (a field), have resolve/read_back/mutate's read-backs call query_with_options(&db, &q, options.clone()), and wrap the route body in crate::query_control::run_query_task(state.config.query_timeout_ms, …) so a shared cancellation handle covers all of a request's root-field queries. This is the same wiring routes/query.rs:573 already uses.

Comment thread fluree-db-graphql/src/runtime.rs Outdated
///
/// `mutations` are the write fields to expose; empty means a read-only schema,
/// which is every tier below 3 and any tier-3 schema that did not opt in.
pub fn build_schema(model: &SchemaModel, mutations: &[MutationField]) -> Result<Schema> {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should-fix (fold in now) — no depth / complexity / aliasing limits on the schema.

Schema::build here is finished with no .limit_depth(_) / .limit_complexity(_); async-graphql defaults both to unlimited. Deep nesting recurses parse_query, selection::walk (selection.rs:123), build_level, and reshape, and the inferred schema has cyclic types (Person.knows: [Person]) so depth is attacker-chosen. Note the first two run in run_graphql (api/graphql.rs:1152,1157) before schema.execute() (:1169), so a schema-level limit_depth alone wouldn't cover them.

Either vector is cheap to exploit: { a: persons(where:…){…} b: persons … } × thousands of aliases (a few hundred KB, within the body limit) → thousands of concurrent queries (each unbounded per blocking #1); or one deeply-nested document → deep recursion before any budget applies.

Add .limit_depth(N).limit_complexity(M) on the builder in build_schema, plus a cheap explicit depth guard in selection::extract/walk (it already threads state for the cyclic-fragment guard, so a depth counter is a small addition). Fixing #1 bounds each query's wall-clock; this bounds the count and the pre-execution recursion. Conservative production defaults, ideally configurable alongside query_timeout_ms.

///
/// Deliberately no clock: `SystemTime::now()` *panics* on
/// `wasm32-unknown-unknown`, and it would compile fine on the way there.
fn new_id() -> String {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Optional — minted subject IRIs are 64-bit; consider widening to avoid a silent create-merge.

new_id mints a 16-hex (64-bit) local id from a per-process AtomicU64 counter hashed under a per-process RandomState. Across a clustered/replicated write deployment each process seeds independently, so uniqueness is birthday-bound (~2³² mints), and a collision makes create_<T> (a Verb::Insert, :178) silently merge the new subject's facts onto an existing one. Low probability, silent outcome.

A 128-bit id (UUIDv4 or ULID) drops the collision risk to negligible and keeps the no-clock property the last commit was after.

if entry.count == 0 {
continue;
}
if policy.is_some_and(|p| crate::cypher_procedures::class_denied(p, &entry.class_sid)) {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Optional — add a test for policy-pruned introspection and the cache bypass.

The filters that keep a denied class/predicate out of the derived schema (:300-314, :594-643) and the cache_key → None that stops a policy view sharing a cached schema (:229-231) are a security boundary with no direct test (no it_graphql* references policy/Deny/default_allow).

An end-to-end test — a policy-restricted view's SDL/introspection omits the denied type, and re-deriving under a different identity does not serve a cached (root) schema — should go red if the pruning is removed. Worth folding in now while the mechanism is fresh.

Comment thread fluree-db-api/src/graphql.rs Outdated
/// over `POST`, so the method says nothing about intent. An unparseable
/// document reads as a query, so the parse error surfaces from the read path
/// where it is the only thing wrong.
pub fn is_mutation(request: &GraphQlRequest) -> bool {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Optional — parse the request document once. is_mutation (:91-95, called from the route) parses the document, then run_graphql (:1152) parses it again and selection::extract walks it. Cheap vs. the query, but the parsed ExecutableDocument could be produced once and threaded, deciding read-vs-write from the same parse.

} else {
ordering
};
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Optional, cosmetic — nested orderBy on large integers loses precision. compare_sort_values compares via a.as_f64().partial_cmp(&b.as_f64()), which misorders Long/Decimal past 2⁵³. It only orders one subject's already-materialized values, so stakes are low; if easy, compare serde_json::Number as i64/decimal before falling back to f64.

# Conflicts:
#	fluree-db-api/Cargo.toml
#	fluree-db-api/src/lib.rs
main's `autotests = false` landed after this branch forked, so the merge
silently stopped auto-discovering tests/graphql_http_integration.rs — the
crate's own harness_coverage guard is what caught it.
…nt limits

GraphQL was the one read surface with no ceiling. `LedgerExecutor::resolve` ran
`Fluree::query`, which is `query_with_options(.., default())` — no cancellation
handle, no timeout — and the route called `state.fluree.graphql` outside
`run_query_task`, so the handle the server installs for every other read could
not reach it. Fuel is no backstop either: the lowered query carries no
`opts.maxFuel`, so it runs untracked.

Thread `QueryExecutionOptions` through `LedgerExecutor` and wrap the route in
`run_query_task`, the same wiring `routes/query.rs` uses. One document's root
fields resolve concurrently, so they share a single handle: a timeout or a
client disconnect cancels the whole fan-out rather than whichever field checks
next. `graphql_with_options` / `graphql_transact_with_options` are the new entry
points; the existing pair delegates with defaults.

A timeout bounds how long one query runs, not how many a document launches, and
does nothing about recursion that happens before execution — so also add depth
and complexity limits. A derived schema is cyclic wherever one class references
another, so nesting depth is the caller's choice. Depth is checked twice:
`selection::walk` counts levels the way async-graphql's `DepthCalculate` does,
because `parse_query` and `selection::extract` run before `schema.execute()` and
a schema-level limit alone would let a deep document recurse through extraction
first. Complexity is left to the schema, where it bounds fields across every
alias and fragment.

Both are baked into the registered schema, which is cached per ledger version,
so they join `RegisteredKey` — otherwise the first request through would fix the
ceiling for every later one. Defaults are 15 and 1000, configurable as
`graphql_max_depth` / `graphql_max_complexity`, with 0 disabling either the way
`query_timeout_ms = 0` disables the timeout.

Also covers the policy boundary, which had no direct test: a denied class or
property is absent from the derived schema rather than present-and-empty, and a
policy view declines the schema cache key so one identity is never served
another's schema. Each new test was checked against a reverted fix.
`async_graphql_parser` counts recursion depth while *building the AST*
(`MAX_RECURSION_DEPTH = 64`), but pest has already descended the grammar
recursively with no limit of its own by the time that counter runs. A document
of a few hundred KB — `{ persons { persons { … } } }` nested ten thousand deep —
overflows the stack inside `parse_query` and aborts the process, which no caller
can catch. Measured: 1,000 levels is refused, 10,000 aborts.

The depth limit added alongside cannot cover this: it runs during
`selection::extract`, after the parse. So guard the raw document instead — an
O(n) scan bounding brace nesting, skipping strings and comments, at the same 64
the parser would itself have enforced. Both entry points check it, `is_mutation`
included: the route parses there first, so it would otherwise be the hole.

Also drop GRAPHQL.md. It was the working design record for this branch and the
design is settled; what outlives it — the gaps a user needs to know about —
moves into docs/query/graphql.md, which is where the two code comments that
referenced it now point.
Three review follow-ups, all cheaper to take before the feature lands than
after.

**Minted IRIs are 128 bits.** `new_id` hashed a process-lifetime counter under
one random seed and emitted 64 bits. Each writer in a clustered deployment seeds
independently, so uniqueness across the cluster was birthday-bound at ~2^32
mints, and the failure is silent: `create_<T>` lowers to an Insert, so a
collision merges the new subject's facts onto whatever already holds that IRI.
Two independent seeds put the bound at ~2^64, which is UUIDv4's. No clock and no
new dependency — `SystemTime::now()` panics on wasm32, which rules out a ULID or
UUIDv7, and the seeds come from the same `RandomState` the 64-bit form already
used. This is the one item here that is genuinely irreversible: a minted IRI is
permanent, so the width cannot be widened for ledgers already written.

**Each document is parsed once.** It was three times: once to decide whether the
operation writes, once to extract the selection tree, and once more inside
async-graphql. `PreparedRequest` does the parse, answers `writes()`, and carries
the document on to `selection::extract` and then to the executor through
`Request::set_parsed_query` — which `execute` uses as-is, with validation, the
depth and complexity limits included, still running on it. This also collapses
the stack guard to one site; it previously had to be repeated wherever a parse
happened, which is exactly the shape a future caller forgets. Replaces the
`is_mutation` free function.

**Nested `orderBy` compares integers exactly.** `compare_sort_values` went
through `as_f64`, which collapses `xsd:long` values past 2^53 onto one float, so
distinct values tied and a stable sort then ordered them by whatever the
hydration happened to produce. Compare as i64/u64 first. Pinned by unit tests on
the comparator rather than end to end: the arrival order masks the tie, and an
integration test that passes with the bug reverted is worse than none.
@bplatz

bplatz commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Thanks — this was a genuinely useful review, and the blocking finding was correct in every particular, including the detail that fuel is no backstop.

All five findings are addressed. I'd originally planned to defer the three optional ones and you'd have been right to push back on that: "don't churn an approved branch" is an argument about shipped code, and none of this has shipped. The minted-IRI one especially — that's a permanent on-disk identifier, so it's the one item here that genuinely cannot be fixed later.

Branch is rebased on main first (two conflicts, both from the wasm-groundwork reshuffle: fluree-db-spatial and pub mod import moved under cfg(not(wasm32))).

#1 — timeout and cancellation · b8ae924c4

LedgerExecutor carries QueryExecutionOptions and every lowered query goes through query_with_options; the route body is wrapped in run_query_task, the same wiring routes/query.rs:573 uses. Root fields resolve concurrently so they share one handle — a timeout or a disconnect cancels the whole fan-out rather than whichever field checks next. New entry points are graphql_with_options / graphql_transact_with_options; the existing pair delegates.

#2 — depth / complexity · b8ae924c4, c32bba2de

graphql_max_depth (15) and graphql_max_complexity (1000), settable by flag, env var and config file, 0 disabling either the way query_timeout_ms = 0 does. Your point about the ordering was the load-bearing one: depth is checked in selection::walk because extraction runs before schema.execute(), and it counts levels exactly as async-graphql's DepthCalculate does so the two agree on which documents are refusable. Complexity is left to the schema. Both are in RegisteredKey — they're baked into the registered schema, which is cached, so otherwise the first request through would fix the ceiling for everyone after.

Chasing this turned up something worse than a slow parse. async_graphql_parser caps recursion at 64 — but only while building the AST, after pest has already descended the grammar unbounded. Measured: 1,000 levels is refused cleanly, 10,000 levels overflows the stack and aborts the process. A request inside the body limit kills the server, and SIGABRT isn't catchable, so a limit that runs after the parse can't help. There's now an O(n) pre-parse scan bounding brace nesting at that same 64 (limits::guard_nesting).

#4 — policy pruning · b8ae924c4

Four tests in it_graphql_policy.rs: a denied class and a denied property are absent from the SDL rather than present-and-empty, a denied class is unqueryable, and a policy view isn't served a cached root schema (checked in both directions).

Minted IRIs · 58c919c26

Now 128 bits, via two independent per-process seeds — collision bound goes from ~2³² mints to ~2⁶⁴, matching UUIDv4. Kept the no-clock property you'd noted from the previous commit: SystemTime::now() panics on wasm32, which rules out a ULID or UUIDv7, so the seeds still come from RandomState. No new dependency.

Double-parse · 58c919c26

It was three parses, not two — is_mutation, run_graphql, and async-graphql's own. PreparedRequest does the parse once, answers writes(), and hands the document to selection::extract and then to the executor via Request::set_parsed_query; execute uses it as-is and validation, limits included, still runs on it. is_mutation is gone. The bigger win is that this collapses the stack guard to a single call site — it previously had to be repeated at every parse, which is exactly the shape someone forgets.

f64 sort · 58c919c26

compare_sort_values compares as i64/u64 before falling back to f64.

Worth recording how this one is tested: the obvious end-to-end test passed with the fix reverted, because the hydration already delivers those values in ascending order so the tie never bites. The real pin is three unit tests on the comparator itself; the integration test stayed but is labelled for what it actually covers. I also wrote and then removed explicit mixed-sign branches — they behave identically to the f64 fallback, since rounding can collapse neighbours but never flip a sign.

One thing the merge broke

main gained autotests = false and the grp_* harnesses after this branch forked, so tests/graphql_http_integration.rs stopped being auto-discovered — all 12 HTTP tests would have gone dark with CI still green. The crate's own harness_coverage guard caught it, which is a nice validation of that guard. Registered in grp_http behind #[cfg(feature = "graphql")] (b8b56e470), plus 3 new HTTP tests for the configured limits.

Verification

Every new test was checked against a reverted fix — I sabotaged each mechanism in turn (options threading, the walk guard, the schema limits, the cache key, both policy filters, the cache decline, the route's config wiring, the stack guard) and confirmed the specific tests fail. The stack guard's proof is the process abort itself.

Two honest caveats:

  • The single-parse change is not measurable at the bench level. graphql_query_overhead still reads ~83.6 µs. A parse is ~5.4 µs so I expected a few percent, but I don't have a clean before-number on the same machine — treat the perf framing as unverified. It stands on the single-guard-site argument regardless.
  • I did not run the full --workspace --all-features suite locally (the target/ dir filled the disk mid-session); CI covers it. Green through c32bba2de, with the last commit queued.

Also dropped GRAPHQL.md — it was the working design record for this branch and the design is settled. The parts a user needs (the deliberate gaps, and now the limits) moved into docs/query/graphql.md, and the two code comments that pointed at it now point there.

@bplatz
bplatz merged commit 9b09c14 into main Sep 3, 2026
15 checks passed
@bplatz
bplatz deleted the feat/graphql-endpoint branch September 3, 2026 02:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:query Query execution, planning, fast paths, overlay, result formatting enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants