Repository navigation
[FEA] GFQL native type system: schemas, inference, validation, and Arrow representation #1046
Description
Activity
- added a commit that references this issue
on May 5, 2026 Follow-on scoping update: the next executable child slices after #1262 are now split and tracked as:\n\n- #1337 — public declarative schema model + stable exports\n- #1338 — schema inference API + typed topology extraction (cross-links #483)\n- #1339 — public schema↔Arrow APIs + plottable boundary enforcement\n\nA consolidated scoping note is now in-repo at ai/docs/gfql/scoping_1046_follow_on_slices.md for execution-order, acceptance criteria, and close-criteria mapping back to #1046.
Meta update: closed PR #1340 because it became a no-op (empty net diff) after removing a throwaway scoping doc file per repo preference.\n\nScoping is preserved via executable child slices instead:\n- #1337 public declarative schema model + stable exports\n- #1338 schema inference API + typed topology extraction\n- #1339 public schema↔Arrow/plottable boundary APIs\n\nSo #1046 continues via those child issues rather than a docs-only throwaway artifact.
Post-#1262 narrowing update: follow-on execution is now explicitly split into four child lanes:\n\n- A #1337 (public declarative schema model + stable exports)\n- B #1338 (schema inference API + typed topology extraction)\n- C #1339 (public schema↔Arrow/plottable boundary APIs)\n- D #1345 (API contract alignment matrix + compatibility policy)\n\nNon-blocking policy: D runs in parallel and is not a start/merge gate for A/B/C. A/B/C each carry scoped acceptance tests + compatibility notes; D hardens cross-lane contract consistency and drift checks.
Execution narrowing landed in PR #1347 (https://github.com/graphistry/pygraphistry/pull/1347):\n- A/B/C follow-on lanes are documented in durable GFQL maintainer docs\n- D lane #1345 is explicitly scoped as non-blocking to A/B/C\n\nPR checks are green.
- added a commit that references this issue
on May 7, 2026 Coordinator schedule update: filed #1464 as the schema tutorial lane, but it is intentionally deferred until after #1338 lands.
Updated order for the schema/type-system track:
- feat(gfql): add public declarative schema model #1457 / GFQL type system follow-on A: public declarative schema model + stable exports #1337 — public declarative schema API and validation behavior.
- GFQL type system follow-on B: schema inference API + typed topology extraction #1338 — schema inference / typed topology extraction.
- GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 — community-facing tutorial: infer, refine, bind, and validate Cypher.
- GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 — Arrow/plottable boundary APIs and enforcement, with tutorial links/extension after the core inference-first tutorial exists.
Rationale: the main tutorial should not make users hand-write every schema. It should show the complete practical workflow starting from existing graph data.
Coordinator schedule update: filed #1465 for
gfql_remote()typed-schema transport.This was not explicitly covered by the prior schema plan. Updated schema/type-system order:
- feat(gfql): add public declarative schema model #1457 / GFQL type system follow-on A: public declarative schema model + stable exports #1337 — local public
GraphSchemadeclaration/binding/preflight. - GFQL type system follow-on B: schema inference API + typed topology extraction #1338 — schema inference / typed topology extraction.
- GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 — stable schema serialization + Arrow/plottable boundary semantics.
- GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 — send bound typed schema through
gfql_remote()/gfql_remote_shape()request envelopes. - GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 — tutorial can mention remote schema behavior only after GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 lands; otherwise keep it local infer/refine/bind/validate.
#1465 does not block #1457. It is a remote transport follow-up once the local schema contract is stable and serializable.
- feat(gfql): add public declarative schema model #1457 / GFQL type system follow-on A: public declarative schema model + stable exports #1337 — local public
Coordinator schedule clarification: typed schema work should proceed as a stack, not as the main active worker pool while #1419 deletion closeout is pending.
Current schema stack remains:
- feat(gfql): add public declarative schema model #1457/GFQL type system follow-on A: public declarative schema model + stable exports #1337 local public schema API.
- GFQL type system follow-on B: schema inference API + typed topology extraction #1338 inference.
- GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 serialization/Arrow/plottable boundary.
- GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465
gfql_remote()schema transport. - GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 tutorial after inference, with remote section only after GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465.
Main project priority is currently #1466 / #1419 mass GFQL deletion audit/closeout.
Coordinator research follow-up recorded locally in
plans/gfql-type-system-research-2026-05-17/report.md.Design direction to preserve:
- Land feat(gfql): add public declarative schema model #1457/GFQL type system follow-on A: public declarative schema model + stable exports #1337 as a narrow, open, Arrow-backed graph schema contract, not a full static graph calculus.
- Treat Cypher labels as predicates over label membership.
- Treat
NodeType/EdgeTypeas trait-like contracts over those predicates. - Use Arrow row fragments as the property type substrate.
- Add a separate follow-on for schema effects: graph-growing calls should be represented as
Graph[S] -> Graph[S + delta].
Key semantic decisions:
(:A:B)/ conjunctive labels should merge trait/property fragments.(:A|B)/ disjunctive labels should branch; guaranteed properties are intersection, admissible properties are union/maybe.- Negative labels refine predicates but should not add positive property guarantees.
- Relationship type disjunction (
[:R|S]) branches; positive relationship-type conjunction should reject or be unsatisfiable.
Open design constraints for follow-ons:
- Conflict/presence policy must distinguish required, optional, maybe-absent, and unknown.
- Topology selectors should eventually move from ambiguous raw label sets toward explicit selectors like
All(...)andAnyOf(...). - First-class graph values should be modeled as
Graph[SchemaSnapshot], especially forGRAPH { ... }, GFQL DAGs, remote execution, notebooks, and long-running update sweeps. - GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 should use a JSON schema payload first, with optional Arrow IPC after GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 settles exact conversion/coercion.
Practical sequence remains:
- feat(gfql): add public declarative schema model #1457/GFQL type system follow-on A: public declarative schema model + stable exports #1337 declared open schema contract.
- GFQL type system follow-on B: schema inference API + typed topology extraction #1338 schema inference/refinement.
- GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 Arrow/plottable boundary enforcement/coercion.
- New schema-effects follow-on for graph-growing calls.
- GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 remote schema payload.
- GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 tutorial after inference.
Typed-schema design follow-up from #1637 / local design receipt:
Recommended architecture to preserve:
- Arrow remains the canonical graph contract for internals, wire protocol, backend storage, and cross-language interoperability.
- Python-native dataclasses, Pydantic models, and typed dataframe/entity wrappers are client adapters derived from that Arrow-backed
GraphSchema. - Do not make Python models the source of truth or the remote wire format.
- For static typing, runtime-generated classes are useful but not sufficient; mypy/pyright users will need optional emitted
.py/.pyiartifacts. - Type narrowing for Python clients should use explicit APIs such as typed frame/entity façades, not rely on static inference from arbitrary dataframe query strings.
New child issues filed:
- GFQL typed schema: Python dataclass/Pydantic adapters for GraphSchema #1642 — Python dataclass/Pydantic adapters <> Arrow-backed
GraphSchema - GFQL typed schema: Python adapters + typed dataframe/entity façades #1643 — Typed dataframe/entity façades over
GraphSchema
Existing related issues:
- GFQL type system follow-on B: schema inference API + typed topology extraction #1338 remains the right home for schema inference / CSV bootstrap proposal flow.
- GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 remains the right home for
gfql_remote()schema transport. - GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 already landed the public schema↔Arrow boundary foundation.
Design slogan:
Arrow is the graph contract. Python models are generated conveniences. Typed dataframe views are explicit façades.
Correction to the typed Python follow-up split:
- GFQL typed schema: Python dataclass/Pydantic adapters for GraphSchema #1642 is closed as superseded.
- GFQL typed schema: Python adapters + typed dataframe/entity façades #1643 now owns the bundled Python adapter + typed dataframe/entity façade design.
Rationale: dataclass/Pydantic adapters should not be a standalone early lane. They should be designed with the ORM-like Python user workflow they enable: typed dataframe views, explicit type narrowing, and typed extracted entities, while Arrow remains the canonical graph contract.
Summary
Add a native type system to GFQL covering schema specification, inference, query validation, and typed data representation (Arrow). This is a foundational capability that would improve correctness, performance, and developer experience across the GFQL stack.
Motivation
Scope
1. Schema specification
Define node and edge schemas declaratively:
Design considerations:
2. Schema inference
Infer schemas from existing graph data:
Design considerations:
3. Query validation against schema
Validate GFQL chains and Cypher queries against a schema:
Design considerations:
4. Arrow representation
Map schema types to Arrow types for efficient columnar storage:
Design considerations:
Architecture questions
Prior art
Suggested approach
Relationship to existing code
AI contributor notes