Skip to content

[FEA] GFQL native type system: schemas, inference, validation, and Arrow representation #1046

Description

@lmeyerov

Summary

Add a native type system to GFQL covering schema specification, inference, query validation, and typed data representation (Arrow). This is a foundational capability that would improve correctness, performance, and developer experience across the GFQL stack.

Motivation

  • Correctness: Catch schema mismatches at query compile time instead of runtime
  • Performance: Arrow-typed columns enable zero-copy GPU transfer and columnar optimization
  • Developer experience: Autocompletion, documentation, and error messages that reference schema
  • Interop: Typed schemas enable code generation for downstream consumers (TypeScript, Rust, etc.)

Scope

1. Schema specification

Define node and edge schemas declaratively:

from graphistry.schema import NodeType, EdgeType, GraphSchema

Person = NodeType("Person", {
    "id": int,
    "name": str,
    "age": Optional[int],
    "scores": list[float],
})

Company = NodeType("Company", {
    "id": int,
    "name": str,
    "founded": datetime,
})

WorksAt = EdgeType("WORKS_AT",
    source=Person,
    destination=Company,
    properties={
        "since": datetime,
        "role": str,
    },
)

schema = GraphSchema(
    node_types=[Person, Company],
    edge_types=[WorksAt],
)

Design considerations:

  • Multi-label: Cypher nodes can have multiple labels (`:Person:Employee`). Schema should support this — a node satisfies a type if it has ALL required labels.
  • Topology constraints: Edge types should declare valid (source_type, destination_type) pairs. `WORKS_AT` can only connect Person → Company.
  • Union types: A property might be `str | int` (heterogeneous data). Support via Python union syntax or explicit `Union[str, int]`.
  • Optional fields: Properties that may be null/missing. Align with `Optional[T]` / `T | None`.
  • Pydantic alignment: Consider using or extending Pydantic models for schema definition, getting validation, serialization, and IDE support for free.

2. Schema inference

Infer schemas from existing graph data:

schema = g.infer_schema()
# Returns GraphSchema with node types derived from label__* columns,
# edge types from relationship type column, property types from DataFrame dtypes

Design considerations:

  • Infer from `label__X` boolean columns (existing GFQL convention)
  • Infer property types from pandas/cudf dtypes → Arrow types
  • Handle mixed-type columns (object dtype) gracefully
  • Detect topology patterns (which edge types connect which node types)
  • Support incremental refinement: infer base schema, then user annotates/overrides

3. Query validation against schema

Validate GFQL chains and Cypher queries against a schema:

schema = GraphSchema(...)
g = g.bind(schema=schema)

# Compile-time validation:
g.gfql("MATCH (p:Person)-[:WORKS_AT]->(c:Company) RETURN p.age, c.nonexistent")
# → SchemaValidationError: Company has no property 'nonexistent'

g.gfql("MATCH (p:Person)-[:WORKS_AT]->(q:Person) RETURN p, q")
# → SchemaValidationError: WORKS_AT edge type requires destination=Company, got Person

Design considerations:

  • Validate at Cypher compile time (in lowering.py) — no runtime cost
  • Validate native GFQL chains via schema-aware `n()` / `e()` constructors
  • Provide helpful error messages referencing the schema definition
  • Optional strict mode vs permissive mode (warn vs error on unknown properties)

4. Arrow representation

Map schema types to Arrow types for efficient columnar storage:

schema.to_arrow_schema()
# Returns pyarrow.Schema with typed fields for each property

# Load/save with enforced types:
g = graphistry.from_arrow(nodes_table, edges_table, schema=schema)
g.to_arrow(schema=schema)  # Validates and casts to schema types

Design considerations:

  • Map Python types → Arrow types: `int → int64`, `str → utf8`, `list[float] → list`, etc.
  • Support cudf/RAPIDS Arrow interop
  • Enable zero-copy roundtrip: Arrow IPC → cudf → GFQL → Arrow IPC
  • Schema evolution: handle missing columns, extra columns, type coercion

Architecture questions

  1. Where does schema live? On the Plottable? As a separate object? Both?
  2. GFQL-first or Cypher-first? If we start at GFQL (schema-aware `n()`/`e()`), Cypher gets validation for free via the existing lowering path. Starting at Cypher requires mapping Cypher types to GFQL types.
  3. Pydantic integration depth: Full Pydantic models (with validation, serialization) vs lightweight dataclasses with Pydantic-style annotations?
  4. Inference vs declaration: Should `infer_schema()` produce the same schema objects as manual declaration? Or separate "inferred" vs "declared" types?
  5. Incremental adoption: How to add schemas to existing untyped graphs without breaking anything?

Prior art

  • Neo4j constraints: `CREATE CONSTRAINT ... REQUIRE (n.prop) IS :: INTEGER` — runtime enforcement
  • Apache AGE: PostgreSQL-based, inherits PostgreSQL type system
  • Kuzu: Built-in schema with typed node/edge tables
  • Pydantic: Python schema validation library — potential building block
  • Apache Arrow Schema: Columnar type system — the target representation
  • GraphQL: Typed schema for API queries — similar schema-first philosophy
  • openCypher type system: `INTEGER`, `FLOAT`, `STRING`, `BOOLEAN`, `LIST`, `MAP`, `PATH`, `NODE`, `RELATIONSHIP`

Suggested approach

  1. Spike: Define `NodeType`, `EdgeType`, `GraphSchema` dataclasses. Implement `infer_schema()` from existing graph data.
  2. Validate: Add schema validation to the Cypher lowering path — check property references and topology constraints at compile time.
  3. Arrow: Map schema to `pyarrow.Schema`, add `to_arrow()` / `from_arrow()` with schema enforcement.
  4. Iterate: Pydantic integration, multi-label support, union types based on real usage patterns.

Relationship to existing code

  • `graphistry/compute/gfql/cypher/lowering.py`: Cypher property references validated here — schema validation plugs in naturally
  • `graphistry/compute/ast.py`: `ASTNode`, `ASTEdge` could carry schema type info
  • `graphistry/Engine.py`: Engine resolution (pandas/cudf) — Arrow bridge point
  • `graphistry/Plottable.py`: Schema could attach here as `._schema`
  • `label__X` convention: Existing multi-label encoding — schema inference reads these

AI contributor notes

  • Repo AI guidance: `AGENTS.md`, `ai/README.md`
  • GFQL architecture: `ai/docs/` has guides for the query pipeline
  • Test patterns: `graphistry/tests/compute/gfql/cypher/test_lowering.py` (600+ tests)
  • The Cypher compiler is pure Python (no pandas dependency) — schema validation can be added without runtime overhead

Activity

  1. lmeyerov commented on May 5, 2026

    @lmeyerov
    ContributorAuthor

    Type-system staging update: T2 is complete via #1302 closed by PR #1308 (merged 2026-05-05, commit a23c34c). Program tracking remains in #1262 with next active lane T3 (#1300).

  2. added 3 commits that reference this issue on May 5, 2026
  3. lmeyerov commented on May 6, 2026

    @lmeyerov
    ContributorAuthor

    Program receipt update: T3.b nullable-helper consolidation (#1309) is complete via #1317 (merged to master at 911bf00).\n\nMeta trackers updated:\n- #1262 (T3.b marked done)\n- #1259 (recently completed + active-lane refresh)

  4. lmeyerov commented on May 6, 2026

    @lmeyerov
    ContributorAuthor

    Progress update: #1320 (validate-only API lane) is now closed via #1321 (merged to master at cd1f3c8).

  5. lmeyerov commented on May 7, 2026

    @lmeyerov
    ContributorAuthor

    Follow-on scoping update: the next executable child slices after #1262 are now split and tracked as:\n\n- #1337 — public declarative schema model + stable exports\n- #1338 — schema inference API + typed topology extraction (cross-links #483)\n- #1339 — public schema↔Arrow APIs + plottable boundary enforcement\n\nA consolidated scoping note is now in-repo at ai/docs/gfql/scoping_1046_follow_on_slices.md for execution-order, acceptance criteria, and close-criteria mapping back to #1046.

  6. lmeyerov commented on May 7, 2026

    @lmeyerov
    ContributorAuthor

    Meta update: closed PR #1340 because it became a no-op (empty net diff) after removing a throwaway scoping doc file per repo preference.\n\nScoping is preserved via executable child slices instead:\n- #1337 public declarative schema model + stable exports\n- #1338 schema inference API + typed topology extraction\n- #1339 public schema↔Arrow/plottable boundary APIs\n\nSo #1046 continues via those child issues rather than a docs-only throwaway artifact.

  7. lmeyerov commented on May 7, 2026

    @lmeyerov
    ContributorAuthor

    Post-#1262 narrowing update: follow-on execution is now explicitly split into four child lanes:\n\n- A #1337 (public declarative schema model + stable exports)\n- B #1338 (schema inference API + typed topology extraction)\n- C #1339 (public schema↔Arrow/plottable boundary APIs)\n- D #1345 (API contract alignment matrix + compatibility policy)\n\nNon-blocking policy: D runs in parallel and is not a start/merge gate for A/B/C. A/B/C each carry scoped acceptance tests + compatibility notes; D hardens cross-lane contract consistency and drift checks.

  8. lmeyerov commented on May 7, 2026

    @lmeyerov
    ContributorAuthor

    Execution narrowing landed in PR #1347 (https://github.com/graphistry/pygraphistry/pull/1347):\n- A/B/C follow-on lanes are documented in durable GFQL maintainer docs\n- D lane #1345 is explicitly scoped as non-blocking to A/B/C\n\nPR checks are green.

  9. added a commit that references this issue on May 7, 2026
  10. lmeyerov commented on May 16, 2026

    @lmeyerov
    ContributorAuthor

    Coordinator schedule update: filed #1464 as the schema tutorial lane, but it is intentionally deferred until after #1338 lands.

    Updated order for the schema/type-system track:

    1. feat(gfql): add public declarative schema model #1457 / GFQL type system follow-on A: public declarative schema model + stable exports #1337 — public declarative schema API and validation behavior.
    2. GFQL type system follow-on B: schema inference API + typed topology extraction #1338 — schema inference / typed topology extraction.
    3. GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 — community-facing tutorial: infer, refine, bind, and validate Cypher.
    4. GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 — Arrow/plottable boundary APIs and enforcement, with tutorial links/extension after the core inference-first tutorial exists.

    Rationale: the main tutorial should not make users hand-write every schema. It should show the complete practical workflow starting from existing graph data.

  11. lmeyerov commented on May 16, 2026

    @lmeyerov
    ContributorAuthor

    Coordinator schedule update: filed #1465 for gfql_remote() typed-schema transport.

    This was not explicitly covered by the prior schema plan. Updated schema/type-system order:

    1. feat(gfql): add public declarative schema model #1457 / GFQL type system follow-on A: public declarative schema model + stable exports #1337 — local public GraphSchema declaration/binding/preflight.
    2. GFQL type system follow-on B: schema inference API + typed topology extraction #1338 — schema inference / typed topology extraction.
    3. GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 — stable schema serialization + Arrow/plottable boundary semantics.
    4. GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 — send bound typed schema through gfql_remote() / gfql_remote_shape() request envelopes.
    5. GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 — tutorial can mention remote schema behavior only after GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 lands; otherwise keep it local infer/refine/bind/validate.

    #1465 does not block #1457. It is a remote transport follow-up once the local schema contract is stable and serializable.

  12. lmeyerov commented on May 16, 2026

    @lmeyerov
    ContributorAuthor

    Coordinator schedule clarification: typed schema work should proceed as a stack, not as the main active worker pool while #1419 deletion closeout is pending.

    Current schema stack remains:

    1. feat(gfql): add public declarative schema model #1457/GFQL type system follow-on A: public declarative schema model + stable exports #1337 local public schema API.
    2. GFQL type system follow-on B: schema inference API + typed topology extraction #1338 inference.
    3. GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 serialization/Arrow/plottable boundary.
    4. GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 gfql_remote() schema transport.
    5. GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 tutorial after inference, with remote section only after GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465.

    Main project priority is currently #1466 / #1419 mass GFQL deletion audit/closeout.

  13. lmeyerov commented on May 17, 2026

    @lmeyerov
    ContributorAuthor

    Coordinator research follow-up recorded locally in plans/gfql-type-system-research-2026-05-17/report.md.

    Design direction to preserve:

    Key semantic decisions:

    • (:A:B) / conjunctive labels should merge trait/property fragments.
    • (:A|B) / disjunctive labels should branch; guaranteed properties are intersection, admissible properties are union/maybe.
    • Negative labels refine predicates but should not add positive property guarantees.
    • Relationship type disjunction ([:R|S]) branches; positive relationship-type conjunction should reject or be unsatisfiable.

    Open design constraints for follow-ons:

    Practical sequence remains:

    1. feat(gfql): add public declarative schema model #1457/GFQL type system follow-on A: public declarative schema model + stable exports #1337 declared open schema contract.
    2. GFQL type system follow-on B: schema inference API + typed topology extraction #1338 schema inference/refinement.
    3. GFQL type system follow-on C: public schema-Arrow APIs + plottable boundary enforcement #1339 Arrow/plottable boundary enforcement/coercion.
    4. New schema-effects follow-on for graph-growing calls.
    5. GFQL remote: send bound typed GraphSchema with gfql_remote requests #1465 remote schema payload.
    6. GFQL schema tutorial: infer, refine, bind, and validate Cypher #1464 tutorial after inference.
  14. lmeyerov commented on May 25, 2026

    @lmeyerov
    ContributorAuthor

    Typed-schema design follow-up from #1637 / local design receipt:

    Recommended architecture to preserve:

    • Arrow remains the canonical graph contract for internals, wire protocol, backend storage, and cross-language interoperability.
    • Python-native dataclasses, Pydantic models, and typed dataframe/entity wrappers are client adapters derived from that Arrow-backed GraphSchema.
    • Do not make Python models the source of truth or the remote wire format.
    • For static typing, runtime-generated classes are useful but not sufficient; mypy/pyright users will need optional emitted .py / .pyi artifacts.
    • Type narrowing for Python clients should use explicit APIs such as typed frame/entity façades, not rely on static inference from arbitrary dataframe query strings.

    New child issues filed:

    Existing related issues:

    Design slogan:

    Arrow is the graph contract. Python models are generated conveniences. Typed dataframe views are explicit façades.

  15. lmeyerov commented on May 25, 2026

    @lmeyerov
    ContributorAuthor

    Correction to the typed Python follow-up split:

    Rationale: dataclass/Pydantic adapters should not be a standalone early lane. They should be designed with the ORM-like Python user workflow they enable: typed dataframe views, explicit type narrowing, and typed extracted entities, while Arrow remains the canonical graph contract.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions