Skip to content

GFQL: no contract for NULL node ids / NULL edge endpoints — production answers it both ways (4 strict xfails) #1995

Description

@lmeyerov

Four strict xfails in graphistry/tests/compute/gfql/test_endpoint_closure_matrix.py record cross-engine disagreements over NULL graph endpoints. They are not four bugs — they are one unanswered contract question, answered both ways inside production code, with the xfail oracles picking the minority answer. Nothing here should be "fixed" until the contract is chosen, because fixing either half in isolation makes an engine self-inconsistent.

The question

Is a NULL node id an identity that a NULL edge endpoint resolves to?

The two answers currently shipping

"NULL never links" — stated in six production comments and implemented in the pandas traversal core:

  • graphistry/compute/hop.py _domain_unique: return series.dropna().unique() — a NULL id can never be in a BFS frontier
  • graphistry/compute/chain.py (unconstrained fast path): "dropna so a NaN node id can't validate a NaN endpoint — .isin treats NaN as matchable but the BFS joins never match NaN<->NaN"
  • graphistry/compute/chain_fast_paths.py ×4: "null ids never link", "a null id/endpoint must not link"
  • graphistry/compute/gfql_fast_paths.py ×3: "NULL endpoints never count: null == null is False on all three engines"

"A NULL endpoint resolves iff the id universe holds a NULL id" — one site, added in #1888 review round 6:

  • graphistry/compute/gfql/lazy/engine/polars/hop_eager.py _keep_edges_with_both_endpoints_resolvable (a_null_id_is_resolvable = resolvable_ids.null_count() > 0), documented in CHANGELOG.

openCypher gives the first answer cover, not the second: a relationship always has two endpoints, and NULL = NULL is UNKNOWN under three-valued logic. There is no ErrorCode covering a NULL node id (E301–E305 are column/binding), nothing normative in docs/source/gfql/spec/language.md, and nothing in the schema validator.

The four cells

Fixture (well-formed, not a broken harness): nodes id = [0.0, 1.0, 2.0, None], edges (0,1) (1,2) (None,2).

test oracle pandas polars
test_the_kept_null_endpoint_edge_also_gets_its_node_row[polars] node set {0,1,2,NULL} {0,1,2,NULL} {0,1,2}
test_the_chain_answers_the_same_null_endpoint_question_as_hop 3 edges 2 2
test_cypher_undirected_count_counts_the_null_endpoint_edge[polars] 6 6 4
test_the_synthesized_chain_is_not_gated_for_a_null_endpoint[pandas,cudf] 3 edges 2 3

Three findings the xfail reasons do not record

  1. The polars hop is self-contradictory. It keeps the (NULL,2) edge (round 6's rule) but its node output all_nodes.join(needed, how="semi") uses polars' default nulls_equal=False, so the kept edge's NULL endpoint gets no node row. The output is not endpoint-closed under either policy. (nulls_equal=True fixes it and is safe on the polars>=1.29 floor.)

  2. pandas answers the same undirected chain two different ways depending on whether the ops are named. _try_chain_fast_path declines on undirected + aliases, so naming routes to the full BFS and changes the value:

    pandas unnamed undirected chain -> [(0,1), (1,2)]
    pandas NAMED   undirected chain -> [(0,1), (1,2), (NULL, 2)]
    

    The fast path's own comment claims parity ("the BFS joins never match NaN<->NaN") — true for forward, false for undirected, because an undirected walk reaches the (NULL,2) edge through its non-null endpoint 2 and never keys on the NULL at all. This is a real defect independent of which policy wins.

  3. pandas is self-inconsistent between its chain and its count on the same pattern: [n(), e_undirected(), n()] yields 2 edges (→ 4 orientations) while MATCH (a)-[x]-(b) RETURN count(*) says 6. And polars' "correct" 3 in the synthesized cell is only half correct — it returns 3 edges but node set {0,1,2}, so it too is not endpoint-closed.

Forcing nulls_equal=True on every polars semi/anti join moves the count 4 → 5, not 6: there is a second, still-unlocated null-blind site in the undirected-orientation materialization.

Why a typed decline is the wrong shape here

The natural place would be a new E3xx for a NULL node id, but the same file already has ~7 passing tests over this exact fixture, and OPTIONAL MATCH legitimately manufactures NULL endpoints — a hard reject would break OPTIONAL MATCH. Declining is only conceivably viable for a bound node table with a NULL id column, never for edge endpoints.

Suggested sequencing

  1. Choose the contract. Recommendation: "a NULL node id is not an identity; nulls never link" — it matches openCypher 3VL and eight of the nine production sites, leaving one dissenter (hop_eager.py's round-6 patch).
  2. Fix the two policy-independent defects: the polars hop's un-backed NULL endpoint, and the pandas named-vs-unnamed undirected divergence.
  3. Rewrite the four oracles as green pins of the chosen contract (_NULL_ENDPOINT_CLOSED → 2 edges, count → 4), and update the CHANGELOG note from design: endpoint-closure contract — five surfaces, five answers for one dangling-edge graph #1888.
  4. Write the contract down in docs/source/gfql/spec/language.md.

cuDF arms unverified (the sweep that produced this ran without a GPU); sites 2 and 4 claim cuDF/pandas parity.

🤖 Generated with Claude Code
https://claude.ai/code/session_01AjbKuKheqDu78oapRT5AYm

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions