Skip to content

feat(cypher): support connected MATCH + OPTIONAL MATCH (#996) - #1021

Merged
lmeyerov merged 2 commits into
masterfrom
fix/issue-996-optional-match-connected
Apr 2, 2026
Merged

lmeyerov merged 2 commits into
masterfrom
fix/issue-996-optional-match-connected

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Apr 2, 2026

Copy link
Copy Markdown
Contributor

Summary

  • feat(cypher): Support MATCH (connected-path) OPTIONAL MATCH ... RETURN queries where the non-optional MATCH is a connected path. Unblocks LDBC SNB Interactive interactive-short-7 (IS7) and any query combining a connected traversal with an optional extension and mixed-alias or CASE-expression projections (Cypher/GFQL cannot lower MATCH + OPTIONAL MATCH reply-author flag queries after a connected MATCH #996).
  • fix(tests): Replace cudf.DataFrame.from_pandas() → cudf.from_pandas() across the GFQL Cypher test suite for RAPIDS 25.02 + 26.02 compatibility (42 call sites).

Architecture

New compilation path _compile_connected_optional_match() in lowering.py:

  1. Detects the 2-clause pattern (connected non-optional + optional) before _merged_match_clause() runs
  2. Lowers each clause independently via lower_match_clause()
  3. Builds a post_join_chain using the same _build_projection_plan + return_() + _lower_order_by_clause infrastructure as the normal path

New execution path _apply_connected_optional_match() in gfql_unified.py:

  1. Runs each chain with rows(binding_ops=...) to produce binding-row tables
  2. Left-outer-joins on shared node alias columns
  3. Delegates RETURN / ORDER BY / SKIP / LIMIT to the standard row pipeline via _chain_dispatch

No hand-rolled expression evaluation — full feature parity with the normal Cypher path because the same row pipeline handles all expression types.

Test plan

  • 19 new connected OPTIONAL MATCH tests covering:
    • Expression breadth: type(), coalesce(), arithmetic, CASE WHEN ... IS NULL
    • Join edge cases: no matches, all match, multi-row per base, empty base, two shared aliases, integer IDs, custom node column name, longer optional chains (2-hop)
    • Post-projection: ORDER BY DESC, SKIP+LIMIT, DISTINCT
  • Full test_lowering.py CPU: 546 passed, 49 skipped
  • Full GFQL GPU on DGX (RAPIDS 26.02): 886 passed, 5 skipped (igraph only), 0 failed
  • Broader suites (test_gfql, test_compute_chain, test_chain_let, test_parser): 286 passed
  • typecheck: success (204 files)
  • lint: all checks passed

🤖 Generated with Claude Code

lmeyerov and others added 2 commits April 1, 2026 23:39
Add a new compilation path for queries where the non-optional MATCH is a
connected path followed by an OPTIONAL MATCH.  The compiler lowers each
clause independently, left-outer-joins binding-row tables on shared node
aliases at runtime, and delegates RETURN / ORDER BY / SKIP / LIMIT to the
standard row pipeline — giving full expression parity with the normal
Cypher path.

Also fix cudf.DataFrame.from_pandas → cudf.from_pandas across the GFQL
Cypher test suite for RAPIDS 25.02 + 26.02 compatibility.

Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
…cudf compat

- Add query.where guard to _is_connected_optional_match_query() so queries
  with WHERE clauses (like TCK match-where6-1) stay on the existing rejection
  path instead of producing wrong results.
- Remove plans/ from tracked files (CI no-plans-in-repo check).

Co-Authored-By: Claude Opus 4.6 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant