Skip to content

feat(cypher): OPTIONAL MATCH + CASE execution for interactive-short-7 - #1061

Merged
lmeyerov merged 7 commits into
masterfrom
feat/issue-996-optional-match-case
Apr 5, 2026
Merged

lmeyerov merged 7 commits into
masterfrom
feat/issue-996-optional-match-case

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Apr 5, 2026

Copy link
Copy Markdown
Contributor

Closes #996

Summary

Implement direct Cypher execution for MATCH ... OPTIONAL MATCH ... RETURN CASE ... patterns needed by LDBC SNB Interactive interactive-short-7.

Current state: compiler rejects these queries with [unsupported-cypher-query] Only node-only pre-binding MATCH clauses are supported before the final connected MATCH in this phase.

Target query shape (IS7)

MATCH (m:Message {id: $messageId})<-[:REPLY_OF]-(c:Comment)-[:HAS_CREATOR]->(p:Person)
OPTIONAL MATCH (m)-[:HAS_CREATOR]->(a:Person)-[r:KNOWS]-(p)
RETURN c.id AS commentId,
    c.content AS commentContent,
    c.creationDate AS commentCreationDate,
    p.id AS replyAuthorId,
    p.firstName AS replyAuthorFirstName,
    p.lastName AS replyAuthorLastName,
    CASE r
        WHEN null THEN false
        ELSE true
    END AS replyAuthorKnowsOriginalMessageAuthor
ORDER BY commentCreationDate DESC, replyAuthorId

Plan

See plans/cypher-996-optional-match-case/plan.md

lmeyerov and others added 2 commits April 4, 2026 20:45
…e_eq__

pandas `Series == None` always returns False; for CASE x WHEN null the
equality must use the null-mask of the non-null side instead.

Adds test for IS7-shape: connected MATCH + OPTIONAL MATCH + CASE r.

Fixes #996

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
@lmeyerov
lmeyerov force-pushed the feat/issue-996-optional-match-case branch from a0b2f13 to 8bb0a2d Compare April 5, 2026 03:45
lmeyerov and others added 2 commits April 4, 2026 20:53
…; fix CHANGELOG

- 7 tests: __cypher_case_eq__ null boundary (null matches null, no-ELSE, mixed series, regression guard)
- 5 tests: ConnectedOptionalMatch left-join (all match, none match, node alias null, ORDER BY opt col, searched CASE regression)
- Fix unused variable lint in node_alias_null test
- CHANGELOG: add Fixed entry for pandas null-comparison bug + Tests entry for IS7 regression coverage

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
@lmeyerov
lmeyerov marked this pull request as ready for review April 5, 2026 04:32
lmeyerov and others added 3 commits April 5, 2026 14:44
Exercises __cypher_case_eq__ null fix on GPU row pipeline via cudf engine.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
…edge alias synthesis

marker.apply(lambda v: ...) fails on cudf due to module pickling error.
notna() is equivalent and works on both pandas and cudf.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
@lmeyerov
lmeyerov merged commit b7843b7 into master Apr 5, 2026
101 checks passed
@lmeyerov
lmeyerov deleted the feat/issue-996-optional-match-case branch April 5, 2026 22:37
lmeyerov added a commit that referenced this pull request Apr 6, 2026
Gap tests:
- safe_map_series: cudf + pd.Series mapping path (Engine.py:399-400)
- hop: missing_mask guard when edge_hop_col=None (hop.py:973)
- hop: empty edge_map_df else branch (hop.py:961-965)

DRY: extract _mk_abc_chain() helper in test_hop.py — 5 label tests refactored

CHANGELOG: fix call-site count (8, not 7); add #977 Tests entry (23 tests);
correct file list (no gfql_unified.py — that fix was in #1061)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
lmeyerov added a commit that referenced this pull request Apr 6, 2026
…1066)

* chore: init branch for #977 cudf SIGSEGV fix

* fix(gfql): safe_map_series — replace cudf-incompatible Series.map(dict) with pandas bridge (#977)

cudf Series.map(dict/Series) triggers numba JIT which SIGSEGVs on RAPIDS 25.02.
Introduce safe_map_series() in Engine.py that bridges through pandas for cudf.
Apply at all 7 affected call sites: hop.py (4), chain.py (1), df_executor.py (2), pipeline.py (1).
Includes DGX smoke test + unit tests for safe_map_series + cudf guard regression tests.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* fix(gfql): guard hop.py edge_map usage against UnboundLocalError (#977)

edge_map is only assigned inside the edge_hop_col branch; line 974
used it unconditionally which raised UnboundLocalError when
edge_hop_col is None. Previously masked by SIGSEGV on 25.02.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* test(gfql): add cudf op bisect script for #977 25.02 SIGSEGV investigation

* test(gfql): bisect2 — minimal step-by-step crash location for #977 hop label_node_hops

* test(gfql): bisect3 — targeted hop label_node_hops cudf crash hunt

* fix(test): bisect3 typo fix

* fix(gfql): safe_map_series — use merge+to_arrow for cudf to avoid numba SIGSEGV on 25.02

* test(gfql): fix smoke test list-return query (unsupported), use simple ORDER BY

* fix(gfql): safe_map_series use to_arrow — CHANGELOG update, DGX-validated on 25.02+26.02 (#977)

* refactor(gfql): simplify safe_map_series — drop sort, drop_duplicates on lookup, cleaner cudf check

* test(gfql): amplify safe_map_series, hop labels, chain, ORDER BY — #977 coverage

* chore: remove temporary bisect scripts used to diagnose #977 SIGSEGV

* fix(types): safe_map_series — use DataframeLike/Union signature per Engine.py conventions

* fix(types): safe_map_series — use lazy_cudf_import + isinstance per Engine.py conventions

* fix(test): remove stray conflict marker from test_lowering.py

<<<<<<< HEAD line was left from incomplete conflict resolution
during rebase; caused SyntaxError in CI test collection.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* test(gfql): fourth amplification round + DRY — #977 coverage

Gap tests:
- safe_map_series: cudf + pd.Series mapping path (Engine.py:399-400)
- hop: missing_mask guard when edge_hop_col=None (hop.py:973)
- hop: empty edge_map_df else branch (hop.py:961-965)

DRY: extract _mk_abc_chain() helper in test_hop.py — 5 label tests refactored

CHANGELOG: fix call-site count (8, not 7); add #977 Tests entry (23 tests);
correct file list (no gfql_unified.py — that fix was in #1061)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* test(gfql): fifth amplification round + DRY — #977 coverage

New tests:
- output_max_hops path (hop.py:834-891 fallback_map try/except + line-1011 post-filter)
- label_seeds=True with implicit starting_nodes (hop.py:843-859 both-branches path)
- numeric (int) node IDs through safe_map_series int→int hop labeling (hop.py:938)
- undirected label_seeds=False clears seed hop (hop.py:987-1000)

DRY: extract _mk_abcd_chain() helper — multi_hop_ordering + output_max_hops refactored

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* test(gfql): extend DGX smoke script with amplification-round cudf paths (#977)

New smoke tests cover:
- safe_map_series + pd.Series mapping (Engine.py:399-400)
- output_max_hops filtering (hop.py:834-891 fallback_map + line-1011 post-filter)
- label_seeds=True seed rows from node_hop_records (hop.py:843-859)
- label_seeds=False + undirected clears seed hop (hop.py:987-1000)

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

* fix(test): smoke script — use to_arrow().to_pylist() for int cudf columns

to_pandas() on integer cudf Series SIGSEGVs on RAPIDS 25.02 (same numba
path as the original #977 bug). String columns are safe. Switch all numeric
hop value reads to to_arrow().to_pylist() to avoid the numba JIT trigger.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>

---------

Co-authored-by: Claude Sonnet 4.6 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cypher/GFQL cannot lower MATCH + OPTIONAL MATCH reply-author flag queries after a connected MATCH

1 participant