Skip to content

GFQL: legitimate .edges() rebind resurrects a STALE adjacency index (silent wrong answers, both engines) + documented recovery fails #1913

Description

@lmeyerov

Highest-severity finding of the amplification campaign — release-blocking. Round-009 lifecycle probe, ~40 scenarios + 585 differential checks, both engines, twice-verified. Scripts in session scratchpad probe-r9-lifecycle/.

1. SILENT-WRONG on both engines: rebind resurrects a stale index

The declared-index identity+fingerprint guard is correct, and get_valid() correctly reports the index stale after a rebind. The chain executor then hands it a new lease on life. compute/chain.py:1220-1223 and its polars twin gfql/lazy/engine/polars/chain.py:1055-1058 re-point the resident index at the new frame unconditionally:

_reg = get_registry(g)
if not _reg.is_empty():
    g = set_registry(g, _reg.rebind_edges(indexed_edges_df))

GfqlIndexRegistry.rebind_edges (gfql/index/registry.py:290-328) validates only the NEW frame's structural fingerprint — (row count, bound column names, engine) — and never checks the index was still valid for the frame it is augmenting. Its own docstring states that precondition; no caller enforces it. After g.edges(other_frame) the identity check has already failed, but the shallow copy re-passes the structural check, so a CSR built over the OLD edges is marked valid for the NEW ones.

E1 = pd.DataFrame({'s':[0,1,2,3,4], 'd':[1,2,3,4,5]})
E2 = pd.DataFrame({'s':[0,3,1,4,2], 'd':[3,0,4,1,5]})   # same row count
g  = gfql_index_edges(graphistry.nodes(N,'id').edges(E1,'s','d'))
g.chain([n({'id':0}), e_forward(), n(), e_forward(), n()])   # warm
gr = g.edges(E2, 's', 'd')                                   # documented rebind
gr.chain([n({'id':0}), e_forward(), n(), e_forward(), n()])

oracle (fresh bind of E2) nodes [0,3] edges [(0,3),(3,0)] — actual [2] / []. Control without an index is correct. At N=4000 with the index confirmed engaged via index_trace: 2-hop [2] vs oracle [0,1582,1707].

Causality proven: monkeypatching rebind_edges to DROP indexes instead of re-pointing makes every wrong answer correct on both engines, nothing else changed.

Precision boundary — note the second line, which is an ordinary user action:

  • same row count, different values → WRONG
  • same edge set, rows permuted (df.sort_values(...)) → WRONG ([2] vs [0,1,2])
  • different row count → safe (fingerprint catches it)
  • E1.assign(w=...), E1.copy(deep=True) → correct, so the optimization's intended shallow-augmentation case is unaffected

Affected: multi-hop and var-length chain on both engines. Single-hop chain, direct hop(), degree facts, node-prop/col-stats indexes and the Cypher fast paths were clean in this sweep.

Aggravating: show_indexes(gr) reports "stale fingerprint (frames rebound since build) -> rebuild" while the same query is being served from those indexes — the diagnostic surface contradicts the executor.

Why existing tests miss it: test_rebind_edges_revalidates_after_shallow_augmentation and test_rebind_edges_drops_index_on_fingerprint_mismatch (tests/compute/gfql/index/test_index.py:929-975) both start from an index that is already valid for the frame being augmented. The already-stale-then-augmented case has no test.

2. Both documented recovery recipes fail

ComputeMixin.gfql.__doc__ (ComputeMixin.py:757-761) says: "After mutating, rebind fresh frames or call gfql_clear_caches()." With a resident adjacency index, neither works (N=4000, in-place edit of both indexed columns, oracle [0,2000,2001,3000]):

recovery pandas polars
gfql_clear_caches() STILL WRONG STILL WRONG
rebind a fresh deep-copied frame STILL WRONG STILL WRONG
drop_index(...) recovered recovered
brand-new nodes().edges() recovered recovered

Two causes: (a) gfql_clear_caches's own docstring says graph-keyed caches are NOT touched — the index registry is exactly such a cache — so the gfql() docstring contradicts it; (b) the "rebind fresh frames" arm is defeated by Finding 1, since an in-place edit always preserves row count. The polars NaN-clean verdict cache honors both recipes correctly; this is specific to the per-graph index registry.

3. Completeness-lock blind spots (latent, no live instance)

test_clear_caches_covers_every_cache.py passes and every GFQL-tree cache is genuinely registered, but discovery scans only gfql/** + gfql_unified.py (missing chain.py, hop.py, *_fast_paths.py, ComputeMixin.py, all on the execution path — a repo-wide AST sweep found no memo there today) and only matches names containing cache/memo (so _OFFENGINE_BRIDGE_WARNED in gfql/call/executor.py:69 and _COST_GATE_FRAC_OVERRIDES in index/cost.py:18 are unregistered process-global state; warning-order and plan-choice effects only, not wrong answers).

4. Raw engine errors on ComputeMixin helpers with polars frames

filter_nodes_by_dict / filter_edges_by_dict → AttributeError: 'DataFrame' object has no attribute 'assign'; prune_self_edges → raw ValueError. Non-GFQL surfaces, listed for completeness.

Suggested fix

Require idx.source_ref is old_edges before re-pointing (drop otherwise), or gate the two chain call sites on _reg.get_valid(kind, g._edges, (src,dst), engine) is not None. Either restores the safe miss-to-scan the registry docstring promises while keeping the intended shallow-augmentation win. Then reconcile the two recovery docstrings.

What was QUIET (strong results)

Zero library-side mutation of caller frames — 52 surfaces × 2 engines, 88 executions, deep before/after snapshots incl. object and per-column .values identity. Parse/compile cache — 126 cross-graph checks on binding/dtype-collision hazards, 0 mismatches. id()-keyed cache recycle hazard — 30,000 iterations with 30,000 confirmed id recycles, 0 wrong answers: weakref.finalize fires before the id can be reused, so the guard is sound rather than lucky. Per-caller setattr channels — all copy/bind first. Concurrency — ~480 interleaved executions across 6 threads matched single-threaded oracles exactly.

Activity

  1. added 5 commits that reference this issue on Aug 15, 2026
  2. added a commit that references this issue on Aug 16, 2026
  3. added 2 commits that reference this issue on Aug 16, 2026
  4. added a commit that references this issue on Aug 19, 2026
  5. lmeyerov commented on Aug 19, 2026

    @lmeyerov
    ContributorAuthor

    Closing: the master audit (vs e6625ed28) verified findings 1–3 fixed by the landed #1914 — rebind returns fresh answers, the deep-copy recovery recipe works, and the contract is pinned (test_rebind_edges_leaves_index_resident_after_in_place_shape_mutation et al). Finding 4 — the polars helper-surface crash family — landed in #1942 (f001cbbc1): prune_self_edges is polars-native (incl. the silent column-selection hazard when rowcount==colcount) and the filter_*_by_dict half fell out of the #1882 fix, both pinned. All four findings served.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions