Highest-severity finding of the amplification campaign — release-blocking. Round-009 lifecycle probe, ~40 scenarios + 585 differential checks, both engines, twice-verified. Scripts in session scratchpad probe-r9-lifecycle/.
1. SILENT-WRONG on both engines: rebind resurrects a stale index
The declared-index identity+fingerprint guard is correct, and get_valid() correctly reports the index stale after a rebind. The chain executor then hands it a new lease on life. compute/chain.py:1220-1223 and its polars twin gfql/lazy/engine/polars/chain.py:1055-1058 re-point the resident index at the new frame unconditionally:
_reg = get_registry(g)
if not _reg.is_empty():
g = set_registry(g, _reg.rebind_edges(indexed_edges_df))
GfqlIndexRegistry.rebind_edges (gfql/index/registry.py:290-328) validates only the NEW frame's structural fingerprint — (row count, bound column names, engine) — and never checks the index was still valid for the frame it is augmenting. Its own docstring states that precondition; no caller enforces it. After g.edges(other_frame) the identity check has already failed, but the shallow copy re-passes the structural check, so a CSR built over the OLD edges is marked valid for the NEW ones.
E1 = pd.DataFrame({'s':[0,1,2,3,4], 'd':[1,2,3,4,5]})
E2 = pd.DataFrame({'s':[0,3,1,4,2], 'd':[3,0,4,1,5]}) # same row count
g = gfql_index_edges(graphistry.nodes(N,'id').edges(E1,'s','d'))
g.chain([n({'id':0}), e_forward(), n(), e_forward(), n()]) # warm
gr = g.edges(E2, 's', 'd') # documented rebind
gr.chain([n({'id':0}), e_forward(), n(), e_forward(), n()])
oracle (fresh bind of E2) nodes [0,3] edges [(0,3),(3,0)] — actual [2] / []. Control without an index is correct. At N=4000 with the index confirmed engaged via index_trace: 2-hop [2] vs oracle [0,1582,1707].
Causality proven: monkeypatching rebind_edges to DROP indexes instead of re-pointing makes every wrong answer correct on both engines, nothing else changed.
Precision boundary — note the second line, which is an ordinary user action:
- same row count, different values → WRONG
- same edge set, rows permuted (
df.sort_values(...)) → WRONG ([2] vs [0,1,2])
- different row count → safe (fingerprint catches it)
E1.assign(w=...), E1.copy(deep=True) → correct, so the optimization's intended shallow-augmentation case is unaffected
Affected: multi-hop and var-length chain on both engines. Single-hop chain, direct hop(), degree facts, node-prop/col-stats indexes and the Cypher fast paths were clean in this sweep.
Aggravating: show_indexes(gr) reports "stale fingerprint (frames rebound since build) -> rebuild" while the same query is being served from those indexes — the diagnostic surface contradicts the executor.
Why existing tests miss it: test_rebind_edges_revalidates_after_shallow_augmentation and test_rebind_edges_drops_index_on_fingerprint_mismatch (tests/compute/gfql/index/test_index.py:929-975) both start from an index that is already valid for the frame being augmented. The already-stale-then-augmented case has no test.
2. Both documented recovery recipes fail
ComputeMixin.gfql.__doc__ (ComputeMixin.py:757-761) says: "After mutating, rebind fresh frames or call gfql_clear_caches()." With a resident adjacency index, neither works (N=4000, in-place edit of both indexed columns, oracle [0,2000,2001,3000]):
| recovery |
pandas |
polars |
gfql_clear_caches() |
STILL WRONG |
STILL WRONG |
| rebind a fresh deep-copied frame |
STILL WRONG |
STILL WRONG |
drop_index(...) |
recovered |
recovered |
brand-new nodes().edges() |
recovered |
recovered |
Two causes: (a) gfql_clear_caches's own docstring says graph-keyed caches are NOT touched — the index registry is exactly such a cache — so the gfql() docstring contradicts it; (b) the "rebind fresh frames" arm is defeated by Finding 1, since an in-place edit always preserves row count. The polars NaN-clean verdict cache honors both recipes correctly; this is specific to the per-graph index registry.
3. Completeness-lock blind spots (latent, no live instance)
test_clear_caches_covers_every_cache.py passes and every GFQL-tree cache is genuinely registered, but discovery scans only gfql/** + gfql_unified.py (missing chain.py, hop.py, *_fast_paths.py, ComputeMixin.py, all on the execution path — a repo-wide AST sweep found no memo there today) and only matches names containing cache/memo (so _OFFENGINE_BRIDGE_WARNED in gfql/call/executor.py:69 and _COST_GATE_FRAC_OVERRIDES in index/cost.py:18 are unregistered process-global state; warning-order and plan-choice effects only, not wrong answers).
4. Raw engine errors on ComputeMixin helpers with polars frames
filter_nodes_by_dict / filter_edges_by_dict → AttributeError: 'DataFrame' object has no attribute 'assign'; prune_self_edges → raw ValueError. Non-GFQL surfaces, listed for completeness.
Suggested fix
Require idx.source_ref is old_edges before re-pointing (drop otherwise), or gate the two chain call sites on _reg.get_valid(kind, g._edges, (src,dst), engine) is not None. Either restores the safe miss-to-scan the registry docstring promises while keeping the intended shallow-augmentation win. Then reconcile the two recovery docstrings.
What was QUIET (strong results)
Zero library-side mutation of caller frames — 52 surfaces × 2 engines, 88 executions, deep before/after snapshots incl. object and per-column .values identity. Parse/compile cache — 126 cross-graph checks on binding/dtype-collision hazards, 0 mismatches. id()-keyed cache recycle hazard — 30,000 iterations with 30,000 confirmed id recycles, 0 wrong answers: weakref.finalize fires before the id can be reused, so the guard is sound rather than lucky. Per-caller setattr channels — all copy/bind first. Concurrency — ~480 interleaved executions across 6 threads matched single-threaded oracles exactly.
Highest-severity finding of the amplification campaign — release-blocking. Round-009 lifecycle probe, ~40 scenarios + 585 differential checks, both engines, twice-verified. Scripts in session scratchpad probe-r9-lifecycle/.
1. SILENT-WRONG on both engines: rebind resurrects a stale index
The declared-index identity+fingerprint guard is correct, and
get_valid()correctly reports the index stale after a rebind. The chain executor then hands it a new lease on life.compute/chain.py:1220-1223and its polars twingfql/lazy/engine/polars/chain.py:1055-1058re-point the resident index at the new frame unconditionally:GfqlIndexRegistry.rebind_edges(gfql/index/registry.py:290-328) validates only the NEW frame's structural fingerprint — (row count, bound column names, engine) — and never checks the index was still valid for the frame it is augmenting. Its own docstring states that precondition; no caller enforces it. Afterg.edges(other_frame)the identity check has already failed, but the shallow copy re-passes the structural check, so a CSR built over the OLD edges is marked valid for the NEW ones.oracle (fresh bind of E2) nodes
[0,3]edges[(0,3),(3,0)]— actual[2]/[]. Control without an index is correct. At N=4000 with the index confirmed engaged viaindex_trace: 2-hop[2]vs oracle[0,1582,1707].Causality proven: monkeypatching
rebind_edgesto DROP indexes instead of re-pointing makes every wrong answer correct on both engines, nothing else changed.Precision boundary — note the second line, which is an ordinary user action:
df.sort_values(...)) → WRONG ([2]vs[0,1,2])E1.assign(w=...),E1.copy(deep=True)→ correct, so the optimization's intended shallow-augmentation case is unaffectedAffected: multi-hop and var-length
chainon both engines. Single-hop chain, directhop(), degree facts, node-prop/col-stats indexes and the Cypher fast paths were clean in this sweep.Aggravating:
show_indexes(gr)reports"stale fingerprint (frames rebound since build) -> rebuild"while the same query is being served from those indexes — the diagnostic surface contradicts the executor.Why existing tests miss it:
test_rebind_edges_revalidates_after_shallow_augmentationandtest_rebind_edges_drops_index_on_fingerprint_mismatch(tests/compute/gfql/index/test_index.py:929-975) both start from an index that is already valid for the frame being augmented. The already-stale-then-augmented case has no test.2. Both documented recovery recipes fail
ComputeMixin.gfql.__doc__(ComputeMixin.py:757-761) says: "After mutating, rebind fresh frames or callgfql_clear_caches()." With a resident adjacency index, neither works (N=4000, in-place edit of both indexed columns, oracle[0,2000,2001,3000]):gfql_clear_caches()drop_index(...)nodes().edges()Two causes: (a)
gfql_clear_caches's own docstring says graph-keyed caches are NOT touched — the index registry is exactly such a cache — so thegfql()docstring contradicts it; (b) the "rebind fresh frames" arm is defeated by Finding 1, since an in-place edit always preserves row count. The polars NaN-clean verdict cache honors both recipes correctly; this is specific to the per-graph index registry.3. Completeness-lock blind spots (latent, no live instance)
test_clear_caches_covers_every_cache.pypasses and every GFQL-tree cache is genuinely registered, but discovery scans onlygfql/**+gfql_unified.py(missingchain.py,hop.py,*_fast_paths.py,ComputeMixin.py, all on the execution path — a repo-wide AST sweep found no memo there today) and only matches names containingcache/memo(so_OFFENGINE_BRIDGE_WARNEDingfql/call/executor.py:69and_COST_GATE_FRAC_OVERRIDESinindex/cost.py:18are unregistered process-global state; warning-order and plan-choice effects only, not wrong answers).4. Raw engine errors on ComputeMixin helpers with polars frames
filter_nodes_by_dict/filter_edges_by_dict→AttributeError: 'DataFrame' object has no attribute 'assign';prune_self_edges→ rawValueError. Non-GFQL surfaces, listed for completeness.Suggested fix
Require
idx.source_ref is old_edgesbefore re-pointing (drop otherwise), or gate the two chain call sites on_reg.get_valid(kind, g._edges, (src,dst), engine) is not None. Either restores the safe miss-to-scan the registry docstring promises while keeping the intended shallow-augmentation win. Then reconcile the two recovery docstrings.What was QUIET (strong results)
Zero library-side mutation of caller frames — 52 surfaces × 2 engines, 88 executions, deep before/after snapshots incl. object and per-column
.valuesidentity. Parse/compile cache — 126 cross-graph checks on binding/dtype-collision hazards, 0 mismatches.id()-keyed cache recycle hazard — 30,000 iterations with 30,000 confirmed id recycles, 0 wrong answers:weakref.finalizefires before the id can be reused, so the guard is sound rather than lucky. Per-caller setattr channels — all copy/bind first. Concurrency — ~480 interleaved executions across 6 threads matched single-threaded oracles exactly.