Skip to content

GFQL index breaks the documented 'same answer, faster' contract: EXISTS self-loop returns wrong rows with an index, declines without #1986

Description

@lmeyerov

Found by a 2026-08-19 audit for suppressed wrong answers. Verified on master 54266cc6f. Pre-existing.

Repro

import polars as pl, graphistry
g = (graphistry
     .nodes(pl.DataFrame({'id': [0,1,2,3]}), 'id')
     .edges(pl.DataFrame({'s': [0,3], 'd': [1,3]}), 's','d'))   # only node 3 self-loops

q = "MATCH (n) WHERE EXISTS { (n)-->(n) } RETURN n.id"
g.gfql(q, engine='polars')                    # NotImplementedError  (honest decline)
g.gfql_index_all().gfql(q, engine='polars')   # [{'n.id': 0}, {'n.id': 3}]

Truth is [3]. Node 0's only edge is 0 -> 1, which is not a self-loop.

So attaching an index converts an honest decline into a wrong answer, and the wrong answer is worse than the decline it replaced.

Why this one matters beyond the bug

It contradicts the contract this repo states in its own words at graphistry/compute/gfql/index/api.py:190:

Fast paths are contracted "same answer, faster"

An index is an optimization. It must never change an answer. Here it changes a decline into wrong rows.

Mechanism

graphistry/compute/gfql/lazy/engine/polars/pattern_apply.py:138-172 answers pattern membership from edge-table CSR keys, skipping:

  • the duplicate-alias check (which is what makes (n)-->(n) a self-loop rather than any outgoing edge), and
  • the node-table intersection.

The NOT EXISTS form has the mirror defect: it drops rows that genuinely satisfy the predicate.

Trigger

polars frames plus a resident index. Without the index the same query declines correctly, so this is reachable only on the indexed path.

Expected

Either the indexed path answers identically to the scan path, or it declines to the scan path. Never a third, wrong answer.

Related

A weaker instance of the same contract gap is already tolerated in a test: graphistry/tests/compute/gfql/index/test_index.py:1298-1307 relaxes an equality assertion to extra <= seed_ids because indexed and scan to_fixed_point disagree (indexed 1957 node rows vs scan 1956 — the indexed result includes the seed, the scan does not). That comment is honest about refusing to encode a bug as expected, but it tolerates the divergence with no strict xfail to flip when fixed.

Activity

  1. added a commit that references this issue on Aug 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions