You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
GFQL: pandas Cypher RETURN row-pipeline is ~O(N) — ~95ms to return 1 row at 500k nodes (polars ~80× faster) #1670
The pandas GFQL Cypher RETURN / row-pipeline path is ~O(N) in the graph size even for a 1-row result. Returning a single matched node from a 500K-node graph takes ~95 ms on pandas vs ~1.2 ms on polars (~80×). The cost is in the RETURN/result-postprocess path, not the node filter (which finds the row in <1 ms) and not indexing.
This surfaced while investigating point-lookup competitiveness (a Ladybug/Kuzu comparison): the "slow point lookup" is really this RETURN-path overhead, so it's worth its own issue independent of the CSR adjacency-index work (#1658).
Scales ~linearly with N (1-row RETURN a.id, pandas): 15.6 ms @ 50K → 100 ms @ 500K.
Both RETURN a (whole entity) and RETURN a.id (scalar) are ~equally slow, so it is not whole-entity rendering specifically — it looks like a fixed O(N) pass in the row pipeline / result_postprocess.
An O(N) pass in the pandas RETURN path (compute/gfql/cypher/result_postprocess.py and/or the row pipeline) that runs regardless of result cardinality.
polars already avoids it (~1 ms), so a candidate fix is to bring the pandas path in line (or steer lookup-heavy Cypher to polars/cuDF).
Impact
Lookup/point-query workloads on pandas. GFQL already wins full scans / counts / traversals; this is the one shape where a tiny result pays a large fixed cost.
Summary
The pandas GFQL Cypher
RETURN/ row-pipeline path is ~O(N) in the graph size even for a 1-row result. Returning a single matched node from a 500K-node graph takes ~95 ms on pandas vs ~1.2 ms on polars (~80×). The cost is in the RETURN/result-postprocess path, not the node filter (which finds the row in <1 ms) and not indexing.This surfaced while investigating point-lookup competitiveness (a Ladybug/Kuzu comparison): the "slow point lookup" is really this RETURN-path overhead, so it's worth its own issue independent of the CSR adjacency-index work (#1658).
Evidence (local CPU, 500K nodes / 2M edges, id = 1..N)
gfql([n({'id': X})])MATCH (a) WHERE a.id = X RETURN a(1 row)MATCH (a) WHERE a.id = X RETURN a.id(1 row)Scales ~linearly with N (1-row
RETURN a.id, pandas): 15.6 ms @ 50K → 100 ms @ 500K.Both
RETURN a(whole entity) andRETURN a.id(scalar) are ~equally slow, so it is not whole-entity rendering specifically — it looks like a fixed O(N) pass in the row pipeline / result_postprocess.Repro
Hypothesis / where to look
compute/gfql/cypher/result_postprocess.pyand/or the row pipeline) that runs regardless of result cardinality.Impact
Lookup/point-query workloads on pandas. GFQL already wins full scans / counts / traversals; this is the one shape where a tiny result pays a large fixed cost.
Notes
plans/gfql-1658-seeded-index/receipts/lookup-finding.md.