Repository navigation
perf(gfql): projection-pushdown into binding materialization — skip unused alias property joins (pandas/cuDF) #1711
Description
Activity
Update (2026-07-06): this is now critical path, and it affects polars/polars-gpu too — not just pandas/cuDF.
I built the lazy binding-build (polars, to enable polars-gpu offload) and benchmarked on dgx. Result:
- q1/q9/q3 on polars-gpu: now win (GPU offload works, ~1.5-1.7×).
- q8 (unfiltered 2-hop count) on polars-gpu: REGRESSED 19× (228ms → 4294ms).
Root cause = exactly this issue: q8 attaches
a.age/b.age/c.ageto the full 10M-path intermediate for acount(*)that needs none of them. On CPU that just wastes ~330ms; on GPU the large unpruned intermediate is catastrophic (4.3s). The de-risk probe (join-chain only, no property joins) was 12ms on GPU — so pruning turns q8 from a 4.3s regression into a ~12ms win.Consequence: projection-pushdown here is the prerequisite for the lazy-GPU work, and it delivers on all four engines (pandas 383→~50ms build, cuDF, polars-cpu 7×, and unblocks polars-gpu). The lazy-build patch is preserved and re-applies cleanly once this lands. Keeping the scope as originally described (thread the referenced
alias.propertysuperset from lowering into the builders; conservative keep-all default). Not shipping the lazy build until this is in — HEAD stays regression-free.DONE + validated on dgx (commits 626e6ed6 projection-pushdown + c381a481 lazy build; branch dev/gfql-cypher-std-conformance).
Result @100k nodes / 1M edges (dgx, safe_run, all values correct): the q8-gpu regression is fixed and it is now the fastest engine.
q8 2-hop count before after polars-gpu 4294ms (lazy-alone regression) 54ms (79×, fastest) polars-cpu 244ms 90ms cuDF 973ms 439ms pandas 4386ms 1122ms The pushdown (skip unused property joins) cut every engine ~2.5-4× on q8, and — paired with the lazy
collect-once-on-target build — keeps the GPU intermediate small sopolars-gpufinally offloads and wins. Conservative gate as described (single-MATCH, no WITH/WHERE, no repeated alias, no collect → attach-all otherwise). Differential parity vs pandas holds (gfql 3044 + compute 4331 green).Remaining as future refinement (non-blocking): per-PROPERTY pruning (currently per-alias), and extending past the no-WHERE gate (filtered queries keep-all, already fast). Closing the benchmark-blocking scope.
- added 20 commits that reference this issue
on Jul 9, 2026 - added a commit that references this issue
on Jul 9, 2026 - added a commit that references this issue
on Aug 15, 2026
Summary
The bindings-row builder attaches every node alias's property columns (
alias.{col}) for every node alias in the pattern, regardless of whether the downstream query references them. For aggregate/count queries that reference few or no properties, this materializes large intermediates for nothing.Evidence (profiled 2026-07-06)
MATCH (a)-[]->(b)-[]->(c) RETURN count(*)on polars @100k nodes / 1M edges:rows(binding_ops)build alone = 382.9ms (96%); group_by+select = 16.9ms.a, b, c, a.node_id, a.age, b.node_id, b.age, c.node_id, c.age— butcount(*)needs none of the.agecolumns.So skipping unreferenced property columns is a ~7× win on q8 (383→~50ms build), and it directly de-risks the pandas full-data OOM concern (57M-path 2-hop binding × unused property columns at 2.42M edges).
Fix (pandas/cuDF)
Thread the set of referenced
alias.propertycolumns from the cypher lowering intorows(binding_ops=..., keep_alias_cols=<set>), and have the builders (_gfql_connected_bindings_row_frame_from_statepandas;binding_rows_polarsfor the non-lazy path) only left-join the kept columns. Bare alias id columns + grouping/join keys always kept.Safety: the keep-set MUST be a conservative superset of referenced columns (walk all downstream ops: RETURN/WHERE/ORDER BY/GROUP BY keys+aggs/WITH stages/hidden-reentry refs). Under-inclusion = wrong answer; over-inclusion = merely slower. Default to keep-all when the set can't be computed confidently.
Scope / relation