Skip to content

perf(gfql): projection-pushdown into binding materialization — skip unused alias property joins (pandas/cuDF) #1711

Description

@lmeyerov

Summary

The bindings-row builder attaches every node alias's property columns (alias.{col}) for every node alias in the pattern, regardless of whether the downstream query references them. For aggregate/count queries that reference few or no properties, this materializes large intermediates for nothing.

Evidence (profiled 2026-07-06)

MATCH (a)-[]->(b)-[]->(c) RETURN count(*) on polars @100k nodes / 1M edges:

  • full query = 399.8ms; the rows(binding_ops) build alone = 382.9ms (96%); group_by+select = 16.9ms.
  • build output cols: a, b, c, a.node_id, a.age, b.node_id, b.age, c.node_id, c.age — but count(*) needs none of the .age columns.
  • the raw join-chain (no property attach) is only ~50ms → ~330ms is wasted unused-property joins.

So skipping unreferenced property columns is a ~7× win on q8 (383→~50ms build), and it directly de-risks the pandas full-data OOM concern (57M-path 2-hop binding × unused property columns at 2.42M edges).

Fix (pandas/cuDF)

Thread the set of referenced alias.property columns from the cypher lowering into rows(binding_ops=..., keep_alias_cols=<set>), and have the builders (_gfql_connected_bindings_row_frame_from_state pandas; binding_rows_polars for the non-lazy path) only left-join the kept columns. Bare alias id columns + grouping/join keys always kept.

Safety: the keep-set MUST be a conservative superset of referenced columns (walk all downstream ops: RETURN/WHERE/ORDER BY/GROUP BY keys+aggs/WITH stages/hidden-reentry refs). Under-inclusion = wrong answer; over-inclusion = merely slower. Default to keep-all when the set can't be computed confidently.

Scope / relation

Activity

  1. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    Update (2026-07-06): this is now critical path, and it affects polars/polars-gpu too — not just pandas/cuDF.

    I built the lazy binding-build (polars, to enable polars-gpu offload) and benchmarked on dgx. Result:

    • q1/q9/q3 on polars-gpu: now win (GPU offload works, ~1.5-1.7×).
    • q8 (unfiltered 2-hop count) on polars-gpu: REGRESSED 19× (228ms → 4294ms).

    Root cause = exactly this issue: q8 attaches a.age/b.age/c.age to the full 10M-path intermediate for a count(*) that needs none of them. On CPU that just wastes ~330ms; on GPU the large unpruned intermediate is catastrophic (4.3s). The de-risk probe (join-chain only, no property joins) was 12ms on GPU — so pruning turns q8 from a 4.3s regression into a ~12ms win.

    Consequence: projection-pushdown here is the prerequisite for the lazy-GPU work, and it delivers on all four engines (pandas 383→~50ms build, cuDF, polars-cpu 7×, and unblocks polars-gpu). The lazy-build patch is preserved and re-applies cleanly once this lands. Keeping the scope as originally described (thread the referenced alias.property superset from lowering into the builders; conservative keep-all default). Not shipping the lazy build until this is in — HEAD stays regression-free.

  2. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    DONE + validated on dgx (commits 626e6ed6 projection-pushdown + c381a481 lazy build; branch dev/gfql-cypher-std-conformance).

    Result @100k nodes / 1M edges (dgx, safe_run, all values correct): the q8-gpu regression is fixed and it is now the fastest engine.

    q8 2-hop count before after
    polars-gpu 4294ms (lazy-alone regression) 54ms (79×, fastest)
    polars-cpu 244ms 90ms
    cuDF 973ms 439ms
    pandas 4386ms 1122ms

    The pushdown (skip unused property joins) cut every engine ~2.5-4× on q8, and — paired with the lazy collect-once-on-target build — keeps the GPU intermediate small so polars-gpu finally offloads and wins. Conservative gate as described (single-MATCH, no WITH/WHERE, no repeated alias, no collect → attach-all otherwise). Differential parity vs pandas holds (gfql 3044 + compute 4331 green).

    Remaining as future refinement (non-blocking): per-PROPERTY pruning (currently per-alias), and extending past the no-WHERE gate (filtered queries keep-all, already fast). Closing the benchmark-blocking scope.

  3. added 20 commits that reference this issue on Jul 9, 2026
  4. added a commit that references this issue on Aug 15, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions