Skip to content

benchmarks/gfql: q1-q9 runner uses dataframe shortcuts (bypasses GFQL engine) + untimed lowercase precompute — unfair vs competitor DBs #1710

Description

@lmeyerov

Summary (benchmark integrity — must fix before publishing "official" numbers)

benchmarks/gfql/graph_benchmark_q1_q9.py (GFQL vs Kuzu/Ladybug/Neo4j/Memgraph on the prrao87 graph-benchmark suite) gives GFQL two unfair advantages vs the competitor databases, which all run their own real Cypher:

1. Dataframe shortcuts bypass the GFQL engine entirely

For q1, q3, q4, q5, q6, q7 (query_variant="standard", the default — _uses_dataframe_shortcut, line 238-243), GFQL does not run its query engine. It runs hand-written pandas/cuDF groupby/merge code (_query{1,3,4,5,6,7}_dataframe_shortcut, plus _query5_polars_shortcut, _query6_cudf_shortcut, etc.). That measures hand-tuned dataframe code, not GFQL — while Kuzu/Neo4j/Memgraph parse+plan+execute real Cypher. Not apples-to-apples.

  • Fix: run real GFQL (g.gfql(<cypher>, engine=...)) for all queries. On pandas/cuDF this works for q1-q9 once the count(<other-alias>) lowering routing is fixed (see the lowering issue; q1 also works today via the count(*) form). On polars, traversal queries NIE (see the polars binding_ops issue) and count(*) is currently wrong (see the P0 count-broadcast issue).

2. Untimed lowercase precompute (both GFQL and the Memgraph runner)

  • graph_benchmark_q1_q9.py precomputes gender_lc/interest_lc at load (untimed) for the toLower() filters in q5/q6/q7. Measured impact when moved inside the timed region: <1ms (negligible) — but for symmetry it should be timed or removed.
  • graph_benchmark_memgraph_q1_q9.py similarly precomputes lowercased columns. The competitors (Kuzu/Ladybug/Neo4j) run raw tolower() inline. For official numbers, no side gets untimed precompute — everyone runs the raw query.

Data-scale caveat (already caught)

/tmp/graph-benchmark-gfql-memgraph = TINY (1k persons/10k edges). Full data (100k/2.42M) = /tmp/graph-benchmark-gfql-memgraph-full (repo data/ symlinks to it; the Docker mount does NOT follow the symlink — mount the target dir directly). An earlier Memgraph q1-q9 run on the tiny set produced invalid ~100× fast numbers.

Acceptance

  • GFQL runs its real engine (no dataframe shortcuts) for every query it reports.
  • No untimed precompute for any system.
  • Every system runs the identical canonical query from the prrao87 repo (neo4j/query.py etc.).
  • Where GFQL can't yet run a query as real GFQL on an engine, the cell is marked honestly (NIE / in-progress), NOT shortcut-faked.

Activity

  1. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    Progress + blocker (2026-07-06): scoped the runner rewrite. GFQL Cypher translations written for all q1-q9 (schema uses node_type/rel as properties → inline {node_type:...}/-[{rel:...}]-> maps). q1/q8/q9 validated as correct pure single-query GFQL. q3/q4 are var-length (need validation). But q2/q5/q6/q7 are blocked: they are the multi-part MATCH...WHERE...WITH <subset> MATCH (alias)... shape, which hits #1712 — a silent wrong answer (subset WITH-carry does not restrict the second MATCH; returns unfiltered counts on pandas AND cuDF). So they cannot run as single-query GFQL until #1712 is fixed or orchestrated as multi-call GFQL.

    Root cause of why this slipped: the current runner shortcuts bypass the GFQL engine, so these shapes had no real-engine correctness tests. Added an xfail(strict)-locked correctness test for the subset-carry shape (commit locks it). Full plan + query translations recorded for resumption. Next: fix/orchestrate #1712, validate q3/q4, then the full-data 4-engine + competitor run.

  2. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    Full real-GFQL correctness map (tested each qN shape on known-answer graphs, 2026-07-06):

    queries real-GFQL status
    q1, q8, q9 ✅ correct (pure single-query GFQL, tested)
    q3, q4 ❌ honest NIE — RETURN c.city, avg(p.age) (grouped projection over TWO aliases) hits the #1273 "one MATCH source alias" boundary
    q2, q5, q6, q7 ❌ silent wrong (#1712) — multi-part subset carry / connected multi-pattern intersection

    Only 3/9 run correctly as real GFQL today. The honest full-data cross-DB benchmark is blocked on two substantial engine gaps: #1712 (q2/q5/q6/q7) and #1273 (q3/q4). Until fixed, GFQL cells for q2–q7 must be marked NIE/in-progress — not shortcut-faked. q1/q8/q9 are publishable now. Recorded the full plan + query translations for resumption.

  3. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    UNBLOCKED (2026-07-07): both engine gaps fixed. #1712 (commit f5d05674) — connected/multi-part shared-alias filter now intersects correctly (pandas+cuDF), so q2/q5/q6/q7 run as real GFQL (expressed as the comma-form with tolower). #1273 (commit 00902f4f) — multi-source grouped aggregate RETURN c.city, avg(p.age) routes to bindings, so q3/q4 run. Verified all q1/q3/q6/q8 shapes correct on pandas+cuDF on known-answer graphs. All q1-q9 now run as real GFQL (min/max multi-source still NIE but not used by q1-q9). The +9 correctness tests fill the coverage the shortcuts had bypassed. Remaining #1710 work is now just the mechanical runner rewrite + full-data cross-DB run + report refresh (no engine blockers).

  4. lmeyerov commented on Jul 7, 2026

    @lmeyerov
    ContributorAuthor

    FULL-DATA REAL-GFQL RESULTS (2026-07-07, dgx, safe_run, all values verified identical across pandas/cudf/polars/polars-gpu):

    The shortcut-faked table was hiding the truth. Real GFQL LOSES all 9 graph-benchmark OLAP queries to the dedicated graph DBs (GFQL-best vs best-DB: q1 1.4×, q2 1.2×, q3 5.9×, q4 3.2×, q5 839×, q6 234×, q7 366×, q8 37×, q9 18×). The old shortcut numbers (q1 pandas ~80ms) were 25-322× faster than real GFQL (q1 pandas 1992ms, q9 71s) — hand-rolled dataframe code, not the engine.

    This is expected and honest: graph-benchmark is heterogeneous OLAP aggregation — full-Cypher-database territory (Kuzu/Neo4j have query planners, column stores, native group-by; GFQL materializes big binding tables). GFQL is NOT the right tool for this suite, and the report must say so. GFQL best OLAP engine = polars (q8 389ms, q9 1424ms). GFQLs real win is the Pokec seeded-traversal suite (#1658 index), a different workload.

    Runner committed; full JSON in plan. Remaining follow-up gaps: q5/q6/q7 crash on polars (comma-form connected-join uses pandas .merge — should be a clean NIE); polars-gpu q9 anomaly (55s); these do not change the conclusion. #1710 core deliverable (honest real-GFQL numbers) DONE; report refresh next.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions