Repository navigation
benchmarks/gfql: q1-q9 runner uses dataframe shortcuts (bypasses GFQL engine) + untimed lowercase precompute — unfair vs competitor DBs #1710
Description
Activity
Progress + blocker (2026-07-06): scoped the runner rewrite. GFQL Cypher translations written for all q1-q9 (schema uses node_type/rel as properties → inline
{node_type:...}/-[{rel:...}]->maps). q1/q8/q9 validated as correct pure single-query GFQL. q3/q4 are var-length (need validation). But q2/q5/q6/q7 are blocked: they are the multi-partMATCH...WHERE...WITH <subset> MATCH (alias)...shape, which hits #1712 — a silent wrong answer (subset WITH-carry does not restrict the second MATCH; returns unfiltered counts on pandas AND cuDF). So they cannot run as single-query GFQL until #1712 is fixed or orchestrated as multi-call GFQL.Root cause of why this slipped: the current runner shortcuts bypass the GFQL engine, so these shapes had no real-engine correctness tests. Added an xfail(strict)-locked correctness test for the subset-carry shape (commit locks it). Full plan + query translations recorded for resumption. Next: fix/orchestrate #1712, validate q3/q4, then the full-data 4-engine + competitor run.
Full real-GFQL correctness map (tested each qN shape on known-answer graphs, 2026-07-06):
queries real-GFQL status q1, q8, q9 ✅ correct (pure single-query GFQL, tested) q3, q4 ❌ honest NIE — RETURN c.city, avg(p.age)(grouped projection over TWO aliases) hits the #1273 "one MATCH source alias" boundaryq2, q5, q6, q7 ❌ silent wrong (#1712) — multi-part subset carry / connected multi-pattern intersection Only 3/9 run correctly as real GFQL today. The honest full-data cross-DB benchmark is blocked on two substantial engine gaps: #1712 (q2/q5/q6/q7) and #1273 (q3/q4). Until fixed, GFQL cells for q2–q7 must be marked NIE/in-progress — not shortcut-faked. q1/q8/q9 are publishable now. Recorded the full plan + query translations for resumption.
UNBLOCKED (2026-07-07): both engine gaps fixed. #1712 (commit f5d05674) — connected/multi-part shared-alias filter now intersects correctly (pandas+cuDF), so q2/q5/q6/q7 run as real GFQL (expressed as the comma-form with tolower). #1273 (commit 00902f4f) — multi-source grouped aggregate
RETURN c.city, avg(p.age)routes to bindings, so q3/q4 run. Verified all q1/q3/q6/q8 shapes correct on pandas+cuDF on known-answer graphs. All q1-q9 now run as real GFQL (min/max multi-source still NIE but not used by q1-q9). The +9 correctness tests fill the coverage the shortcuts had bypassed. Remaining #1710 work is now just the mechanical runner rewrite + full-data cross-DB run + report refresh (no engine blockers).FULL-DATA REAL-GFQL RESULTS (2026-07-07, dgx, safe_run, all values verified identical across pandas/cudf/polars/polars-gpu):
The shortcut-faked table was hiding the truth. Real GFQL LOSES all 9 graph-benchmark OLAP queries to the dedicated graph DBs (GFQL-best vs best-DB: q1 1.4×, q2 1.2×, q3 5.9×, q4 3.2×, q5 839×, q6 234×, q7 366×, q8 37×, q9 18×). The old shortcut numbers (q1 pandas ~80ms) were 25-322× faster than real GFQL (q1 pandas 1992ms, q9 71s) — hand-rolled dataframe code, not the engine.
This is expected and honest: graph-benchmark is heterogeneous OLAP aggregation — full-Cypher-database territory (Kuzu/Neo4j have query planners, column stores, native group-by; GFQL materializes big binding tables). GFQL is NOT the right tool for this suite, and the report must say so. GFQL best OLAP engine = polars (q8 389ms, q9 1424ms). GFQLs real win is the Pokec seeded-traversal suite (#1658 index), a different workload.
Runner committed; full JSON in plan. Remaining follow-up gaps: q5/q6/q7 crash on polars (comma-form connected-join uses pandas .merge — should be a clean NIE); polars-gpu q9 anomaly (55s); these do not change the conclusion. #1710 core deliverable (honest real-GFQL numbers) DONE; report refresh next.
- added 7 commits that reference this issue
on Jul 9, 2026
Summary (benchmark integrity — must fix before publishing "official" numbers)
benchmarks/gfql/graph_benchmark_q1_q9.py(GFQL vs Kuzu/Ladybug/Neo4j/Memgraph on the prrao87 graph-benchmark suite) gives GFQL two unfair advantages vs the competitor databases, which all run their own real Cypher:1. Dataframe shortcuts bypass the GFQL engine entirely
For q1, q3, q4, q5, q6, q7 (
query_variant="standard", the default —_uses_dataframe_shortcut, line 238-243), GFQL does not run its query engine. It runs hand-written pandas/cuDFgroupby/merge code (_query{1,3,4,5,6,7}_dataframe_shortcut, plus_query5_polars_shortcut,_query6_cudf_shortcut, etc.). That measures hand-tuned dataframe code, not GFQL — while Kuzu/Neo4j/Memgraph parse+plan+execute real Cypher. Not apples-to-apples.g.gfql(<cypher>, engine=...)) for all queries. On pandas/cuDF this works for q1-q9 once thecount(<other-alias>)lowering routing is fixed (see the lowering issue; q1 also works today via thecount(*)form). On polars, traversal queries NIE (see the polars binding_ops issue) and count(*) is currently wrong (see the P0 count-broadcast issue).2. Untimed lowercase precompute (both GFQL and the Memgraph runner)
graph_benchmark_q1_q9.pyprecomputesgender_lc/interest_lcat load (untimed) for the toLower() filters in q5/q6/q7. Measured impact when moved inside the timed region: <1ms (negligible) — but for symmetry it should be timed or removed.graph_benchmark_memgraph_q1_q9.pysimilarly precomputes lowercased columns. The competitors (Kuzu/Ladybug/Neo4j) run rawtolower()inline. For official numbers, no side gets untimed precompute — everyone runs the raw query.Data-scale caveat (already caught)
/tmp/graph-benchmark-gfql-memgraph= TINY (1k persons/10k edges). Full data (100k/2.42M) =/tmp/graph-benchmark-gfql-memgraph-full(repodata/symlinks to it; the Docker mount does NOT follow the symlink — mount the target dir directly). An earlier Memgraph q1-q9 run on the tiny set produced invalid ~100× fast numbers.Acceptance
neo4j/query.pyetc.).