Skip to content

perf(gfql): memoize node dtype reads per frame (#2029) - #2030

Merged
lmeyerov merged 1 commit into
masterfrom
perf/gfql-node-dtypes-memo
Sep 5, 2026
Merged

lmeyerov merged 1 commit into
masterfrom
perf/gfql-node-dtypes-memo

Conversation

@lmeyerov

@lmeyerov lmeyerov commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

Fixes #2029.

_compile_string_query keys the compile cache on the node dtypes, which forces _LazyNodeDtypes to materialize on every string query; _read_node_dtypes then runs pd.api.types.infer_dtype over the full values of every object column (the string-content gate for predicate pushdown). On the LDBC SNB SF0.1 node table as pyg-bench binds it (327,588 rows × 32 columns, 27 object), that scan was 305 ms of a 319 ms 3-call profile; the seeded typed-hop fast path it gates ran in 12 ms.

Change: _read_node_dtypes memoizes per node frame — keyed on (id(frame), engine), valid while a weakref still resolves to the same object and the (length, columns) fingerprint matches (the resident indexes' identity contract), bounded to 32 entries, registered with cache_registry so gfql_clear_caches() empties it. No semantic change: the gate still scans on the first read of any frame.

Measured on the real SF0.1 graph (pandas, index resident, warm median of 5):

query before after
official IS5 (MATCH (m:Message {id})-[:HAS_CREATOR]->(p:Person) RETURN p.id, p.firstName, p.lastName) 105.9 ms 1.95 ms
official IS1 (two-alias, 8-column projection; no fast path) 126.3 ms 22.4 ms
MATCH (p:Person {id}) RETURN p.firstName, p.lastName 115.6 ms 13.7 ms

polars is unaffected (no object-dtype inference). Tests: memo hit does not rescan, rebound/grown frame rescans, per-engine entries, gfql_clear_caches() empties it, end-to-end string queries. Lint/type-hygiene/comment guards pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1

The compile-cache key materializes the lazy node dtypes on every string
query, and reading them scans the values of every object column (the
string-content gate). On the SNB SF0.1 node table (327k rows, 27 object
columns) that was 105 ms of a 106 ms seeded lookup on pandas; the fast
path itself is 2 ms. Reads are now memoized per node frame keyed on
object identity (weakref) plus a length/columns fingerprint, bounded,
registered with the cache registry so gfql_clear_caches() empties it.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
@lmeyerov

lmeyerov commented Sep 5, 2026

Copy link
Copy Markdown
Contributor Author

looks fine, bigger thing is locking in expected rough latency (pyg-bench?) to avoid overall latency regressions again, at least on a few common basic query shapes of varying complexity like as happened here

@lmeyerov
lmeyerov merged commit a64b566 into master Sep 5, 2026
78 checks passed
@lmeyerov
lmeyerov deleted the perf/gfql-node-dtypes-memo branch September 5, 2026 05:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf(gfql): Cypher compile scans every object column of the node table on each call (pandas point lookups 100+ ms on SNB)

1 participant