Repository navigation
perf(gfql): memoize node dtype reads per frame (#2029) - #2030
Merged
Merged
Conversation
lmeyerov
force-pushed
the
perf/gfql-node-dtypes-memo
branch
from
September 4, 2026 21:08
d2faba1 to
3dd6bdf
Compare
The compile-cache key materializes the lazy node dtypes on every string query, and reading them scans the values of every object column (the string-content gate). On the SNB SF0.1 node table (327k rows, 27 object columns) that was 105 ms of a 106 ms seeded lookup on pandas; the fast path itself is 2 ms. Reads are now memoized per node frame keyed on object identity (weakref) plus a length/columns fingerprint, bounded, registered with the cache registry so gfql_clear_caches() empties it. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1
lmeyerov
force-pushed
the
perf/gfql-node-dtypes-memo
branch
from
September 4, 2026 21:19
3dd6bdf to
f0a40db
Compare
This was referenced Sep 4, 2026
Contributor
Author
|
looks fine, bigger thing is locking in expected rough latency (pyg-bench?) to avoid overall latency regressions again, at least on a few common basic query shapes of varying complexity like as happened here |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #2029.
_compile_string_querykeys the compile cache on the node dtypes, which forces_LazyNodeDtypesto materialize on every string query;_read_node_dtypesthen runspd.api.types.infer_dtypeover the full values of everyobjectcolumn (the string-content gate for predicate pushdown). On the LDBC SNB SF0.1 node table as pyg-bench binds it (327,588 rows × 32 columns, 27 object), that scan was 305 ms of a 319 ms 3-call profile; the seeded typed-hop fast path it gates ran in 12 ms.Change:
_read_node_dtypesmemoizes per node frame — keyed on(id(frame), engine), valid while a weakref still resolves to the same object and the (length, columns) fingerprint matches (the resident indexes' identity contract), bounded to 32 entries, registered withcache_registrysogfql_clear_caches()empties it. No semantic change: the gate still scans on the first read of any frame.Measured on the real SF0.1 graph (pandas, index resident, warm median of 5):
MATCH (m:Message {id})-[:HAS_CREATOR]->(p:Person) RETURN p.id, p.firstName, p.lastName)MATCH (p:Person {id}) RETURN p.firstName, p.lastNamepolars is unaffected (no object-dtype inference). Tests: memo hit does not rescan, rebound/grown frame rescans, per-engine entries,
gfql_clear_caches()empties it, end-to-end string queries. Lint/type-hygiene/comment guards pass.🤖 Generated with Claude Code
https://claude.ai/code/session_01WwMmVFo44ADiRRj5cxh7i1