Repository navigation
feat(layout): replace external treemap layout dep with built-in implementation - #1059
Merged
Merged
Conversation
…rized implementation - Add graphistry/layout/gib/_squarify.py: pure Python + numpy squarified treemap layout algorithm (normalize_sizes + squarify + internal helpers) - Wire treemap.py to use the built-in implementation - Remove third-party dep from setup.py and mypy.ini - Add graphistry/tests/layout/test_treemap.py: 356 unit tests covering normalize_sizes, squarify geometry invariants, strip-loop boundary conditions, float-accumulation stress, numpy array inputs, and treemap() end-to-end integration; cross-validated against reference Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
np.sum() has ~3µs per-call overhead on small lists vs Python's sum() at ~100ns. Switching to sum() restores performance parity with the removed external dep (1.00-1.03x across n=2-500 partitions, verified on local + dgx-spark). Also removes the now-unnecessary numpy import from _squarify.py. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Replace four per-row dict .map() lookups in partitioned_layout.py with two vectorized DataFrame merges (normalize + global positioning), giving O(nodes) work in a single join instead of 4× serial Python dict scans. Also fix treemap.py to compute partition_ids once outside the comprehension (was calling reset_index() inside the loop). Add benchmarks/layout/treemap.py with pre-resident data design: - DataFrame built once before timing loop (measures only treemap() call) - CPU=pandas in memory, GPU=cuDF on device (when available) - Algorithm-only sweep (ref vs built-in) + E2E sweep (cpu vs gpu) - RESULTS.md written alongside script Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
E2E numbers with pre-resident data (pandas in-memory vs cuDF on-device): - CPU baseline: ~250µs (3× faster than pre-vectorization ~800µs) - GPU crossover: ~50k total nodes; at 500k nodes GPU is 1.30× faster - Algorithm-only: 1.00–1.02× parity with reference across n=2–500 Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
…LTS.md - partitioned_layout.py: replace pd.DataFrame(partition_offsets) with df_cons(engine)(partition_offsets) so the offsets table is cuDF-native on the GPU path, not a pandas frame merged into a cuDF DataFrame - benchmarks/layout/.gitignore: exclude RESULTS.md (machine-specific output) - git rm benchmarks/layout/RESULTS.md (not source, belongs in .gitignore) Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Rename offsets columns to ox/oy/odx/ody before merging so columns that only exist on the right side (dx, dy) don't silently drop their suffix. Previously dx/dy came through unsuffixed (no collision) while x/y got _local/_offset suffixes, causing KeyError: 'dx_offset' at runtime. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
graphistry/layout/gib/_squarify.pypartitioned_layout.py: replaced 4× per-row dict.map()lookups with a single DataFrame merge for normalize + global-positioning stepstreemap()andgroup_in_a_box_layout()are unchanged from the caller's perspectivegraphistry/tests/layout/test_treemap.py) covering geometry invariants, strip-loop boundary conditions, float-accumulation stress, numpy array inputs, and end-to-endtreemap()integrationbenchmarks/layout/treemap.pywith CPU+GPU comparison; results inbenchmarks/layout/RESULTS.mdPerformance
Benchmarked with pre-resident data (pandas in memory for CPU, cuDF on-device for GPU), 200–500 repeated measurements, median reported.
Algorithm only (normalize + layout kernel, no DataFrame):
Built-in impl matches reference within noise (1.00–1.02×).
E2E treemap() — CPU (pandas) vs GPU (cuDF), dgx-spark, Rapids 26.02:
CPU E2E baseline: ~250µs (down from ~800µs pre-vectorization). GPU crossover at ~50k total nodes.
Validation
test_gib_cudfandtest_gib_cudf_with_partitionsboth greenTest plan
python3.10 -m pytest graphistry/tests/layout/test_treemap.py graphistry/tests/layout/test_gib.py— 353 passedbenchmarks/layout/treemap.pyrun locally (CPU) and on dgx-spark (CPU+GPU)🤖 Generated with Claude Code