Repository navigation
Use huge pages for large buffers on Linux - #3937
Merged
Merged
Conversation
Buffers are allocated with plain new T[], which never calls madvise(MADV_HUGEPAGE). With transparent_hugepage=madvise, the default on many systems and common on HPC clusters where users cannot change it, Scipp thus ran on 4 kByte pages while NumPy in the same process used huge pages, since it madvises arrays of 4 MByte and more. libtbbmalloc_proxy and TBB_MALLOC_USE_HUGE_PAGES do not help: tbbmalloc never calls madvise and considers huge pages unavailable unless they are reserved explicitly or the kernel setting is `always`. Request huge pages for buffers of 4 MByte and more, matching NumPy's threshold. Smaller ranges are left alone since the kernel rounds up to the 2 MByte huge page size. SCIPP_MADVISE_HUGEPAGE=0 disables the request. See https://github.com/orgs/scipp/discussions/3936.
MridulS
approved these changes
Jul 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
It was always been possible to use Scipp with hugepages (see https://scipp.github.io/user-guide/tips-tricks-and-anti-patterns.html#Using-HugePages), but it required user intervention and adequate system settings, which required root permissions.
Scipp allocates its buffers with plain
new T[], which never callsmadvise(MADV_HUGEPAGE). On systems withtransparent_hugepage=madvise— the default on many distributions, and the case users hit on HPC clusters where they cannot change the setting — Scipp therefore runs on 4 kByte pages, while NumPy in the same process uses huge pages, since it madvises every array of 4 MByte and more. Settingtransparent_hugepage=alwayswas the only documented workaround, and it needs root.Neither
libtbbmalloc_proxy.sonorTBB_MALLOC_USE_HUGE_PAGEShelp: tbbmalloc has nomadvisecall at all and treats huge pages as unavailable unless they are reserved explicitly or the kernel setting isalways. The claim in our docs that themadvisesetting "is not sufficient for Scipp" was accurate, but the reason was on our side, and it costs up to 1.7x on memory-bound operations.Buffers of 4 MByte and more now request huge pages, matching NumPy's threshold. Below that the kernel would round up to the 2 MByte huge page size and waste memory.
SCIPP_MADVISE_HUGEPAGE=0disables it, which may be needed on systems suffering from memory fragmentation: with the commondefrag=madvisesetting the kernel compacts memory synchronously when a page of an advised range is first touched, which can stall the allocating thread.Reported in https://github.com/orgs/scipp/discussions/3936.
Benchmarks
Operations on 1 GByte buffers, best of 3, on a system with
transparent_hugepage=madvise,defrag=madvise, glibc 2.35, kernel 6.8:a * b + ahistof 1e8 events into 1e6 binsBoth columns are the same build, the first one run with
SCIPP_MADVISE_HUGEPAGE=0, i.e., with the behavior before this change.The script used for this, and the raw output including a comparison with
libtbbmalloc_proxy.soandGLIBC_TUNABLES=glibc.malloc.hugetlb=1, are in https://gist.github.com/SimonHeybrock/132292e373059e68f6ac970524caeb23.Test plan
The added unit test allocates a large buffer and checks
VmFlagsin/proc/self/smapsfor thehgflag, i.e., that the kernel recorded the advice. This works independently of whether the system can actually provide huge pages, so it is meaningful in containers, and it skips on kernels built without huge page support.🤖 Generated with Claude Code