Repository navigation
[improve][test] Add host and per-container CPU and perf counter sampling, default bookie cache sizes and bottleneck analysis guidance to the performance tests - #26749
Merged
Conversation
…thread CPU analysis Add an "Eliminating bottlenecks" mode to the performance tests' agent guide: find the limiting stage (a busy serial thread, blocked time, run-queue time, a saturated resource outside the broker), compare cost per unit of work, screen one change at a time quickly, check large unbatched entry sizes, add metrics to check assumptions, and record where the limit moved after each change. Document per-thread CPU analysis: split the measurement recording's CPU samples by thread with the converter's --threads option and rank threads, thread pools and the busiest thread's frames with DuckDB's quack_flamegraph. Assisted-by: Claude Code (claude-opus-5-5)
… during performance runs Write host-io.csv next to host-stats.csv: once per second, the busy and I/O-wait share of all CPUs from /proc/stat, and each physical disk's read and write MB/s and busy share from /proc/diskstats. All the containers of a run share the host, so this shows whether a run is limited by the host's CPUs or by its storage, which the bookies share and which their own metrics can't show. Assisted-by: Claude Code (claude-opus-5-5)
…default cache sizes The test image's run-bookie.sh sets dbStorage_writeCacheMaxSizeMb and dbStorage_readAheadCacheMaxSizeMb to 16 MB unless they are set. With a 16 MB write cache, a bookie taking large entries flushes it about 17 times per second and throttles adds while it does (bookie_throttled_write): at 128 KB unbatched entries the max-rate scenario reached 2.1k msg/s, and 3.5k msg/s with 256 MB. Set both caches to a quarter of the bookies' direct memory, BookKeeper's default, in the cluster memory configurations. Assisted-by: Claude Code (claude-opus-5-5)
…erations in the agent guide Assisted-by: Claude Code (claude-opus-5-5)
11 tasks
… perf counters in performance runs Once no thread is saturated, the limit of a run is usually the host's CPUs, shared by the broker, the bookies and the workload clients, and a run didn't show how each container used them. Add two per-second samplers and a Containers section to the run report: - container-stats.csv, sampled from the host without privileges: each container's CPUs used, from its cgroup's cpu.stat, and its threads' voluntary and involuntary context switches, from /proc/<tid>/status. The container's cgroup is found from its main process's /proc/<pid>/cgroup entry, with the PID from Docker's container inspect. - perf-stat.csv, from `perf stat -a --for-each-cgroup` in a privileged sidecar container built from Alpine's perf package: task-clock, context switches, CPU migrations, cycles and instructions per container. The counts are exact and cheap, so they suit unprofiled runs. --no-perf-stat (-Pperformance.perfStat=false) turns it off; counters the host doesn't provide to containers are left empty. The Containers section shows, for the measurement, each container's CPUs, CPU seconds per million messages, voluntary and involuntary switches per second, CPU migrations per second, the clock rate and the instructions per cycle (IPC). async-profiler and jonoffcpu don't provide these as counts: async-profiler samples one event with stacks at a time, and jonoffcpu records sampled off-CPU intervals. Assisted-by: Claude Code (claude-opus-5-5)
…aries, JSON for agents, and VM-based Docker engines - perf-stat.csv also counts page faults, last-level cache references and misses, L1 data cache load misses and branch misses: perf's generic events, which count without multiplexing on common x86 and Arm cores. The report derives the misses per thousand instructions (MPKI) and the last-level cache miss rate from them. - The Containers section starts with the host's CPU busy and I/O-wait share and disk throughput from host-io.csv, splits the containers into a CPU use table and a CPU efficiency table, and adds a row for all containers. - container-summary.json holds the section's numbers for scripts and AI agents; the agent guide points to it. - The perf sidecar runs in the Docker engine host's PID and cgroup namespaces and finds each container's cgroup from the PID of Docker's container inspect, so it works in the Linux VM of Docker Desktop or OrbStack, on x86-64 and arm64. When the launcher's host doesn't see the containers' cgroups, the container stats come through the sidecar; when perf can't count, the sidecar still serves them. Assisted-by: Claude Code (claude-opus-5-5)
…he perf stat sidecar host-io.csv now describes the Docker engine's host. When the engine runs in a Linux VM, such as Docker Desktop's or OrbStack's on macOS, the launcher reads /proc/stat, /proc/diskstats and /sys/block in the VM through the privileged sidecar, which now starts before the cluster. It tells the two cases apart by comparing the kernels' boot IDs. host-stats.csv (temperatures, frequencies, throttling) stays the launcher host's own. Assisted-by: Claude Code (claude-opus-5-5)
merlimat
approved these changes
Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
A first round of bottleneck hunting with the performance tests (the
iot-telemetry-max-ratescenario with 128-byte, 8 KB and 128 KB unbatched entries) ran into three gaps in the tooling:bookie_throttled_write) while their DbLedgerStorage write cache flushed. The cause was the test image'srun-bookie.sh, which setsdbStorage_writeCacheMaxSizeMbanddbStorage_readAheadCacheMaxSizeMbto 16 MB unless they are set. With BookKeeper's default of a quarter of the direct memory (256 MB for the high-memory configuration's bookies), the 128 KB scenario reached 3.5–3.8k msg/s instead of 2.1–2.6k (+45–66 %, single screening runs), and the 8 KB scenario 48–49k msg/s instead of 28–34k.Modifications
host-io.csv: the launcher's newHostIoSamplerwrites, once per second, the busy and I/O-wait share of all CPUs from/proc/stat, and each physical disk's read and write MB/s and busy share from/proc/diskstats(block devices under/sys/blockwith adevicelink). They describe the Docker engine's host: when the engine runs in a Linux VM, such as Docker Desktop's or OrbStack's on macOS, the launcher reads these files in the VM through the privileged perf stat sidecar below, and it tells the two apart by comparing the kernels' boot IDs.host-stats.csv's temperatures, frequencies and throttle counters stay the launcher host's own. It is documented next tohost-stats.csvindocs/run-reports.mdand in the agent guide's list of run inputs.container-stats.csv,perf-stat.csvandcontainer-summary.json, and a Containers section in the run report. Once per second: each container's CPUs used (its cgroup'scpu.stat) and its threads' voluntary and involuntary context switches (/proc/<tid>/status); and exactperf stat -a --for-each-cgroupcounts per container: task-clock, context switches, CPU migrations, page faults, cycles, instructions, last-level cache references and misses, L1 data cache load misses and branch misses. The section starts with the host's CPU busy and I/O-wait share and disk throughput, then shows each container and all of them: CPUs, CPU seconds per million messages, voluntary and involuntary switches, migrations and page faults per second, and in a second table the clock rate, instructions per cycle (IPC), last-level cache, L1 data cache and branch misses per thousand instructions (MPKI) and the last-level cache miss rate.container-summary.jsonholds the same numbers for scripts and AI agents, and the agent guide points to it.perfpackage, runs in the Docker engine host's PID and cgroup namespaces and finds each container's cgroup from the PID of Docker's container inspect, so it works on a Linux host as well as in the Linux VM of Docker Desktop or OrbStack, on x86-64 and arm64. The events are perf's generic ones, which the kernel maps to the CPU's own and which count without multiplexing on common x86 and Arm cores; a count that the CPU or VM doesn't provide is left empty. On a Linux host the launcher reads the container stats from its own/procwithout privileges; elsewhere the sidecar serves them, also when perf can't count.--no-perf-stat(-Pperformance.perfStat=false) turns the sidecar off.host-io.csv, profiling the client workloads, comparing broker, managed-ledger and bookie latencies to locate a queue, checking the test setup's defaults, and diagnosing bookies, which the launcher doesn't profile.docs/analyzing-profiles.mdshows how to split a measurement recording's CPU samples by thread with the converter's--threadsoption, and how to rank threads and thread pools and break down the busiest thread with DuckDB'squack_flamegraph.Verifying this change
This change added tests and can be verified as follows:
HostIoSamplerTest: disk discovery, the rates computed from two readings, and reading/proccounters intohost-io.csv.ContainerStatsSamplerTest: finding a container's cgroup, and CPU and context-switch rates with threads that start and exit between readings.PerfStatSidecarTest: convertingperf statinterval output into rows per container, including counters perf reports as not supported, parsing the sidecar's container counters, and splitting the engine host's/proc/statand/proc/diskstatsoutput.ContainerStatsReportTest: the host line, both tables with the all-containers row, the derived clock rate, IPC and MPKI, andcontainer-summary.json.iot-telemetry-max-rateruns on a Linux x86-64 host, one reading the container stats from the host's/procand one through the sidecar (as with a VM-based engine), wrote all files and the Containers section with matching numbers. A repeat of the pair after moving the host I/O sampling to the sidecar again matched: 84.8 % and 85.6 % host CPU busy, 67 and 63 MB/s disk writes, the broker at 5.8 CPUs and IPC 0.73. In the sidecar run,--procfspointed at an empty directory, sohost-io.csvcould only come from the sidecar; for example, the broker used 6.0–6.1 CPUs (45 of the 97 CPU seconds per million messages of all containers) at an IPC of 0.69–0.70 with 28 L1 data cache misses per thousand instructions, 16–17k CPU migrations and 34k page faults per second../gradlew quickCheckand the performance modules' tests (common,tools,launcher,report-tool,metrics: 200 tests) pass locally. The sampler and the cache settings were used in the runs described above.Does this pull request potentially affect one of the following parts:
If the box was checked, please highlight the changes
The changes are limited to
tests/performance.This PR was prepared with AI assistance (Claude Code) and reviewed by a human contributor.