Skip to content

[improve][test] Add host and per-container CPU and perf counter sampling, default bookie cache sizes and bottleneck analysis guidance to the performance tests - #26749

Merged
merlimat merged 7 commits into
masterfrom
lh-improve-perf-bottleneck-tooling
Sep 30, 2026
Merged

merlimat merged 7 commits into
masterfrom
lh-improve-perf-bottleneck-tooling

Conversation

@lhotari

@lhotari lhotari commented Sep 29, 2026 •

Copy link
Copy Markdown
Member

Motivation

A first round of bottleneck hunting with the performance tests (the iot-telemetry-max-rate scenario with 128-byte, 8 KB and 128 KB unbatched entries) ran into three gaps in the tooling:

  • Once no thread is saturated, the limit is often the host itself: the broker, the bookies and the 500-producer / 20-consumer workloads share its CPUs and its disk. Nothing in a run's output showed the host's CPU utilization or disk throughput, nor how each container used the CPUs: its CPU time, context switches, CPU migrations, clock rate and instructions per cycle.
  • Large entries were limited by the bookies throttling adds (bookie_throttled_write) while their DbLedgerStorage write cache flushed. The cause was the test image's run-bookie.sh, which sets dbStorage_writeCacheMaxSizeMb and dbStorage_readAheadCacheMaxSizeMb to 16 MB unless they are set. With BookKeeper's default of a quarter of the direct memory (256 MB for the high-memory configuration's bookies), the 128 KB scenario reached 3.5–3.8k msg/s instead of 2.1–2.6k (+45–66 %, single screening runs), and the 8 KB scenario 48–49k msg/s instead of 28–34k.
  • The agent guide describes validating one change, but not how to hunt for the limiting stage quickly across iterations, and the flame graphs merge all threads, which hides a serial stage such as a topic's managed-ledger thread.

Modifications

  • host-io.csv: the launcher's new HostIoSampler writes, once per second, the busy and I/O-wait share of all CPUs from /proc/stat, and each physical disk's read and write MB/s and busy share from /proc/diskstats (block devices under /sys/block with a device link). They describe the Docker engine's host: when the engine runs in a Linux VM, such as Docker Desktop's or OrbStack's on macOS, the launcher reads these files in the VM through the privileged perf stat sidecar below, and it tells the two apart by comparing the kernels' boot IDs. host-stats.csv's temperatures, frequencies and throttle counters stay the launcher host's own. It is documented next to host-stats.csv in docs/run-reports.md and in the agent guide's list of run inputs.
  • container-stats.csv, perf-stat.csv and container-summary.json, and a Containers section in the run report. Once per second: each container's CPUs used (its cgroup's cpu.stat) and its threads' voluntary and involuntary context switches (/proc/<tid>/status); and exact perf stat -a --for-each-cgroup counts per container: task-clock, context switches, CPU migrations, page faults, cycles, instructions, last-level cache references and misses, L1 data cache load misses and branch misses. The section starts with the host's CPU busy and I/O-wait share and disk throughput, then shows each container and all of them: CPUs, CPU seconds per million messages, voluntary and involuntary switches, migrations and page faults per second, and in a second table the clock rate, instructions per cycle (IPC), last-level cache, L1 data cache and branch misses per thousand instructions (MPKI) and the last-level cache miss rate. container-summary.json holds the same numbers for scripts and AI agents, and the agent guide points to it.
    • A privileged sidecar container, built from Alpine's perf package, runs in the Docker engine host's PID and cgroup namespaces and finds each container's cgroup from the PID of Docker's container inspect, so it works on a Linux host as well as in the Linux VM of Docker Desktop or OrbStack, on x86-64 and arm64. The events are perf's generic ones, which the kernel maps to the CPU's own and which count without multiplexing on common x86 and Arm cores; a count that the CPU or VM doesn't provide is left empty. On a Linux host the launcher reads the container stats from its own /proc without privileges; elsewhere the sidecar serves them, also when perf can't count. --no-perf-stat (-Pperformance.perfStat=false) turns the sidecar off.
    • async-profiler and jonoffcpu don't provide these as counts: async-profiler samples one event with stacks at a time, and jonoffcpu records sampled off-CPU intervals.
  • Bookie caches: the cluster memory configurations set both DbLedgerStorage caches to a quarter of the bookies' direct memory (256 MB high-memory, 128 MB medium and low), BookKeeper's own default.
  • Agent guide: a new Eliminating bottlenecks section describes the iterative mode: find the limiting stage (a busy serial thread, blocked time, run-queue time, a saturated resource outside the broker), compare cost per unit of work, screen one change at a time, check large unbatched entry sizes, add metrics to check assumptions, and record where the limit moved. It also records lessons: recognizing a host-CPU limit with host-io.csv, profiling the client workloads, comparing broker, managed-ledger and bookie latencies to locate a queue, checking the test setup's defaults, and diagnosing bookies, which the launcher doesn't profile.
  • Per-thread CPU: docs/analyzing-profiles.md shows how to split a measurement recording's CPU samples by thread with the converter's --threads option, and how to rank threads and thread pools and break down the busiest thread with DuckDB's quack_flamegraph.

Verifying this change

  • Make sure that the change passes the CI checks.

This change added tests and can be verified as follows:

  • HostIoSamplerTest: disk discovery, the rates computed from two readings, and reading /proc counters into host-io.csv.
  • ContainerStatsSamplerTest: finding a container's cgroup, and CPU and context-switch rates with threads that start and exit between readings.
  • PerfStatSidecarTest: converting perf stat interval output into rows per container, including counters perf reports as not supported, parsing the sidecar's container counters, and splitting the engine host's /proc/stat and /proc/diskstats output.
  • ContainerStatsReportTest: the host line, both tables with the all-containers row, the derived clock rate, IPC and MPKI, and container-summary.json.
  • Two iot-telemetry-max-rate runs on a Linux x86-64 host, one reading the container stats from the host's /proc and one through the sidecar (as with a VM-based engine), wrote all files and the Containers section with matching numbers. A repeat of the pair after moving the host I/O sampling to the sidecar again matched: 84.8 % and 85.6 % host CPU busy, 67 and 63 MB/s disk writes, the broker at 5.8 CPUs and IPC 0.73. In the sidecar run, --procfs pointed at an empty directory, so host-io.csv could only come from the sidecar; for example, the broker used 6.0–6.1 CPUs (45 of the 97 CPU seconds per million messages of all containers) at an IPC of 0.69–0.70 with 28 L1 data cache misses per thousand instructions, 16–17k CPU migrations and 34k page faults per second.
  • ./gradlew quickCheck and the performance modules' tests (common, tools, launcher, report-tool, metrics: 200 tests) pass locally. The sampler and the cache settings were used in the runs described above.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

The changes are limited to tests/performance.

This PR was prepared with AI assistance (Claude Code) and reviewed by a human contributor.

…thread CPU analysis

Add an "Eliminating bottlenecks" mode to the performance tests' agent guide: find the limiting stage (a busy serial
thread, blocked time, run-queue time, a saturated resource outside the broker), compare cost per unit of work, screen
one change at a time quickly, check large unbatched entry sizes, add metrics to check assumptions, and record where
the limit moved after each change.

Document per-thread CPU analysis: split the measurement recording's CPU samples by thread with the converter's
--threads option and rank threads, thread pools and the busiest thread's frames with DuckDB's quack_flamegraph.

Assisted-by: Claude Code (claude-opus-5-5)
… during performance runs

Write host-io.csv next to host-stats.csv: once per second, the busy and I/O-wait share of all CPUs from /proc/stat,
and each physical disk's read and write MB/s and busy share from /proc/diskstats. All the containers of a run share
the host, so this shows whether a run is limited by the host's CPUs or by its storage, which the bookies share and
which their own metrics can't show.

Assisted-by: Claude Code (claude-opus-5-5)
…default cache sizes

The test image's run-bookie.sh sets dbStorage_writeCacheMaxSizeMb and dbStorage_readAheadCacheMaxSizeMb to 16 MB
unless they are set. With a 16 MB write cache, a bookie taking large entries flushes it about 17 times per second
and throttles adds while it does (bookie_throttled_write): at 128 KB unbatched entries the max-rate scenario reached
2.1k msg/s, and 3.5k msg/s with 256 MB. Set both caches to a quarter of the bookies' direct memory, BookKeeper's
default, in the cluster memory configurations.

Assisted-by: Claude Code (claude-opus-5-5)
…erations in the agent guide

Assisted-by: Claude Code (claude-opus-5-5)
… perf counters in performance runs

Once no thread is saturated, the limit of a run is usually the host's CPUs, shared by the broker, the bookies and the
workload clients, and a run didn't show how each container used them. Add two per-second samplers and a Containers
section to the run report:

- container-stats.csv, sampled from the host without privileges: each container's CPUs used, from its cgroup's
  cpu.stat, and its threads' voluntary and involuntary context switches, from /proc/<tid>/status. The container's
  cgroup is found from its main process's /proc/<pid>/cgroup entry, with the PID from Docker's container inspect.
- perf-stat.csv, from `perf stat -a --for-each-cgroup` in a privileged sidecar container built from Alpine's perf
  package: task-clock, context switches, CPU migrations, cycles and instructions per container. The counts are exact
  and cheap, so they suit unprofiled runs. --no-perf-stat (-Pperformance.perfStat=false) turns it off; counters the
  host doesn't provide to containers are left empty.

The Containers section shows, for the measurement, each container's CPUs, CPU seconds per million messages,
voluntary and involuntary switches per second, CPU migrations per second, the clock rate and the instructions per
cycle (IPC). async-profiler and jonoffcpu don't provide these as counts: async-profiler samples one event with
stacks at a time, and jonoffcpu records sampled off-CPU intervals.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari lhotari changed the title [improve][test] Add host CPU and disk sampling, BookKeeper default bookie caches and bottleneck analysis guidance to the performance tests [improve][test] Add host and per-container CPU and perf counter sampling, default bookie cache sizes and bottleneck analysis guidance to the performance tests Sep 29, 2026
…aries, JSON for agents, and VM-based Docker engines

- perf-stat.csv also counts page faults, last-level cache references and misses, L1 data cache load misses and
  branch misses: perf's generic events, which count without multiplexing on common x86 and Arm cores. The report
  derives the misses per thousand instructions (MPKI) and the last-level cache miss rate from them.
- The Containers section starts with the host's CPU busy and I/O-wait share and disk throughput from host-io.csv,
  splits the containers into a CPU use table and a CPU efficiency table, and adds a row for all containers.
- container-summary.json holds the section's numbers for scripts and AI agents; the agent guide points to it.
- The perf sidecar runs in the Docker engine host's PID and cgroup namespaces and finds each container's cgroup from
  the PID of Docker's container inspect, so it works in the Linux VM of Docker Desktop or OrbStack, on x86-64 and
  arm64. When the launcher's host doesn't see the containers' cgroups, the container stats come through the sidecar;
  when perf can't count, the sidecar still serves them.

Assisted-by: Claude Code (claude-opus-5-5)
…he perf stat sidecar

host-io.csv now describes the Docker engine's host. When the engine runs in a Linux VM,
such as Docker Desktop's or OrbStack's on macOS, the launcher reads /proc/stat,
/proc/diskstats and /sys/block in the VM through the privileged sidecar, which now
starts before the cluster. It tells the two cases apart by comparing the kernels'
boot IDs. host-stats.csv (temperatures, frequencies, throttling) stays the launcher
host's own.

Assisted-by: Claude Code (claude-opus-5-5)
@merlimat
merlimat merged commit 1f62aec into master Sep 30, 2026
44 checks passed
@merlimat
merlimat deleted the lh-improve-perf-bottleneck-tooling branch September 30, 2026 14:39
@lhotari lhotari added this to the 5.0.0 milestone Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants