Skip to content

[feat][test] Add a Docker-based performance testing framework with profiling, metrics and AI agent support - #26714

Merged
lhotari merged 120 commits into
masterfrom
lh-use-jonoffcpu-profiler
Sep 28, 2026
Merged

lhotari merged 120 commits into
masterfrom
lh-use-jonoffcpu-profiler

Conversation

@lhotari

@lhotari lhotari commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Motivation

This PR adds a complete Docker-based performance testing framework to Pulsar, with agent guides that let AI agents run AI-assisted investigations of performance bottlenecks and optimize them. Finding and fixing a bottleneck is a loop of many steps: run a realistic scenario, read its reports, profiles and metrics, form a hypothesis, change the code or a setting, and compare the change with its baseline over repeated runs, since a single run's result isn't reliable. Without an agent, most of that loop is manual work that few contributors have time for. tests/performance/AGENTS.md, which the repository's AGENTS.md routes to, tells an agent how to check and configure the host, run and profile scenarios, read the reports, profiles and metrics, compare revisions, and report a result with its evidence, so that a contributor can hand an agent an experiment and review what it found.

Pulsar had no practical way to measure a change's effect on a realistic workload during development, with a feedback loop quick enough to iterate on. Today's AI agents need such a quick feedback loop to make progress: a run takes minutes on a developer's machine, without a separate benchmark environment, so an agent can try a change, measure it and move on to the next one. The framework runs scenarios that model real-world use cases, currently IoT telemetry, on a Pulsar cluster in Docker containers, with producers and consumers that check every message's delivery, ordering and duplicates. Every run writes a report with its settings, provenance, throughput and backlog, and latency with charts rendered from the HDR histograms of the publish and end-to-end latencies.

Profiling uses jonoffcpu, a profiler for both on-CPU and off-CPU time. For on-CPU time, its bundled async-profiler samples the CPU, allocations and locks into a JDK Flight Recorder (JFR) recording, which also collects the JVM's own events, such as garbage collections. For off-CPU time, an eBPF program in the Linux kernel records when threads stopped running, and jonoffcpu correlates each interval with the Java stack that async-profiler recorded. A broker's throughput limits are often waits: a thread blocked on a monitor, a lock or a queue hand-off. Those waits don't show up in CPU samples, and async-profiler's wall-clock mode samples idle threads as well, which buries the few waits that matter. jonoffcpu ranks the blocked time by the method that waited, so contention such as a dispatcher monitor or an executor queue lock stands out directly.

The metrics stack comes out of the box: VictoriaMetrics and Grafana, configured with Pulsar's dashboards, run with Docker Compose. During each run they collect the metrics of the brokers, bookies and ZooKeeper, and keep them across runs for later analysis, next to each run's JFR recordings and off-CPU profiles. The run's timeline is added to Grafana as annotations: the gateways' start, the end of the warmup, and the gateways' and the applications' finish, so that the dashboards show where the measurement period is. The run report includes images of Grafana panels over the run, such as the publish and delivery rates, the backlog, the JVM's heap and GC, and the bookies' add latency, each linked to Grafana.

A run's results also depend on the host's state: a CPU that runs at turbo frequencies slows down as it heats up and throttles, and daemons that tune power management, such as thermald, change the settings during a run. The reports show the host's temperature and throttling, and scripts configure the host for performance testing and restore it afterwards.

Together, the run reports, the JFR recordings, the off-CPU profiles and the metrics give an agent the evidence it needs to find a bottleneck, and to show that a change improves Pulsar.

Modifications

Profiling:

  • Profile scenarios through the jonoffcpu 0.8.0 agent (io.github.jonoffcpu:jonoffcpu-agent, Apache-2.0): its bundled async-profiler records CPU and allocation samples into the JFR recording, and its capture records the off-CPU intervals under the sampling policy of each profiled component's offCpuOptions in the scenario's profiling section. After the run the launcher correlates each capture in-process with jonoffcpu-correlator.
  • Render four off-CPU flame graphs per recording, where a box's width is off-CPU time: all off-CPU time, the blocked time without the idle waits (Netty event loops in epollWait, executor workers waiting for a task, JDK and HotSpot service threads), and both again with each stack starting at its first Pulsar or BookKeeper frame (^org\.apache\.(pulsar|bookkeeper)\.), so the same code reached from different thread pools joins into one tree; stacks without such a frame are left out of those two and counted apart. The idle-wait patterns (offcpu-idle-waits.txt) and the BookKeeper and Pulsar frames that only dispatch work (offcpu-dispatch-hide.txt, hidden before rooting) are copied into each output directory.
  • The correlator also writes an analysis digest per recording, jonoffcpu-summary.md, which ranks the blocked time by the application method that waited, by where threads entered the application and by application method, with the idle waits left out.
  • Render the standard async-profiler views (cpu, wall, alloc, lock) of each measurement recording with jonoffcpu-jfr-converter, each only when its event is in the recording's async-profiler options, as a flame graph, split by thread and as a heatmap over time.
  • Add dockerBuildWolfi tasks that build the Pulsar image and the java-test-image from Wolfi under a <tag>-wolfi tag. The profiler's native libraries need glibc, and on Alpine every native frame is the unsymbolized musl loader. The launcher's profile task uses the Wolfi image by default; the Alpine images are unchanged.
  • The integration-test containers can run the component JVM as another user (PULSAR_PROCESS_USER, applied to the supervisord configuration), so that a profiled JVM runs as root and may load the eBPF programs. The existing kernel-tuning task also clears kernel.unprivileged_bpf_disabled.
  • PulsarClusterSpec and PulsarContainer accept the jonoffcpu agent JAR and sampling options.

Reports, in a new module :tests:performance:report-tool that the launcher calls once a run has finished:

  • The run report, README.md in the run directory with its HTML page index.html, for every run, profiled or not: the run's settings with links to the scenario file and its resolved configuration; where, by whom and from which commit the run was made (host, user, project directory, git branch and commit, Pulsar version, also in run-info.json); links to each profile's reports (first the jonoffcpu report (off-CPU summary), which is the off-CPU digest, then the profile report and the flame graphs) and the JFR recordings; correctness; producer and delivered throughput; latency percentiles for publish and for each consumer application; and the backlog and publish and dispatch rates, sampled once per second from the broker's topic stats.
  • The host's CPU, sampled once per second from Linux's sysfs files into host-stats.csv: the package and hottest core temperature, the mean and lowest core frequency, the kernel's thermal throttle counters and the fastest fan. The run report sums it up in one line of the settings table and has a Host section with the values at the start and during the measurement, a bold note when the CPU throttled during the measurement, and temperature and frequency charts, so that a run slowed down by a hot host can be told apart.
  • Charts: latency by percentile and the maximum latency per logged interval, as HistogramLogAnalyzer plots them, drawn with XChart as PNG, with publish and each application as separate lines; throughput and backlog over time. Each chart names the branch, commit and run time in a small footer.
  • A profile report per profiled component, README.md and index.html in its directory: its files (digest, JFR recordings, capture stream, patterns), the off-CPU flame graphs with their totals, and the views. It names the profiled process, the broker or a Pulsar client, says what jonoffcpu's off-CPU flame graphs and async-profiler's flame graphs each show, lists the blocked time first, and keeps the details in a collapsed section.
  • Each Markdown report, and the digest, is also rendered to a self-contained HTML page with commonmark-java, with links rewritten between the pages and descriptive link texts; the digest's page abbreviates Java package names as the flame graphs do.
  • Each latency log also gets its percentile distribution as a .hgrm file, HdrHistogram's percentile output format, which HdrHistogram's plotFiles.html reads.

Performance Launcher's IoT scenario reporting:

image


jonoff summary report for application's blocking callsites:

image


Off-cpu flamegraph example for application's blocking callsites:

image


Run output layout:

  • Every run gets its own directory, <reports root>/<yyyy-MM-dd>/<branch>/<name>/<MM-dd-HH-mm-ss>/. The reports root is build/performance unless -Pperformance.reportsDir sets it (also in ~/.gradle/gradle.properties); the name is the scenario file name, the scenario's output.name, or --name; --output still writes to an exact directory. The launcher scenarios drop their fixed output.directory.
  • README.md and index.html are the run report, so that a directory of runs can be served over HTTP or pushed to a GitHub repository. ./gradlew :tests:performance:report-tool:serveReports, or tests/performance/serve-reports.py without Gradle, serves a reports root on the loopback interface by default (for an SSH tunnel from another machine), showing the YAML, CSV, HDR logs, logs and Markdown as text. The launcher prints each report's URL on that server, with a configurable base URL. tests/performance/docs/run-reports.md shows an SSH tunnel, also as a ~/.ssh/config host, that forwards the reports server, Grafana and VictoriaMetrics from another machine.
  • Each consumer application's outputs are in a directory named after its subscription (iot-application-0/ …), as the report names the application; container logs are saved as container.log.txt.

Workload:

  • The producer and consumer applications log their latency every second (HdrLatencyRecorder), so that the HDR logs show latency over the run; merged, the intervals give the same distribution as before.
  • --cooldown-temperature <°C> (or -Pperformance.cooldownTemperature, also in ~/.gradle/gradle.properties) makes the launcher wait for the CPU package to cool down to that temperature before starting the cluster, and again after the warmup rounds. For the second wait the producer serves two endpoints with the JDK's built-in HTTP server, which the launcher reaches through the port Testcontainers maps on the host: GET /measurement/ready?waitMillis=<ms> answers as soon as every warmup round has been received, and POST /measurement/start starts the measurement. --cooldown-timeout bounds each wait, the workloads' timeouts are extended by it, and the run report says how long each wait took. Runs that follow each other, as in an A/B comparison, then start from comparable thermal conditions. A wait of 10 s or more before the measurement is cut out of the throughput and backlog charts, whose time axis then breaks, with a small mark, and shows the real time before the break.
  • workloads.iotTelemetry.gateways.env and workloads.iotTelemetry.applications.env set environment variables of the producer (gateways) and the consumer (applications) containers, as cluster.brokers.env and cluster.bookies.env do for the brokers and the bookies, for example GLIBC_TUNABLES. A JAVA_TOOL_OPTIONS given there is appended to the launcher's JVM options, which keeps the heap settings and the profiling agent.

Metrics, in a new module :tests:performance:metrics (described in tests/performance/docs/metrics.md):

  • A metrics stack, VictoriaMetrics and Grafana with Grafana's image renderer, which Docker Compose runs from tests/performance/metrics/compose.yaml as the project pulsar-performance-metrics: ./gradlew :tests:performance:metrics:up starts it in the background, and ./gradlew :tests:performance:metrics:down stops it. metrics:up prints the links to Grafana and to VictoriaMetrics' web UI, vmui, for browsing the metrics and for PromQL queries. Its volumes and the network that the runs connect to are external to the project, created with the docker CLI, so that the metrics, the dashboards, the annotations and Grafana's settings outlive the stack.
  • When Grafana's volume is empty, a setup step provisions the VictoriaMetrics data source and the 19 dashboards of the Apache Pulsar Helm chart's pulsar dashboards, pinned to a commit, each with an annotation query for the runs. Grafana runs the slim image with the Prometheus data source plugin, has a fixed admin password, and lets anyone view the dashboards; the stack is published on the loopback interface by default.
  • A run collects its metrics by default, into the running stack or into one that it starts for itself and stops at its end (--no-metrics or -Pperformance.metrics=false collects none). Its brokers, bookies and ZooKeeper join the metrics network, and VictoriaMetrics scrapes them with the run's directory path as the cluster label, and the job and kubernetes_pod_name labels that the dashboards filter on. The scenario's metrics.intervalSeconds, 5 by default, sets the scrape interval and the brokers' and bookies' stats periods, which have to match it since the stats start over at each period.
  • At the end of a run, the launcher adds its events to Grafana as annotations (the gateways' start, the end of the warmup, the gateways' finish and the applications' finish), renders eight panels of the Messaging, JVM and BookKeeper dashboards as PNG images into the run report, with links to the panels and the dashboards over the run, and writes metrics.json, with the run's label selector, time range, jobs, events, and the URLs, credentials and data source with which scripts and agents query VictoriaMetrics and Grafana. A run whose metrics fail goes on without them.

Grafana
Integrated and pre-configured to collect metrics during the test runs
image

VictoriaMetrics Explore
Metrics backend for Grafana, exploring metrics example:
image

VictoriaMetrics PromQL
Metrics backend for Grafana, PromQL example:
image

Other additions:

  • Heap dumps of the broker, the gateways and the applications on request (heapDumps in a scenario): on OutOfMemoryError, at the start or the end, at given times, periodically and at the highest heap usage, optionally gzip-compressed. tests/performance/docs/heap-dumps.md describes them, and tests/performance/docs/analyzing-profiles.md analyzing them with jafar-shell.
  • The launcher keeps what it prints on the console in console.log.txt, and deletes launcher.log, which holds the containers' logs, when a run succeeds (--keep-launcher-log keeps it).
  • -Pperformance.clusterPulsarImage runs the cluster on a released Pulsar image, to compare the checkout with a release.
  • The topic stats sampling reaches every broker of a cluster with several brokers.

Host setup, in tests/performance/environment (described in its README):

  • scripts/configure-perf-test-environment.sh, run as root. install installs TuneD on Debian based distros, disables its dynamic tuning, installs the performance-testing TuneD profile, limits the size of Docker's container logs in /etc/docker/daemon.json while keeping its other settings, and leaves the TuneD daemon disabled. start checks that the host is on AC power and warns when the disk is nearly full, stops thermald and, on Pop!_OS, com.system76.PowerDaemon.service, activates and verifies the profile, and skips the :tests:integration:tuneKernelPerfEvents task in the user's ~/.gradle/gradle.properties. stop switches TuneD to the balanced profile (RESTORE_PROFILE), which allows power saving, stops TuneD, applies the system's configured dirty page limits and swappiness again, starts the stopped daemons again and removes the Gradle property.
  • The performance-testing profile is based on latency-performance and adds: turbo disabled, which fixes the CPU frequency at the base frequency with latency-performance's min_perf_pct=100; vm.swappiness=1 and NUMA balancing disabled, keeping the host's own dirty page limits; the none I/O scheduler, the performance ACPI platform profile and NVMe power state transitions disabled; and the perf event and BPF settings for profiling, the NMI watchdog disabled and the Transparent Huge Pages settings for -XX:+UseTransparentHugePages, which are left in place when the profile is deactivated.
  • The README describes the setup, including a sudoers rule for running start and stop without a password.
  • tuneKernelPerfEvents sets the THP defrag mode to madvise instead of defer, so that -XX:+AlwaysPreTouch gets huge pages at startup instead of depending on memory fragmentation and khugepaged. The inttest.asyncprofiler.skipPerfEventTuning property skips the task only when it is empty or true, so that a later value in gradle.properties can override an earlier one.

Tests:

  • The launcher and report tool tests under tests/performance, including the ones that existed before this change, use AssertJ assertions instead of TestNG's Assert, as the other tests/performance modules already do.

  • Document the performance tests in tests/performance/README.md, a tutorial, with reference pages in tests/performance/docs and tests/performance/scenarios/docs, and guide AI agents with tests/performance/AGENTS.md, which has a quick reference of the commands and points agents to the jonoffcpu report, which is text, for blocked time, and to the collapsed stacks for the full call trees; point CONTRIBUTING.md to it.

Verifying this change

  • Make sure that the change passes the CI checks.

This change added tests and can be verified as follows:

  • The report tool's tests cover the run and profile reports, their numbers, charts, links and HTML pages (RunReportTest, ProfileReportTest, MarkdownPagesTest, HdrHistogramRendererTest, TimeSeriesRendererTest), the run's provenance (RunInfoTest) and the rendered async-profiler views (JfrFlamegraphViewsTest); the launcher's tests cover the run directory layout and index links (RunDirectoryTest); HdrLatencyRecorderTest covers the per-second latency logs. WorkloadEnvironmentTest covers the producer and consumer containers' environment variables.
  • The host script passes bash -n and shellcheck -S warning.
  • ./gradlew :tests:performance:launcher:profile --args='--scenario tests/performance/scenarios/iot-telemetry-high-rate.yaml --extends configs/profile-broker --extends configs/profile-gateways' completes with correct delivery (all messages, no duplicates, no ordering violations) and writes the reports, off-CPU, flame-graph and digest outputs described in the README. Profiling needs a Docker engine whose kernel has BTF, and privileged containers. It works on Linux and on macOS, and async-profiler, jonoffcpu's off-CPU profiling and JFR were also tested on macOS arm64 with the OrbStack Docker engine; Linux x86_64, configured with tests/performance/environment, is recommended for measurements, since it is Pulsar's main target platform, and dedicated hardware has no noisy neighbours and less thermal and power throttling and CPU frequency variance.
  • The tooling was tested on Pop!_OS 24.04 (Ubuntu-based) Linux, x86_64, with 32 GB of RAM and Docker Engine, and on macOS on an Apple M3 Max with 36 GB of RAM and OrbStack, with a 20 GB memory limit for OrbStack. Docker Desktop and Podman Desktop are untested. 32 GB of RAM on the host is recommended, although testing may be possible with less.
  • The metrics module's and the launcher's tests cover the scrape configuration, the Grafana links and render URLs, the metrics settings and the run report's Metrics section (ScrapeConfigTest, MetricsSettingsTest, RunReportTest). Runs with the metrics stack started by metrics:up, and started by the run itself, collected the brokers', bookies' and ZooKeeper's metrics, added the annotations, rendered the panels into the run report and kept the data in the volumes across restarts of the stack.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

Dependencies, used only by the performance launcher, its report tool and the integration-test profiling support, and not part of the Pulsar distribution or images: io.github.jonoffcpu:jonoffcpu-agent, jonoffcpu-correlator and jonoffcpu-jfr-converter 0.8.0 (Apache-2.0); org.commonmark:commonmark, commonmark-ext-gfm-tables and commonmark-ext-heading-anchor 0.30.0 (BSD-2-Clause); and org.knowm.xchart:xchart 4.0.4 (Apache-2.0), used for PNG only, so that none of its optional dependencies, such as the LGPL VectorGraphics2D behind its SVG export, is pulled in.

This PR was prepared with AI assistance (Claude Code) and reviewed by a human contributor.

@lhotari
lhotari marked this pull request as draft September 26, 2026 09:56
@lhotari
lhotari force-pushed the lh-use-jonoffcpu-profiler branch from a782d92 to ab6fd45 Compare September 26, 2026 11:33
@lhotari
lhotari marked this pull request as ready for review September 26, 2026 11:33
@lhotari
lhotari force-pushed the lh-use-jonoffcpu-profiler branch from ab6fd45 to aeb35ff Compare September 26, 2026 12:02
@lhotari
lhotari marked this pull request as draft September 26, 2026 12:31
…off-CPU agent

Motivation

The profiled IoT scenario only captured async-profiler CPU samples. Off-CPU
time (scheduler waits, I/O, lock parking) is invisible there, yet it is where
most broker latency goes.

Modifications

- Add `JonoffcpuAgent` (jonoffcpu 0.4.0), which mounts the agent JAR into a
  profiled container, runs it privileged as root with a Docker-managed tracefs
  volume at /sys/kernel/tracing (advised for Docker Desktop), and writes the
  agent's YAML config with the required `sampling:` block.
- Let the test image's run scripts switch supervisord's user through
  PULSAR_PROCESS_USER so the profiled JVM keeps root; a non-root container
  process has no effective capabilities even when privileged.
- Profiled runs use the Alpine test image by default, since the agent bundles
  musl and glibc natives; add a Wolfi-based variant (`dockerBuildWolfi`),
  selected with -Pinttest.testImageVariant=wolfi.
- Add `OffCpuFlamegraphs`, which runs the jonoffcpu correlator over each
  recording's capture stream cut to the measurement window without the
  row-level audit files and with a 10 ms synthetic JFR quantum, and renders
  the collapsed stacks in-process with the jonoffcpu-jfr-converter dependency
  (`--units µs`).
- Relax unprivileged_bpf_disabled in tuneKernelPerfEvents; pass the agent
  path through a provider so the profile task is configuration cache
  compatible, and give it an explicit 4 GB heap.
- Record only blocked intervals of 100 µs and above, sampled in proportion to
  their length below 10 ms, and drop the lock event from the profiled
  scenario.
- Document the profiling flow in tests/performance.

Assisted-by: Claude Code (Opus 5.5)
Motivation

Over 99 % of a broker's off-CPU time is threads waiting for work: Netty event
loops in epollWait, executor workers waiting for a task, JDK and HotSpot
service threads. That hides the waits worth optimizing, and long package
names leave little room in a frame for the class and method. On the Alpine
image, native frames are all the unsymbolized /lib/ld-musl-x86_64.so.1, so
the JVM's own threads cannot be told apart.

Modifications

- Render two slices of the correlator's stack profile with
  `stacks --package-names abbreviate`: offcpu with every interval, and
  offcpu-no-idle, which excludes intervals whose Java, native or kernel stack
  holds an idle-wait frame (OffCpuFlamegraphs.IDLE_WAIT_FRAMES). Each has a
  collapsed file, a JSON summary accounting for the excluded time, and an
  HTML flame graph.
- The idle-wait frames name the wait itself rather than the thread's run
  loop, so contention while running a task stays visible. On a broker run
  they leave 50 s of 8,519 s: monitor and lock contention, GC phases and
  safepoints.
- Profile on the glibc-based Wolfi test image by default, where native frames
  are symbolized; -Pinttest.testImageVariant=alpine selects Alpine.

Assisted-by: Claude Code (Opus 5.5)
Move the idle-wait frame patterns from a Java list to the launcher resource
offcpu-idle-waits.txt, passed to jonoffcpu stacks with --exclude-from and copied
into each output directory so a run records the patterns it used. Also exclude
threads switched out inside async-profiler's signal handler, and render stacks
with package names dropped.

Assisted-by: Claude Code (Opus 5.5)
…ecordings

After a profiled run, render the cpu, wall, alloc and lock views of each
measurement recording into <recording>-flamegraphs/ (HTML, per-thread HTML and
collapsed stacks), each only when its event is in the recording's async-profiler
options as recorded in the agent configuration. Keep full names in the off-CPU
slices' collapsed files so that scripts and diff tools can classify frames by
package; only the flame graphs drop package names.

Assisted-by: Claude Code (Opus 5.5)
Describe the jonoffcpu profiling flow in one section of the performance README:
requirements, the files a profiled run writes, and how to find what to optimize
(the idle-free off-CPU graph, ranking blocked time by application frame with
DuckDB, stack profile slices, run comparisons), plus which JFR events Jafar
should query. Point CONTRIBUTING.md and the IoT scenario notes to it.

Assisted-by: Claude Code (Opus 5.5)
jonoffcpu moved from github.com/lhotari/jonoffcpu to github.com/jonoffcpu/jonoffcpu.
Its Maven group is now io.github.jonoffcpu and the correlator's package is
io.github.jonoffcpu.correlator. Upgrade to 0.6.0 with the new coordinates,
update the correlator import and point the documentation links to the new
repository.

Assisted-by: Claude Code (claude-opus-5-5)
Assisted-by: Claude Code (claude-opus-5-5)
… heatmaps for profiled runs

Adopt the jonoffcpu 0.7.0 reporting in the performance launcher:

- Correlate with --idle-from and the launcher's idle-wait patterns, so the
  off-CPU digest jonoffcpu-summary.md ranks the busy time and leaves out the
  same waits for work as offcpu-no-idle. Drop --quantum-ns: 0.7.0 no longer
  writes the synthetic JFR.
- Render two more off-CPU slices with stacks --root-at ^org\.apache\.:
  offcpu-app-root and offcpu-no-idle-app-root start each stack at its first
  Pulsar or BookKeeper frame, so the same code reached from different thread
  pools or event loops forms one tree.
- Abbreviate package names in the off-CPU flame graphs instead of dropping
  them, and highlight the application's frames there and in the JFR views.
- Render a heatmap of each JFR view (<view>-heatmap.html).
- Write profile-report.md into each profiled directory, such as
  broker-profile/: the run, and for each recording links to the digest, the
  off-CPU flame graphs with their totals, and the JFR views and heatmaps.
- Update the performance README: what a run writes, and top and top
  --baseline instead of the DuckDB query and the synthetic-JFR diff.

Assisted-by: Claude Code (claude-opus-5-5)
…ed backlog, rendered to HTML

After each run the launcher left summaries and HDR latency logs, but no readable
result, and the Markdown profile reports linked to HTML flame graphs that most
Markdown viewers do not follow.

- Every run writes run-report.md: the scenario settings, correctness per
  application, producer and delivered throughput, the seconds consumers were
  still draining, and publish and end-to-end latency percentiles with the
  latency distribution chart from HdrHistogramRenderer.
- A sampler polls the workload topics' stats once per second while the
  producers run and the consumers drain, into topic-stats.csv: each
  subscription's backlog and the published and dispatched message counters.
  The report adds the sampled maximum backlog, the per-second publish and
  dispatch rates (median and minimum within the measurement), and two charts
  over time, throughput and backlog, as wide as the latency chart.
- Each Markdown report, the run and profile reports and the jonoffcpu off-CPU
  digests, gets an HTML page rendered with commonmark-java (BSD-2-Clause) and
  a small inlined stylesheet. Links to other Markdown reports lead to their
  pages, and absolute paths inside the run directory become relative.

Assisted-by: Claude Code (claude-opus-5-5)
…dle in off-CPU profiles

The broker's pulsar-web reserved threads block in
ReservedThreadExecutor$ReservedThread.waitForTask until Jetty hands them a
job. The wait was missing from the idle list, so it made up 37 of the 41
seconds of busy off-CPU time on a Key_Shared 500x20 broker profile.

Assisted-by: Claude Code (claude-opus-5-5)
commonmark-java renders headings without an id, so in-page links of the
Markdown reports, such as the off-CPU digest's #where-the-time-went, led
nowhere in the HTML pages. Add its heading-anchor extension, which gives
each heading a GitHub-style id.

Assisted-by: Claude Code (claude-opus-5-5)
…h its profiles, bound the backlog maximum to the measurement, tidy the latency chart

Assisted-by: Claude Code (claude-opus-5-5)
…, end the backlog maximum at the producers' finish

Assisted-by: Claude Code (claude-opus-5-5)
…-CPU digest and app-root flame graphs, abbreviate Java names on the digest page

Assisted-by: Claude Code (claude-opus-5-5)
…branch and name, and record its provenance

Assisted-by: Claude Code (claude-opus-5-5)
…, and leave stacks without an application frame out of the app-root flame graphs

Assisted-by: Claude Code (claude-opus-5-5)
…of its report, and save container logs as text

Assisted-by: Claude Code (claude-opus-5-5)
jonoffcpu 0.8.0 bundles an async-profiler that aligns its clock with the JVM's JFR clock on arm64, so that the
profiler's events line up with the JVM's in whole-recording readers and in the off-CPU correlation. It replaces the
snapshot from the local Maven repository, which is removed.

Assisted-by: Claude Code (claude-opus-5-5)
…ed Pulsar

A performance run could only test the checkout, so comparing it with a release meant building that release's
revision. -Pperformance.clusterPulsarImage=<image>, such as apachepulsar/pulsar:4.0.13 or apachepulsar/pulsar:latest,
builds the test image on that Alpine-based Pulsar image, pulling it first, and runs ZooKeeper, the configuration
store, the bookies and the brokers on it. The gateways and the applications stay on the checkout's test image, and so
does the Pulsar client.

The launcher reads the brokers' version once the cluster has started. The run report's title and the charts' footer
lead with it, and name the checkout's branch and commit as the clients', since the client affects the results when
the clients are the bottleneck. run-info.json has the image and the version, and the runs go under pulsar-<tag>
instead of the branch. PulsarClusterSpec gets a cluster image, which the cluster's ZooKeeper, configuration store,
bookie, broker and proxy containers use.

Assisted-by: Claude Code (claude-opus-5-5)
…nce testing README

The README's comparison step now shows -Pperformance.clusterPulsarImage with the latest release and with a
particular one, and the comparison page's example runs both. With latest, the runs of different releases share the
pulsar-latest directory, while the reports name the version that the brokers reported.

Assisted-by: Claude Code (claude-opus-5-5)
…rio needs, and a sustainable high-rate scenario

The IoT telemetry scenarios had inconsistent results with large backlogs and delays. In the full topology,
iot-telemetry.yaml, 1,000 messages per second had a publish p50 of about 400 ms and an end-to-end p99 of seconds,
and in the high-rate scenario the publish p50 was about 2 s and applications fell behind by millions of messages.

- The medium-memory configuration's broker gets a 4 GB heap instead of 2 GB. The full topology's 20 applications
  with 100 pods on 30 topics are 60,000 Key_Shared consumers, whose consistent hash rings, with 100 points per
  consumer, and consumer state leave about 2 GB of live objects: with a 2 GB heap the broker spent its CPU in
  garbage collection. The configuration now needs about 11 GB of memory available to Docker instead of 9 GB.
- iot-telemetry-high-rate.yaml publishes 30,000 messages per second instead of as fast as possible, and its
  gateways keep at most 10,000 messages in flight instead of 100,000. Without a rate limit, on a host whose CPU runs
  at a fixed base frequency, the gateways published faster than the applications received, their backlogs grew, and
  an application that fell behind could stall for tens of seconds; the 100,000 messages in flight alone made the
  publish latency about 2 s.
- The docs describe the memory and the high-rate scenario, and that warmup.messages works with a rate too.

Measured twice each with the performance testing environment active: iot-telemetry.yaml has a publish p50 of 1.4
ms and an end-to-end p99 of 47 ms with backlogs of at most 128 messages, iot-telemetry-high-rate.yaml delivers
30,000 messages per second with an end-to-end p99 of at most 40 ms and backlogs under 3,000 messages, and
iot-telemetry-small.yaml, which needed no change, has an end-to-end p99 of 11 ms.

Assisted-by: Claude Code (claude-opus-5-5)
…mand, and analyze them with jafar-shell

Finding what fills a component's heap, such as the broker's in the full IoT topology, needed heap dumps taken by
hand inside the containers, at the right time, and the JVMs in the containers write files that only their own user
can read.

- A scenario's heapDumps section asks for heap dumps of the broker, the gateways and the applications: when the JVM
  runs out of memory (onOutOfMemoryError), when the gateways start (atStart), at given seconds after that
  (atSeconds), periodically (everySeconds), at the highest heap usage (atPeakUsage), and after every application has
  received every message (atEnd, the broker only). configs/heap-dumps-broker.yaml adds the broker's peak, end and
  out-of-memory dumps to any scenario with --extends.
- The launcher writes the dumps with jcmd GC.heap_dump, as the user that the JVM runs as, one at a time, into
  heap-dumps/<component> in the run directory, which each container mounts, and lists them in
  heap-dumps/heap-dumps.csv. atPeakUsage samples the heap usage every 5 s and keeps one dump, of a heap at least half
  full and 10 % above the previous peak dump. onOutOfMemoryError adds -XX:+HeapDumpOnOutOfMemoryError to the JVM's
  JAVA_TOOL_OPTIONS. When the run ends, also after a failure, a one-off container makes the dumps readable and the
  launcher prints where they are; the run report lists them in a Heap dumps section.
- docs/heap-dumps.md describes the settings, the files and the cost: a dump stops the JVM, so a run with scheduled
  dumps is for finding what holds the memory, not for measuring.
- docs/analyzing-profiles.md describes jafar-shell, the newer version of jfr-shell, which reads JFR recordings and
  heap dumps with its JfrPath and HdumpPath queries, also from standard input for scripts and agents, and links its
  heap dump quick start. Its heap dump support isn't in a release yet, so the docs show how to build and install it.

Assisted-by: Claude Code (claude-opus-5-5)
…n with the current memory settings

The performance testing README's example output of iot-telemetry.yaml showed publish latencies of hundreds of
milliseconds, from a run whose broker heap was too small. It now shows the lines of a run with the 4 GB broker heap.

Assisted-by: Claude Code (claude-opus-5-5)
…ter honored numProxies

HealthCheckTest set numProxies to 0, which PulsarCluster ignored until it started the proxy only when the spec
asks for one. Without the proxy, testZooKeeperDown's healthcheck hangs past the test's 300 s timeout, in CI and
locally, while with the proxy it passes as it did before. The test now has the default proxy again, the topology
it always ran with.

Assisted-by: Claude Code (claude-opus-5-5)
…-compressed

A scenario's heapDumps.gzipLevel, from 1 to 9, has the JVMs write every heap dump compressed, as .hprof.gz: the
launcher's dumps with jcmd GC.heap_dump -gz=<level>, and the dumps on OutOfMemoryError with
-XX:HeapDumpGzipLevel=<level>, which JDK 17 and later have, so also the Java 21 of the Pulsar 4 images. Without it
the dumps stay uncompressed .hprof files. At level 1, a small broker's dumps came to about a quarter of their size in
the time that its uncompressed dumps took.

jafar-shell reads uncompressed dumps only, so the docs show decompressing a dump first with gunzip -k.

Assisted-by: Claude Code (claude-opus-5-5)
… run directory, and delete launcher.log of successful runs

launcher.log holds Testcontainers' and the Pulsar containers' logs, which make it large, and a successful run
rarely needs it. The launcher still writes it during the run, so that it can be followed, and a failed run keeps it,
but a successful run deletes it at the end, after stopping the logging, which closes the file. --keep-launcher-log,
or -Pperformance.keepLauncherLog, keeps it.

What the launcher prints on the console goes also to console.log.txt in the run directory, which every run keeps
and the run report links: the resolved scenario, the phases, the progress lines, the heap dumps and the failure.

Assisted-by: Claude Code (claude-opus-5-5)
…rver, with a configurable base URL

When the tests run on another host, a report's URL that opens with a click in the terminal saves finding it through
the server's directory listings. The launcher prints the URL of the run report and of each profile report beside its
file, as the base URL with the report's path in the reports root appended, without a directory's index.html.

-Pperformance.reportsServer.baseUrl sets the base URL, such as http://192.168.1.123:8000/. Without it, it is
http://<bind address>:<port>/ of serveReports, and with the bind address 0.0.0.0, the IPv4 address of the host's
first network interface, by index, that is up, other than a loopback or link-local one; a host with several
interfaces sets the base URL instead. serveReports prints the same base URL when it starts.

Assisted-by: Claude Code (claude-opus-5-5)
…ctoriaMetrics and show them on Grafana's Pulsar dashboards

A run's report shows the workload's view of throughput, latency and backlog, but not what the brokers, bookies and
ZooKeeper went through meanwhile, and runs couldn't be compared on those metrics afterwards.

A new module, :tests:performance:metrics, runs VictoriaMetrics and Grafana, with Grafana's image renderer, with Docker
Compose from its compose file, as the project pulsar-performance-metrics. ./gradlew :tests:performance:metrics:up
starts it in the background, and it keeps running, also across restarts of Docker, until
./gradlew :tests:performance:metrics:down stops it. Its volumes and the network that the runs connect to are external,
created with the docker CLI, so that stopping the stack keeps the metrics, the dashboards, the annotations and
Grafana's settings. Grafana's setup, when its volume is empty, provisions the VictoriaMetrics data source and the 19
dashboards of the Apache Pulsar Helm chart, pinned to a commit, with an annotation query for the runs. Grafana has the
fixed admin password pulsar-performance, and viewing needs no login; the slim Grafana image installs the Prometheus
plugin.

A run collects its metrics by default, -Pperformance.metrics=false or --no-metrics doesn't: into the running stack,
or into the stack that it starts for itself and stops at its end. Its brokers, bookies and ZooKeeper join the metrics
network under their container names, and the run's scrape configuration goes into VictoriaMetrics, with the run's
directory path in the reports root as the cluster label, and the jobs and instances that the dashboards filter on.
The scenario's metrics.intervalSeconds, 5 by default, sets the scrape interval, and the broker's and the bookies'
stats periods, which have to match it since the stats start over at each period: statsUpdateFrequencyInSecs,
statsUpdateInitialDelayInSecs, managedLedgerStatsPeriodSeconds and managedLedgerPrometheusStatsLatencyRolloverSeconds,
and the bookies' prometheusStatsLatencyRolloverSeconds.

At the end of a run, the launcher adds its events to Grafana as annotations: the gateways' start, the end of the
warmup, the gateways' finish and the applications' finish. It renders eight panels of the Messaging, JVM and
BookKeeper dashboards as PNG images into the run directory, which the run report's Metrics section shows with links
to the panels and the dashboards over the run, at -Pperformance.metrics.grafanaUrl. metrics.json has the run's
selector, time range, jobs, events, and the URLs, credentials and data source with which scripts and agents query
VictoriaMetrics and Grafana. A run without metrics, or whose metrics fail, goes on, and its report has no Metrics
section.

Assisted-by: Claude Code (claude-opus-5-5)
…p dumps, sampling and metrics stack

A review of the branch found these, each confirmed before it was fixed:

- Profiling a released Pulsar 4.x broker failed: the jonoffcpu agent's JVM options had
  --sun-misc-unsafe-memory-access=allow, which Java 21 rejects. The agent no longer passes it; Pulsar's scripts add
  it on Java 23 and later.
- The broker's heap dump on OutOfMemoryError went to /var/log/pulsar: the test image's scripts put
  -XX:HeapDumpPath=/var/log/pulsar on the command line, which overrode JAVA_TOOL_OPTIONS. The broker's options go to
  PULSAR_EXTRA_OPTS, which come after it.
- jcmd inherited a profiled workload's JAVA_TOOL_OPTIONS and started the profiling agent in its own JVM; it runs
  without them.
- With several brokers, the topic stats of the topics that another broker owns redirected to names that the host
  can't resolve, and weren't sampled. The sampler has a client per broker, and samples each topic through its owner.
- A metrics stack whose start failed part way kept running; the failed start is stopped too.
- The workload startup's 60 second stall timeout wasn't checked while the container kept logging without progress.
- Heap dumps due after the workload's end delayed the end of the run by up to 10 minutes; they are cancelled.
- The IoT telemetry reference gave the full topology's broker a 2 GiB heap, not 4 GiB, and the README's setting of
  the reports root failed when ~/.gradle didn't exist.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari
lhotari marked this pull request as ready for review September 27, 2026 18:43
@lhotari
lhotari marked this pull request as draft September 27, 2026 18:44
…bring the performance tests' guides up to date

The run report's Profiles section named each profile with a link to its profile report, and the off-CPU digest was
one more click away. Each profile now has its reports a click away: the profile report, the jonoffcpu report, which
is the digest, and the off-CPU, CPU and allocation flame graphs that it has.

The README tutorial's steps were run as written and pass. Its example output now has the metrics lines that a run
prints, and the agent guide has a quick reference of the commands: the launcher's options with --help, running and
profiling a scenario, the reports root, serving the reports and starting and stopping the metrics stack. Its advice
on charts in pull requests uses HTML images with a width, linked to the uploaded files, which Markdown images in a
table's columns don't display properly.

The guides say which platforms the tooling was tested on: Linux and macOS both work, including profiling with
async-profiler, jonoffcpu and JDK Flight Recorder, tested on Pop!_OS 24.04 x86_64 with Docker Engine and on macOS
arm64 with OrbStack; Docker Desktop and Podman Desktop are untested. Linux x86_64, configured for performance
testing, is recommended for measurements, with 32 GB of RAM on the host.

Assisted-by: Claude Code (claude-opus-5-5)
…ictoriaMetrics, and link vmui's views

The SSH tunnel for reaching the reports server of another machine now forwards Grafana's and VictoriaMetrics' ports
too, and docs/run-reports.md shows a ~/.ssh/config host that opens the same tunnel with `ssh <host>-pulsar-perf`.
The metrics docs and `metrics:up` link VictoriaMetrics' web UI both for browsing the metrics (vmui/?#/metrics) and
for PromQL queries (vmui).

Assisted-by: Claude Code (claude-opus-5-5)
… the forwarding

The tunnel is silent while it works. The guide now says what it prints when a forwarded service isn't running on
the remote machine (channel open failed, connect failed) or a local port is already in use (bind: Address already in
use, Could not request local forwarding.), and that -v logs each forwarded port and each connection through the
tunnel, as checked in OpenSSH 9.6p1's channels.c and ssh.c.

Assisted-by: Claude Code (claude-opus-5-5)
…off-CPU summary)

The run report's Profiles section and the profile report both name the off-CPU digest "jonoffcpu report (off-CPU
summary)". In the Profiles table it comes first among a profile's links, in bold. In the profile report's files it is
bold, and its description starts with a bold "Start here" link to it. The README and the guide to analyzing profiles
say that the off-CPU flame graphs show the same blocked time as call trees, to inspect visually which code paths lead
to the blocking methods.

Assisted-by: Claude Code (claude-opus-5-5)
…phs show, and name the blocked time as such

The reports described the off-CPU flame graphs by how they were filtered rather than by what they show, and didn't
tell them apart from async-profiler's flame graphs. jonoffcpu produces both: the off-CPU flame graphs from its
off-CPU capture, where a box's width is off-CPU time, and the CPU, allocation and other flame graphs from
async-profiler's samples in the JFR recording.

- The run report's Profiles section introduces both kinds for the profiled process, the broker or a Pulsar client,
  names its column "Blocked time" and links the blocked time flame graph.
- The profile report names the process it profiled. Its "Off-CPU flame graphs: where threads waited" section says
  what the flame graphs show and which of them matter for throughput, and lists the blocked time first: "Without
  idle waits" is now "Blocked time", and "from the application's first frame" is "from where threads entered Pulsar
  or BookKeeper code". A collapsed section explains the details: how idle waits are recognized, how the flame graphs
  are rooted, and that box widths are the observed time of the recorded waits, which the sampling policy can bias
  against short waits.
- Its "async-profiler flame graphs" section says where their samples come from and what each view shows.
- The jonoffcpu report's description in the profile report states its value: the methods where the process's threads
  blocked, on what and for how long, with the blocked time flame graphs for the code paths that lead to them.
- The flame graph pages' titles and the docs use the same names.

Assisted-by: Claude Code (claude-opus-5-5)
…d stacks for blocked time

The agent guide says to look first at a profile's jonoffcpu report, a text digest of where the Pulsar broker's or
the Pulsar client's threads blocked, and to read the full call trees from the collapsed stacks beside it, since the
flame graphs' HTML pages need a browser to render.

Assisted-by: Claude Code (claude-opus-5-5)
…riptions

- The off-CPU section no longer contrasts its widths with async-profiler's sample counts: the allocation and lock
  views weigh bytes and time. Each async-profiler view states what its widths count.
- The lock view includes thread parking, such as idle executor workers, not only waits to enter Java monitors.
- `--weights estimated` corrects for the capture's sampling, but waits shorter than the sampling policy's minimum
  aren't recorded at all.

Assisted-by: Claude Code (claude-opus-5-5)
…board that it is on

A panel's image opened the panel on its own in Grafana. Below each panel, the run report now also links the
dashboard that the panel is on, with the run's cluster, the dashboard's variables and the run's time range, so that
the panel can be seen among the others. metrics.json records each panel's dashboard and its URL.

Assisted-by: Claude Code (claude-opus-5-5)
…the README

The README's table of tested Docker engines described the Linux host's testing more briefly than the macOS one's.
Both rows now list runs, profiling with async-profiler, jonoffcpu and JDK Flight Recorder, and metrics, which were
tested on both.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari lhotari changed the title [improve][test] Profile performance scenarios off-CPU with jonoffcpu [feat][test] Add a Docker-based performance testing framework with profiling, metrics and AI agent support Sep 27, 2026
@lhotari
lhotari marked this pull request as ready for review September 27, 2026 22:26
@lhotari
lhotari merged commit d3978ac into master Sep 28, 2026
71 of 73 checks passed
@lhotari
lhotari deleted the lh-use-jonoffcpu-profiler branch October 1, 2026 00:02
@lhotari lhotari added this to the 5.0.0 milestone Oct 1, 2026
Radiancebobo pushed a commit to Radiancebobo/pulsar that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants