Repository navigation
[feat][test] Add a Docker-based performance testing framework with profiling, metrics and AI agent support - #26714
Merged
Merged
Conversation
lhotari
added this pull request to stack #26718
September 25, 2026 19:28
This was referenced Sep 25, 2026
dao-jun
approved these changes
Sep 26, 2026
lhotari
marked this pull request as draft
September 26, 2026 09:56
lhotari
force-pushed
the
lh-use-jonoffcpu-profiler
branch
from
September 26, 2026 11:33
a782d92 to
ab6fd45
Compare
lhotari
marked this pull request as ready for review
September 26, 2026 11:33
lhotari
force-pushed
the
lh-use-jonoffcpu-profiler
branch
from
September 26, 2026 12:02
ab6fd45 to
aeb35ff
Compare
lhotari
marked this pull request as draft
September 26, 2026 12:31
…off-CPU agent Motivation The profiled IoT scenario only captured async-profiler CPU samples. Off-CPU time (scheduler waits, I/O, lock parking) is invisible there, yet it is where most broker latency goes. Modifications - Add `JonoffcpuAgent` (jonoffcpu 0.4.0), which mounts the agent JAR into a profiled container, runs it privileged as root with a Docker-managed tracefs volume at /sys/kernel/tracing (advised for Docker Desktop), and writes the agent's YAML config with the required `sampling:` block. - Let the test image's run scripts switch supervisord's user through PULSAR_PROCESS_USER so the profiled JVM keeps root; a non-root container process has no effective capabilities even when privileged. - Profiled runs use the Alpine test image by default, since the agent bundles musl and glibc natives; add a Wolfi-based variant (`dockerBuildWolfi`), selected with -Pinttest.testImageVariant=wolfi. - Add `OffCpuFlamegraphs`, which runs the jonoffcpu correlator over each recording's capture stream cut to the measurement window without the row-level audit files and with a 10 ms synthetic JFR quantum, and renders the collapsed stacks in-process with the jonoffcpu-jfr-converter dependency (`--units µs`). - Relax unprivileged_bpf_disabled in tuneKernelPerfEvents; pass the agent path through a provider so the profile task is configuration cache compatible, and give it an explicit 4 GB heap. - Record only blocked intervals of 100 µs and above, sampled in proportion to their length below 10 ms, and drop the lock event from the profiled scenario. - Document the profiling flow in tests/performance. Assisted-by: Claude Code (Opus 5.5)
Motivation Over 99 % of a broker's off-CPU time is threads waiting for work: Netty event loops in epollWait, executor workers waiting for a task, JDK and HotSpot service threads. That hides the waits worth optimizing, and long package names leave little room in a frame for the class and method. On the Alpine image, native frames are all the unsymbolized /lib/ld-musl-x86_64.so.1, so the JVM's own threads cannot be told apart. Modifications - Render two slices of the correlator's stack profile with `stacks --package-names abbreviate`: offcpu with every interval, and offcpu-no-idle, which excludes intervals whose Java, native or kernel stack holds an idle-wait frame (OffCpuFlamegraphs.IDLE_WAIT_FRAMES). Each has a collapsed file, a JSON summary accounting for the excluded time, and an HTML flame graph. - The idle-wait frames name the wait itself rather than the thread's run loop, so contention while running a task stays visible. On a broker run they leave 50 s of 8,519 s: monitor and lock contention, GC phases and safepoints. - Profile on the glibc-based Wolfi test image by default, where native frames are symbolized; -Pinttest.testImageVariant=alpine selects Alpine. Assisted-by: Claude Code (Opus 5.5)
Move the idle-wait frame patterns from a Java list to the launcher resource offcpu-idle-waits.txt, passed to jonoffcpu stacks with --exclude-from and copied into each output directory so a run records the patterns it used. Also exclude threads switched out inside async-profiler's signal handler, and render stacks with package names dropped. Assisted-by: Claude Code (Opus 5.5)
…ecordings After a profiled run, render the cpu, wall, alloc and lock views of each measurement recording into <recording>-flamegraphs/ (HTML, per-thread HTML and collapsed stacks), each only when its event is in the recording's async-profiler options as recorded in the agent configuration. Keep full names in the off-CPU slices' collapsed files so that scripts and diff tools can classify frames by package; only the flame graphs drop package names. Assisted-by: Claude Code (Opus 5.5)
Describe the jonoffcpu profiling flow in one section of the performance README: requirements, the files a profiled run writes, and how to find what to optimize (the idle-free off-CPU graph, ranking blocked time by application frame with DuckDB, stack profile slices, run comparisons), plus which JFR events Jafar should query. Point CONTRIBUTING.md and the IoT scenario notes to it. Assisted-by: Claude Code (Opus 5.5)
jonoffcpu moved from github.com/lhotari/jonoffcpu to github.com/jonoffcpu/jonoffcpu. Its Maven group is now io.github.jonoffcpu and the correlator's package is io.github.jonoffcpu.correlator. Upgrade to 0.6.0 with the new coordinates, update the correlator import and point the documentation links to the new repository. Assisted-by: Claude Code (claude-opus-5-5)
Assisted-by: Claude Code (claude-opus-5-5)
… heatmaps for profiled runs Adopt the jonoffcpu 0.7.0 reporting in the performance launcher: - Correlate with --idle-from and the launcher's idle-wait patterns, so the off-CPU digest jonoffcpu-summary.md ranks the busy time and leaves out the same waits for work as offcpu-no-idle. Drop --quantum-ns: 0.7.0 no longer writes the synthetic JFR. - Render two more off-CPU slices with stacks --root-at ^org\.apache\.: offcpu-app-root and offcpu-no-idle-app-root start each stack at its first Pulsar or BookKeeper frame, so the same code reached from different thread pools or event loops forms one tree. - Abbreviate package names in the off-CPU flame graphs instead of dropping them, and highlight the application's frames there and in the JFR views. - Render a heatmap of each JFR view (<view>-heatmap.html). - Write profile-report.md into each profiled directory, such as broker-profile/: the run, and for each recording links to the digest, the off-CPU flame graphs with their totals, and the JFR views and heatmaps. - Update the performance README: what a run writes, and top and top --baseline instead of the DuckDB query and the synthetic-JFR diff. Assisted-by: Claude Code (claude-opus-5-5)
…ed backlog, rendered to HTML After each run the launcher left summaries and HDR latency logs, but no readable result, and the Markdown profile reports linked to HTML flame graphs that most Markdown viewers do not follow. - Every run writes run-report.md: the scenario settings, correctness per application, producer and delivered throughput, the seconds consumers were still draining, and publish and end-to-end latency percentiles with the latency distribution chart from HdrHistogramRenderer. - A sampler polls the workload topics' stats once per second while the producers run and the consumers drain, into topic-stats.csv: each subscription's backlog and the published and dispatched message counters. The report adds the sampled maximum backlog, the per-second publish and dispatch rates (median and minimum within the measurement), and two charts over time, throughput and backlog, as wide as the latency chart. - Each Markdown report, the run and profile reports and the jonoffcpu off-CPU digests, gets an HTML page rendered with commonmark-java (BSD-2-Clause) and a small inlined stylesheet. Links to other Markdown reports lead to their pages, and absolute paths inside the run directory become relative. Assisted-by: Claude Code (claude-opus-5-5)
…dle in off-CPU profiles The broker's pulsar-web reserved threads block in ReservedThreadExecutor$ReservedThread.waitForTask until Jetty hands them a job. The wait was missing from the idle list, so it made up 37 of the 41 seconds of busy off-CPU time on a Key_Shared 500x20 broker profile. Assisted-by: Claude Code (claude-opus-5-5)
commonmark-java renders headings without an id, so in-page links of the Markdown reports, such as the off-CPU digest's #where-the-time-went, led nowhere in the HTML pages. Add its heading-anchor extension, which gives each heading a GitHub-style id. Assisted-by: Claude Code (claude-opus-5-5)
…h its profiles, bound the backlog maximum to the measurement, tidy the latency chart Assisted-by: Claude Code (claude-opus-5-5)
…, end the backlog maximum at the producers' finish Assisted-by: Claude Code (claude-opus-5-5)
…-CPU digest and app-root flame graphs, abbreviate Java names on the digest page Assisted-by: Claude Code (claude-opus-5-5)
…branch and name, and record its provenance Assisted-by: Claude Code (claude-opus-5-5)
…, and leave stacks without an application frame out of the app-root flame graphs Assisted-by: Claude Code (claude-opus-5-5)
…of its report, and save container logs as text Assisted-by: Claude Code (claude-opus-5-5)
jonoffcpu 0.8.0 bundles an async-profiler that aligns its clock with the JVM's JFR clock on arm64, so that the profiler's events line up with the JVM's in whole-recording readers and in the off-CPU correlation. It replaces the snapshot from the local Maven repository, which is removed. Assisted-by: Claude Code (claude-opus-5-5)
…ed Pulsar A performance run could only test the checkout, so comparing it with a release meant building that release's revision. -Pperformance.clusterPulsarImage=<image>, such as apachepulsar/pulsar:4.0.13 or apachepulsar/pulsar:latest, builds the test image on that Alpine-based Pulsar image, pulling it first, and runs ZooKeeper, the configuration store, the bookies and the brokers on it. The gateways and the applications stay on the checkout's test image, and so does the Pulsar client. The launcher reads the brokers' version once the cluster has started. The run report's title and the charts' footer lead with it, and name the checkout's branch and commit as the clients', since the client affects the results when the clients are the bottleneck. run-info.json has the image and the version, and the runs go under pulsar-<tag> instead of the branch. PulsarClusterSpec gets a cluster image, which the cluster's ZooKeeper, configuration store, bookie, broker and proxy containers use. Assisted-by: Claude Code (claude-opus-5-5)
…nce testing README The README's comparison step now shows -Pperformance.clusterPulsarImage with the latest release and with a particular one, and the comparison page's example runs both. With latest, the runs of different releases share the pulsar-latest directory, while the reports name the version that the brokers reported. Assisted-by: Claude Code (claude-opus-5-5)
…rio needs, and a sustainable high-rate scenario The IoT telemetry scenarios had inconsistent results with large backlogs and delays. In the full topology, iot-telemetry.yaml, 1,000 messages per second had a publish p50 of about 400 ms and an end-to-end p99 of seconds, and in the high-rate scenario the publish p50 was about 2 s and applications fell behind by millions of messages. - The medium-memory configuration's broker gets a 4 GB heap instead of 2 GB. The full topology's 20 applications with 100 pods on 30 topics are 60,000 Key_Shared consumers, whose consistent hash rings, with 100 points per consumer, and consumer state leave about 2 GB of live objects: with a 2 GB heap the broker spent its CPU in garbage collection. The configuration now needs about 11 GB of memory available to Docker instead of 9 GB. - iot-telemetry-high-rate.yaml publishes 30,000 messages per second instead of as fast as possible, and its gateways keep at most 10,000 messages in flight instead of 100,000. Without a rate limit, on a host whose CPU runs at a fixed base frequency, the gateways published faster than the applications received, their backlogs grew, and an application that fell behind could stall for tens of seconds; the 100,000 messages in flight alone made the publish latency about 2 s. - The docs describe the memory and the high-rate scenario, and that warmup.messages works with a rate too. Measured twice each with the performance testing environment active: iot-telemetry.yaml has a publish p50 of 1.4 ms and an end-to-end p99 of 47 ms with backlogs of at most 128 messages, iot-telemetry-high-rate.yaml delivers 30,000 messages per second with an end-to-end p99 of at most 40 ms and backlogs under 3,000 messages, and iot-telemetry-small.yaml, which needed no change, has an end-to-end p99 of 11 ms. Assisted-by: Claude Code (claude-opus-5-5)
…mand, and analyze them with jafar-shell Finding what fills a component's heap, such as the broker's in the full IoT topology, needed heap dumps taken by hand inside the containers, at the right time, and the JVMs in the containers write files that only their own user can read. - A scenario's heapDumps section asks for heap dumps of the broker, the gateways and the applications: when the JVM runs out of memory (onOutOfMemoryError), when the gateways start (atStart), at given seconds after that (atSeconds), periodically (everySeconds), at the highest heap usage (atPeakUsage), and after every application has received every message (atEnd, the broker only). configs/heap-dumps-broker.yaml adds the broker's peak, end and out-of-memory dumps to any scenario with --extends. - The launcher writes the dumps with jcmd GC.heap_dump, as the user that the JVM runs as, one at a time, into heap-dumps/<component> in the run directory, which each container mounts, and lists them in heap-dumps/heap-dumps.csv. atPeakUsage samples the heap usage every 5 s and keeps one dump, of a heap at least half full and 10 % above the previous peak dump. onOutOfMemoryError adds -XX:+HeapDumpOnOutOfMemoryError to the JVM's JAVA_TOOL_OPTIONS. When the run ends, also after a failure, a one-off container makes the dumps readable and the launcher prints where they are; the run report lists them in a Heap dumps section. - docs/heap-dumps.md describes the settings, the files and the cost: a dump stops the JVM, so a run with scheduled dumps is for finding what holds the memory, not for measuring. - docs/analyzing-profiles.md describes jafar-shell, the newer version of jfr-shell, which reads JFR recordings and heap dumps with its JfrPath and HdumpPath queries, also from standard input for scripts and agents, and links its heap dump quick start. Its heap dump support isn't in a release yet, so the docs show how to build and install it. Assisted-by: Claude Code (claude-opus-5-5)
…n with the current memory settings The performance testing README's example output of iot-telemetry.yaml showed publish latencies of hundreds of milliseconds, from a run whose broker heap was too small. It now shows the lines of a run with the 4 GB broker heap. Assisted-by: Claude Code (claude-opus-5-5)
…ter honored numProxies HealthCheckTest set numProxies to 0, which PulsarCluster ignored until it started the proxy only when the spec asks for one. Without the proxy, testZooKeeperDown's healthcheck hangs past the test's 300 s timeout, in CI and locally, while with the proxy it passes as it did before. The test now has the default proxy again, the topology it always ran with. Assisted-by: Claude Code (claude-opus-5-5)
…-compressed A scenario's heapDumps.gzipLevel, from 1 to 9, has the JVMs write every heap dump compressed, as .hprof.gz: the launcher's dumps with jcmd GC.heap_dump -gz=<level>, and the dumps on OutOfMemoryError with -XX:HeapDumpGzipLevel=<level>, which JDK 17 and later have, so also the Java 21 of the Pulsar 4 images. Without it the dumps stay uncompressed .hprof files. At level 1, a small broker's dumps came to about a quarter of their size in the time that its uncompressed dumps took. jafar-shell reads uncompressed dumps only, so the docs show decompressing a dump first with gunzip -k. Assisted-by: Claude Code (claude-opus-5-5)
… run directory, and delete launcher.log of successful runs launcher.log holds Testcontainers' and the Pulsar containers' logs, which make it large, and a successful run rarely needs it. The launcher still writes it during the run, so that it can be followed, and a failed run keeps it, but a successful run deletes it at the end, after stopping the logging, which closes the file. --keep-launcher-log, or -Pperformance.keepLauncherLog, keeps it. What the launcher prints on the console goes also to console.log.txt in the run directory, which every run keeps and the run report links: the resolved scenario, the phases, the progress lines, the heap dumps and the failure. Assisted-by: Claude Code (claude-opus-5-5)
…rver, with a configurable base URL When the tests run on another host, a report's URL that opens with a click in the terminal saves finding it through the server's directory listings. The launcher prints the URL of the run report and of each profile report beside its file, as the base URL with the report's path in the reports root appended, without a directory's index.html. -Pperformance.reportsServer.baseUrl sets the base URL, such as http://192.168.1.123:8000/. Without it, it is http://<bind address>:<port>/ of serveReports, and with the bind address 0.0.0.0, the IPv4 address of the host's first network interface, by index, that is up, other than a loopback or link-local one; a host with several interfaces sets the base URL instead. serveReports prints the same base URL when it starts. Assisted-by: Claude Code (claude-opus-5-5)
…ctoriaMetrics and show them on Grafana's Pulsar dashboards A run's report shows the workload's view of throughput, latency and backlog, but not what the brokers, bookies and ZooKeeper went through meanwhile, and runs couldn't be compared on those metrics afterwards. A new module, :tests:performance:metrics, runs VictoriaMetrics and Grafana, with Grafana's image renderer, with Docker Compose from its compose file, as the project pulsar-performance-metrics. ./gradlew :tests:performance:metrics:up starts it in the background, and it keeps running, also across restarts of Docker, until ./gradlew :tests:performance:metrics:down stops it. Its volumes and the network that the runs connect to are external, created with the docker CLI, so that stopping the stack keeps the metrics, the dashboards, the annotations and Grafana's settings. Grafana's setup, when its volume is empty, provisions the VictoriaMetrics data source and the 19 dashboards of the Apache Pulsar Helm chart, pinned to a commit, with an annotation query for the runs. Grafana has the fixed admin password pulsar-performance, and viewing needs no login; the slim Grafana image installs the Prometheus plugin. A run collects its metrics by default, -Pperformance.metrics=false or --no-metrics doesn't: into the running stack, or into the stack that it starts for itself and stops at its end. Its brokers, bookies and ZooKeeper join the metrics network under their container names, and the run's scrape configuration goes into VictoriaMetrics, with the run's directory path in the reports root as the cluster label, and the jobs and instances that the dashboards filter on. The scenario's metrics.intervalSeconds, 5 by default, sets the scrape interval, and the broker's and the bookies' stats periods, which have to match it since the stats start over at each period: statsUpdateFrequencyInSecs, statsUpdateInitialDelayInSecs, managedLedgerStatsPeriodSeconds and managedLedgerPrometheusStatsLatencyRolloverSeconds, and the bookies' prometheusStatsLatencyRolloverSeconds. At the end of a run, the launcher adds its events to Grafana as annotations: the gateways' start, the end of the warmup, the gateways' finish and the applications' finish. It renders eight panels of the Messaging, JVM and BookKeeper dashboards as PNG images into the run directory, which the run report's Metrics section shows with links to the panels and the dashboards over the run, at -Pperformance.metrics.grafanaUrl. metrics.json has the run's selector, time range, jobs, events, and the URLs, credentials and data source with which scripts and agents query VictoriaMetrics and Grafana. A run without metrics, or whose metrics fail, goes on, and its report has no Metrics section. Assisted-by: Claude Code (claude-opus-5-5)
…p dumps, sampling and metrics stack A review of the branch found these, each confirmed before it was fixed: - Profiling a released Pulsar 4.x broker failed: the jonoffcpu agent's JVM options had --sun-misc-unsafe-memory-access=allow, which Java 21 rejects. The agent no longer passes it; Pulsar's scripts add it on Java 23 and later. - The broker's heap dump on OutOfMemoryError went to /var/log/pulsar: the test image's scripts put -XX:HeapDumpPath=/var/log/pulsar on the command line, which overrode JAVA_TOOL_OPTIONS. The broker's options go to PULSAR_EXTRA_OPTS, which come after it. - jcmd inherited a profiled workload's JAVA_TOOL_OPTIONS and started the profiling agent in its own JVM; it runs without them. - With several brokers, the topic stats of the topics that another broker owns redirected to names that the host can't resolve, and weren't sampled. The sampler has a client per broker, and samples each topic through its owner. - A metrics stack whose start failed part way kept running; the failed start is stopped too. - The workload startup's 60 second stall timeout wasn't checked while the container kept logging without progress. - Heap dumps due after the workload's end delayed the end of the run by up to 10 minutes; they are cancelled. - The IoT telemetry reference gave the full topology's broker a 2 GiB heap, not 4 GiB, and the README's setting of the reports root failed when ~/.gradle didn't exist. Assisted-by: Claude Code (claude-opus-5-5)
lhotari
marked this pull request as ready for review
September 27, 2026 18:43
lhotari
marked this pull request as draft
September 27, 2026 18:44
…bring the performance tests' guides up to date The run report's Profiles section named each profile with a link to its profile report, and the off-CPU digest was one more click away. Each profile now has its reports a click away: the profile report, the jonoffcpu report, which is the digest, and the off-CPU, CPU and allocation flame graphs that it has. The README tutorial's steps were run as written and pass. Its example output now has the metrics lines that a run prints, and the agent guide has a quick reference of the commands: the launcher's options with --help, running and profiling a scenario, the reports root, serving the reports and starting and stopping the metrics stack. Its advice on charts in pull requests uses HTML images with a width, linked to the uploaded files, which Markdown images in a table's columns don't display properly. The guides say which platforms the tooling was tested on: Linux and macOS both work, including profiling with async-profiler, jonoffcpu and JDK Flight Recorder, tested on Pop!_OS 24.04 x86_64 with Docker Engine and on macOS arm64 with OrbStack; Docker Desktop and Podman Desktop are untested. Linux x86_64, configured for performance testing, is recommended for measurements, with 32 GB of RAM on the host. Assisted-by: Claude Code (claude-opus-5-5)
…ictoriaMetrics, and link vmui's views The SSH tunnel for reaching the reports server of another machine now forwards Grafana's and VictoriaMetrics' ports too, and docs/run-reports.md shows a ~/.ssh/config host that opens the same tunnel with `ssh <host>-pulsar-perf`. The metrics docs and `metrics:up` link VictoriaMetrics' web UI both for browsing the metrics (vmui/?#/metrics) and for PromQL queries (vmui). Assisted-by: Claude Code (claude-opus-5-5)
… the forwarding The tunnel is silent while it works. The guide now says what it prints when a forwarded service isn't running on the remote machine (channel open failed, connect failed) or a local port is already in use (bind: Address already in use, Could not request local forwarding.), and that -v logs each forwarded port and each connection through the tunnel, as checked in OpenSSH 9.6p1's channels.c and ssh.c. Assisted-by: Claude Code (claude-opus-5-5)
…off-CPU summary) The run report's Profiles section and the profile report both name the off-CPU digest "jonoffcpu report (off-CPU summary)". In the Profiles table it comes first among a profile's links, in bold. In the profile report's files it is bold, and its description starts with a bold "Start here" link to it. The README and the guide to analyzing profiles say that the off-CPU flame graphs show the same blocked time as call trees, to inspect visually which code paths lead to the blocking methods. Assisted-by: Claude Code (claude-opus-5-5)
…phs show, and name the blocked time as such The reports described the off-CPU flame graphs by how they were filtered rather than by what they show, and didn't tell them apart from async-profiler's flame graphs. jonoffcpu produces both: the off-CPU flame graphs from its off-CPU capture, where a box's width is off-CPU time, and the CPU, allocation and other flame graphs from async-profiler's samples in the JFR recording. - The run report's Profiles section introduces both kinds for the profiled process, the broker or a Pulsar client, names its column "Blocked time" and links the blocked time flame graph. - The profile report names the process it profiled. Its "Off-CPU flame graphs: where threads waited" section says what the flame graphs show and which of them matter for throughput, and lists the blocked time first: "Without idle waits" is now "Blocked time", and "from the application's first frame" is "from where threads entered Pulsar or BookKeeper code". A collapsed section explains the details: how idle waits are recognized, how the flame graphs are rooted, and that box widths are the observed time of the recorded waits, which the sampling policy can bias against short waits. - Its "async-profiler flame graphs" section says where their samples come from and what each view shows. - The jonoffcpu report's description in the profile report states its value: the methods where the process's threads blocked, on what and for how long, with the blocked time flame graphs for the code paths that lead to them. - The flame graph pages' titles and the docs use the same names. Assisted-by: Claude Code (claude-opus-5-5)
…d stacks for blocked time The agent guide says to look first at a profile's jonoffcpu report, a text digest of where the Pulsar broker's or the Pulsar client's threads blocked, and to read the full call trees from the collapsed stacks beside it, since the flame graphs' HTML pages need a browser to render. Assisted-by: Claude Code (claude-opus-5-5)
…riptions - The off-CPU section no longer contrasts its widths with async-profiler's sample counts: the allocation and lock views weigh bytes and time. Each async-profiler view states what its widths count. - The lock view includes thread parking, such as idle executor workers, not only waits to enter Java monitors. - `--weights estimated` corrects for the capture's sampling, but waits shorter than the sampling policy's minimum aren't recorded at all. Assisted-by: Claude Code (claude-opus-5-5)
…board that it is on A panel's image opened the panel on its own in Grafana. Below each panel, the run report now also links the dashboard that the panel is on, with the run's cluster, the dashboard's variables and the run's time range, so that the panel can be seen among the others. metrics.json records each panel's dashboard and its URL. Assisted-by: Claude Code (claude-opus-5-5)
…the README The README's table of tested Docker engines described the Linux host's testing more briefly than the macOS one's. Both rows now list runs, profiling with async-profiler, jonoffcpu and JDK Flight Recorder, and metrics, which were tested on both. Assisted-by: Claude Code (claude-opus-5-5)
lhotari
marked this pull request as ready for review
September 27, 2026 22:26
10 tasks
Radiancebobo
pushed a commit
to Radiancebobo/pulsar
that referenced
this pull request
Oct 8, 2026
…ofiling, metrics and AI agent support (apache#26714)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
This PR adds a complete Docker-based performance testing framework to Pulsar, with agent guides that let AI agents run AI-assisted investigations of performance bottlenecks and optimize them. Finding and fixing a bottleneck is a loop of many steps: run a realistic scenario, read its reports, profiles and metrics, form a hypothesis, change the code or a setting, and compare the change with its baseline over repeated runs, since a single run's result isn't reliable. Without an agent, most of that loop is manual work that few contributors have time for.
tests/performance/AGENTS.md, which the repository'sAGENTS.mdroutes to, tells an agent how to check and configure the host, run and profile scenarios, read the reports, profiles and metrics, compare revisions, and report a result with its evidence, so that a contributor can hand an agent an experiment and review what it found.Pulsar had no practical way to measure a change's effect on a realistic workload during development, with a feedback loop quick enough to iterate on. Today's AI agents need such a quick feedback loop to make progress: a run takes minutes on a developer's machine, without a separate benchmark environment, so an agent can try a change, measure it and move on to the next one. The framework runs scenarios that model real-world use cases, currently IoT telemetry, on a Pulsar cluster in Docker containers, with producers and consumers that check every message's delivery, ordering and duplicates. Every run writes a report with its settings, provenance, throughput and backlog, and latency with charts rendered from the HDR histograms of the publish and end-to-end latencies.
Profiling uses jonoffcpu, a profiler for both on-CPU and off-CPU time. For on-CPU time, its bundled async-profiler samples the CPU, allocations and locks into a JDK Flight Recorder (JFR) recording, which also collects the JVM's own events, such as garbage collections. For off-CPU time, an eBPF program in the Linux kernel records when threads stopped running, and jonoffcpu correlates each interval with the Java stack that async-profiler recorded. A broker's throughput limits are often waits: a thread blocked on a monitor, a lock or a queue hand-off. Those waits don't show up in CPU samples, and async-profiler's wall-clock mode samples idle threads as well, which buries the few waits that matter. jonoffcpu ranks the blocked time by the method that waited, so contention such as a dispatcher monitor or an executor queue lock stands out directly.
The metrics stack comes out of the box: VictoriaMetrics and Grafana, configured with Pulsar's dashboards, run with Docker Compose. During each run they collect the metrics of the brokers, bookies and ZooKeeper, and keep them across runs for later analysis, next to each run's JFR recordings and off-CPU profiles. The run's timeline is added to Grafana as annotations: the gateways' start, the end of the warmup, and the gateways' and the applications' finish, so that the dashboards show where the measurement period is. The run report includes images of Grafana panels over the run, such as the publish and delivery rates, the backlog, the JVM's heap and GC, and the bookies' add latency, each linked to Grafana.
A run's results also depend on the host's state: a CPU that runs at turbo frequencies slows down as it heats up and throttles, and daemons that tune power management, such as thermald, change the settings during a run. The reports show the host's temperature and throttling, and scripts configure the host for performance testing and restore it afterwards.
Together, the run reports, the JFR recordings, the off-CPU profiles and the metrics give an agent the evidence it needs to find a bottleneck, and to show that a change improves Pulsar.
Modifications
Profiling:
io.github.jonoffcpu:jonoffcpu-agent, Apache-2.0): its bundled async-profiler records CPU and allocation samples into the JFR recording, and its capture records the off-CPU intervals under the sampling policy of each profiled component'soffCpuOptionsin the scenario'sprofilingsection. After the run the launcher correlates each capture in-process withjonoffcpu-correlator.epollWait, executor workers waiting for a task, JDK and HotSpot service threads), and both again with each stack starting at its first Pulsar or BookKeeper frame (^org\.apache\.(pulsar|bookkeeper)\.), so the same code reached from different thread pools joins into one tree; stacks without such a frame are left out of those two and counted apart. The idle-wait patterns (offcpu-idle-waits.txt) and the BookKeeper and Pulsar frames that only dispatch work (offcpu-dispatch-hide.txt, hidden before rooting) are copied into each output directory.jonoffcpu-summary.md, which ranks the blocked time by the application method that waited, by where threads entered the application and by application method, with the idle waits left out.cpu,wall,alloc,lock) of each measurement recording withjonoffcpu-jfr-converter, each only when its event is in the recording's async-profiler options, as a flame graph, split by thread and as a heatmap over time.dockerBuildWolfitasks that build the Pulsar image and the java-test-image from Wolfi under a<tag>-wolfitag. The profiler's native libraries need glibc, and on Alpine every native frame is the unsymbolized musl loader. The launcher'sprofiletask uses the Wolfi image by default; the Alpine images are unchanged.PULSAR_PROCESS_USER, applied to the supervisord configuration), so that a profiled JVM runs as root and may load the eBPF programs. The existing kernel-tuning task also clearskernel.unprivileged_bpf_disabled.PulsarClusterSpecandPulsarContaineraccept the jonoffcpu agent JAR and sampling options.Reports, in a new module
:tests:performance:report-toolthat the launcher calls once a run has finished:README.mdin the run directory with its HTML pageindex.html, for every run, profiled or not: the run's settings with links to the scenario file and its resolved configuration; where, by whom and from which commit the run was made (host, user, project directory, git branch and commit, Pulsar version, also inrun-info.json); links to each profile's reports (first the jonoffcpu report (off-CPU summary), which is the off-CPU digest, then the profile report and the flame graphs) and the JFR recordings; correctness; producer and delivered throughput; latency percentiles for publish and for each consumer application; and the backlog and publish and dispatch rates, sampled once per second from the broker's topic stats.host-stats.csv: the package and hottest core temperature, the mean and lowest core frequency, the kernel's thermal throttle counters and the fastest fan. The run report sums it up in one line of the settings table and has a Host section with the values at the start and during the measurement, a bold note when the CPU throttled during the measurement, and temperature and frequency charts, so that a run slowed down by a hot host can be told apart.README.mdandindex.htmlin its directory: its files (digest, JFR recordings, capture stream, patterns), the off-CPU flame graphs with their totals, and the views. It names the profiled process, the broker or a Pulsar client, says what jonoffcpu's off-CPU flame graphs and async-profiler's flame graphs each show, lists the blocked time first, and keeps the details in a collapsed section..hgrmfile, HdrHistogram's percentile output format, which HdrHistogram's plotFiles.html reads.Performance Launcher's IoT scenario reporting:
jonoff summary report for application's blocking callsites:
Off-cpu flamegraph example for application's blocking callsites:
Run output layout:
<reports root>/<yyyy-MM-dd>/<branch>/<name>/<MM-dd-HH-mm-ss>/. The reports root isbuild/performanceunless-Pperformance.reportsDirsets it (also in~/.gradle/gradle.properties); the name is the scenario file name, the scenario'soutput.name, or--name;--outputstill writes to an exact directory. The launcher scenarios drop their fixedoutput.directory.README.mdandindex.htmlare the run report, so that a directory of runs can be served over HTTP or pushed to a GitHub repository../gradlew :tests:performance:report-tool:serveReports, ortests/performance/serve-reports.pywithout Gradle, serves a reports root on the loopback interface by default (for an SSH tunnel from another machine), showing the YAML, CSV, HDR logs, logs and Markdown as text. The launcher prints each report's URL on that server, with a configurable base URL.tests/performance/docs/run-reports.mdshows an SSH tunnel, also as a~/.ssh/confighost, that forwards the reports server, Grafana and VictoriaMetrics from another machine.iot-application-0/…), as the report names the application; container logs are saved ascontainer.log.txt.Workload:
HdrLatencyRecorder), so that the HDR logs show latency over the run; merged, the intervals give the same distribution as before.--cooldown-temperature <°C>(or-Pperformance.cooldownTemperature, also in~/.gradle/gradle.properties) makes the launcher wait for the CPU package to cool down to that temperature before starting the cluster, and again after the warmup rounds. For the second wait the producer serves two endpoints with the JDK's built-in HTTP server, which the launcher reaches through the port Testcontainers maps on the host:GET /measurement/ready?waitMillis=<ms>answers as soon as every warmup round has been received, andPOST /measurement/startstarts the measurement.--cooldown-timeoutbounds each wait, the workloads' timeouts are extended by it, and the run report says how long each wait took. Runs that follow each other, as in an A/B comparison, then start from comparable thermal conditions. A wait of 10 s or more before the measurement is cut out of the throughput and backlog charts, whose time axis then breaks, with a small mark, and shows the real time before the break.workloads.iotTelemetry.gateways.envandworkloads.iotTelemetry.applications.envset environment variables of the producer (gateways) and the consumer (applications) containers, ascluster.brokers.envandcluster.bookies.envdo for the brokers and the bookies, for exampleGLIBC_TUNABLES. AJAVA_TOOL_OPTIONSgiven there is appended to the launcher's JVM options, which keeps the heap settings and the profiling agent.Metrics, in a new module
:tests:performance:metrics(described intests/performance/docs/metrics.md):tests/performance/metrics/compose.yamlas the projectpulsar-performance-metrics:./gradlew :tests:performance:metrics:upstarts it in the background, and./gradlew :tests:performance:metrics:downstops it.metrics:upprints the links to Grafana and to VictoriaMetrics' web UI, vmui, for browsing the metrics and for PromQL queries. Its volumes and the network that the runs connect to are external to the project, created with the docker CLI, so that the metrics, the dashboards, the annotations and Grafana's settings outlive the stack.pulsardashboards, pinned to a commit, each with an annotation query for the runs. Grafana runs the slim image with the Prometheus data source plugin, has a fixed admin password, and lets anyone view the dashboards; the stack is published on the loopback interface by default.--no-metricsor-Pperformance.metrics=falsecollects none). Its brokers, bookies and ZooKeeper join the metrics network, and VictoriaMetrics scrapes them with the run's directory path as theclusterlabel, and thejobandkubernetes_pod_namelabels that the dashboards filter on. The scenario'smetrics.intervalSeconds, 5 by default, sets the scrape interval and the brokers' and bookies' stats periods, which have to match it since the stats start over at each period.metrics.json, with the run's label selector, time range, jobs, events, and the URLs, credentials and data source with which scripts and agents query VictoriaMetrics and Grafana. A run whose metrics fail goes on without them.Grafana

Integrated and pre-configured to collect metrics during the test runs
VictoriaMetrics Explore

Metrics backend for Grafana, exploring metrics example:
VictoriaMetrics PromQL

Metrics backend for Grafana, PromQL example:
Other additions:
heapDumpsin a scenario): onOutOfMemoryError, at the start or the end, at given times, periodically and at the highest heap usage, optionally gzip-compressed.tests/performance/docs/heap-dumps.mddescribes them, andtests/performance/docs/analyzing-profiles.mdanalyzing them with jafar-shell.console.log.txt, and deleteslauncher.log, which holds the containers' logs, when a run succeeds (--keep-launcher-logkeeps it).-Pperformance.clusterPulsarImageruns the cluster on a released Pulsar image, to compare the checkout with a release.Host setup, in
tests/performance/environment(described in its README):scripts/configure-perf-test-environment.sh, run as root.installinstalls TuneD on Debian based distros, disables its dynamic tuning, installs theperformance-testingTuneD profile, limits the size of Docker's container logs in/etc/docker/daemon.jsonwhile keeping its other settings, and leaves the TuneD daemon disabled.startchecks that the host is on AC power and warns when the disk is nearly full, stops thermald and, on Pop!_OS,com.system76.PowerDaemon.service, activates and verifies the profile, and skips the:tests:integration:tuneKernelPerfEventstask in the user's~/.gradle/gradle.properties.stopswitches TuneD to thebalancedprofile (RESTORE_PROFILE), which allows power saving, stops TuneD, applies the system's configured dirty page limits and swappiness again, starts the stopped daemons again and removes the Gradle property.performance-testingprofile is based onlatency-performanceand adds: turbo disabled, which fixes the CPU frequency at the base frequency withlatency-performance'smin_perf_pct=100;vm.swappiness=1and NUMA balancing disabled, keeping the host's own dirty page limits; thenoneI/O scheduler, theperformanceACPI platform profile and NVMe power state transitions disabled; and the perf event and BPF settings for profiling, the NMI watchdog disabled and the Transparent Huge Pages settings for-XX:+UseTransparentHugePages, which are left in place when the profile is deactivated.startandstopwithout a password.tuneKernelPerfEventssets the THP defrag mode tomadviseinstead ofdefer, so that-XX:+AlwaysPreTouchgets huge pages at startup instead of depending on memory fragmentation and khugepaged. Theinttest.asyncprofiler.skipPerfEventTuningproperty skips the task only when it is empty ortrue, so that a later value ingradle.propertiescan override an earlier one.Tests:
The launcher and report tool tests under
tests/performance, including the ones that existed before this change, use AssertJ assertions instead of TestNG'sAssert, as the othertests/performancemodules already do.Document the performance tests in
tests/performance/README.md, a tutorial, with reference pages intests/performance/docsandtests/performance/scenarios/docs, and guide AI agents withtests/performance/AGENTS.md, which has a quick reference of the commands and points agents to the jonoffcpu report, which is text, for blocked time, and to the collapsed stacks for the full call trees; pointCONTRIBUTING.mdto it.Verifying this change
This change added tests and can be verified as follows:
RunReportTest,ProfileReportTest,MarkdownPagesTest,HdrHistogramRendererTest,TimeSeriesRendererTest), the run's provenance (RunInfoTest) and the rendered async-profiler views (JfrFlamegraphViewsTest); the launcher's tests cover the run directory layout and index links (RunDirectoryTest);HdrLatencyRecorderTestcovers the per-second latency logs.WorkloadEnvironmentTestcovers the producer and consumer containers' environment variables.bash -nandshellcheck -S warning../gradlew :tests:performance:launcher:profile --args='--scenario tests/performance/scenarios/iot-telemetry-high-rate.yaml --extends configs/profile-broker --extends configs/profile-gateways'completes with correct delivery (all messages, no duplicates, no ordering violations) and writes the reports, off-CPU, flame-graph and digest outputs described in the README. Profiling needs a Docker engine whose kernel has BTF, and privileged containers. It works on Linux and on macOS, and async-profiler, jonoffcpu's off-CPU profiling and JFR were also tested on macOS arm64 with the OrbStack Docker engine; Linux x86_64, configured withtests/performance/environment, is recommended for measurements, since it is Pulsar's main target platform, and dedicated hardware has no noisy neighbours and less thermal and power throttling and CPU frequency variance.ScrapeConfigTest,MetricsSettingsTest,RunReportTest). Runs with the metrics stack started bymetrics:up, and started by the run itself, collected the brokers', bookies' and ZooKeeper's metrics, added the annotations, rendered the panels into the run report and kept the data in the volumes across restarts of the stack.Does this pull request potentially affect one of the following parts:
If the box was checked, please highlight the changes
Dependencies, used only by the performance launcher, its report tool and the integration-test profiling support, and not part of the Pulsar distribution or images:
io.github.jonoffcpu:jonoffcpu-agent,jonoffcpu-correlatorandjonoffcpu-jfr-converter0.8.0 (Apache-2.0);org.commonmark:commonmark,commonmark-ext-gfm-tablesandcommonmark-ext-heading-anchor0.30.0 (BSD-2-Clause); andorg.knowm.xchart:xchart4.0.4 (Apache-2.0), used for PNG only, so that none of its optional dependencies, such as the LGPL VectorGraphics2D behind its SVG export, is pulled in.This PR was prepared with AI assistance (Claude Code) and reviewed by a human contributor.