Skip to content

[improve][test] Improve performance experiments, analysis tooling and agent guidance - #26715

Merged
lhotari merged 22 commits into
masterfrom
lh-key-shared-500x20-scenario
Sep 28, 2026
Merged

lhotari merged 22 commits into
masterfrom
lh-key-shared-500x20-scenario

Conversation

@lhotari

@lhotari lhotari commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Motivation

Follow up on #26714, which introduced the Docker-based performance testing framework. Make the framework easier to use for repeatable development experiments and AI-assisted performance analysis: explain how to run a controlled comparison, find and interpret the evidence, and use the analysis tools without hunting for JARs. Add a maximum-rate IoT telemetry scenario for focused dispatcher and managed-ledger investigations.

Modifications

Experiment and analysis guidance

  • Expand tests/performance/AGENTS.md into a working reference for the experiment loop: establish a baseline, profile, form a hypothesis, change one thing, and compare repeated runs. Document commands, outputs, correctness checks, platform limitations, and the evidence to include in a handoff.
  • Map run reports, metrics, JFR recordings, off-CPU profiles, and heap dumps to the tools that read them. Clarify measurement windows, normalization, profile weights, and the distinction between diagnostic profiles and unprofiled measurements.
  • Document automated analysis of collapsed (also called folded) stacktrace files using DuckDB's quack_flamegraph community extension. Include parameterized SQL for stacks, method coverage, call edges, and attribution to application frames, plus JSON and CSV output for agents and scripts.
  • Link the analysis examples from the performance README and the integration test agent guide. Describe jfr-converter, flameshow, and Inferno as additional ways to inspect or render collapsed stacks.

Analysis commands

  • Add :tests:performance:report-tool:runJonoffcpuCorrelator and :tests:performance:report-tool:runJfrConverter. Both use the existing pinned dependencies, accept arbitrary CLI options through --args, and run without starting a cluster or requiring a separate tool installation.
  • Document repository-relative paths and the shared performance.profile.maxHeapSize setting, and replace manual JAR commands with the Gradle tasks.

Host validation

  • Check Docker disk usage from a disposable container so validation also works when Docker runs in a VM, including on macOS. Keep host tuning checks Linux-specific.
  • Return documented exit-code bits for insufficient disk space, unavailable Docker or unreadable disk usage, and failed host configuration checks. Explain the disk-check image, its possible download, and the scope of the check.

Maximum-rate scenario

  • Add tests/performance/scenarios/iot-telemetry-max-rate.yaml, the standalone launcher's counterpart of the legacy runner's key-shared-500x20 workload, named in the IoT domain's terms.
  • Run 500 gateways against one topic, consumed by one application with 20 Key_Shared pods. Use no rate limit, up to 100,000 messages in flight, unbatched 128-byte messages, one million warmup messages, and four million measured messages.
  • Use single-copy ledgers (ensemble, write quorum, and ack quorum of 1). This is a focused dispatch experiment, not capacity guidance for a durable deployment.
  • Document the scenario and use the existing profiling overlays; no separate profiling scenario is needed.

Usage

Run the maximum-rate scenario without profiling:

./gradlew :tests:performance:launcher:run \
  --args='--scenario tests/performance/scenarios/iot-telemetry-max-rate.yaml'

Profile the broker and gateways using the same scenario:

./gradlew :tests:performance:launcher:profile \
  --args='--scenario tests/performance/scenarios/iot-telemetry-max-rate.yaml --extends configs/profile-broker --extends configs/profile-gateways'

Explore the analysis tools' options:

./gradlew -q :tests:performance:report-tool:runJonoffcpuCorrelator --args='top --help'
./gradlew -q :tests:performance:report-tool:runJfrConverter --args='--help'

See tests/performance/AGENTS.md for the experiment workflow and tests/performance/docs/analyzing-profiles.md for analysis examples.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

Performance test tooling, scenarios, host setup scripts, and documentation only; no production Pulsar behavior changes.

This PR was prepared with AI assistance (Claude Code and Codex).

@lhotari
lhotari added this pull request to stack #26718 September 25, 2026 19:28
@lhotari
lhotari marked this pull request as draft September 25, 2026 19:29
@lhotari
lhotari marked this pull request as ready for review September 25, 2026 19:30
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch 2 times, most recently from fc3e32c to 03f3632 Compare September 25, 2026 22:22
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch 2 times, most recently from 99a7595 to ffefc57 Compare September 26, 2026 12:02
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch from ffefc57 to bf8cc19 Compare September 26, 2026 12:32
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch from bf8cc19 to 5c7a932 Compare September 26, 2026 17:19
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch 3 times, most recently from 3f902cb to dc2fc42 Compare September 26, 2026 20:18
@lhotari
lhotari removed this pull request from stack #26718 September 26, 2026 20:22
@lhotari
lhotari added this pull request to stack #26722 September 26, 2026 20:22
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch 2 times, most recently from 793a55f to 2d4b341 Compare September 27, 2026 12:30
…files

Assisted-by: Claude Code (claude-opus-5-5)
…red scenario settings

The performance scenarios' cluster, IoT workload and profiling settings were restructured, and the launcher rejects
settings that no longer exist, so the Key_Shared 500x20 scenarios failed to start.

- iot-key-shared-500x20.yaml sets the brokers' single-copy ledgers under cluster.brokers.env, and the measured
  messages, the payload size, unbatched producers and one application with 20 pods with the nested workload
  settings.
- iot-key-shared-500x20-profile.yaml extends the scenario with the broker and gateway profile files, which profile
  the same components as before.

Assisted-by: Claude Code (claude-opus-5-5)
…the high-rate scenario got a rate limit

iot-telemetry-high-rate.yaml, which the Key_Shared 500x20 scenario extends, now publishes 30,000 messages per second
with at most 10,000 in flight. The Key_Shared 500x20 scenario measures the broker's dispatch at saturation, so it sets
the settings it inherited before: no rate limit, and up to 100,000 messages in flight.

Assisted-by: Claude Code (claude-opus-5-5)
@lhotari
lhotari force-pushed the lh-key-shared-500x20-scenario branch from beb39f3 to 1d34213 Compare September 28, 2026 07:23
@lhotari
lhotari marked this pull request as draft September 28, 2026 07:52
…-rate in the IoT domain's terms

Rename iot-key-shared-500x20.yaml to iot-telemetry-max-rate.yaml, after its rate: 0, and describe it in the
domain's terms: the high-rate scenario's 500 gateways publishing to one topic without a rate limit, consumed by one
application with 20 pods. The single-copy ledgers are explained as writing each entry once while the three bookies
share the ledgers, which balances their write load.

Remove iot-key-shared-500x20-profile.yaml: profiling adds configs/profile-broker and configs/profile-gateways with
--extends, as the other scenarios do. List the scenario in the scenarios' README and the IoT telemetry guide, and
point the legacy runner's key-shared-500x20 to it.

Assisted-by: Claude Code (claude-opus-5-5)
…s' agent quick reference

The host check is optional and applies only to Linux hosts, where it checks the configuration that keeps the
run-to-run variance low, so it goes last instead of first. The metrics stack's row gives the addresses of Grafana
and VictoriaMetrics.

Assisted-by: Claude Code (claude-opus-5-5)
…gent guide names

Link the README's "Read the report" section, the comparing-revisions, run-reports and analyzing-profiles pages,
which the guide named by their paths, and the jafar-perf plugin and the Jafar MCP server.

Assisted-by: Claude Code (claude-opus-5-5)
…, with a map of a run directory's inputs

Separate the guide's concerns into sections with tables: the checks before running, the platforms and what their
results are good for, the rules for using a run's results, the inputs for analysis in a run directory, the
analysis tools, and tuning experiments.

The inputs map every file of a run, a profiled component and a heap dump to what it has and what reads it,
including the off-CPU and async-profiler collapsed stacks, the flame graphs, the JFR recordings, the jonoffcpu
capture and stack profile, the latency logs, the sampled stats, metrics.json and the heap dumps. The tools table
lists what each analysis tool reads, what it is for and where to get it, linking the page that describes it.

Assisted-by: Claude Code (claude-opus-5-5)
… Docker's disk on macOS, with an exit code per kind of failure

"validate" checks that Docker is available and that Docker's disk has space on every operating system, and the
host's configuration on Linux only. Where Docker's data directory isn't on the host, as with Docker in a virtual
machine on macOS, it reads the disk's usage in a container, whose root file system is on the same disk as Docker's
data; df -P replaces GNU df's --output.

Its exit code tells the kinds of failed checks apart, as the sum of 2 for Docker's disk, 4 for Docker being
unavailable and 8 for the host's configuration, so that scripts and agents can act on each. The docs list the exit
codes, and no longer say that validate runs without root.

Assisted-by: Claude Code (claude-opus-5-5)
… guide as a table

List each host operating system and Docker engine with whether it was tested, whether off-CPU profiling works on
it, what configure-perf-test-environment.sh validate checks there and what its results are good for, following the
tested Docker engines in the README.

Assisted-by: Claude Code (claude-opus-5-5)
…ance tests' agent quick reference

Assisted-by: Claude Code (claude-opus-5-5)
…figure-perf-test-environment.sh

Always read how full Docker's disk is with df in a container, whose root file system is on the same disk as Docker's
data and the bookies' ledgers, instead of from Docker's data directory on the host, which isn't there when Docker
runs in a virtual machine. docker info only names the data directory in the output.

Assisted-by: Claude Code (claude-opus-5-5)
…exit code as a bit mask, and link it from the agent guide

Describe the exit code as a bit mask in which each kind of failed check sets its bit, naming each bit by its
position and its value: bit 1 (2) for Docker's disk, bit 2 (4) for Docker being unavailable and bit 3 (8) for the
host's configuration. The table has a heading of its own, Exit codes, and the agent guide refers to it when validate
fails, instead of explaining the bits itself.

Assisted-by: Claude Code (claude-opus-5-5)
Describe the experiment loop, command outputs, artifact and tool inputs, and reproducible evidence handoff. Clarify profiling tasks, measurement validity, and metric interpretation.

Assisted-by: Codex
Clarify unlimited-rate prerequisites, metric windows, profiling and platform constraints, and optional analysis tooling. Add cool-down guidance and distinguish duplicate warnings from failed delivery checks.

Assisted-by: Codex
Expose the correlator and JFR converter CLIs using pinned dependencies and document arbitrary argument passthrough.

Assisted-by: Codex
Describe jfr-converter, flameshow, Inferno and DuckDB quack_flamegraph with usage examples.

Assisted-by: Codex
Move SQL examples into the profile analysis guide, parameterize profile and package filters, and add coverage, call-edge and leaf attribution queries.

Assisted-by: Codex
Link DuckDB and quack_flamegraph from the integration agent guide, clarify collapsed and folded stack terminology, and document JSON and CSV output for automation.

Assisted-by: Codex
Make automated collapsed stack analysis discoverable, improve documentation flow and interpretation guidance, and add command-line quack_flamegraph examples.

Assisted-by: Codex
@lhotari lhotari changed the title [improve][test] Add a launcher scenario for 500 producers and a 20-member Key_Shared subscription [improve][test] Add an IoT telemetry scenario for maximum-rate measurements Sep 28, 2026
Bring the performance experiment guidance, analysis tools and host validation improvements into the maximum-rate scenario follow-up to #26714.

Assisted-by: Codex
@lhotari lhotari changed the title [improve][test] Add an IoT telemetry scenario for maximum-rate measurements [improve][test] Improve performance experiments, analysis tooling and agent guidance Sep 28, 2026
@lhotari
lhotari marked this pull request as ready for review September 28, 2026 13:51
@lhotari
lhotari merged commit f925a25 into master Sep 28, 2026
52 of 55 checks passed
@lhotari
lhotari deleted the lh-key-shared-500x20-scenario branch October 1, 2026 00:02
@lhotari lhotari added this to the 5.0.0 milestone Oct 1, 2026
Radiancebobo pushed a commit to Radiancebobo/pulsar that referenced this pull request Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants