Skip to content

Lead the README with what jonoffcpu is for, and run CI's JVM tests once - #32

Merged
lhotari merged 10 commits into
mainfrom
readme-value-proposition
Sep 25, 2026
Merged

lhotari merged 10 commits into
mainfrom
readme-value-proposition

Conversation

@lhotari

@lhotari lhotari commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

The README opened with a one-line description, then about 500 lines of theory, kernel settings and Docker before the quick start. It never said:

  • why someone would pick jonoffcpu
  • how it relates to async-profiler's own lock profiling
  • where the project is going

CI ran the JVM tests four times, once in every architecture × C-library job, and built every native bundle even for a README typo.

README and docs/

The README is now an overview of about 400 lines, in the order a newcomer asks:

  • What jonoffcpu is: a profiling toolset that runs async-profiler and extends it with kernel-measured off-CPU profiling.
  • Six value points:
    • exact durations
    • one recording with every source, from an async-profiler fork that tracks upstream master
    • a small observer effect
    • waiting for work kept apart from blocked time
    • outputs made for automation
    • glibc/musl on x86-64 and arm64
  • What is off-CPU profiling: a thread is running, waiting for a CPU, or blocked, with the scheduler's timeline diagram.
  • Why jonoffcpu:
    • in a system like Apache Pulsar, millions of events per second flow through hundreds of threads;
    • the waits that matter are a small fraction of the off-CPU time;
    • bottlenecks differ between usage scenarios;
    • every change has to be compared with the previous run in each scenario, which does not scale without automated performance analysis. Solving that is the core of the vision in the next section.
  • Where jonoffcpu is going:
    • the Apache Pulsar origin (linking tests/performance);
    • extracting what matters from the JFR, async-profiler and eBPF data of one run;
    • normalizing stacks and detecting application and library boundaries from the data;
    • integrating with tools such as Jafar and jafar-perf-box rather than competing with them.
  • How a recording works, with the architecture diagram:
    • jfrsync records the JDK's Flight Recorder events, async-profiler's samples and the signal samples into one JFR file;
    • JDK Mission Control and async-profiler's converter read that file as usual;
    • the correlation stream sits beside it.
  • Complementing async-profiler and JFR:
    • what lock=, nativelock=, JFR's thresholded wait events and wall-clock profiling cover;
    • what only scheduler-measured off-CPU time sees: I/O, queues, futures, native futexes, run-queue delay;
    • a question → tool table.
  • Quick start: prerequisites and a one-line kernel check, one sysctl command, then four steps from the JARs to the digest, the flame graph and JDK Mission Control.
  • Documentation index, then Project status: pre-1.0.

The details moved, mostly verbatim, to:

  • docs/off-cpu-profiling.md
  • docs/how-it-works.md
  • docs/setup.md
  • docs/capture.md
  • docs/analysis.md
  • docs/automation.md
  • docs/building.md

Across the README and the docs pages, concepts link at first use to their common definitions: Wikipedia, the HotSpot glossary, JEPs, the JDK API docs, and the kernel docs and man pages. Examples:

  • process states, preemption, the run queue
  • futexes, page faults, safepoints
  • BTF, CO-RE, the BPF ring buffer
  • the sampling policies as Poisson/PPS and Bernoulli sampling
  • the population estimate as the Horvitz–Thompson estimator

docs/recording.md is new and covers:

  • what jfrsync adds, and which JDK events it turns off when async-profiler records the same kind;
  • the events by source;
  • choosing the options;
  • JDK Mission Control, flame graphs of the other events, and the jfr tool.

The release tag and the Maven coordinates that release.yml rewrites stay in the README.

Fixed while moving:

  • Removed the pre-1.0 "earlier releases" compatibility notes, here and in OFFLINE.md.
  • --waiting-from / --hide-from defaults are now given as preset:*.
  • file= in asyncProfilerOptions is optional and may be relative.
  • Added the agent keys the options table left out.
  • The retention figure is now OFFLINE.md's.
  • Export is described once instead of three times.
  • --on-limit help listed a "coarsen" step that doesn't exist; the help snapshots are updated.

Documentation check

DocumentationTest replaces the README check in CommandLineTest. It runs as readmeTest over the repository's Markdown, which a typed CommandLineArgumentProvider declares as its inputs. It checks that:

  • no document repeats a heading;
  • the options tables and correlator sections name only real options, with the parser's defaults;
  • every java -jar jonoffcpu-correlator.jar … command, in a code block or inline, uses only its own subcommand's options;
  • every relative link, image and anchor resolves.

Test categories and CI

Two lifecycle tasks:

  • jvmCheck (root): formatting and every test that needs only a JDK. It runs on any OS and never builds a native bundle.
  • :jonoffcpu-agent:nativeTest: the agent's integration tests (host-native and privileged-container) for the selected bundles.

check is unchanged.

build-and-verify.yml (shared with releases):

  1. Unit tests: cargo fmt, DuckDB, jvmCheck. Runs once.
  2. The four native jobs: after that, each builds its bundle and runs only nativeTest.
  3. Package combined artifacts: also runs the agent's packaged-JAR test, whose JAR embeds all four bundles.

ci.yml:

  • changes (dorny/paths-filter v4.0.3, pinned, with .github/changes-filter.yaml) marks a change docs-only when every file is Markdown outside src/ or under docs/.
  • build-and-verify is skipped for a docs-only change.
  • Documentation check runs readmeTest when the change touches documentation.
  • All checks passed accepts exactly those two skips: the build for a docs-only change, and the documentation check for a change without docs. Its name is unchanged, so branch protection needs no edit.
  • The reusable workflow's caller is renamed "Build and verify", since it now runs the unit tests too.

Compatibility

  • README anchors: links from outside the repository to the old section anchors now land on the README without reaching the section. The content is under docs/.
  • CI latency: the native jobs now start after the JVM job, which adds its duration to CI's latency. In exchange, each run skips three redundant runs of the JVM tests.

Verification

  • jvmCheck: passes, also with --rerun-tasks. The second run reuses the configuration cache, and --dry-run lists no native build.
  • Negative docs check: a deliberately broken link, anchor and option each fail with a message naming it.
  • nativeTest: passes on x86-64 for glibc on the host and for musl in containers, including every privileged end-to-end test.
  • Agent packaging: :jonoffcpu-agent:packagedJarTest and verifyRuntimeJar pass.
  • Workflows: actionlint (the pinned image) and zizmor report no findings.
  • To confirm on this PR: the docs-only skip.

The README opened with a one-line description and 500 lines of theory,
kernel settings and Docker before the quick start. It never said why
someone would pick jonoffcpu, how it relates to async-profiler's own
lock profiling, or where the project is going, and it had drifted from
the code.

The README is now an overview of about 300 lines: what jonoffcpu is
(a toolset that runs async-profiler and extends it with kernel-measured
off-CPU profiling), its value in six points, how off-CPU profiling
complements async-profiler's lock, native-lock and wall-clock modes, a
quick start, a documentation index, the direction (automated
performance optimization and tuning, from its Apache Pulsar origin,
integrating with tools such as Jafar and jafar-perf-box), and the
pre-1.0 status. The details moved to docs/: off-CPU concepts, how it
works, host setup, capture configuration, analysis, AI agents and SQL,
and building and testing. The release tag and the Maven coordinates
that release.yml rewrites stay in the README.

Fixed while moving: the pre-1.0 compatibility notes ("earlier releases")
here and in OFFLINE.md, the --waiting-from and --hide-from defaults
(preset:*), file= in asyncProfilerOptions (optional, may be relative),
the agent keys the options table left out, the correlator's retention
figure (now OFFLINE.md's), export described three times, and a help
text that listed a "coarsen" step --on-limit degrade does not have.

DocumentationTest replaces the README check in CommandLineTest and runs
as readmeTest over the repository's Markdown, which a typed argument
provider declares as its inputs: the options tables and sections name
only correlator options with the parser's defaults, every correlator
command uses only its own subcommand's options, and every relative link
and anchor resolves.

CI ran :jonoffcpu-agent:check and :jonoffcpu-correlator:check in all
four architecture and C-library jobs, so the JVM tests ran four times,
and a documentation change built every native bundle. Two lifecycle
tasks now name the test categories: jvmCheck (formatting and every test
that needs only a JDK, on any OS) and :jonoffcpu-agent:nativeTest (the
agent's integration tests, which need the selected bundles). CI runs
jvmCheck once, with cargo fmt and DuckDB, then the four native jobs run
only nativeTest, and the package job adds the agent's packaged-JAR test.
A changes job (dorny/paths-filter v4.0.3 and .github/changes-filter.yaml)
skips the build for a documentation-only change, a documentation check
job always runs readmeTest, and "All checks passed" accepts exactly
that skip.
The README's "Why off-CPU profiling" opened with a request-latency
example and put the architecture diagram next to it. It now reads in
the order a newcomer asks:

- What is off-CPU profiling: a thread is running, waiting for a CPU, or
  blocked; the scheduler's timeline diagram; what a CPU profiler misses
  and how jonoffcpu pairs the kernel's interval with the Java stack.
- Why jonoffcpu: in a system like Apache Pulsar, millions of events per
  second flow through hundreds of threads, the waits that matter are a
  small fraction of the off-CPU time, and every change has to be
  compared with the previous run, which needs automation.
- Where jonoffcpu is going, moved up: extracting what matters from the
  JFR, async-profiler and eBPF data of one run, normalizing stacks and
  detecting application and library boundaries from the data, and
  integrating with tools such as Jafar.
- How a recording works, with the architecture diagram: jfrsync records
  the JDK's own Flight Recorder events, async-profiler's samples and the
  signal samples into one JFR file, which JDK Mission Control and the
  converter read as usual, beside the correlation stream.
- Complementing async-profiler and JFR, which now includes JFR's own
  thresholded wait events.

docs/recording.md describes the recording: what jfrsync adds and which
JDK events it turns off when async-profiler records the same kind, the
events by source, choosing the options, JDK Mission Control, flame
graphs of the other events (moved from docs/analysis.md), the jfr tool
and the correlation stream. The quick start records lock contention
too and points at JDK Mission Control. The off-CPU concepts page uses
examples closer to a messaging system than a JDBC query, and the
repository layout links the async-profiler submodule to its branch on
GitHub.
The documentation check ran for every change, independently of Detect
changes. It now needs Detect changes and runs when the change touches
documentation; a change without documentation skips it, since the unit
tests run the same readmeTest in jvmCheck. "All checks passed" accepts
that skip as well as the build's skip for a documentation-only change.

The jobs are renamed for what they are: "Unit tests" (was "Formatting
and JVM tests"), "Package combined artifacts" (was "Package combined
Java artifacts"), and "Build and verify" for the reusable workflow's
caller (was "Native build and verification"), which now runs the unit
tests too.
The README's three thread states hid two kinds of blocking: a thread
waiting for work to arrive, and a thread held by the OS or the JVM (a
page fault that reads from disk, a safepoint). They are now rows of a
table that says, for each state, how jonoffcpu records it: not off-CPU
time, run-queue time of reason runnable or preempted (or after a blocked
interval's wakeup), or reason blocked, which the analysis sets apart as
waiting or ranks as blocked. The off-CPU concepts page says the same
under "Why the thread left the CPU", with links for page faults and
safepoints.
The README and the docs pages now link each concept, where it is
introduced, to its common definition:
- the scheduler's process states (running, ready, blocked), context
  switches, preemption, the run queue, futexes and page faults
- eBPF, CO-RE, BTF, tracepoints, the BPF ring buffer, signals,
  capabilities and the kernel's sysctl, schedstat, delay-accounting and
  CFS bandwidth-control docs
- the HotSpot glossary's safepoints, JIT compilers, JNI, TLABs and
  garbage collection, JFR (JEP 328) and hidden classes (JEP 371)
- the event loops, thread pools, locks and monitors that threads wait on
- the admission policies as Poisson sampling with probability
  proportional to size and Bernoulli sampling, and the population
  estimate as the Horvitz-Thompson estimator, an inverse probability
  weighting

The off-CPU concepts page also maps Java's Thread.State onto the
kernel's view: a blocking socket read is RUNNABLE to Java and blocked to
the kernel.

The README's "Why jonoffcpu" gives the Apache Pulsar example without
numbers, says that bottlenecks differ between usage scenarios, and that
profiling, optimizing and weighing trade-offs across many scenarios does
not scale without automated performance analysis. "Where jonoffcpu is
going" makes solving that the core of the vision. A table row says
"Preempted by the kernel while running", and every mention of the
project says Apache Pulsar.

The previous commit left the old "Where jonoffcpu is going" behind when
it moved the section up. DocumentationTest now checks that no document
repeats a heading, which fails on that README.
…e CPUs

Its thread pools are sized from Runtime.getRuntime().availableProcessors(),
often twice that, so a broker runs roughly 50 to 200 threads depending on
the CPU count, not hundreds.
"Compare across runs and scenarios" now says what jonoffcpu already
supports (top --baseline per unit of work, export with run labels and
metadata) and where the rest can come from: the generic parts of Apache
Pulsar's performance scenarios and community contributions, such as
load generator integrations and test report generators, as subprojects
of the jonoffcpu organization.

"Integrate, don't compete" becomes "Complement existing tooling, and
plug into it": the recording as an ordinary JFR file, the outputs as
documented messages, integration through other tools' plugins, and
coding agent skills and plugins that know the jonoffcpu workflow, as
jafar-perf-box does for Jafar. The section ends by inviting ideas and
contributions through issues.
"Using the artifacts as libraries" says that the artifacts are on Maven
Central under the group id io.github.jonoffcpu, linking the namespace,
and adds a Maven example beside the Gradle one. The Maven example keeps
the version in one jonoffcpu.version property, which the release job's
README update now rewrites and checks along with the Gradle coordinates
and the download tag.
"Measured, not estimated" said that each recorded wait is exact, but not
that a capture usually records only a sample of them to keep the
observer effect small, so the totals are estimates. The first value
point now links the two: the eBPF program times each interval, a
capture records a duration-weighted sample, each recorded wait carries
its exact duration, reason, sleeping and run-queue split and the
threshold it was drawn against, and the thresholds turn the sample into
unbiased (Horvitz-Thompson) estimates of the total off-CPU time. The
observer-effect point keeps the kernel-side mechanics and no longer
calls the estimate exact.

The intro, "Complementing async-profiler and JFR", the recording page
and the off-CPU concepts page likewise say that every kind of wait is
covered and each sampled wait is recorded with its exact duration,
rather than that every wait is recorded.
@lhotari
lhotari merged commit cc482da into main Sep 25, 2026
10 checks passed
@lhotari
lhotari deleted the readme-value-proposition branch September 25, 2026 13:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant