Skip to content

Add top and summarize: ranked tables and an analysis digest - #12

Merged
lhotari merged 1 commit into
mainfrom
top-and-digest
Sep 23, 2026
Merged

lhotari merged 1 commit into
mainfrom
top-and-digest

Conversation

@lhotari

@lhotari lhotari commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Third PR of the 0.5.0 correlator stack (GitHub stack #15; spec: top-and-digest.md). Based on #11 (export).

What changes

  • top ranks where off-CPU time went. It works on a --profile or any --collapsed-input and shares the selection, filter and transform options of stacks.
    • Intervals are idle when a frame of any stack matches --idle/--idle-from (e.g. preset:jvm-idle) and busy otherwise. Idle time gets its own table instead of disappearing.
    • --by boundary (default): deepest --app frame + blocker, the entry frame of the wait machinery below it (--machinery-from, default preset:jvm-wait-machinery). Busy time without an application frame is broken down by thread pool.
    • --by self|method|class|package|pool rank the other ways.
    • Columns: seconds, share, intervals, estimated seconds and the sleeping/run-queue split when available, dominant reason, heaviest caller.
    • Totals add up, and an over-exclusion check counts idle entries that still waited on a lock or monitor.
    • --format md|json|csv. JSON is the source and includes the reproduce command; md and csv are rendered from it.
  • top --baseline compares two runs by application boundary, per unit of work with --units/--baseline-units. It warns when shares are length-biased (sampled runs without estimates) and when the runs' native symbolization differs. It refuses observed weights when both runs have estimates.
  • summarize and the digest jonoffcpu-summary.{md,json} (schema v1):
    • Contents: capture coverage and losses from the report, where the time went, the busy, pool and idle tables, the heaviest transformed stacks, and a reproduce command for each table.
    • The Markdown is rendered from the JSON.
    • Correlation writes the digest by default (--summary-output false skips it). A failure goes into the report's digest.error and never fails the correlation.
  • The packaged smoke now also checks that the report names the digest.

Verification

  • ./gradlew spotlessCheck :jonoffcpu-correlator:check -PjonoffcpuFixtures=… passes.
  • Packaged agent smoke passed on x86-64 glibc (1,496 matched samples) with the digest check.
  • The worked example is reproduced exactly (new checks in testFixtureAcceptance):
    • Totals: all 2,791 / 308,777 / 8,519.334 s; idle 660 / 301,389; busy 2,131 / 7,388 / 48.968 s; no application frame 132 / 237 / 28.374 s; over-exclusion 18 / 18 / 0.047 s.
    • The top 10 boundary + blocker rows with seconds and intervals, e.g. internalConsumerFlow + C2 Runtime complete_monitor_locking 11.982 s / 4,245.
    • Pools ZDriverMinor 22.517 s (88) and ZDriverMajor 5.527 s (27).
    • Idle [no application frame] 7,687.8 s, internalTakeAll 231.2 s, and so on.
  • Comparison example (Wolfi vs Alpine, 5 M messages each): 2.034/2.396, 0.735/0.989, 0.091/0.161, 0.116/0.124, 0.064/0.101 s per M; shares 61.4/58.2 %, …; busy application time 16.574 vs 20.593 s; busy total 2,019.121 vs 48.968 s; both warnings printed.
  • top agrees with stacks --leaf-at '^org\.apache\.' on all 57 boundaries.
  • Digest on the fixture: 15,841 bytes of Markdown (< 16 KB), heaviest stacks 78 lines at a mean depth of 4.1.
  • TopTest covers the spec's unit cases:
    • the boundary past interleaved library frames
    • the pool table, and an idle entry that is not counted as busy
    • --by method counting recursion once
    • --weights estimated refused without an estimate, and the --baseline warning
    • md, json and csv agree
    • the digest is byte-stable and its Markdown is rendered from its JSON

Choices the spec left open

  • The over-exclusion "lock-acquire frame" is ^java\.util\.concurrent\.locks\..*\.(lock|acquire)\w*$|complete_monitor_locking. Of the definitions tried, it is the one that yields the spec's 18 / 18 / 0.047 s.
  • Baseline rows are keyed by boundary only, not boundary + blocker. The spec's comparison numbers (e.g. isDuplicateNormal 0.161 s/M = 0.805 s) only match that way.
  • The "unresolved-native share" warning compares the share of busy time in stacks with an unsymbolized /lib/… frame.
  • The digest uses --canonical-names, so callers read $$Lambda.run. Its heaviest stacks use --package-names drop, like the spec's … example, which keeps the Markdown under 16 KB. Both options appear in its reproduce commands.
  • top --by boundary without --app is a usage error pointing to --by self. summarize, like the correlation default, falls back to --by self after --collapse-leaf.
  • Spec 5's accountedLoss is shown in the digest's capture section when the report has it (Account for counted sequence contention in population estimates #9).

@lhotari
lhotari added this pull request to stack #14 September 23, 2026 20:16
@lhotari
lhotari removed this pull request from stack #14 September 23, 2026 20:17
@lhotari
lhotari added this pull request to stack #15 September 23, 2026 20:22
Base automatically changed from export-for-sql to main September 23, 2026 20:36
Deciding what to optimize from a capture took a DuckDB session and six
scripts in the Apache Pulsar analysis: attribute each wait to the deepest
application frame and the lock below it, set idle waits apart without
losing them, break the rest down by thread pool, and compare runs per unit
of work. The correlator now does that itself.

- top ranks a profile (or any collapsed file) with the selection, filter
  and transform options of stacks. An interval is idle when a frame of any
  of its stacks matches --idle/--idle-from, busy otherwise; idle time gets
  its own table instead of disappearing. --by boundary (the default) keys
  busy rows by the deepest --app frame and its blocker, the entry frame of
  the wait machinery below it (--machinery-from, default
  preset:jvm-wait-machinery), and breaks busy time without an application
  frame down by thread pool. --by self, method, class, package and pool
  rank otherwise. Every total adds up, and an over-exclusion check counts
  idle entries that still waited on a lock or monitor.
- Columns: seconds, share, intervals, estimated seconds and the
  sleeping/run-queue split when the profile has them, the dominant reason
  and the heaviest caller. --format md|json|csv; JSON is the source and
  names the command that reproduces it.
- top --baseline compares two runs by application boundary, per unit of
  work with --units/--baseline-units. It warns when observed time is
  length-biased by sampling and when the runs' native symbolization
  differs, and refuses observed weights when both runs have estimates.
- summarize writes jonoffcpu-summary.json and .md: the capture's coverage
  and losses from the report, where the time went, the busy, pool and
  idle tables, the heaviest transformed stacks, and each table's
  reproduce command. The Markdown is rendered from the JSON. Correlation
  writes it by default (--summary-output false skips it); a failure to
  write it is recorded in the report's digest object and never fails the
  correlation.

On the Pulsar broker fixtures every number of the spec's worked example
and comparison is reproduced exactly: 2,131 busy entries, 7,388 intervals,
48.968 s, of which 28.374 s has no application frame (ZDriverMinor 22.517 s,
ZDriverMajor 5.527 s); internalConsumerFlow on C2 Runtime
complete_monitor_locking 11.982 s / 4,245 at the top; an over-exclusion of
18 entries, 0.047 s; and Alpine against Wolfi 2.034 vs 2.396 s per million
messages. top's boundary rows agree with stacks --leaf-at on all 57
boundaries, and the digest's Markdown is 15,841 bytes.
@lhotari
lhotari merged commit dea8da8 into main Sep 23, 2026
4 checks passed
@lhotari
lhotari deleted the top-and-digest branch September 24, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant