Skip to content

Migrate tests to JUnit Jupiter and bring them under the CI time budget - #19

Merged
lhotari merged 1 commit into
mainfrom
junit-jupiter-tests
Sep 24, 2026
Merged

lhotari merged 1 commit into
mainfrom
junit-jupiter-tests

Conversation

@lhotari

@lhotari lhotari commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

Why

CI's "Build and run checks" step spent most of its time in one fixture: StreamingCorrelatorTest.scale took 158 s of
a 200 s class on both architectures. The tests also ran on a homegrown main-based framework (FixtureExec,
FixtureSteps, copied check helpers), and the agent's end-to-end checks were Python scripts outside Gradle.

What changes

Test framework. JUnit Jupiter 6.1.3, AssertJ 3.27.7, Awaitility 4.3.0 and the Testcontainers 2.0.5 BOM, versioned
in the catalog and wired through jonoffcpu.java-conventions. Every fixture is now a JUnit test, with table-driven
cases as @ParameterizedTest, the end-to-end classes as @ParameterizedClass over the selected C libraries, and every
wait bounded through Awaitility.

Unit and integration suites. src/test holds pure unit tests that run anywhere with a JDK; src/integrationTest
holds the rest, routed by tag:

Tag Runs in
host-native integrationTest on a Linux host with a selected libc, or containerIntegrationTest<Platform>: the JUnit Console Launcher in the pinned Corretto image of each selected libc, all JARs flat on -cp lib/* (default off Linux, so macOS works; CI's musl jobs use it)
privileged-container integrationTest, through Testcontainers: packaged smoke, shutdown, abrupt exit and async-profiler-first-stop, replacing the four jonoffcpu-agent/tools/run-*.py scripts
packaged-jar packagedJarTest, against the shaded JAR alone
scale scaleTest, in a JVM capped at the heap it proves

Time budget.

  • Scale test: 2M rows and 158 s → 60k rows, exactly 200 distinct stacks and about 9 s. The full size is still
    -PscaleRows=2000000 -PscaleHeap=1g.
  • Degradation ladder: 200k rows → 10k rows. This needed a new public Limits.watermarkRows, with a hidden
    --watermark-rows option; the default stays 65,536.
  • A cold check of every module (all tasks rerun) takes about 70 s locally, for glibc and for musl in container mode.

Deterministic fixtures. The scale test's old "≥ 100 distinct stacks" floor was only met through JIT variants: the
real fan-out was 64, JFR's default stack depth. Now:

  • the fixture's methods stay interpreted,
  • each test class gets a fresh JVM,
  • the scale test raises JFR's stack depth,
  • the ladder budgets a measured share of the fixture's peak retention, because JFR can reorder a few events under load.

CI. The end-to-end tests run inside check, so the separate Python smoke step is gone.

Docs. New CODING.md with the code and test conventions, linked from AGENTS.md and README.md. The agent README
and OFFLINE.md are updated.

Verification

  • ./gradlew :jonoffcpu-agent:check :jonoffcpu-correlator:check :jonoffcpu-jfr-converter:check --rerun-tasks passes
    on x86-64 with -PnativeLibcs=glibc, and with -PnativeLibcs=musl -PintegrationTestsInContainer=true.
  • The fixture-sensitive classes passed repeated runs, alone and inside a full check.
  • spotlessCheck is clean.
  • Not yet verified: arm64 (left to CI), the DuckDB-missing skip path, and FixtureAcceptanceTest against real
    recordings.

Replace the main-based fixtures (FixtureExec, FixtureSteps and per-class
check helpers) with JUnit Jupiter 6.1.3, AssertJ, Awaitility and
Testcontainers, all versioned in the catalog, and split each module into
a `test` suite of pure unit tests and an `integrationTest` suite for the
native bundle, packaged JARs, Docker and external tools, routed by tag:

- host-native tests run in the test JVM on a Linux host, or through the
  JUnit Console Launcher in a pinned Corretto container of each selected
  C library (the default off Linux), with every JAR flat on `-cp lib/*`
- privileged-container tests replace the agent's Python smoke, shutdown
  and interruption scripts with Testcontainers classes that run the
  packaged agent and correlator end to end against the host kernel
- packaged-jar tests run against the shaded JAR alone; the scale test
  runs in a JVM of its own under the heap cap it proves

Scale the slow scenarios down: the scale test drops from 2M rows and
158 s to 60k rows, 200 real distinct stacks and about 9 s (full size
behind -PscaleRows/-PscaleHeap), and the degradation ladder runs on 10k
rows through a new Limits.watermarkRows (hidden --watermark-rows) with a
budget measured from the fixture. Fixtures are made deterministic: the
fixture's frames stay interpreted, each test class gets a fresh JVM, and
JFR's stack depth covers the scale test's recursion.

CI runs the end-to-end tests inside `check` (musl jobs in container
mode) instead of a separate smoke step. CODING.md documents the code and
test conventions and is linked from AGENTS.md and README.md.
@lhotari
lhotari merged commit 8f6cb7a into main Sep 24, 2026
5 checks passed
@lhotari
lhotari deleted the junit-jupiter-tests branch September 24, 2026 13:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant