Repository navigation
Add CI (Azure Pipelines + CodeQL) with build/test/coverage, and test-bench fixes needed to get there - #2
Closed
Thrameos wants to merge 4 commits into
Closed
Add CI (Azure Pipelines + CodeQL) with build/test/coverage, and test-bench fixes needed to get there#2Thrameos wants to merge 4 commits into
Thrameos wants to merge 4 commits into
Conversation
- tools/run_tests.sh, run_tests.bat: fix a dead CLS_PATH variable -- the JOGL jars were built into a classpath string that was never actually passed to java (only $OUT_PATH was), so any OSR-based JUnit test was silently unrunnable via the official script (NoClassDefFoundError: com/jogamp/opengl/awt/GLCanvas). - TestSetupExtension: add headless-safe CefApp command-line flags (disable GPU/software-rasterizer/Vulkan/on-device-model probing). Without them, a GPU-process crash during browser teardown on a display-less CI agent triggers a slow (multi-minute) fallback probe that blows past any reasonable test timeout. - TestFrameTest, DisplayHandlerTest: convert from TestFrame's default windowed (non-OSR) browser mode to OSR. Windowed close never fires onBeforeClose after native window disposal -- reproduced identically in two independent headless environments, matches upstream java-cef#364 -- so these tests could not complete in any automated environment before this change. - DisplayHandlerTest: onTitleChange/onAddressChange can legitimately fire more than once with the same value (observed a second onTitleChange call from inside CefBrowser_N.close() itself during OSR teardown). The prior assertFalse(gotCallback_) assumed exactly one call; made idempotent instead. - DragDataTest.createFile(): fixed a wrong assertion -- CefDragData.getFileNames() returns the display names passed to addFile(), not the paths (see cef_drag_data.h, which has a separate GetFilePaths() for that). Confirmed via CEF's own header comments. Verified: the full JUnit suite (12 tests) now runs to completion with zero failures via tools/run_tests.sh, in both Release and Debug builds, under both a real X server and a from-scratch Xvfb+window-manager+dbus setup -- the exact combination a real headless CI agent needs. This is the first time the full suite has completed cleanly. Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01BLqFiEqAPbHc3MifgpNjA8
Applies --coverage compile/link flags to the jcef target only (not the vendored libcef_dll_wrapper). Linux/GCC-Clang only for now, mirroring jpype's own CMakeLists.txt AND NOT WIN32 pattern for the same option -- MSVC needs a different toolchain not wired up here. Forces a Debug build when enabled, since gcov coverage data is unreliable against optimized code. Co-Authored-By: Claude Sonnet 5 <[email protected]> Claude-Session: https://claude.ai/code/session_01BLqFiEqAPbHc3MifgpNjA8
Linux-only for now (see plan/findings.md for staging rationale -- no
verified-working recipe for headless Windows/macOS CEF test execution
yet; proving the pattern on one platform first beats a slow first push).
Windows/macOS matrix expansion is an explicit fast-follow.
- .azure/build.yml: two jobs, Test (fast, no instrumentation) and
Coverage (gcov for native/*.cpp + JaCoCo for org.cef.*, both uploaded
to Codecov). Both cache the downloaded CEF binary distribution
(~615MB compressed) keyed on CEF_VERSION + platform, parsed out of
CMakeLists.txt, so a version bump naturally invalidates the cache.
- .azure/scripts/{cache-cef,build,test,coverage}.yml: templates.
- .github/workflows/codeql.yml: lightweight Java + C++ static analysis,
same split as jpype's own CI (Azure for the real build matrix, GitHub
Actions for cheap GitHub-native checks).
Known limitation, documented inline and in plan/findings.md: the
Coverage job's test-run step is marked continueOnError. CEF's native
shutdown reliably crashes in Debug builds (which coverage requires) --
DCHECK failed: all_.empty() in browser_context.cc -- killing the process
before gcov's atexit flush can write any .gcda data. Not a regression
from this change; reproduced with a single test class and as few as 2
browser lifecycles. Not attempting a native-code fix here per project
direction (upstream is under-resourced -- log/track, don't take on
fixing everything).
Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01BLqFiEqAPbHc3MifgpNjA8
This was referenced Aug 29, 2026
Closed
…ore crash
- .azure/scripts/jdk.yml: pin an explicit JDK (21) instead of relying on
a hosted image's default. Found locally that a too-new JDK produces
bytecode JaCoCo 0.8.12 can't parse (class file major version 69 vs.
its support ceiling) -- a real version-skew risk this avoids.
- .azure/build.yml: reference it from every job; add BuildWindows/
BuildMacOS (build-only, no test execution -- see file header for why).
Unverified: no Windows/macOS machine in this dev environment to
actually run them on.
- CoverageTestHelper.{cpp,java}: coverage-only test/CI infrastructure,
not part of the public API or any normal build (only compiled in when
ENABLE_COVERAGE is on, see native/CMakeLists.txt). Explicitly flushes
gcov data via __gcov_dump() right before the known-crashing
CefApp.dispose() call (java-cef#4), so a coverage CI run still
captures real numbers for everything that ran instead of losing all
of it to the crash. A no-op (caught UnsatisfiedLinkError) on any
non-coverage build. This is a mitigation, not a fix, for #4 -- root-
causing that crash is separate follow-up work.
Verified locally end-to-end: rebuilt with -DENABLE_COVERAGE=ON, ran the
full suite (still crashes on shutdown as before, unrelated to this
change -- see #4), confirmed 72 .gcda files are now written (zero
without this change) and gcovr produces a real, non-trivial coverage
report from them (27% overall, full per-file breakdown) despite the
crash.
Co-Authored-By: Claude Sonnet 5 <[email protected]>
Claude-Session: https://claude.ai/code/session_01BLqFiEqAPbHc3MifgpNjA8
3 tasks done
Owner
Author
|
Superseded by PR #35's GitHub Actions CI, which is now merged to master. Closing without merging -- no functional loss, master now has real CI (build/test/coverage/CodeQL) where it had none when this PR was opened. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #7. Also documents/references #3, #4, #5 (not fixed by this PR -- see plan/findings.md and the "Known limitation" section below).
Summary
Stacked on #1 (this repo currently has no CI at all and thin test coverage on
both sides of the JNI boundary -- see
plan/findings.md, not included in this PRsince it's gitignored). This PR gets a real, green, hang-free CI baseline in place
before the planned Panama/Silk JNI-boundary strangler-fig migration starts.
Test-bench fixes (prerequisite -- the suite couldn't complete before these)
tools/run_tests.sh/run_tests.bat: fixed a deadCLS_PATHvariable that meantthe JOGL jars were never actually on the classpath -- any OSR-based JUnit test was
silently unrunnable via the official script.
TestSetupExtension: added headless-safe CefApp flags (GPU/Vulkan/on-device-modeldisabled) -- without them a GPU-process crash on teardown triggers a multi-minute
fallback probe on a display-less CI agent.
TestFrameTest,DisplayHandlerTest: converted fromTestFrame's defaultwindowed browser mode to OSR. Windowed close never fires
onBeforeCloseafternative window disposal -- reproduced in two independent headless environments,
matches upstream java-cef#364 -- so these tests could not complete in any
automated environment before this change.
DisplayHandlerTest:onTitleChange/onAddressChangecan legitimately fire morethan once with the same value; made the assertions idempotent instead of assuming
exactly one call.
DragDataTest.createFile(): fixed a wrong assertion --CefDragData.getFileNames()returns display names, not paths (confirmed againstcef_drag_data.h, which has a separateGetFilePaths()).Result: the full JUnit suite (12 tests) now runs to completion with zero
failures, in both Release and Debug builds, under both a real X server and a
from-scratch Xvfb+window-manager+dbus setup (the combination a real headless CI
agent needs). First time this has ever completed cleanly.
CI (Linux-only for now -- see plan/findings.md for staging rationale)
.azure/build.yml+ templates:Testjob (fast) andCoveragejob (gcov fornative/*.cpp+ JaCoCo fororg.cef.*, uploaded to Codecov). Both cache thedownloaded CEF binary distribution (~615MB compressed) keyed on
CEF_VERSION+platform.
CMakeLists.txt/native/CMakeLists.txt: newENABLE_COVERAGEoption, applies--coverageto thejceftarget only..github/workflows/codeql.yml: lightweight Java + C++ static analysis.Coverage crash: mitigated (real numbers now), root cause tracked separately (#4)
CEF's native shutdown reliably crashes in Debug builds (required for gcov) with
DCHECK failed: all_.empty()inbrowser_context.cc-- reproduced with a singletest class and as few as 2 browser lifecycles, a general Debug-build defect (filed
as #4), not something introduced here. Individual test results are unaffected (crash
is in final process teardown, after every test already passed).
Tried increasing the shutdown message-pump count -- did not help, reverted (no
unproven change kept). What actually works:
CoverageTestHelper.{cpp,java}(new,coverage-only test/CI infra, not part of the public API, only compiled in when
ENABLE_COVERAGEis on) explicitly flushes gcov data via__gcov_dump()immediately before the known-crashing shutdown call. Verified: 72
.gcdafiles nowwritten (zero without this),
gcovrproduces a real report from them -- 27%overall, full per-file breakdown. The crash still happens every time; this rescues
the coverage data instead of losing all of it.
continueOnError: trueon that stepso it doesn't hard-block the pipeline. Root-causing the crash itself is left to a
follow-up branch per #4, not this PR.
Also added
.azure/scripts/jdk.yml: pin JDK 21 explicitly (found locally that a too-new JDKproduces bytecode JaCoCo 0.8.12 can't parse -- a real version-skew risk this
avoids), referenced from every job.
BuildWindows/BuildMacOSjobs: build-only (native + Java compile, no testexecution -- none of the hangs/crashes found this session are build-time issues).
Unverified -- no Windows/macOS machine in this dev environment to run them on.
Test plan
(the real-CI-equivalent setup).
cmake -DENABLE_COVERAGE=ONconfigures and builds cleanly.caused by any change in this PR (reproduces with a single test class).
pipeline YAML is structure-checked against jpype's known-working templates,
not executed. First real run happens once the pipeline is linked to this fork.
🤖 Generated with Claude Code
https://claude.ai/code/session_01BLqFiEqAPbHc3MifgpNjA8