Repository navigation
Debug/coverage build (ENABLE_COVERAGE=ON) hits multiple DCHECK/FATAL crashes not present in Release #9
Description
Activity
Update: refined diagnosis + workaround that unblocks a real gcovr number
While starting Track B (getting a real native coverage number, tracked in
plan/roadmap.md), re-confirmed all three crashes above still reproduce on the current CEF version, plus found more precise repros and one new crash signature.Crash 1 & 2 are the same root cause, not two separate ones
Bisected via
--select-method: every single test method inCefPrintSettingsTestcrashes individually, both as the very first native call in a fresh process (crash 1's exact repro:print_settings_cpptoc.cc:452] CefPrintSettings_0_CppToC called with invalid version -1) and after browser tests have already run (crash 2:printing/units.cc:13] DCHECK failed: new_unit > 0). It is not specific tosetPrinterPrintableArea'sDimension/Rectanglehandling as originally guessed above — literally any call toCefPrintSettings.create()crashes in a Debug build, regardless of what runs before it. The two different crash messages appear to be two different failure points hit on the way to the same underlying problem (the print-settings native object's version/unit state is never correctly initialized in Debug builds).New:
CefRequestContext.getGlobalContext()crashes even with zero preference-accessor callsBisecting
CefRequestContextTestthe same way: all three of its test methods crash individually, includingglobalContextIsGlobal(), which only callsgetGlobalContext(),isGlobal(), andgetHandler()— no preference accessors at all. The crash is a different DCHECK than the one documented above for crash 3:FATAL:cef/libcef/browser/request_context_impl.cc:117] DCHECK failed: false. context not valid(vs. the originally-documented
request_context_impl.cc:242] DCHECK failed: false. called on invalid thread, which is specific to the preference accessors and still separately reproduces too). SoCefRequestContext.getGlobalContext()itself is unreliable in a fresh Debug process, not just its preference-accessor methods off-thread.New:
CefPostDataElement.setToFile("")hits a Debug-onlyCHECK(notDCHECK)Found via a new unit test added this session (
CefPostDataTest.elementSetToEmptyFilePathLeavesElementEmpty, which passes cleanly in Release — confirmed the element's type correctly staysPDE_TYPE_EMPTYthere). In the Debug/coverage build the same call crashes:FATAL:.../libcef_dll/ctocpp/post_data_element_ctocpp.cc:69] Check failed: !fileName.empty().This one's a plain
CHECK, not aDCHECK— worth double-checking it doesn't also affect Release builds under different circumstances, sinceCHECKis nominally not compiled out. It didn't reproduce in Release here, but the mechanism isn't obvious from the JCEF side.Workaround used to unblock Track B
Excluding the two known-bad classes (
CefPrintSettingsTest,CefRequestContextTest) plus the newly-found one (CefPostDataTest) via the JUnit console launcher's--exclude-classname, the rest of the suite now runs to completion against the coverage build (confirmed viaTestSetupExtension.close()firing, which JUnit5 only invokes after the full test plan finishes) and flushes real.gcdadata before the already-documented shutdown crash (browser_context.cc:44] DCHECK failed: all_.empty(), tracked separately). Got a realgcovrnumber this way: 44% (3838/8680 lines) onnative/, up from the original 27% baseline — saved toplan/coverage-native-current.txt(gitignored working notes, not in this PR).Not attempting a native-side fix for any of these three in this session — logging the refined diagnosis here per this fork's usual practice, since a real fix needs CEF/Chromium internals familiarity per the original "Suggested next steps" above.
Update 2: a 6th crash signature, found via new CefBrowser API coverage
Continuing coverage work (plan/roadmap.md Phase 4), added `CefBrowserApiTest.java` -- a broad sweep of `CefBrowser` API methods (frame enumeration, zoom, `find`/`stopFinding`, `viewSource`, `replaceMisspelling`, `executeJavaScript`, `createScreenshot`, etc.) against a live browser. Passes cleanly and reliably in the Release build (confirmed in the full suite, 139/139), but crashes the Debug/coverage build:
```
FATAL:mojo/public/cpp/bindings/lib/interface_endpoint_client.cc:538] DCHECK failed: !has_pending_responders().
```Preceded by repeated `SetError: {code=4, message="MEDIA_ELEMENT_ERROR: Media load rejected by URL safety check"}` from `html_media_element.cc` -- unexpected, since no test content includes a
<video>/<audio>element; this looks like some internal Chromium probe unrelated to the test's own page content. Not yet isolated to one specific method in the class (my best guess by inspection is `viewSource()`, which triggers a real navigation to a `view-source:` URL and a renderer swap, or `find()`/`stopFinding()`'s find-in-page mojo interface -- unconfirmed).Workaround: added `CefBrowserApiTest` to the same `--exclude-classname` list already used for the three classes above, which unblocked a fresh full `gcovr` run: 46% (4006/8680 lines) on `native/`, up from the 44% reported in the previous comment. Updated report saved to `plan/coverage-native-current.txt` (gitignored working notes).
Not attempting a fix or further isolation this session -- logging per this fork's usual practice.
Summary
While expanding unit test coverage (tracked in #5) and trying to get a real native
gcovrnumber against theENABLE_COVERAGEDebug build, several distinct nativecrashes surfaced that do not reproduce in the normal Release build. All were
found incidentally while writing straightforward value-object unit tests -- none
involve unusual API usage. This is on top of the already-tracked shutdown crash in
#4; these are separate, earlier-in-the-run crashes.
Environment: local sandbox,
cmake -DENABLE_COVERAGE=ON(forces Debug perCMakeLists.txt), Linux, JDK 25 (Temurin), JCEF at the version pinned by thisfork's current
CMakeLists.txt.Crash 1:
CefPrintSettingscreated as the first native object in a fresh processRepro: run a JUnit class that only calls
CefPrintSettings.create()(no browsercreated first) as the only test class in the process. Same shape as an earlier
finding that
CefDragData.create()alone in a fresh process hits an analogouscrash -- something about a CEF API struct's embedded version field only being
correctly initialized after a real browser has been created first, in Debug builds.
Crash 2:
CefPrintSettingsTestas part of the full suite (browser tests ran first)Different crash than (1) under otherwise-similar conditions:
Not yet isolated to the exact test method (11 in that class);
setPrinterPrintableArea'sDimension/Rectanglehandling is the likely suspect by inspection but unconfirmed.Crash 3:
CefRequestContextpreference accessors called off the browser UI threadCefRequestContext.hasPreference/getPreference/canSetPreference/setPreferencearedocumented (Javadoc) as: "This method must be called on the browser process UI
thread, otherwise it will always return false [or null]." That description promises
graceful degradation off-thread, but in a Debug build the underlying native call
DCHECK-crashes instead of returning the documented default. Repro: call any of
these four methods from a plain JUnit test thread (not the CEF UI thread) against
CefRequestContext.getGlobalContext().Impact
All three make it very difficult to get a real, complete native coverage number
via
gcovragainst the Debug/ENABLE_COVERAGEbuild -- the process crashespartway through the suite, before all tests run and before final coverage flush.
Crash 3 in particular looks like a real behavioral bug independent of coverage
tooling: the documented "off-thread calls return false/null" contract is violated
in Debug builds.
Suggested next steps
Debug builds, or (more likely the actual intent) the DCHECK should be relaxed to
match the documented graceful-degradation behavior.
why Debug-only DCHECKs fire on otherwise-ordinary API usage; not yet root-caused
from the JCEF side.
Found via
Writing
CefPrintSettingsTest.javaandCefRequestContextTest.java(new unittests, coverage-expansion effort tracked in #5). All three logged here rather than
fixed, per this fork's general practice of logging real bugs found while building
out CI/coverage rather than chasing every one immediately.