Repository navigation
Deterministic DCHECK(all_.empty()) in browser_context.cc during final CefApp shutdown (Debug build) #23
Description
Activity
Found and fixed one genuine, confirmed leak that trips this exact DCHECK:
CefRequestContextHandlerTest.handlerServesContentThroughARequestContext()created a non-globalCefRequestContextviaCefRequestContext.createContext(...)and never calleddispose()on it -- its underlying nativeCefBrowserContextwas still registered in CEF's internal tracking (browser_context.cc'sImplManager::all_) whenCefShutdown()ran. Fixed by disposing it afterawaitCompletion()returns.However, the DCHECK still reproduces deterministically even with that fix in place. Root-caused one level deeper by reading CEF's own source (
~/devel/cef/libcef/browser/browser_context.cc):CefBrowserContext::Shutdown()assertsrequest_context_set_.empty()before deregistering itself from the globalImplManager. That means someCefRequestContextis still referencing aCefBrowserContextat shutdown time -- either a second leak I haven't found yet, or (more likely, given this suite creates and disposes many browsers across ~180 tests sharing oneCefApp/global context) an async-teardown-ordering race: the last browser's own close sequence (and whateverCefRequestContextreference it implicitly holds) may not have fully unwound by the timeCefApp.shutdown()'sN_Shutdown()call runs.Our
Context::Shutdown()already pumpsCefDoMessageLoopWork()10 times before callingCefShutdown()(confirmedwindowless_rendering_enabled=trueby default in this test harness, soexternal_message_pump_is true and that pump loop does run) -- but 10 iterations apparently isn't enough to let the last browser's context references settle. Also applied a related, independently-useful fix: madeContext::Shutdown()idempotent (a static already-shut-down guard, matching another JetBrains fix in the same area) -- not the fix for this specific DCHECK, but cheap defensive hardening regardless.Not yet resolved. Next step for whoever picks this up: instrument/log
CefBrowserContext's constructor/Shutdown()(or therequest_context_set_add/remove calls) to identify exactly which context is still outstanding at the DCHECK, or try increasing/replacing the pump-before-shutdown loop with an actual wait-for-idle signal.This does NOT block coverage measurement -- confirmed the crash happens only after all
.gcdafiles have already flushed, sogcovrnumbers remain trustworthy despite the non-zero exit code.- added a commit that references this issue
on Aug 30, 2026 Update: two real leaks fixed, deeper trigger found, still not resolved (commit
6d80f9e)Fixed two genuine static-reference leaks:
CefRequestContext_N.globalInstanceandCefCookieManager_N.globalInstanceeach correctly memoize the native wrapper forgetGlobalContext()/getGlobalManager()(avoiding a freshAddRefper call), but neither static field was ever cleared — a static field is a permanent GC root, so the one persistentAddRef'd reference to CEF's globalCefRequestContext/CefBrowserContextsurvived indefinitely, pastCefShutdown().CefBrowser_N.getRequestContext()falls back toCefRequestContext.getGlobalContext()for every browser created without an explicit context, so this is populated almost immediately in normal use.Added
CefRequestContext.disposeGlobalContext()/CefCookieManager.disposeGlobalManager(), wired intoCefApp.shutdown()beforeN_Shutdown(). Verified real and correct (169/169 passing, no regressions) — but not sufficient alone: theall_.empty()DCHECK still reproduces in the Debug/coverage build with this fix applied.Found what looks like the actual trigger while chasing it further: immediately before the DCHECK, the log always shows a Mojo message-validation failure —
Invalid message: VALIDATION_ERROR_DESERIALIZATION_FAILED Message ... rejected by interface content.mojom.NavigationClient— followed by a separate
FATAL: DCHECK failed: ptr_.atbase/memory/scoped_refptr.h:292, in a different PID than the main browser process (a child/renderer process). This is consistent with an abnormally-terminated renderer skipping its ownCefBrowserContextcleanup and leaving it permanently registered in CEF'sImplManager— exactly what the later shutdown-timeall_.empty()DCHECK detects.This trigger is itself intermittent/timing-sensitive, like #22 — reproduced in 3/3 quick runs but did not reproduce within a 30s window on another run. Tried to pin down exactly which test triggers it (one run showed it consistently right after browser id=17's
onAfterCreated, immediately followingCefDialogHandlerTest'sfile_dialog.htmlbrowser at id=16) but JUnit's console reporter buffers all output until the whole run completes, so a crash mid-run loses attribution.Not resolved. This reads as a real Chromium/Mojo-internal race under Debug-build IPC message validation, plausibly related to this suite's high browser create/destroy churn (~180 short-lived browsers in quick succession), rather than a JCEF-specific coding mistake — but that's not confirmed.
Next step for whoever picks this up: get a live stack trace for the
scoped_refptr.h:292crash itself (the same technique that cracked #22 — a luckyhs_err_pid*.log/core dump would show exactly which Mojo interface/message triggers it), or add aTestExecutionListenerthat flushes each test's name to stdout the instant it starts (not buffered like the built-in tree reporter) so a crash mid-run still attributes cleanly to one test class.- added 7 commits that reference this issue
on Aug 30, 2026 3 remaining items
- added 15 commits that reference this issue
on Sep 1, 2026
Summary
A deterministic (not intermittent, unlike #22)
DCHECKfires at the very end of a full Debug/ENABLE_COVERAGEtest-suite run, during finalCefApp/CEF shutdown, after every individual test has already completed successfully:Immediately preceded in the log by:
Reproducibility
100% deterministic across 4 consecutive full-suite runs this session (
--select-package tests.junittests, same four classes excluded by name as always for unrelated reasons -- see #9/#16). Every run: all tests pass, all.gcdacoverage data flushes successfully (confirmed via file count/mtime), and then this DCHECK fires during the final CEF shutdown sequence, afterCefApp's ownshutdown()call.Impact
Blocks getting a clean-exit Debug/
ENABLE_COVERAGErun (the process always core-dumps at the very end), though it does not lose coverage data -- unlike #20/#21, this happens afterCoverageTestHelper.flush()has already run for every test, sogcovrnumbers measured against the resulting.gcdafiles are trustworthy. Does not reproduce in the Release build (no DCHECKs compiled in there, and the equivalent code path may simply be skippedNDEBUG-style).Likely cause (not confirmed)
libcef/browser/browser_context.cc:44'sall_presumably tracks liveCefBrowserContextinstances CEF-side; the DCHECK asserts all of them were destroyed beforeCefShutdown()runs. This suite creates and disposes aCefRequestContextobject in a few tests (CefRequestContextTest-- currently excluded from coverage runs for an unrelated crash reason,CefRequestContextHandlerTest,CefCookieAccessFilterTest) -- worth checking whether one of those leaves aCefRequestContext-backedCefBrowserContextalive past its owning browser's close, or whetherCefApp's singleton-teardown ordering in this test harness (TestSetupExtension, shared across the whole suite) disposes things in the wrong order relative to CEF's own expectations.Not attempted
No isolation/bisection attempted yet (which specific test class holds the outstanding
CefBrowserContext) -- flagging now since it's now confirmed deterministic (a good sign for whoever picks it up:--select-class/--exclude-classnamebisection should converge quickly on the reproducer, since it happens every time rather than intermittently like #22).Found via
Retrying Track B's (#5)
ENABLE_COVERAGEDebug-build measurement after fixing the crashes tracked in #20/#21 -- this DCHECK is what's now hit at the tail end of every clean run, instead of #22's intermittent SIGSEGV (#22 did not reproduce in these 4 runs).