Skip to content

Deterministic DCHECK(all_.empty()) in browser_context.cc during final CefApp shutdown (Debug build) #23

Description

@Thrameos

Summary

A deterministic (not intermittent, unlike #22) DCHECK fires at the very end of a full Debug/ENABLE_COVERAGE test-suite run, during final CefApp/CEF shutdown, after every individual test has already completed successfully:

[FATAL:cef/libcef/browser/browser_context.cc:44] DCHECK failed: all_.empty().

Immediately preceded in the log by:

shutdown on Thread[#38,AWT-EventQueue-0,6,main]
[ERROR:base/process/kill_posix.cc:32] waitpid(...): No child processes (10)

Reproducibility

100% deterministic across 4 consecutive full-suite runs this session (--select-package tests.junittests, same four classes excluded by name as always for unrelated reasons -- see #9/#16). Every run: all tests pass, all .gcda coverage data flushes successfully (confirmed via file count/mtime), and then this DCHECK fires during the final CEF shutdown sequence, after CefApp's own shutdown() call.

Impact

Blocks getting a clean-exit Debug/ENABLE_COVERAGE run (the process always core-dumps at the very end), though it does not lose coverage data -- unlike #20/#21, this happens after CoverageTestHelper.flush() has already run for every test, so gcovr numbers measured against the resulting .gcda files are trustworthy. Does not reproduce in the Release build (no DCHECKs compiled in there, and the equivalent code path may simply be skipped NDEBUG-style).

Likely cause (not confirmed)

libcef/browser/browser_context.cc:44's all_ presumably tracks live CefBrowserContext instances CEF-side; the DCHECK asserts all of them were destroyed before CefShutdown() runs. This suite creates and disposes a CefRequestContext object in a few tests (CefRequestContextTest -- currently excluded from coverage runs for an unrelated crash reason, CefRequestContextHandlerTest, CefCookieAccessFilterTest) -- worth checking whether one of those leaves a CefRequestContext-backed CefBrowserContext alive past its owning browser's close, or whether CefApp's singleton-teardown ordering in this test harness (TestSetupExtension, shared across the whole suite) disposes things in the wrong order relative to CEF's own expectations.

Not attempted

No isolation/bisection attempted yet (which specific test class holds the outstanding CefBrowserContext) -- flagging now since it's now confirmed deterministic (a good sign for whoever picks it up: --select-class/--exclude-classname bisection should converge quickly on the reproducer, since it happens every time rather than intermittently like #22).

Found via

Retrying Track B's (#5) ENABLE_COVERAGE Debug-build measurement after fixing the crashes tracked in #20/#21 -- this DCHECK is what's now hit at the tail end of every clean run, instead of #22's intermittent SIGSEGV (#22 did not reproduce in these 4 runs).

Activity

  1. Thrameos commented on Aug 30, 2026

    @Thrameos
    OwnerAuthor

    Found and fixed one genuine, confirmed leak that trips this exact DCHECK: CefRequestContextHandlerTest.handlerServesContentThroughARequestContext() created a non-global CefRequestContext via CefRequestContext.createContext(...) and never called dispose() on it -- its underlying native CefBrowserContext was still registered in CEF's internal tracking (browser_context.cc's ImplManager::all_) when CefShutdown() ran. Fixed by disposing it after awaitCompletion() returns.

    However, the DCHECK still reproduces deterministically even with that fix in place. Root-caused one level deeper by reading CEF's own source (~/devel/cef/libcef/browser/browser_context.cc): CefBrowserContext::Shutdown() asserts request_context_set_.empty() before deregistering itself from the global ImplManager. That means some CefRequestContext is still referencing a CefBrowserContext at shutdown time -- either a second leak I haven't found yet, or (more likely, given this suite creates and disposes many browsers across ~180 tests sharing one CefApp/global context) an async-teardown-ordering race: the last browser's own close sequence (and whatever CefRequestContext reference it implicitly holds) may not have fully unwound by the time CefApp.shutdown()'s N_Shutdown() call runs.

    Our Context::Shutdown() already pumps CefDoMessageLoopWork() 10 times before calling CefShutdown() (confirmed windowless_rendering_enabled=true by default in this test harness, so external_message_pump_ is true and that pump loop does run) -- but 10 iterations apparently isn't enough to let the last browser's context references settle. Also applied a related, independently-useful fix: made Context::Shutdown() idempotent (a static already-shut-down guard, matching another JetBrains fix in the same area) -- not the fix for this specific DCHECK, but cheap defensive hardening regardless.

    Not yet resolved. Next step for whoever picks this up: instrument/log CefBrowserContext's constructor/Shutdown() (or the request_context_set_ add/remove calls) to identify exactly which context is still outstanding at the DCHECK, or try increasing/replacing the pump-before-shutdown loop with an actual wait-for-idle signal.

    This does NOT block coverage measurement -- confirmed the crash happens only after all .gcda files have already flushed, so gcovr numbers remain trustworthy despite the non-zero exit code.

  2. Thrameos commented on Aug 30, 2026

    @Thrameos
    OwnerAuthor

    Update: two real leaks fixed, deeper trigger found, still not resolved (commit 6d80f9e)

    Fixed two genuine static-reference leaks: CefRequestContext_N.globalInstance and CefCookieManager_N.globalInstance each correctly memoize the native wrapper for getGlobalContext()/getGlobalManager() (avoiding a fresh AddRef per call), but neither static field was ever cleared — a static field is a permanent GC root, so the one persistent AddRef'd reference to CEF's global CefRequestContext/CefBrowserContext survived indefinitely, past CefShutdown(). CefBrowser_N.getRequestContext() falls back to CefRequestContext.getGlobalContext() for every browser created without an explicit context, so this is populated almost immediately in normal use.

    Added CefRequestContext.disposeGlobalContext()/CefCookieManager.disposeGlobalManager(), wired into CefApp.shutdown() before N_Shutdown(). Verified real and correct (169/169 passing, no regressions) — but not sufficient alone: the all_.empty() DCHECK still reproduces in the Debug/coverage build with this fix applied.

    Found what looks like the actual trigger while chasing it further: immediately before the DCHECK, the log always shows a Mojo message-validation failure —

    Invalid message: VALIDATION_ERROR_DESERIALIZATION_FAILED
    Message ... rejected by interface content.mojom.NavigationClient
    

    — followed by a separate FATAL: DCHECK failed: ptr_. at base/memory/scoped_refptr.h:292, in a different PID than the main browser process (a child/renderer process). This is consistent with an abnormally-terminated renderer skipping its own CefBrowserContext cleanup and leaving it permanently registered in CEF's ImplManager — exactly what the later shutdown-time all_.empty() DCHECK detects.

    This trigger is itself intermittent/timing-sensitive, like #22 — reproduced in 3/3 quick runs but did not reproduce within a 30s window on another run. Tried to pin down exactly which test triggers it (one run showed it consistently right after browser id=17's onAfterCreated, immediately following CefDialogHandlerTest's file_dialog.html browser at id=16) but JUnit's console reporter buffers all output until the whole run completes, so a crash mid-run loses attribution.

    Not resolved. This reads as a real Chromium/Mojo-internal race under Debug-build IPC message validation, plausibly related to this suite's high browser create/destroy churn (~180 short-lived browsers in quick succession), rather than a JCEF-specific coding mistake — but that's not confirmed.

    Next step for whoever picks this up: get a live stack trace for the scoped_refptr.h:292 crash itself (the same technique that cracked #22 — a lucky hs_err_pid*.log/core dump would show exactly which Mojo interface/message triggers it), or add a TestExecutionListener that flushes each test's name to stdout the instant it starts (not buffered like the built-in tree reporter) so a crash mid-run still attributes cleanly to one test class.

  3. added 7 commits that reference this issue on Aug 30, 2026
  4. Thrameos commented on Aug 31, 2026

    @Thrameos
    OwnerAuthor

    Same root cause as #4, fixed in commit b7baff5 -- see #4's closing comment for the full writeup. Closing as duplicate.

  5. 3 remaining items

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions