Skip to content

CefBrowserApiDebugSafeTest deterministically crashes the process (mojo interface_endpoint_client.cc:538 DCHECK) in the Debug/coverage build #27

Description

@Thrameos

Summary

java/tests/junittests/CefBrowserApiDebugSafeTest.java (two tests:
executeJavaScriptAndLoadRequestDoNotThrow, createScreenshotReturnsARealImage)
deterministically crashes the whole JUnit process before any @Test method
even starts, when run against the ENABLE_COVERAGE/coverage-instrumented
Debug build -- reproduces both in isolation and as part of a small
multi-class run.

FATAL:mojo/public/cpp/bindings/lib/interface_endpoint_client.cc:538]
DCHECK failed: !has_pending_responders()

preceded by several Exception in thread "AWT-EventQueue-0" with no visible
stack trace.

Not root-caused. Currently @Disabled per this project's standing strategy
(disable flaky/broken ported tests rather than chase the underlying native
crash inline).

Reproduction

Run CefBrowserApiDebugSafeTest alone (or as part of a small 6-class
--select-class run alongside 5 other independently-reliable classes)
against a coverage-instrumented Debug build.

Impact

browser.loadRequest(), browser.executeJavaScript() + real request-object
loading, and browser.createScreenshot() all lose their only existing test
coverage while this stays disabled.

Activity

  1. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Status update, not closing -- currently not reproducing, root cause not found.

    Initial suspicion was that this was actually the same bug as #31 (a CefClient.onGotFocus() infinite-recursion StackOverflowError, since this issue's own description -- "several Exception in thread AWT-EventQueue-0 with no visible stack trace" -- matches that bug's signature exactly). Tested directly: rebuilt the ENABLE_COVERAGE Debug build fresh and ran CefBrowserApiDebugSafeTest 3x in isolation and once in a 6-class combined run (matching both repro conditions this issue describes) -- passes cleanly every time. Then re-tested with CefClient.java reverted to its pre-#31-fix state specifically, to isolate whether #31's fix was responsible: also passes cleanly 3/3, refuting that hypothesis directly.

    So this currently doesn't reproduce at all, with or without #31's fix, under either of this issue's own stated repro conditions. Possibly stale (fixed by something else, unrelated to #31, in the time since this was filed), possibly environment/timing-dependent and just didn't fire today. Leaving open since root cause is still unconfirmed either way -- if this resurfaces, the mojo interface_endpoint_client.cc:538 DCHECK signature is the thing to chase directly next time, not the AWT-EventQueue-0 storm (that's #31's, now fixed, and evidently a red herring here).

  2. Thrameos commented on Sep 2, 2026

    @Thrameos
    OwnerAuthor

    Root-caused, via the sibling class CefBrowserApiTest (same interface_endpoint_client.cc:538 DCHECK failed: !has_pending_responders() signature). Used the jpype gdb technique (handle SIGSEGV nostop noprint pass, launch-under-gdb) to catch it live and get a real backtrace:

    N_Close(force=true)                                         [native/CefBrowser_N.cpp:1515]
     -> CefBrowserHostCToCpp::CloseBrowser(true)
     -> AlloyBrowserHostImpl::CloseContents()
     -> RenderProcessHostImpl::FastShutdownIfPossible() -> FastShutdown()
     -> RenderProcessHostImpl::ProcessDied()
     -> SiteInstanceGroup::RenderProcessExited()
     -> RenderFrameHostImpl::RenderProcessGone() -> RenderFrameDeleted()
     -> WebContentsImpl::RenderFrameDeleted()
     -> CefBrowserContentsDelegate::RenderFrameDeleted()
     -> CefBrowserInfo::RemoveFrame()
     -> CefFrameHostImpl::Detach() -> DetachRenderFrame()
     -> mojo::Remote<RenderFrame>::ResetWithReason()
     -> InterfaceEndpointClient::CloseWithReason() -> PassHandle()  [DCHECK fires here]
    

    FastShutdown() kills the renderer process without waiting for in-flight IPC to settle. If a mojo call to the renderer (in this case, browser.find()/browser.viewSource()) still has a pending responder when the browser is force-closed right after, PassHandle()'s !has_pending_responders() DCHECK fires -- compiled out in Release, live in Debug/coverage builds, hence why this never showed up outside ENABLE_COVERAGE.

    Underlying gap: JCEF has no CefFindHandler binding at all (no onFindResult), so there's no way for Java code to actually wait for find()'s async completion before closing -- unlike CEF's own find_handler_unittest.cc (~/devel/cef/tests/ceftests/), which always waits for OnFindResult's finalUpdate before calling StopFinding()+closing. That's the real, complete fix (a new find_handler.cpp/CefFindHandler Java interface) but is its own scoped feature addition, not a quick fix here.

    Mitigation applied (CefBrowserApiTest.java, commit follows): give the CEF UI thread's message pump ~300ms to drain the pending response before closing. Verified 5/5 clean isolated runs plus the 6-class batch (including CefBrowserApiDebugSafeTest) clean under ENABLE_COVERAGE. Same fix class as #4/#23's finalizer-guard work (c8df5fc) and matches the general pattern in this codebase: shutdown proceeding without draining a pending in-flight op.

    Given this class isn't currently reproducing per the last comment, no change needed here -- but if it resurfaces, apply the same settle-before-close mitigation (or better, wait on a real CefFindHandler/request-completion signal once that binding exists) rather than re-investigating from scratch.

  3. Thrameos commented on Sep 2, 2026

    @Thrameos
    OwnerAuthor

    Related: #32 (JCEF had no CefFindHandler binding) is now fixed in 135dc98. `CefBrowserApiTest` (the sibling class hitting the identical DCHECK signature, root-caused in 0d5890b) now waits on the real `onFindResult(finalUpdate=true)` before closing instead of a blind settle delay -- worth trying against `CefBrowserApiDebugSafeTest`'s two methods here if they also call `find()`/`viewSource()` before closing.

    🤖 Generated with Claude Code

  4. Thrameos commented on Sep 4, 2026

    @Thrameos
    OwnerAuthor

    Confirmed fixed via a real run (the thing this issue was still waiting on). CefBrowserApiDebugSafeTest was already un-@Disabled from the shared-browser (Tier 1) harness migration, which fixed a cold-start hazard -- a different mechanism from CefBrowserApiTest's settle-period/onFindResult mitigation discussed above, and not needed here.

    3/3 clean isolated runs (--select-class tests.junittests.CefBrowserApiDebugSafeTest, exit 0, no DCHECK/abort) plus a clean batched run alongside CefBrowserApiTest and CefDisplayHandlerTest (6/6 tests passing, exit 0).

    Details: plan/tasks/20260903-11-issue27-browserapidebugsafe-mojo-dcheck.md (archived).

    🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions