Repository navigation
CefBrowser.getDevToolsClient()/executeDevToolsMethod() and browser.print() can hang indefinitely, unrecoverable #12
Description
Activity
Partial progress on repro 2 (`browser.print()` hang), same investigative technique that fully resolved #17 and #18 -- but this one didn't turn out as clean.
Found a real, confirmed re-entrancy hazard: the original repro called `terminateTest()` (which begins browser/window teardown) from inside `onPrintDialog()`, before returning `false`. Per CEF's own source (`libcef/browser/printing/print_dialog_linux.cc`'s `ShowDialog()`: `if (!handler_->OnPrintDialog(...)) { callback_impl->Disconnect(); OnPrintCancel(); }`), the native print pipeline expects to synchronously continue its own cleanup after this Java call returns -- tearing the browser down while still inside that call stack is a plausible source of exactly this kind of hang.
Fix attempted: defer teardown to `onPrintReset()` instead (called from `CefPrintDialogLinux`'s own destructor, after the pipeline has already unwound cleanly).
Honest result: genuinely intermittent, not a clean fix. One isolated run recovered cleanly via the normal 30s test watchdog (no `SIGKILL` needed) -- but the very next run of the identical test needed a hard `SIGKILL` past a 45s external timeout, with zero diagnostic output (added `System.out.println` to every `CefPrintHandler` callback -- none fired that time, meaning it can hang somewhere before even `onPrintStart`). This points to a second, independent, timing-dependent hang further upstream in CEF's own print pipeline, unrelated to the re-entrancy issue.
New lead, not confirmed: this environment does have a real CUPS daemon running (`systemctl status cups` shows it active), but zero printers configured (`/printers/` returns 404). A print pipeline blocking/retrying while waiting for printer enumeration or default-printer resolution that never resolves is a plausible trigger, but not verified -- would need native-side instrumentation (gdb/strace on the hung process) to confirm, which wasn't attempted given the intermittent nature makes attaching mid-hang unreliable.
Kept the `onPrintReset()`-based version in the fork (strictly safer than the original -- eliminates one guaranteed-hang code path) but left `@Disabled` on, since this remains a genuine unrecoverable-hang risk, just a narrower one than before. Not closing this out.
- added a commit that references this issue
on Aug 30, 2026 Made concrete progress reading CEF's own internal test suite (`~/devel/cef/tests/ceftests/devtools_message_unittest.cc`) for repro 1 (the DevTools hang), though not a full fix.
Isolated exactly where the hang is: `browser.getDevToolsClient()` itself -- which internally calls `addDevToolsMessageObserver()`, constructing a native `DevToolsMessageObserver`/`CefRegistration` -- is a separate, confirmed-safe code path. It's specifically `executeDevToolsMethod()` that hangs unrecoverably. Confirmed via a standalone diagnostic across 3 repeated isolated runs (~2s each, no `SIGKILL` needed). Landed this as real new coverage (`CefRegistration_N.cpp` 0% -> 100%) in `CefDevToolsRegistrationTest.java`, sidestepping the hang entirely rather than leaving that surface untested.
Found a real, separate API gap while comparing against CEF's own C++ interface: our Java `CefDevToolsMessageObserver` (package-private, used internally by `CefDevToolsClient`) only exposes `onDevToolsMethodResult`/`onDevToolsEvent` -- it's missing `OnDevToolsAgentAttached`/`OnDevToolsAgentDetached`, which the real C++ `CefDevToolsMessageObserver` has. CEF's own test explicitly waits on/checks agent-attachment state (`EXPECT_EQ(1, attached_ct_)`). Whoever picks up the hang next should start here: add the missing attach/detach callbacks to JCEF's Java wrapper, then gate `executeDevToolsMethod()` calls on `OnDevToolsAgentAttached` actually firing first rather than just "after page load" (which is what every attempt so far, including CEF's own test, has assumed is sufficient -- but CEF's test never explicitly waits for attachment before its first `ExecuteMethod` call either, so this is a hypothesis, not confirmed to be the fix).
Other technique details compared and ruled out as differentiators: CEF's test also calls `ExecuteDevToolsMethod` with an explicit non-zero `message_id` rather than JCEF's hardcoded `0` -- but `0` is documented as valid ("assigned automatically"), and JCEF's own promise-completion wiring (`CefBrowser_N.java`'s `executeDevToolsMethod`) correctly handles both cases, so this isn't it either.
Update: the
executeDevToolsMethod()hang repro is root-caused and fixed (commit3feac1d)Root cause: the hang was specific to calling
"Browser.getVersion"— a browser-domain DevTools command. A standalone diagnostic that polled the newCefDevToolsClient.isAgentAttached()signal from a plain (non-AWT) thread found:- The DevTools agent attaches near-instantly (~2ms after the first message is sent), exactly matching
OnDevToolsAgentAttached's documented "occurs in response to the first message sent while detached" behavior. Attachment was never the blocked step. "Browser.getVersion"then never receives a matchingOnDevToolsMethodResultcallback. Confirmed this isn't a message-id race: explicitly assigning our own non-zero incrementing message IDs (matching the technique CEF's owntests/ceftests/devtools_message_unittest.ccuses) made no difference.- Substituting
"Page.enable"— a page-domain command, the exact onedevtools_message_unittest.ccexercises — completes in ~2ms with zero hang, reproduced across 6 runs. - Correction to an earlier note on this issue: the AWT/CEF UI thread is not stuck during this hang. The test harness's watchdog can cleanly force-close the browser (
windowClosing/doClose/onBeforeCloseall fire normally) even while theBrowser.getVersioncall is left permanently pending — only that oneCompletableFuturenever completes.
Likely underlying cause: JCEF forces
CEF_RUNTIME_STYLE_ALLOYfor all browsers, and this CEF version's Browser-domain DevTools command handler may not be fully wired for an Alloy-runtime-style, off-screen-renderedCefBrowserHost. Page-domain commands, tied to the renderer/page rather than the browser-process-level Browser domain, are unaffected.Fix landed:
CefDevToolsClientTestnow callsPage.enableinstead ofBrowser.getVersion,@Disabledremoved, and it assertsisAgentAttached()transitions false → true across the round-trip. Full suite is back to green on this path.The other repro under this issue (
browser.print()hang) is unrelated and still open/unresolved.- The DevTools agent attaches near-instantly (~2ms after the first message is sent), exactly matching
Summary
While expanding unit test coverage (tracked in #5, #9), found two distinct API calls that can hang the JVM process indefinitely in the normal Release build (not just the Debug/coverage build issue #9 tracks) -- and the hang is severe enough that it survives past what would normally be a safe recovery point.
Environment: local sandbox, Release build, Linux, JDK 25 (Temurin), JCEF at the version pinned by this fork's current
CMakeLists.txt. Both reproduced via a JUnit test harness that has its own 15-30s per-test watchdog (ajava.util.Timer-based forced browser close, running viaSwingUtilities.invokeLateron the AWT/CEF UI thread) -- normally sufficient to recover from a slow/stuck test. Neither hang was recoverable by that watchdog.Repro 1:
CefBrowser.getDevToolsClient().executeDevToolsMethod(...)Called from
CefLifeSpanHandler.onAfterCreated()against a freshly created OSR browser. The returnedCompletableFuturenever completes (success or failure) and the process never becomes responsive again. Confirmed via an isolated run wrapped in a hard externaltimeout -k 5 45(SIGTERM then SIGKILL) -- evenSIGKILLfallback was needed, i.e. it wasn't just slow, something was genuinely stuck. The harness's own watchdog (scheduled to force-close the browser at 30s viaSwingUtilities.invokeLater) never got a chance to run, suggesting the AWT/CEF UI thread itself may be blocked, not just the DevTools call's own completion.Repro 2:
CefBrowser.print()Same symptom: the call to
browser.print()(fromCefLoadHandler.onLoadingStateChange()once the page finishes loading) hangs the process indefinitely,onPrintDialogreturningfalsedoes not "cancel immediately" as documented. Isolated the same way -- neededSIGKILLvia a hardtimeoutwrapper, the harness's own 30s watchdog never fired. Likely CEF's print pipeline blocks at the native level without a real printer backend available in this headless environment (no real X11 display/printer daemon), but if so that's a robustness gap worth flagging: an API documented to support immediate cancellation via afalsereturn shouldn't be able to hang the whole process instead.Impact
Both are real public JCEF APIs (
CefBrowser.getDevToolsClient()/CefDevToolsClient.executeDevToolsMethod(),CefBrowser.print()) that any embedding application could call. An unrecoverable hang (not just a slow response) in either is a correctness/robustness issue independent of this fork's own test-coverage tooling -- unlike #9, this is not specific toENABLE_COVERAGE/Debug builds.Not attempted
No native-side investigation done this session -- both were isolated at the Java-API level (confirmed via
--select-class/--select-methodJUnit runs under a hard external timeout wrapper) and then the corresponding test code was reverted/deleted rather than left hanging in the suite. Logged here per this fork's usual practice of tracking real findings rather than chasing every one immediately.Found via
Writing
CefDevToolsClientTest.javaand an addition toCefPrintHandlerTest.java(new unit tests, coverage-expansion effort tracked in #5). Both deleted/reverted after confirming the hang; see this fork'scoverage/phase1-value-objects-phase2-handlersbranch history for the revert commits.