Repository navigation
Release build: SIGSEGV in libc.so.6 during JVM shutdown after all tests pass #10
Description
Activity
Status update, not closing -- root cause still not confirmed.
Two unrelated bugs were fixed this session (a
CefClient.onGotFocus()infinite-recursion StackOverflowError -- see the new issue for it -- and #28/part of #21's null-parameterCHECK()class), plus a stalejcef_buildRelease native library (unbuilt since a much earlier session) got rebuilt fresh. After that:tools/run_tests.sh linux64 Releasehas now completed clean 3 times in a row (184/185, 183/185, 185/186 -- the 1-2 failures each time are a different, already-documented, pre-existing intermittent UI-timing flake, not a crash).Neither fix plausibly explains this by mechanism (neither touches glibc malloc/free or the mojo/message-router code this crash's neighborhood implicates). Best unconfirmed guess: the stale native lib was the actual confound all along, not a real fix landing. Not trusting this yet -- a fresh session should keep re-running the full suite a number of times before treating this as resolved. Also relevant: an
LD_PRELOADmalloc/free double-free interposer built this session (tools_native/malloc_trace/) found zero double-frees scoped to this project's own native code during an earlier (still-crashing) run, which if anything points toward a heap/stack overflow (an out-of-bounds write) rather than a double-free as the mechanism, if this does reproduce again.Major update: found a minimal, fully deterministic repro (6/6 across two different, otherwise-unrelated test classes) -- not closing, root cause still unknown, but this should make it much easier to chase down.
While fixing #16 (made
TestSetupExtension.warmUpBrowserProcess()-- create and close one throwaway browser -- run unconditionally instead of only for the isolated leak-sweep case), discovered that any test class run in complete isolation (--select-class, guaranteeing no other browser was ever created in the process) now crashes with this exact signature (SIGSEGV,libc.so.6+0x91651,SEGV_MAPERR) after JUnit prints a clean "all tests successful" summary, matching this issue precisely.Confirmed the crash has nothing to do with the test body --
CefCommandLineTest(which touches zero CEF value objects beyond whatTestSetupExtensionitself creates) crashes identically, 3/3. So the minimal repro is now just: one real browser create+dispose cycle (the warmup), then normalCefAppshutdown, run alone in the process. No message router, no multi-browser churn, no specific test content required.Counter-intuitively, the full suite (many more browsers, much more churn) does NOT reproduce this -- 187/187 clean with the same fix applied. So whatever the trigger is, it needs isolation (a short-lived, low-volume process), not scale -- the opposite of what "high churn causes corruption" would predict.
Two fresh
hs_err_pid*.logfiles from this exact repro are in the repo root (hs_err_pid285391.log,hs_err_pid285938.log) for whoever picks this up next.Per the project's established methodology for this bug (minimum repro -> pure-C++ equivalent -> tracer/gdb to isolate the divergence): the next step is building a pure-C++ version of exactly this shape (
CefInitialize()-> create+close one browser via the external-pump/windowless_rendering_enabled=truepattern -- seetools_native/leak_probe.cc's own comment on why it deliberately avoids that mode -- ->CefShutdown()) to see if the crash reproduces with no JNI/JVM in the loop at all.Full diagnostic ladder run against the minimal repro from the previous comment. Root cause still not found, but this is the most thoroughly characterized this bug has ever been:
- Real JCEF via JNI, isolated run: crashes 6/6 (two different test classes -- one,
CefCommandLineTest, touches zero CEF value objects beyond what test setup itself creates, ruling out test-body content). - Same, under gdb (
handle SIGSEGV/SIGBUS nostop noprint pass, launched-under): suppressed, 0/1 -- process completes cleanly, exits before a backtrace could be taken. Same ptrace-suppression property already known for the oldModifiedUtf8Testrepro, now confirmed on a much simpler one too. - Same, under a trace-instrumented build (compile-time/env-var only, not a debugger): crashes 3/3, not suppressed. Trace tail shows
CefShutdown()returning andContext's own destructor completing cleanly before the crash -- rules out anything in JCEF's own tracked native teardown path. - Same, under an
LD_PRELOADmalloc/free double-free interposer (scoped to this project's and CEF's own shipped native libraries): crashes 3/3, not suppressed, but zero double-frees found. Rules out a double-free reachable through tracked code; points toward either a heap/stack overflow (invisible to a malloc/free interposer) or something in an untracked module. - Pure C++, no JNI/JVM at all, deliberately mirroring
native/context.cpp's exact config (windowless_rendering_enabled=true,external_message_pump=true, manual pump loop) plus, after direct comparison againstnative/CefBrowser_N.cpp, its exact browser-creation call shape (asyncCreateBrowser(),CEF_RUNTIME_STYLE_ALLOY): survives, 0/3, confirmed both before and after fixing those two call-shape divergences.
Conclusion: a JVM is required to trigger this -- it's not just JCEF's CEF-API call sequence. Leading untested candidate: the pure-C++ repro runs everything from a single thread; real JCEF runs inside a JVM process with 20-30+ other threads (GC, JIT, reference handler, etc.), and glibc's malloc uses per-thread arenas -- a real, unverified variable, not a hand-wave. Next step: test whether thread population specifically matters (spawn the pure-C++ repro's CEF calls from a non-main thread first; then a many-idle-threads version if that alone isn't it).
New tool this round:
tools_native/issue10_repro/(the pure-C++ repro above, kept in-tree for whoever picks this up next).- Real JCEF via JNI, isolated run: crashes 6/6 (two different test classes -- one,
Follow-up: tested the threading hypothesis directly. Neither variant explains it either.
Added `ISSUE10_REPRO_MODE` to `issue10_repro`: `"thread"` runs every CEF UI-thread-affine call from one consistently-used spawned thread instead of `main()` (this does NOT violate CEF's actual threading contract -- it's "one consistent thread for every UI-thread call," not literally `main()`; real JCEF itself uses `AWT-EventQueue-0`, never the JVM's actual main thread, so this mirrors that exactly). `"busythreads"` additionally spawns 24 idle background threads that never touch any CEF API, just churn small heap allocations every 5ms for the whole run, to test whether mere thread-count/glibc-arena pressure matters.
Both survive, 3/3 each. So neither "not the process's initial thread" nor "many concurrent threads doing generic heap churn" reproduces it. Doesn't rule out the JVM's specific allocation pattern (JIT compilation, class loading, GC scanning are qualitatively different from uniform 4KB churn) or something JNI-specific entirely -- those are the remaining candidates.
At this point the pure-C++ repro has ruled out: CEF-API call order/shape, running off the main thread, and generic multi-thread heap pressure. Whatever's left is either JNI-specific (worth trying `issue10_repro` routed through JCEF's actual `client_handler.cpp`/`life_span_handler.cpp` classes instead of its own minimal stub client next) or specific to the JVM's own allocator behavior in a way generic busy threads don't capture.
Major update. Per the user's own hypothesis and direct suggestion ("the general assumption is that it is a free after shutdown problem in which some reference outlives the shutdown and then benignly gets cleared... I think you can more easily just call the finalizer in the C++ after to see what happens"):
Several JCEF Java classes have
finalize()methods doing a real native dispose call if the object was never explicitly disposed (CefBrowser_N.finalize()callsclose(true)i.e.GetHost()->CloseBrowser(true);CefFrame_N,CefRequest,CefResponse,CefPostData(Element), and several others have the same shape). Java finalizers run on the JVM's own Finalizer thread with no ordering guarantee relative toCefApp.dispose()/CefShutdown().Added
ISSUE10_REPRO_LATE_CLOSE=1toissue10_repro: skip the normal explicit close, hold the browser alive acrossCefShutdown(), then callCloseBrowser(true)immediately after -- the exact call a late finalizer would make, deterministic instead of GC-timing-dependent.Result: in a trace-instrumented Debug build, this reproduces the ORIGINAL
DCHECK failed: all_.empty()atbrowser_context.cc:44-- in pure C++, zero JNI/JVM -- 3/3. The trace confirms it fires insideCefShutdown()itself (the post-shutdown close call's own trace line never even prints -- leaving the browser open at shutdown time is sufficient on its own).In Release, the identical scenario survives, 3/3 -- but the trace shows why:
CefShutdown()itself force-closes the still-open browser as part of its own internal teardown (OnBeforeClosefires synchronously duringCefShutdown(), before it returns). So the DebugDCHECKasserts a state Release's subsequent code path already handles gracefully for one straggler browser. My post-shutdown close call lands on an already-closed browser -- a safe no-op.So "one browser never closed before shutdown" is now a confirmed, decisive trigger for the original DCHECK -- a new path to the same known crash, not the already-fixed
CefRequestContext_Nleak. It does not by itself reproduce Release's SIGSEGV, though. Leading theory now: a genuine RACE, not sequential ordering (which is ruled out) -- a finalizer'sclose()landing concurrently withCefShutdown()'s own internal force-close of the same object.CloseBrowser()checksCEF_CURRENTLY_ON_UIT()and posts a task to the UI thread if not -- a task posted to that task runner at the exact momentCefShutdown()is tearing it down is a plausible concrete mechanism a purely sequential test can't reach. Next step: a concurrent-thread variant ofissue10_repro(a second thread callingCloseBrowser()timed to race againstCefShutdown(), not called from the same thread sequentially after).Shipped a real fix for the confirmed mechanism (commit c8df5fc) -- but it does NOT close this issue. Read to the end.
22 Java classes have
finalize()methods that call into native code if never explicitly disposed (browsers, frames, requests, responses, post-data values, drag-data, registrations, message routers, ~10 callback types). Two-part fix, both gated on the sameContext::GetInstance() == nullcheckCefApp.cpp'sN_DoMessageLoopWorkalready used:SetCefForJNIObjectHelper::Release(CefBaseRefCounted*)(jni_scoped_helpers.h) -- the single choke point ~20 of these classes'dispose()paths already funnel through -- guarded once, centrally.- A new
JNI_REQUIRE_CEF_ALIVE_OR_RETURNmacro (jni_util.h) at each of the ~9 individual native entry points whosefinalize()calls a semantic method directly (CefBrowser_N'sclose(true), 8 callback classes'cancel()/Continue()/Failure()).
Validated two ways: full Release suite passes clean, 187/187, no regression. And directly: a scratch test that creates a
CefRequest, never disposes it, and callsdispose()on it immediately after a realCefAppshutdown (deterministic simulation of a late finalizer) now prints "SURVIVED" -- the guard is genuinely reached and works.Important, not closing this issue: in that same scratch-test run, immediately after "SURVIVED post-shutdown dispose()" printed, the process still crashed with the identical
libc.so.6+0x91651signature. SinceCefCommandLineTest(creates zero finalizable objects of its own) already crashes identically in complete isolation, this confirms a SECOND, independent mechanism produces the same crash signature, untouched by this fix. The late-finalizer bug was real and is now fixed, but it was never the sole explanation for this issue's baseline crash. The "isolated single-class run crashes, full suite doesn't" asymmetry is still completely unexplained -- that's the next thing to chase, likely via a concurrent-thread (CloseBrowser()racingCefShutdown()) variant ofissue10_repro, or by finding what's different about full-suite-scale runs that avoids whatever this is.- added 4 commits that reference this issue
on Sep 1, 2026 Resumed per plan. Two things this session:
1. Re-confirmed #10 still reproduces at current HEAD (post c8df5fc/find-handler/frame-handler/permission-handler commits) --
CefCommandLineTestalone, freshly rebuilt Debug/trace tree, crashes 3/3 with the identicalSIGSEGV,libc.so.6+0x91651,SEGV_MAPERRsignature already on file. (First attempt hit a stale-native-lib false alarm --jcef_build_tracepredated the CefFindHandler/CefFrameHandler/CefPermissionHandler JNI additions, throwingUnsatisfiedLinkErrorduringcleanupBrowser(). A fullninjarebuild fixed that; not a real finding, noting it so nobody re-chases it.)2. Decisive new lead: this is a wild instruction-fetch fault, not heap-metadata corruption inside malloc/free as previously characterized. Checked the actual signal-context registers in the hs_err logs, not just the "Problematic frame" summary line:
RIP(the real faulting instruction pointer, from the OSucontext) is identical tosi_addrin every captured crash (this session's 3 fresh ones, and the pre-existing 2026-08-31 log) -- e.g.RIP=0x000071e58043ab10=si_addr=0x000071e58043ab10.- That address is nowhere near libc's actual mapped range in the same log (e.g. RIP
0x71e58043ab10vs libc mapped at0x71e658600000-0x71e658628000-- off by more than 4×10^11, not an in-library data offset). ERR=0x14on the page fault = I/D bit set (instruction fetch) + user-mode -- this is CPU trying to execute code at an unmapped address, not a read/write fault inside a malloc/free routine.- The
libc.so.6+0x91651/Instructions: (pc=0x...691651)shown in "Problematic frame" is a different address than the real RIP -- consistently, across every log checked. That's hs_err's post-fault frame-pointer/return-address guess, not where the crash actually happened. The earlier "disassembly matches a classic freelist-unlink pattern" read was against this wrong address.
So: something is corrupting a code pointer (vtable slot, function pointer, or return address) that later gets called/returned into, landing on garbage -- not a heap-header stomp inside glibc's allocator internals. This significantly narrows the search: look for a stale/dangling
CefRefPtr/vtable call surviving past its object's real lifetime (classic UAF-then-virtual-call shape), rather than any allocator-internals theory.Could not get a core dump to go further (
core_patternpipes to/wsl-capture-crash, which doesn't exist in this sandbox -- cores are silently dropped). gdb live-attach is already known to suppress this Heisenbug (documented earlier in this thread), so a real backtrace needs either a working core-dump path, orrr/reverse-debugging, or a hand-rolled SIGSEGV handler inissue10_reprothat prints a raw backtrace viabacktrace()/backtrace_symbols()before dying -- the last one is cheap and doesn't require ptrace, so it's the natural next step and shouldn't perturb the timing the way gdb does.Also added
ISSUE10_REPRO_RACE_CLOSEmode totools_native/issue10_repro/(a second thread callingCloseBrowser(true)timed to raceCefShutdown(), the one "late close" variant not yet tried) -- ran 20x in Debug and 20x in Release. All 20 Debug runs hit the already-knownall_.empty()DCHECK immediately (expected -- that DCHECK fires synchronously atCefShutdown()entry if the browser is still open at all, regardless of race timing, so this mode doesn't actually exercise the race window against that specific check). All 20 Release runs survived. So this variant hasn't found a new repro path yet; not pursuing it further right now given the more promising in-hand-backtrace lead above.🤖 Generated with Claude Code
Architectural hypothesis (from the user, and it fits the RIP/si_addr evidence above well):
CefShutdown()/CloseBrowser()force a synchronous teardown while async work is still in flight, stranding it -- a queued continuation later firing into torn-down state, rather than the shutdown itself corrupting the heap.Context::Shutdown()(native/context.cpp:276-327) does a fixed, unconditional 10xCefDoMessageLoopWork()pump then callsCefShutdown()-- no tracking of which async operations are actually still outstanding, no wait-for-real-completion-signal, no explicit pass marking live references dead before CEF's own teardown proceeds. Already directly tested that pumping more (200x over ~1s, 100x budget) doesn't help -- consistent with this theory: the fix isn't more blind waiting, it's an actual drain signal tied to specific outstanding work, not time.This matches the two already-fixed instances of this exact bug shape: the
CefMessageRouterpersistent-query-callback leak (a router removed without canceling its still-pending callback) and the late-finalizer guard (c8df5fc, a reactive per-call-site null check rather than a preventive drain).Audit surface -- every
CefPostTask/CefPostDelayedTasksite innative/, none of whichContext::Shutdown()currently tracks or drains:CefMessageRouter_N.cpp:76,105,CefCookieManager_N.cpp:128,150,CefURLRequest_N.cpp:90,util_linux.cpp:105(windowed-close's delayed task), and a dozen sites inCefBrowser_N.cpp(zoom, window size/bounds/parent, frame rate, view-source, async destroy). Any one of these still queued onTID_UI(or its capturedCefRefPtr/raw pointer already dangling) whenCefShutdown()tears down that task runner is a concrete match for the wild-jump fault.Not yet attempted, the natural next test: instrument how many tasks are actually queued on
TID_UIwhenContext::Shutdown()begins pumping, across several runs -- if reliably non-zero on a crash and zero on a survival, that's a direct link. Real fix would be phased: stop accepting new async work, drain (a low-priority sentinel task posted last, waited on -- CEF doesn't expose queue depth directly) or explicitly cancel/tombstone every outstanding op, only then callCefShutdown().🤖 Generated with Claude Code
Ran the derisking probe for the shutdown-phasing theory above (plan preserved in this repo's
plan/roadmap.md, 2026-09-03 entry, for continuity). Result: negative for this specific OSR crash.Added
SetCefForJNIObjectHelper::GetLiveObjectCount()-- an atomic counter at the sameAddRef()/Release()choke point the finalizer guard (c8df5fc) uses -- logged viaJCEF_TRACEinContext::Shutdown()right before the pump and right beforeCefShutdown(). This measures "how many CEF-backed Java wrapper objects does JCEF think are still alive" at the exact moment of the crash-triggering call.8/8 crashing
CefCommandLineTest-alone runs showedlive_object_count=0at both checkpoints. Sanity-checked the instrumentation itself is real, not just stuck at zero: a batch run includingCefRequestTestshowedlive_object_count=4with 70CEF_ADDREFevents logged.So: whatever's stranded at shutdown for this crash isn't a JNI-wrapped CEF object this choke point can see. Either it's purely Chromium-internal (and CEF's own
CefMainRunner::StartShutdownOnUIThread()drain, which already exists precisely for this hazard, is supposed to handle that class of thing), or the mechanism for this specific crash isn't "outstanding async work" at all. Not building the quiesce-loop fix for the OSR path on this evidence -- leaving the counter in place since it's harmless and may be useful signal elsewhere.The windowed-close delayed-task hazard (
util_linux.cpp'sDestroyCefBrowser(), a real, separate, untested risk for windowed-mode shutdown specifically) is unaffected by this result and still worth fixing on its own merits.Next real lead, per the RIP/si_addr wild-jump evidence from the previous comment: a working core-dump path for a genuine postmortem backtrace (this sandbox's
core_patternpipes to a nonexistent/wsl-capture-crash, so that's currently blocked here) -- that would finally show what's actually being called through.🤖 Generated with Claude Code
Follow-up per direct feedback: "if [a leaked CefRequestContext] were true it should show up in the tracer" -- checked directly rather than continuing to speculate.
Tracer confirms no leak: this exact crashing
CefCommandLineTest-alone run showsCEF_ADDREF/CEF_RELEASEperfectly balanced, 52/52 (plusCEF_REFPTR_ADDREF/CEF_REFPTR_RELEASE1/1 at theGetGlobalContext()site). So JCEF is not holding an unreleasedCefRefPtr-tracked reference to anything, globalCefRequestContextincluded -- consistent with the earlierModifiedUtf8Testbalance finding, now confirmed for this repro too.But there's a more precise finding from reading real CEF source (
~/devel/cef/libcef/browser/browser_context.cc): theall_.empty()DCHECK'sImplManager::all_tracksCefBrowserContext*instances, added/removed viaAddImpl()/RemoveImpl().CefBrowserContext::Shutdown()(which does theRemoveImpl()) assertsrequest_context_set_.empty()and is driven entirely by Chromium's own internalProfile/refcount teardown machinery -- not by anything that goes through JCEF's C++ code or itsCefRefPtr/JNI-association tracer at all. That's why the tracer shows nothing wrong: this isn't a reference JCEF holds or leaks, it's a Chromium-internal object lifetime JCEF has zero visibility into.Sharper point, from this exact captured trace:
Shutdown() CefShutdown() returnedandContext::~Context() EXITboth fire cleanly in this run (no DCHECK this time -- confirms the Heisenbug property, sometimes it fires insideCefShutdown(), sometimes it doesn't and the crash comes later). Since the crash happens in a window where none of JCEF's own instrumented native code is executing at all, the race isn't "CefShutdown()'s own call graph outraces itself" -- it's thatCefShutdown()returning to the caller doesn't guarantee CEF's background threads (ThreadPool workers, IO thread) have actually finished tearing down. Something there is still mid-flight when the JVM proceeds to its own exit/library-unload sequence, landing on a corrupted callable -- matching the earlier RIP/si_addr wild-jump evidence exactly.This also explains, retroactively, why the already-tried "pump
CefDoMessageLoopWork()200x over ~1s before callingCefShutdown()" experiment didn't help: that only services the UI-thread's own external-pump queue, which has no ability to wait on independent OS threads that don't route through it.Implication: if the actual race is after
CefShutdown()returns, in threadsContext::Shutdown()never touches, then no pre-CefShutdown()fix on JCEF's side (a wait-loop, a quiesce check, more pumping) can address it -- the fix would have to be post-CefShutdown(), or isn't reachable from JCEF's side at all. Next step: check whether CEF's public API offers any documented way to wait for background-thread shutdown completion afterCefShutdown()returns (nothing obvious ininclude/cef_app.hfrom a first pass) -- if not, this may be a genuine upstream CEF/Chromium bug rather than something fixable innative/context.cpp.🤖 Generated with Claude Code
Follow-up per direct feedback: "you can track active browsers and open contexts... prevent new ones from spawning and apply actions from Java to bring to zero FIRST, then call shutdown." Checked whether that already happens, found a real gap in it, tested the fix for the gap directly -- still doesn't help.
That architecture already exists.
CefApp.java/CefClient.java:createClient()rejects new clients onceSHUTTING_DOWN;dispose()synchronously drives every openCefClientto close all its browsers (browser.close(true), which a comment there confirms "synchronously triggersonBeforeClose()"), and only calls nativeshutdown()onceclients_is empty;CefApp.shutdown()(CefApp.java:509) even explicitly callsCefRequestContext.disposeGlobalContext()beforeN_Shutdown()-- added specifically for this exact DCHECK, per its own comment citing this issue by number.Found a real, concrete gap in that existing fix.
disposeGlobalContext()only does anything if Java code explicitly calledCefRequestContext.getGlobalContext()at some point (populating a static cache). A browser created with no explicit context (native/CefBrowser_N.cpp:1021-1042passesnullptrstraight toCefBrowserHost::CreateBrowser()) never populates that cache -- CEF silently uses its own internal global context instead. That's exactly this whole investigation's minimal-repro path (the warmup browser), so the existing fix is a silent no-op here.Tested the fix for this gap directly, not just theorized: temporarily forced the Java-side cache to populate before shutdown, recompiled, reran
CefCommandLineTestalone 7x. Trace confirmed the mechanism engaged exactly as intended -- the context is created, then explicitly released viadisposeGlobalContextNative(), beforeContext::Shutdown()'s pump even starts. Still crashed 7/7, identically. Reverted the experimental change.Conclusion: dropping JCEF's own Java-visible reference(s) to the global context isn't sufficient.
CefBrowserContext::Shutdown()(real CEF source,browser_context.cc:212) requires ALLCefRequestContextImplreferences gone -- there's at least one more JCEF never holds or sees, almost certainly the browser's own internalRenderProcessHost/WebContentsteardown chain.browser.close(true)'s synchronousOnBeforeClose()signals "the browser told you it's closing," not "every Chromium-internal object tied to it is released" -- that teardown continues asynchronously regardless of what JCEF calls or waits on.Every JCEF-side lever tried this investigation has now been tested directly and refuted: more pumping before
CefShutdown()(200x/~1s, no change), JCEF's own live-object count (reliably 0 on crashing runs), and now forcing the global context's Java-side reference to be dropped before shutdown (no change). The remaining reference(s) are internal to Chromium's own teardown chain, with no public API surface to observe, wait on, or force to completion.This is now reasonably strong evidence that, in this specific OSR manifestation, #10 isn't fixable by reordering or adding JCEF-side calls before
CefShutdown(). Two paths forward: file this exact minimal, reproducible sequence upstream with CEF/Chromium, or try a JCEF-side mitigation that absorbs the race rather than provably closing it (an idle grace-period sleep -- no pumping, no other activity -- after browser close and/or afterCefShutdown()returns; meaningfully different from the already-refuted "pump more" experiment since it doesn't contend with straggling background threads for CPU/scheduling). Neither attempted yet.🤖 Generated with Claude Code
Update: found and fixed a genuine, independent bug while regression-testing the last idea in this thread -- not #10 itself, but caught mid-investigation and worth recording here since it was blocking any full-suite validation.
A real hang (not a crash): a full-suite run stalled 37+ minutes with zero output growth.
jstackshowed the main thread parked forever inTestSetupExtension.close()'scountdown_.await()(unbounded -- no timeout, unlike every other latch-wait in this codebase), whileAWT-EventQueue-0sat completely idle. Root cause:CefRequestHandlerCoverageTest's deliberatechrome://crashnavigation kills its browser's renderer; closing that browser later during the suite-wide finalCefApp.dispose()pass does not reliably fire a realOnBeforeClose()for a browser whose renderer already died. SoCefClient.cleanupBrowser()never sees its browser list empty out,clientWasDisposed()never fires,CefApp's client set never empties, and nativeshutdown()is never invoked at all (confirmed: its ownprintlnnever appeared in the log).Fixed by generalizing the existing windowed-close (X11
WM_DELETE_WINDOW) fallback -- a bounded, idempotency-guarded synthesizedOnBeforeClose()-- to the OSR force-close path too (util::ScheduleOnBeforeCloseFallback(),native/util.h/util_linux.cpp, called fromCefBrowser_N.cpp'sN_Close()), plus a defensive 30s bound onTestSetupExtension.close()'s previously-unbounded wait. Verified: the exact repro now completes cleanly instead of hanging, and a full-suite run goes from a 37+ minute hang down to ~44s with only one pre-existing, already-documented flaky test remaining.Separately, a negative result worth recording: tried making
CefClient.createBrowser()always explicitly resolve toCefRequestContext.getGlobalContext()instead of passingnullthrough to CEF (so JCEF'sdisposeGlobalContext()fix is never a silent no-op). Confirmed via trace it engaged correctly and properly released the exact context object the browser used -- but it brokeCefRequestContextTest.globalContextIsGlobal()deterministically (isGlobal()returned false), a real regression, most likely re-triggering the same "GetGlobalContext()doesn't guarantee the same pointer per call" hazard a prior fix in this exact area was built around, just at higher frequency. Reverted -- not committed. Still doesn't explain #10's actual crash either way (see the last two comments above).🤖 Generated with Claude Code
This should be investigated as a use after free as it smells like a stale object vtable jump.
- added 4 commits that reference this issue
on Sep 4, 2026 New, more diagnosable lead (2026-09-05)
Fixing an unrelated crash (a JCEF test misusing
CefDownloadHandler.onBeforeDownload()'s
return contract — seeplan/tasks/20260905-26-download-shelf-check-crash.mdon the
coverage/phase1-value-objects-phase2-handlersbranch) let a full-suite run get further
than before, and it now hits a different, cleanly reproducible (2/2 identical) native
crash in the same shutdown window as this issue — a real null-pointer dereference, not
the earlierlibc.so.6wild-fetch signature.Repro (Release build, plain test job, no gdb/coverage build needed —
hs_err_pid*.log
is sufficient):tools/run_tests.sh linux64 Release --config debugPrint=true \ --exclude-tag leak-sweep --exclude-tag process-isolatedCrash:
SEGV_MAPERR,si_addr=0x150(null-pointer-plus-small-member-offset — a
classic nullCefRefPtrdereference), duringCefApp.dispose()'sCefShutdown()call:C [libcef.so+0x2e7ecb5] CefBrowserInfo::RemoveFrame(content::RenderFrameHost*)+0x145 C [libcef.so+0x2e6d116] non-virtual thunk to CefBrowserContentsDelegate::RenderFrameDeleted(...) C [libcef.so+0x5eaed15] WebContentsObserverList::NotifyObservers<...>(...) C [libcef.so+0x5ec2b63] content::WebContentsImpl::RenderFrameDeleted(...) C [libcef.so+0x5cd3230] content::RenderFrameHostImpl::~RenderFrameHostImpl() C [libcef.so+0x5d1d35a] content::RenderFrameHostManager::~RenderFrameHostManager() C [libcef.so+0x5c1e567] content::FrameTreeNode::~FrameTreeNode() C [libcef.so+0x5cd4af0] content::RenderFrameHostImpl::ResetChildren() C [libcef.so+0x5c1c2d1] content::FrameTree::Shutdown() C [libcef.so+0x5ea1faf] content::WebContentsImpl::~WebContentsImpl() C [libcef.so+0x725d30c] tabs::TabModel::~TabModel() C [libcef.so+0x7261a81] TabStripModel::SendDetachWebContentsNotifications(...) C [libcef.so+0x7268758] TabStripModel::CloseTabs(...) C [libcef.so+0x7267fe4] TabStripModel::CloseAllTabs() C [libcef.so+0xd5d4314] Browser::~Browser() C [libcef.so+0xd5f3675] BrowserManagerService::Shutdown() C [libcef.so+0x9f69f46] ProfileImpl::~ProfileImpl() C [libcef.so+0x9f666ff] ProfileDestroyer::DestroyOriginalProfileNow(...) C [libcef.so+0x9f7305d] ProfileManager::~ProfileManager() C [libcef.so+0x9dff129] BrowserProcessImpl::StartTearDown() ... ChromeBrowserMainParts::PostMainMessageLoopRun() -> CefMainRunner::Shutdown() -> CefShutdown()Root cause, read from real CEF source (
libcef/browser/browser_info.cc/.h, CEF
commit82195616d, matching this project'sCEF_VERSION146.0.10+g8219561):CefBrowserInfo::RemoveFrame()callsbrowser_->request_context()->...wherebrowser_
is aCefRefPtr<CefBrowserHostBase>member.CefBrowserInfo::BrowserDestroyed()nulls
browser_, documented as "Always called after SetClosing and WebContentsDestroyed" — the
assumption being thatWebContentsDestroyed()'s ownRemoveAllFrames()call already
cleared every frame by then. But the realcontent::WebContentsImplC++ object is not
actually destroyed until later — it's kept alive by theTabStripModel/Browser/Profile
object graph until process-wide teardown, which only runs synchronously inside
CefShutdown()itself (BrowserProcessImpl::StartTearDown()→ProfileImpl::~ProfileImpl()
→ ... →~WebContentsImpl()→FrameTree::Shutdown()). When that destructor chain tears
down any remainingRenderFrameHostImpls, it callsRemoveFrame()a second time — on a
CefBrowserInfowhosebrowser_is already null, with no null-guard for this case.This looks like a real CEF-side lifetime-ordering bug (JCEF's browser-close nulls
browser_before the underlying Chromium object is actually destroyed), not something
fixable from JCEF's side — no lever in JCEF's public API reaches this internal teardown
ordering. Worth root-causing further or reporting upstream to CEF with this exact
reproducible sequence, since it's now a clean, deterministic repro rather than the
harder-to-diagnose wild-fetch signature originally filed here.Full write-up:
plan/tasks/20260903-01-gh10-shutdown-sigsegv.md's 2026-09-05 addendum.
Summary
Running the full JUnit suite (
tools/run_tests.sh linux64 Release) consistentlycrashes with
SIGSEGVduring process shutdown, after JUnit has already printed aclean "112 tests successful" summary. This is a different signature from #4 (which is
a
DCHECK failed: all_.empty()abort specific to Debug/coverage builds) -- this oneis a Release-build
SIGSEGVinlibc.so.6, no DCHECK/assertion message, and the mainjavaprocess itself is the one that crashes (not ajcef_helpersubprocess).Repro
Suite reports
112 tests successful, then:An
hs_err_pid<pid>.logis written to the repo root with:Reproduced identically across multiple separate runs (different pids), always after
all tests already completed successfully -- so it does not affect correctness of test
results, only shutdown cleanliness (and means CI would need to tolerate the non-zero
exit / not treat it as a suite failure).
Environment
-DCMAKE_BUILD_TYPE=Release), not theENABLE_COVERAGEDebug buildcovered by Debug build: native shutdown crashes with DCHECK failed: all_.empty() in browser_context.cc #4
Impact
Low urgency -- doesn't affect test results, but pollutes CI logs/exit codes with a
crash and leaves an
hs_err_pid*.logbehind every run. Worth root-causing eventually(likely in
CefApp.dispose()'s native teardown path) but out of scope for the currentcoverage-expansion push.