Skip to content

Release build: SIGSEGV in libc.so.6 during JVM shutdown after all tests pass #10

Description

@Thrameos

Summary

Running the full JUnit suite (tools/run_tests.sh linux64 Release) consistently
crashes with SIGSEGV during process shutdown, after JUnit has already printed a
clean "112 tests successful" summary. This is a different signature from #4 (which is
a DCHECK failed: all_.empty() abort specific to Debug/coverage builds) -- this one
is a Release-build SIGSEGV in libc.so.6, no DCHECK/assertion message, and the main
java process itself is the one that crashes (not a jcef_helper subprocess).

Repro

tools/run_tests.sh linux64 Release --config debugPrint=false

Suite reports 112 tests successful, then:

tools/run_tests.sh: line 44: <pid> Aborted (core dumped) LD_PRELOAD=libcef.so java ...

An hs_err_pid<pid>.log is written to the repo root with:

# A fatal error has been detected by the Java Runtime Environment:
# SIGSEGV (0xb) at pc=..., pid=..., tid=...
...
Current thread is native thread
Native frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)
C  [libc.so.6+0x91651]

Reproduced identically across multiple separate runs (different pids), always after
all tests already completed successfully -- so it does not affect correctness of test
results, only shutdown cleanliness (and means CI would need to tolerate the non-zero
exit / not treat it as a suite failure).

Environment

Impact

Low urgency -- doesn't affect test results, but pollutes CI logs/exit codes with a
crash and leaves an hs_err_pid*.log behind every run. Worth root-causing eventually
(likely in CefApp.dispose()'s native teardown path) but out of scope for the current
coverage-expansion push.

Activity

  1. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Status update, not closing -- root cause still not confirmed.

    Two unrelated bugs were fixed this session (a CefClient.onGotFocus() infinite-recursion StackOverflowError -- see the new issue for it -- and #28/part of #21's null-parameter CHECK() class), plus a stale jcef_build Release native library (unbuilt since a much earlier session) got rebuilt fresh. After that: tools/run_tests.sh linux64 Release has now completed clean 3 times in a row (184/185, 183/185, 185/186 -- the 1-2 failures each time are a different, already-documented, pre-existing intermittent UI-timing flake, not a crash).

    Neither fix plausibly explains this by mechanism (neither touches glibc malloc/free or the mojo/message-router code this crash's neighborhood implicates). Best unconfirmed guess: the stale native lib was the actual confound all along, not a real fix landing. Not trusting this yet -- a fresh session should keep re-running the full suite a number of times before treating this as resolved. Also relevant: an LD_PRELOAD malloc/free double-free interposer built this session (tools_native/malloc_trace/) found zero double-frees scoped to this project's own native code during an earlier (still-crashing) run, which if anything points toward a heap/stack overflow (an out-of-bounds write) rather than a double-free as the mechanism, if this does reproduce again.

  2. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Major update: found a minimal, fully deterministic repro (6/6 across two different, otherwise-unrelated test classes) -- not closing, root cause still unknown, but this should make it much easier to chase down.

    While fixing #16 (made TestSetupExtension.warmUpBrowserProcess() -- create and close one throwaway browser -- run unconditionally instead of only for the isolated leak-sweep case), discovered that any test class run in complete isolation (--select-class, guaranteeing no other browser was ever created in the process) now crashes with this exact signature (SIGSEGV, libc.so.6+0x91651, SEGV_MAPERR) after JUnit prints a clean "all tests successful" summary, matching this issue precisely.

    Confirmed the crash has nothing to do with the test body -- CefCommandLineTest (which touches zero CEF value objects beyond what TestSetupExtension itself creates) crashes identically, 3/3. So the minimal repro is now just: one real browser create+dispose cycle (the warmup), then normal CefApp shutdown, run alone in the process. No message router, no multi-browser churn, no specific test content required.

    Counter-intuitively, the full suite (many more browsers, much more churn) does NOT reproduce this -- 187/187 clean with the same fix applied. So whatever the trigger is, it needs isolation (a short-lived, low-volume process), not scale -- the opposite of what "high churn causes corruption" would predict.

    Two fresh hs_err_pid*.log files from this exact repro are in the repo root (hs_err_pid285391.log, hs_err_pid285938.log) for whoever picks this up next.

    Per the project's established methodology for this bug (minimum repro -> pure-C++ equivalent -> tracer/gdb to isolate the divergence): the next step is building a pure-C++ version of exactly this shape (CefInitialize() -> create+close one browser via the external-pump/windowless_rendering_enabled=true pattern -- see tools_native/leak_probe.cc's own comment on why it deliberately avoids that mode -- -> CefShutdown()) to see if the crash reproduces with no JNI/JVM in the loop at all.

  3. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Full diagnostic ladder run against the minimal repro from the previous comment. Root cause still not found, but this is the most thoroughly characterized this bug has ever been:

    1. Real JCEF via JNI, isolated run: crashes 6/6 (two different test classes -- one, CefCommandLineTest, touches zero CEF value objects beyond what test setup itself creates, ruling out test-body content).
    2. Same, under gdb (handle SIGSEGV/SIGBUS nostop noprint pass, launched-under): suppressed, 0/1 -- process completes cleanly, exits before a backtrace could be taken. Same ptrace-suppression property already known for the old ModifiedUtf8Test repro, now confirmed on a much simpler one too.
    3. Same, under a trace-instrumented build (compile-time/env-var only, not a debugger): crashes 3/3, not suppressed. Trace tail shows CefShutdown() returning and Context's own destructor completing cleanly before the crash -- rules out anything in JCEF's own tracked native teardown path.
    4. Same, under an LD_PRELOAD malloc/free double-free interposer (scoped to this project's and CEF's own shipped native libraries): crashes 3/3, not suppressed, but zero double-frees found. Rules out a double-free reachable through tracked code; points toward either a heap/stack overflow (invisible to a malloc/free interposer) or something in an untracked module.
    5. Pure C++, no JNI/JVM at all, deliberately mirroring native/context.cpp's exact config (windowless_rendering_enabled=true, external_message_pump=true, manual pump loop) plus, after direct comparison against native/CefBrowser_N.cpp, its exact browser-creation call shape (async CreateBrowser(), CEF_RUNTIME_STYLE_ALLOY): survives, 0/3, confirmed both before and after fixing those two call-shape divergences.

    Conclusion: a JVM is required to trigger this -- it's not just JCEF's CEF-API call sequence. Leading untested candidate: the pure-C++ repro runs everything from a single thread; real JCEF runs inside a JVM process with 20-30+ other threads (GC, JIT, reference handler, etc.), and glibc's malloc uses per-thread arenas -- a real, unverified variable, not a hand-wave. Next step: test whether thread population specifically matters (spawn the pure-C++ repro's CEF calls from a non-main thread first; then a many-idle-threads version if that alone isn't it).

    New tool this round: tools_native/issue10_repro/ (the pure-C++ repro above, kept in-tree for whoever picks this up next).

  4. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Follow-up: tested the threading hypothesis directly. Neither variant explains it either.

    Added `ISSUE10_REPRO_MODE` to `issue10_repro`: `"thread"` runs every CEF UI-thread-affine call from one consistently-used spawned thread instead of `main()` (this does NOT violate CEF's actual threading contract -- it's "one consistent thread for every UI-thread call," not literally `main()`; real JCEF itself uses `AWT-EventQueue-0`, never the JVM's actual main thread, so this mirrors that exactly). `"busythreads"` additionally spawns 24 idle background threads that never touch any CEF API, just churn small heap allocations every 5ms for the whole run, to test whether mere thread-count/glibc-arena pressure matters.

    Both survive, 3/3 each. So neither "not the process's initial thread" nor "many concurrent threads doing generic heap churn" reproduces it. Doesn't rule out the JVM's specific allocation pattern (JIT compilation, class loading, GC scanning are qualitatively different from uniform 4KB churn) or something JNI-specific entirely -- those are the remaining candidates.

    At this point the pure-C++ repro has ruled out: CEF-API call order/shape, running off the main thread, and generic multi-thread heap pressure. Whatever's left is either JNI-specific (worth trying `issue10_repro` routed through JCEF's actual `client_handler.cpp`/`life_span_handler.cpp` classes instead of its own minimal stub client next) or specific to the JVM's own allocator behavior in a way generic busy threads don't capture.

  5. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Major update. Per the user's own hypothesis and direct suggestion ("the general assumption is that it is a free after shutdown problem in which some reference outlives the shutdown and then benignly gets cleared... I think you can more easily just call the finalizer in the C++ after to see what happens"):

    Several JCEF Java classes have finalize() methods doing a real native dispose call if the object was never explicitly disposed (CefBrowser_N.finalize() calls close(true) i.e. GetHost()->CloseBrowser(true); CefFrame_N, CefRequest, CefResponse, CefPostData(Element), and several others have the same shape). Java finalizers run on the JVM's own Finalizer thread with no ordering guarantee relative to CefApp.dispose()/CefShutdown().

    Added ISSUE10_REPRO_LATE_CLOSE=1 to issue10_repro: skip the normal explicit close, hold the browser alive across CefShutdown(), then call CloseBrowser(true) immediately after -- the exact call a late finalizer would make, deterministic instead of GC-timing-dependent.

    Result: in a trace-instrumented Debug build, this reproduces the ORIGINAL DCHECK failed: all_.empty() at browser_context.cc:44 -- in pure C++, zero JNI/JVM -- 3/3. The trace confirms it fires inside CefShutdown() itself (the post-shutdown close call's own trace line never even prints -- leaving the browser open at shutdown time is sufficient on its own).

    In Release, the identical scenario survives, 3/3 -- but the trace shows why: CefShutdown() itself force-closes the still-open browser as part of its own internal teardown (OnBeforeClose fires synchronously during CefShutdown(), before it returns). So the Debug DCHECK asserts a state Release's subsequent code path already handles gracefully for one straggler browser. My post-shutdown close call lands on an already-closed browser -- a safe no-op.

    So "one browser never closed before shutdown" is now a confirmed, decisive trigger for the original DCHECK -- a new path to the same known crash, not the already-fixed CefRequestContext_N leak. It does not by itself reproduce Release's SIGSEGV, though. Leading theory now: a genuine RACE, not sequential ordering (which is ruled out) -- a finalizer's close() landing concurrently with CefShutdown()'s own internal force-close of the same object. CloseBrowser() checks CEF_CURRENTLY_ON_UIT() and posts a task to the UI thread if not -- a task posted to that task runner at the exact moment CefShutdown() is tearing it down is a plausible concrete mechanism a purely sequential test can't reach. Next step: a concurrent-thread variant of issue10_repro (a second thread calling CloseBrowser() timed to race against CefShutdown(), not called from the same thread sequentially after).

  6. Thrameos commented on Sep 1, 2026

    @Thrameos
    OwnerAuthor

    Shipped a real fix for the confirmed mechanism (commit c8df5fc) -- but it does NOT close this issue. Read to the end.

    22 Java classes have finalize() methods that call into native code if never explicitly disposed (browsers, frames, requests, responses, post-data values, drag-data, registrations, message routers, ~10 callback types). Two-part fix, both gated on the same Context::GetInstance() == null check CefApp.cpp's N_DoMessageLoopWork already used:

    • SetCefForJNIObjectHelper::Release(CefBaseRefCounted*) (jni_scoped_helpers.h) -- the single choke point ~20 of these classes' dispose() paths already funnel through -- guarded once, centrally.
    • A new JNI_REQUIRE_CEF_ALIVE_OR_RETURN macro (jni_util.h) at each of the ~9 individual native entry points whose finalize() calls a semantic method directly (CefBrowser_N's close(true), 8 callback classes' cancel()/Continue()/Failure()).

    Validated two ways: full Release suite passes clean, 187/187, no regression. And directly: a scratch test that creates a CefRequest, never disposes it, and calls dispose() on it immediately after a real CefApp shutdown (deterministic simulation of a late finalizer) now prints "SURVIVED" -- the guard is genuinely reached and works.

    Important, not closing this issue: in that same scratch-test run, immediately after "SURVIVED post-shutdown dispose()" printed, the process still crashed with the identical libc.so.6+0x91651 signature. Since CefCommandLineTest (creates zero finalizable objects of its own) already crashes identically in complete isolation, this confirms a SECOND, independent mechanism produces the same crash signature, untouched by this fix. The late-finalizer bug was real and is now fixed, but it was never the sole explanation for this issue's baseline crash. The "isolated single-class run crashes, full suite doesn't" asymmetry is still completely unexplained -- that's the next thing to chase, likely via a concurrent-thread (CloseBrowser() racing CefShutdown()) variant of issue10_repro, or by finding what's different about full-suite-scale runs that avoids whatever this is.

  7. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Resumed per plan. Two things this session:

    1. Re-confirmed #10 still reproduces at current HEAD (post c8df5fc/find-handler/frame-handler/permission-handler commits) -- CefCommandLineTest alone, freshly rebuilt Debug/trace tree, crashes 3/3 with the identical SIGSEGV, libc.so.6+0x91651, SEGV_MAPERR signature already on file. (First attempt hit a stale-native-lib false alarm -- jcef_build_trace predated the CefFindHandler/CefFrameHandler/CefPermissionHandler JNI additions, throwing UnsatisfiedLinkError during cleanupBrowser(). A full ninja rebuild fixed that; not a real finding, noting it so nobody re-chases it.)

    2. Decisive new lead: this is a wild instruction-fetch fault, not heap-metadata corruption inside malloc/free as previously characterized. Checked the actual signal-context registers in the hs_err logs, not just the "Problematic frame" summary line:

    • RIP (the real faulting instruction pointer, from the OS ucontext) is identical to si_addr in every captured crash (this session's 3 fresh ones, and the pre-existing 2026-08-31 log) -- e.g. RIP=0x000071e58043ab10 = si_addr=0x000071e58043ab10.
    • That address is nowhere near libc's actual mapped range in the same log (e.g. RIP 0x71e58043ab10 vs libc mapped at 0x71e658600000-0x71e658628000 -- off by more than 4×10^11, not an in-library data offset).
    • ERR=0x14 on the page fault = I/D bit set (instruction fetch) + user-mode -- this is CPU trying to execute code at an unmapped address, not a read/write fault inside a malloc/free routine.
    • The libc.so.6+0x91651 / Instructions: (pc=0x...691651) shown in "Problematic frame" is a different address than the real RIP -- consistently, across every log checked. That's hs_err's post-fault frame-pointer/return-address guess, not where the crash actually happened. The earlier "disassembly matches a classic freelist-unlink pattern" read was against this wrong address.

    So: something is corrupting a code pointer (vtable slot, function pointer, or return address) that later gets called/returned into, landing on garbage -- not a heap-header stomp inside glibc's allocator internals. This significantly narrows the search: look for a stale/dangling CefRefPtr/vtable call surviving past its object's real lifetime (classic UAF-then-virtual-call shape), rather than any allocator-internals theory.

    Could not get a core dump to go further (core_pattern pipes to /wsl-capture-crash, which doesn't exist in this sandbox -- cores are silently dropped). gdb live-attach is already known to suppress this Heisenbug (documented earlier in this thread), so a real backtrace needs either a working core-dump path, or rr/reverse-debugging, or a hand-rolled SIGSEGV handler in issue10_repro that prints a raw backtrace via backtrace()/backtrace_symbols() before dying -- the last one is cheap and doesn't require ptrace, so it's the natural next step and shouldn't perturb the timing the way gdb does.

    Also added ISSUE10_REPRO_RACE_CLOSE mode to tools_native/issue10_repro/ (a second thread calling CloseBrowser(true) timed to race CefShutdown(), the one "late close" variant not yet tried) -- ran 20x in Debug and 20x in Release. All 20 Debug runs hit the already-known all_.empty() DCHECK immediately (expected -- that DCHECK fires synchronously at CefShutdown() entry if the browser is still open at all, regardless of race timing, so this mode doesn't actually exercise the race window against that specific check). All 20 Release runs survived. So this variant hasn't found a new repro path yet; not pursuing it further right now given the more promising in-hand-backtrace lead above.

    🤖 Generated with Claude Code

  8. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Architectural hypothesis (from the user, and it fits the RIP/si_addr evidence above well): CefShutdown()/CloseBrowser() force a synchronous teardown while async work is still in flight, stranding it -- a queued continuation later firing into torn-down state, rather than the shutdown itself corrupting the heap.

    Context::Shutdown() (native/context.cpp:276-327) does a fixed, unconditional 10x CefDoMessageLoopWork() pump then calls CefShutdown() -- no tracking of which async operations are actually still outstanding, no wait-for-real-completion-signal, no explicit pass marking live references dead before CEF's own teardown proceeds. Already directly tested that pumping more (200x over ~1s, 100x budget) doesn't help -- consistent with this theory: the fix isn't more blind waiting, it's an actual drain signal tied to specific outstanding work, not time.

    This matches the two already-fixed instances of this exact bug shape: the CefMessageRouter persistent-query-callback leak (a router removed without canceling its still-pending callback) and the late-finalizer guard (c8df5fc, a reactive per-call-site null check rather than a preventive drain).

    Audit surface -- every CefPostTask/CefPostDelayedTask site in native/, none of which Context::Shutdown() currently tracks or drains: CefMessageRouter_N.cpp:76,105, CefCookieManager_N.cpp:128,150, CefURLRequest_N.cpp:90, util_linux.cpp:105 (windowed-close's delayed task), and a dozen sites in CefBrowser_N.cpp (zoom, window size/bounds/parent, frame rate, view-source, async destroy). Any one of these still queued on TID_UI (or its captured CefRefPtr/raw pointer already dangling) when CefShutdown() tears down that task runner is a concrete match for the wild-jump fault.

    Not yet attempted, the natural next test: instrument how many tasks are actually queued on TID_UI when Context::Shutdown() begins pumping, across several runs -- if reliably non-zero on a crash and zero on a survival, that's a direct link. Real fix would be phased: stop accepting new async work, drain (a low-priority sentinel task posted last, waited on -- CEF doesn't expose queue depth directly) or explicitly cancel/tombstone every outstanding op, only then call CefShutdown().

    🤖 Generated with Claude Code

  9. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Ran the derisking probe for the shutdown-phasing theory above (plan preserved in this repo's plan/roadmap.md, 2026-09-03 entry, for continuity). Result: negative for this specific OSR crash.

    Added SetCefForJNIObjectHelper::GetLiveObjectCount() -- an atomic counter at the same AddRef()/Release() choke point the finalizer guard (c8df5fc) uses -- logged via JCEF_TRACE in Context::Shutdown() right before the pump and right before CefShutdown(). This measures "how many CEF-backed Java wrapper objects does JCEF think are still alive" at the exact moment of the crash-triggering call.

    8/8 crashing CefCommandLineTest-alone runs showed live_object_count=0 at both checkpoints. Sanity-checked the instrumentation itself is real, not just stuck at zero: a batch run including CefRequestTest showed live_object_count=4 with 70 CEF_ADDREF events logged.

    So: whatever's stranded at shutdown for this crash isn't a JNI-wrapped CEF object this choke point can see. Either it's purely Chromium-internal (and CEF's own CefMainRunner::StartShutdownOnUIThread() drain, which already exists precisely for this hazard, is supposed to handle that class of thing), or the mechanism for this specific crash isn't "outstanding async work" at all. Not building the quiesce-loop fix for the OSR path on this evidence -- leaving the counter in place since it's harmless and may be useful signal elsewhere.

    The windowed-close delayed-task hazard (util_linux.cpp's DestroyCefBrowser(), a real, separate, untested risk for windowed-mode shutdown specifically) is unaffected by this result and still worth fixing on its own merits.

    Next real lead, per the RIP/si_addr wild-jump evidence from the previous comment: a working core-dump path for a genuine postmortem backtrace (this sandbox's core_pattern pipes to a nonexistent /wsl-capture-crash, so that's currently blocked here) -- that would finally show what's actually being called through.

    🤖 Generated with Claude Code

  10. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Follow-up per direct feedback: "if [a leaked CefRequestContext] were true it should show up in the tracer" -- checked directly rather than continuing to speculate.

    Tracer confirms no leak: this exact crashing CefCommandLineTest-alone run shows CEF_ADDREF/CEF_RELEASE perfectly balanced, 52/52 (plus CEF_REFPTR_ADDREF/CEF_REFPTR_RELEASE 1/1 at the GetGlobalContext() site). So JCEF is not holding an unreleased CefRefPtr-tracked reference to anything, global CefRequestContext included -- consistent with the earlier ModifiedUtf8Test balance finding, now confirmed for this repro too.

    But there's a more precise finding from reading real CEF source (~/devel/cef/libcef/browser/browser_context.cc): the all_.empty() DCHECK's ImplManager::all_ tracks CefBrowserContext* instances, added/removed via AddImpl()/RemoveImpl(). CefBrowserContext::Shutdown() (which does the RemoveImpl()) asserts request_context_set_.empty() and is driven entirely by Chromium's own internal Profile/refcount teardown machinery -- not by anything that goes through JCEF's C++ code or its CefRefPtr/JNI-association tracer at all. That's why the tracer shows nothing wrong: this isn't a reference JCEF holds or leaks, it's a Chromium-internal object lifetime JCEF has zero visibility into.

    Sharper point, from this exact captured trace: Shutdown() CefShutdown() returned and Context::~Context() EXIT both fire cleanly in this run (no DCHECK this time -- confirms the Heisenbug property, sometimes it fires inside CefShutdown(), sometimes it doesn't and the crash comes later). Since the crash happens in a window where none of JCEF's own instrumented native code is executing at all, the race isn't "CefShutdown()'s own call graph outraces itself" -- it's that CefShutdown() returning to the caller doesn't guarantee CEF's background threads (ThreadPool workers, IO thread) have actually finished tearing down. Something there is still mid-flight when the JVM proceeds to its own exit/library-unload sequence, landing on a corrupted callable -- matching the earlier RIP/si_addr wild-jump evidence exactly.

    This also explains, retroactively, why the already-tried "pump CefDoMessageLoopWork() 200x over ~1s before calling CefShutdown()" experiment didn't help: that only services the UI-thread's own external-pump queue, which has no ability to wait on independent OS threads that don't route through it.

    Implication: if the actual race is after CefShutdown() returns, in threads Context::Shutdown() never touches, then no pre-CefShutdown() fix on JCEF's side (a wait-loop, a quiesce check, more pumping) can address it -- the fix would have to be post-CefShutdown(), or isn't reachable from JCEF's side at all. Next step: check whether CEF's public API offers any documented way to wait for background-thread shutdown completion after CefShutdown() returns (nothing obvious in include/cef_app.h from a first pass) -- if not, this may be a genuine upstream CEF/Chromium bug rather than something fixable in native/context.cpp.

    🤖 Generated with Claude Code

  11. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Follow-up per direct feedback: "you can track active browsers and open contexts... prevent new ones from spawning and apply actions from Java to bring to zero FIRST, then call shutdown." Checked whether that already happens, found a real gap in it, tested the fix for the gap directly -- still doesn't help.

    That architecture already exists. CefApp.java/CefClient.java: createClient() rejects new clients once SHUTTING_DOWN; dispose() synchronously drives every open CefClient to close all its browsers (browser.close(true), which a comment there confirms "synchronously triggers onBeforeClose()"), and only calls native shutdown() once clients_ is empty; CefApp.shutdown() (CefApp.java:509) even explicitly calls CefRequestContext.disposeGlobalContext() before N_Shutdown() -- added specifically for this exact DCHECK, per its own comment citing this issue by number.

    Found a real, concrete gap in that existing fix. disposeGlobalContext() only does anything if Java code explicitly called CefRequestContext.getGlobalContext() at some point (populating a static cache). A browser created with no explicit context (native/CefBrowser_N.cpp:1021-1042 passes nullptr straight to CefBrowserHost::CreateBrowser()) never populates that cache -- CEF silently uses its own internal global context instead. That's exactly this whole investigation's minimal-repro path (the warmup browser), so the existing fix is a silent no-op here.

    Tested the fix for this gap directly, not just theorized: temporarily forced the Java-side cache to populate before shutdown, recompiled, reran CefCommandLineTest alone 7x. Trace confirmed the mechanism engaged exactly as intended -- the context is created, then explicitly released via disposeGlobalContextNative(), before Context::Shutdown()'s pump even starts. Still crashed 7/7, identically. Reverted the experimental change.

    Conclusion: dropping JCEF's own Java-visible reference(s) to the global context isn't sufficient. CefBrowserContext::Shutdown() (real CEF source, browser_context.cc:212) requires ALL CefRequestContextImpl references gone -- there's at least one more JCEF never holds or sees, almost certainly the browser's own internal RenderProcessHost/WebContents teardown chain. browser.close(true)'s synchronous OnBeforeClose() signals "the browser told you it's closing," not "every Chromium-internal object tied to it is released" -- that teardown continues asynchronously regardless of what JCEF calls or waits on.

    Every JCEF-side lever tried this investigation has now been tested directly and refuted: more pumping before CefShutdown() (200x/~1s, no change), JCEF's own live-object count (reliably 0 on crashing runs), and now forcing the global context's Java-side reference to be dropped before shutdown (no change). The remaining reference(s) are internal to Chromium's own teardown chain, with no public API surface to observe, wait on, or force to completion.

    This is now reasonably strong evidence that, in this specific OSR manifestation, #10 isn't fixable by reordering or adding JCEF-side calls before CefShutdown(). Two paths forward: file this exact minimal, reproducible sequence upstream with CEF/Chromium, or try a JCEF-side mitigation that absorbs the race rather than provably closing it (an idle grace-period sleep -- no pumping, no other activity -- after browser close and/or after CefShutdown() returns; meaningfully different from the already-refuted "pump more" experiment since it doesn't contend with straggling background threads for CPU/scheduling). Neither attempted yet.

    🤖 Generated with Claude Code

  12. Thrameos commented on Sep 3, 2026

    @Thrameos
    OwnerAuthor

    Update: found and fixed a genuine, independent bug while regression-testing the last idea in this thread -- not #10 itself, but caught mid-investigation and worth recording here since it was blocking any full-suite validation.

    A real hang (not a crash): a full-suite run stalled 37+ minutes with zero output growth. jstack showed the main thread parked forever in TestSetupExtension.close()'s countdown_.await() (unbounded -- no timeout, unlike every other latch-wait in this codebase), while AWT-EventQueue-0 sat completely idle. Root cause: CefRequestHandlerCoverageTest's deliberate chrome://crash navigation kills its browser's renderer; closing that browser later during the suite-wide final CefApp.dispose() pass does not reliably fire a real OnBeforeClose() for a browser whose renderer already died. So CefClient.cleanupBrowser() never sees its browser list empty out, clientWasDisposed() never fires, CefApp's client set never empties, and native shutdown() is never invoked at all (confirmed: its own println never appeared in the log).

    Fixed by generalizing the existing windowed-close (X11 WM_DELETE_WINDOW) fallback -- a bounded, idempotency-guarded synthesized OnBeforeClose() -- to the OSR force-close path too (util::ScheduleOnBeforeCloseFallback(), native/util.h/util_linux.cpp, called from CefBrowser_N.cpp's N_Close()), plus a defensive 30s bound on TestSetupExtension.close()'s previously-unbounded wait. Verified: the exact repro now completes cleanly instead of hanging, and a full-suite run goes from a 37+ minute hang down to ~44s with only one pre-existing, already-documented flaky test remaining.

    Separately, a negative result worth recording: tried making CefClient.createBrowser() always explicitly resolve to CefRequestContext.getGlobalContext() instead of passing null through to CEF (so JCEF's disposeGlobalContext() fix is never a silent no-op). Confirmed via trace it engaged correctly and properly released the exact context object the browser used -- but it broke CefRequestContextTest.globalContextIsGlobal() deterministically (isGlobal() returned false), a real regression, most likely re-triggering the same "GetGlobalContext() doesn't guarantee the same pointer per call" hazard a prior fix in this exact area was built around, just at higher frequency. Reverted -- not committed. Still doesn't explain #10's actual crash either way (see the last two comments above).

    🤖 Generated with Claude Code

  13. Thrameos commented on Sep 4, 2026

    @Thrameos
    OwnerAuthor

    This should be investigated as a use after free as it smells like a stale object vtable jump.

  14. Thrameos commented on Sep 5, 2026

    @Thrameos
    OwnerAuthor

    New, more diagnosable lead (2026-09-05)

    Fixing an unrelated crash (a JCEF test misusing CefDownloadHandler.onBeforeDownload()'s
    return contract — see plan/tasks/20260905-26-download-shelf-check-crash.md on the
    coverage/phase1-value-objects-phase2-handlers branch) let a full-suite run get further
    than before, and it now hits a different, cleanly reproducible (2/2 identical) native
    crash
    in the same shutdown window as this issue — a real null-pointer dereference, not
    the earlier libc.so.6 wild-fetch signature.

    Repro (Release build, plain test job, no gdb/coverage build needed — hs_err_pid*.log
    is sufficient):

    tools/run_tests.sh linux64 Release --config debugPrint=true \
        --exclude-tag leak-sweep --exclude-tag process-isolated

    Crash: SEGV_MAPERR, si_addr=0x150 (null-pointer-plus-small-member-offset — a
    classic null CefRefPtr dereference), during CefApp.dispose()'s CefShutdown() call:

    C  [libcef.so+0x2e7ecb5]  CefBrowserInfo::RemoveFrame(content::RenderFrameHost*)+0x145
    C  [libcef.so+0x2e6d116]  non-virtual thunk to CefBrowserContentsDelegate::RenderFrameDeleted(...)
    C  [libcef.so+0x5eaed15]  WebContentsObserverList::NotifyObservers<...>(...)
    C  [libcef.so+0x5ec2b63]  content::WebContentsImpl::RenderFrameDeleted(...)
    C  [libcef.so+0x5cd3230]  content::RenderFrameHostImpl::~RenderFrameHostImpl()
    C  [libcef.so+0x5d1d35a]  content::RenderFrameHostManager::~RenderFrameHostManager()
    C  [libcef.so+0x5c1e567]  content::FrameTreeNode::~FrameTreeNode()
    C  [libcef.so+0x5cd4af0]  content::RenderFrameHostImpl::ResetChildren()
    C  [libcef.so+0x5c1c2d1]  content::FrameTree::Shutdown()
    C  [libcef.so+0x5ea1faf]  content::WebContentsImpl::~WebContentsImpl()
    C  [libcef.so+0x725d30c]  tabs::TabModel::~TabModel()
    C  [libcef.so+0x7261a81]  TabStripModel::SendDetachWebContentsNotifications(...)
    C  [libcef.so+0x7268758]  TabStripModel::CloseTabs(...)
    C  [libcef.so+0x7267fe4]  TabStripModel::CloseAllTabs()
    C  [libcef.so+0xd5d4314]  Browser::~Browser()
    C  [libcef.so+0xd5f3675]  BrowserManagerService::Shutdown()
    C  [libcef.so+0x9f69f46]  ProfileImpl::~ProfileImpl()
    C  [libcef.so+0x9f666ff]  ProfileDestroyer::DestroyOriginalProfileNow(...)
    C  [libcef.so+0x9f7305d]  ProfileManager::~ProfileManager()
    C  [libcef.so+0x9dff129]  BrowserProcessImpl::StartTearDown()
    ...  ChromeBrowserMainParts::PostMainMessageLoopRun() -> CefMainRunner::Shutdown() -> CefShutdown()
    

    Root cause, read from real CEF source (libcef/browser/browser_info.cc/.h, CEF
    commit 82195616d, matching this project's CEF_VERSION 146.0.10+g8219561):

    CefBrowserInfo::RemoveFrame() calls browser_->request_context()->... where browser_
    is a CefRefPtr<CefBrowserHostBase> member. CefBrowserInfo::BrowserDestroyed() nulls
    browser_, documented as "Always called after SetClosing and WebContentsDestroyed" — the
    assumption being that WebContentsDestroyed()'s own RemoveAllFrames() call already
    cleared every frame by then. But the real content::WebContentsImpl C++ object is not
    actually destroyed until later — it's kept alive by the TabStripModel/Browser/Profile
    object graph until process-wide teardown, which only runs synchronously inside
    CefShutdown() itself (BrowserProcessImpl::StartTearDown() → ProfileImpl::~ProfileImpl()
    → ... → ~WebContentsImpl() → FrameTree::Shutdown()). When that destructor chain tears
    down any remaining RenderFrameHostImpls, it calls RemoveFrame() a second time — on a
    CefBrowserInfo whose browser_ is already null, with no null-guard for this case.

    This looks like a real CEF-side lifetime-ordering bug (JCEF's browser-close nulls
    browser_ before the underlying Chromium object is actually destroyed), not something
    fixable from JCEF's side — no lever in JCEF's public API reaches this internal teardown
    ordering. Worth root-causing further or reporting upstream to CEF with this exact
    reproducible sequence, since it's now a clean, deterministic repro rather than the
    harder-to-diagnose wild-fetch signature originally filed here.

    Full write-up: plan/tasks/20260903-01-gh10-shutdown-sigsegv.md's 2026-09-05 addendum.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions