Skip to content

[Bug]: flatpak run aborts with "Extension ... has invalid merge-dirs" since 1.18.1 #6783

Description

@CleoMenezesJr

Checklist

  • I agree to follow the Code of Conduct that this project adheres to.
  • I have searched the issue tracker for a bug that matches the one I want to file, without success.
  • If this is an issue with a particular app, I have tried filing it in the appropriate issue tracker for the app (e.g. under https://github.com/flathub/) and determined that it is an issue with Flatpak itself.
  • This issue is not a report of a security vulnerability (see here if you need to report a security issue).

Flatpak version

1.18.1

What Linux distribution are you using?

Fedora Silverblue

Linux distribution version

Fedora 44

What architecture are you using?

x86_64

How to reproduce

  1. Install an app on the org.freedesktop.Platform//25.08 runtime, for example org.gnome.Epiphany//stable from Flathub. The GL extension org.freedesktop.Platform.GL.default//25.08 on this system comes from the gnome-nightly remote (Mesa 26.1.5).
  2. Launch the app a few times, killing it between runs:
    flatpak kill org.gnome.Epiphany
    timeout 15 flatpak run --branch=stable --command=epiphany org.gnome.Epiphany
    
  3. The failure shows up roughly 3 out of 5 runs. When it does, the launch stops with:
    erro: Extension org.freedesktop.Platform.GL.default has invalid merge-dirs
    ** (epiphany:2): ERROR **: Connection: failed to receive credentials: Era esperado ler apenas um byte para receber credenciais, mas foi lido zero byte
    

The app then dies with SIGABRT. I ran it 5 times and the pattern held: whenever "invalid merge-dirs" appears, the app crashes; when it doesn't, everything works.

Expected Behavior

A failure while opening an optional GL extension's merge-dirs should not abort the whole launch. Missing merge-dirs were silently ignored before flatpak 1.18.1, so it should either start without the extension or print a warning and continue.

Actual Behavior

flatpak run aborts the launch with Extension org.freedesktop.Platform.GL.default has invalid merge-dirs whenever glnx_chaseat() on a merge-dir fails with an errno other than ENOENT/ENOTDIR (for example ELOOP, EXDEV, EACCES or EMFILE). I believe this is a regression from the security fix for GHSA-w69g-9x8j-7p8f:

  • Commit ead0ca6b (main) / ad044fc (1.18 branch), "run: Use RESOLVE_BENEATH for host-side extension file access"
  • Before that, glnx_dirfd_iterator_init_at silently ignored failures on extension merge-dirs
  • Now glnx_chaseat(files_dfd, ext->merge_dirs[i], GLNX_CHASE_RESOLVE_BENEATH | GLNX_CHASE_MUST_BE_DIRECTORY) fails fatally with "Extension %s has invalid merge-dirs" (common/flatpak-run.c, flatpak_run_add_extension_args, lines 246-267)

The GL extension ships its merge-dirs as symlinks (OpenCL -> lib/OpenCL, glvnd -> share/glvnd, vdpau -> lib/vdpau), and with RESOLVE_BENEATH those now fail on transient conditions (fd exhaustion, a race with an extension update, ELOOP/EXDEV) instead of being ignored.

The side effect on WebKitGTK apps is a hard crash: when the GL mount fails, the WebKitWebProcess dies before the credential handshake, and the UI process aborts via g_error() in IPC::Connection::remoteProcessID() (ConnectionGLib.cpp:541), ending in SIGABRT.

Additional Information

  • Reproduced with the rpm package flatpak-1.18.1-1.fc44 on Fedora 44, Wayland, locale pt_BR.UTF-8.
  • Same result when launching from the desktop file (flatpak run --branch=stable --arch=x86_64 --command=epiphany --file-forwarding org.gnome.Epiphany @@u %U @@).
  • Control test: launching without --command (flatpak run --branch=stable org.gnome.Epiphany) worked fine in my tests. The difference seems to be timing rather than the flag itself.
  • I could not find any existing issue or fix: nothing in the tracker mentions "invalid merge-dirs", and no commit after ad044fc touches merge_dirs in flatpak-run.c. There is no 1.18.2 release either.
  • The GL extension deployment looks intact on disk (files/lib with dri/, gbm/, vulkan/, libEGL_mesa.so, and so on), so this does not look like a broken installation.

Activity

  1. swick commented on Aug 18, 2026

    @swick
    Collaborator

    Tried to reproduce this but failed. There is another issue with Epiphany I'm running into after a successful sandbox setup:

    (epiphany:2): Gtk-CRITICAL **: 13:44:56.687: Unable to register the application: GDBus.Error:org.freedesktop.DBus.Error.NameHasNoOwner: Could not activate remote peer 'org.a11y.atspi.Registry': unit failed
    

    Can you get the actual error that causes invalid merge-dirs? I could just turn it into a warning, but I don't understand what's going on here.

  2. CleoMenezesJr commented on Aug 18, 2026

    @CleoMenezesJr
    Author

    Got it, i traced the launch with a tiny LD_PRELOAD shim that logs every openat2/openat/readlinkat/statx with path, result and errno, then looped until it broke again (crash on attempt 2 of 8, SIGABRT).

    the failing call is right in the middle of the merge-dirs loop, between lib/gbm and the crash. Verbatim from the trace (crash attempt 2 of 8. strerror output is localized to pt_BR since that's my system locale):

    openat2(egl/egl_external_platform.d) -> -1 errno=2 (Arquivo ou diretório inexistente)
    openat2(OpenCL/vendors) -> 26 errno=2 (Arquivo ou diretório inexistente)
    openat2(lib/dri) -> 26 errno=0 (Sucesso)
    openat2(lib/d3d) -> -1 errno=2 (Arquivo ou diretório inexistente)
    openat2(lib/gbm) -> 26 errno=2 (Arquivo ou diretório inexistente)
    openat2(vulkan/explicit_layer.d) -> -1 errno=11 (Recurso temporariamente indisponível)
    

    (that last one is EAGAIN, "Resource temporarily unavailable")

    errno=11 (EAGAIN) isn't in the ENOENT/ENOTDIR whitelist in flatpak_run_add_extension_args, so the launch dies with "invalid merge-dirs". it showed up twice in ~4000 traced syscalls (same call again at the second occurrence), which matches the intermittency (roughly 3 of 5 launches failing earlier).

    also ran a tiny program doing the exact same openat2 with RESOLVE_BENEATH | RESOLVE_NO_MAGICLINKS on every merge-dir of the installed extension: always succeeds in steady state (8 open, 2 legit ENOENT). so it's transient, not a broken install.

    the at-spi error doesn't show up here, so no idea, probably unrelated.

  3. swick commented on Aug 18, 2026

    @swick
    Collaborator

    from man openat2()

           EAGAIN how.resolve contains either RESOLVE_IN_ROOT or
                  RESOLVE_BENEATH, and the kernel could not ensure that a
                  ".." component didn't escape (due to a race condition or
                  potential attack).  The caller may choose to retry the
                  openat2() call.
    

    So this doesn't appear to be entirely hypothetical then. Should be easy to fix in libglnx. Can you try to build with this:

    diff --git ./glnx-chase.c ../glnx-chase.c
    index 2251cb7..49e606a 100644
    --- ./glnx-chase.c
    +++ ../glnx-chase.c
    @@ -679,17 +679,20 @@ glnx_chaseat_full (int                 dirfd,
             openat2_resolve |= RESOLVE_NO_SYMLINKS;
     
           how = (struct open_how) {
             .flags = openat2_flags,
             .mode = 0,
             .resolve = openat2_resolve,
           };
     
    -      fd = openat2 (dirfd, path, &how, sizeof (how));
    +      do
    +        fd = openat2 (dirfd, path, &how, sizeof (how));
    +      while (fd < 0 && errno == EAGAIN);
    +
           if (fd < 0)
             {
               /* If the syscall is not implemented (ENOSYS) or blocked by
                * seccomp (EPERM), we need to fall back to the manual path chasing
                * via open_tree. */
               if (G_IN_SET (errno, ENOSYS, EPERM))
                 can_openat2 = FALSE;
               else
  4. swick commented on Aug 18, 2026

    @swick
    Collaborator

    You can also try to add GLNX_CHASE_DEBUG_NO_OPENAT2 to glnx_chaseat(files_dfd, ext->merge_dirs[i], GLNX_CHASE_RESOLVE_BENEATH | GLNX_CHASE_MUST_BE_DIRECTORY) which should work around the issue.

  5. CleoMenezesJr commented on Aug 18, 2026

    @CleoMenezesJr
    Author

    tested your patch and it works. instead of building flatpak i ran the launch with an LD_PRELOAD shim that does the same retry on openat2 (same logic as your diff), then looped 8 times: 0 crashes.

    two real EAGAINs happened during the run and the retry absorbed both, so it's hitting the actual kernel race, not just luck.

    the only difference from your diff: my shim capped the retries at 1000, yours loops forever on EAGAIN. doesn't matter in practice, the race resolves in a few iterations.

    opened a PR with the fix as you wrote it: #6786

  6. added a commit that references this issue on Aug 18, 2026
    db79045
  7. added a commit that references this issue on Aug 18, 2026
    fd0df8d
  8. added a commit that references this issue on Aug 27, 2026
    a53f8eb
  9. tanty commented on Sep 7, 2026

    @tanty

    I started experiencing this mid-July with

    $ flatpak --user run org.gnome.Epiphany.Devel

    Also with Canary.

    Eventually, I also had this same error spit in the logs.

    Reproducible with Debian Trixie, which packs Flatpak 1.16.6.

    I managed to avoid this problem by downgrading to Flatpak 1.14.10

    $ apt install flatpak=1.14.10-1~deb12u2
    $ apt-mark hold flatpak
  10. smcv commented on Sep 7, 2026

    @smcv
    Collaborator

    I managed to avoid this problem by downgrading to Flatpak 1.14.10

    This is re-introducing several serious security vulnerabilities, so please avoid that. (Or, at least, don't run any apps that you don't 100% trust with the older Flatpak version!)

    Reproducible with Debian Trixie, which packs Flatpak 1.16.6.

    Can you reproduce the problem under strace?

    If you find that strace reports openat2(something) returning EAGAIN, then the fix for this that is present in 1.18.2 could be backported to 1.16.x.

  11. mcatanzaro commented on Sep 8, 2026

    @mcatanzaro
    Collaborator

    This should indeed be fixed by 1.18.2. Distros that have manually backported the fix that introduced this regression -- I'm not sure which commit it was, but presumably one of the various CVE fixes -- will need to backport the fix for this issue as well.

    @swick it looks like this issue was supposed to be closed when landing the fix, but that didn't happen because it was in a subtree. I will close it now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions