Skip to content

extra-data mechanism not suitable for multi-gigabyte downloads #6828

Description

@night199uk

Checklist

  • I agree to follow the Code of Conduct that this project adheres to.
  • I have searched the issue tracker for a bug that matches the one I want to file, without success.
  • If this is an issue with a particular app, I have tried filing it in the appropriate issue tracker for the app (e.g. under https://github.com/flathub/) and determined that it is an issue with Flatpak itself.
  • This issue is not a report of a security vulnerability (see here if you need to report a security issue).

Flatpak version

Flatpak 1.19.1

What Linux distribution are you using?

Fedora Silverblue

Linux distribution version

44

What architecture are you using?

x86_64

How to reproduce

This is an upstream bug in glib g_compute_checksum_for_bytes and g_compute_checksum_for_data.

I'm creating this flatpak issue as a tracker since flatpak makes heavy used of g_compute_checksum_for_bytes internally.

Create a flatpak with an extra-data blob which is >8GB.

The checksum verification will fail; sha256sum on the extra-data blob gives a different result to Flatpak.

Expected Behavior

The sha256sum for extra-data blobs >8GB should match.

Actual Behavior

The sha256sum for extra-data blobs >8GB do not match.

Additional Information

I opened:
GNOME/glib#4052
to track the upstream issue.

Activity

  1. smcv commented on Sep 17, 2026

    @smcv
    Collaborator

    Create a flatpak with an extra-data blob which is >8GB.

    What app needs this, and why?

    If Flatpak is loading the entire extra-data blob into memory at once (which it must be, if it's submitting > 8G to g_compute_checksum_for_bytes()), then the extra-data mechanism is clearly not designed for dealing with this much data at a time!

    If it makes sense to have extra-data that's multiple gigabytes, then Flatpak shouldn't load the whole thing into memory at once; it should iterate through a stream (file or pipe or something), feeding some smaller, more reasonable amount of data into the checksum machinery.

    Or if not that, you should interpret this as a sign that the extra-data mechanism isn't designed for data blobs this large, and the app you're (presumably) trying to create should do something else, not extra-data.

  2. night199uk commented on Sep 18, 2026

    @night199uk
    Author

    Practically, the extra-data mechanism is being used to Flatpak package apps for which we don't have distribution rights for the binaries. i.e. commercial software. It allows the user to install software via Flatpak, however the software binaries are resolved at install time: the repo does not need to store the commercial binaries and it is transparent to the Flatpak user.

    See Flatpark for examples of using extra-data to package commercial software:
    https://flatpark.org/

    Specifically - what I am working on - we use extra-data to download the 11GB installater .zip for DaVinci Resolve. This is how the software is packaged by Blackmagic and we have no control of that.

    Thus, I think it is critical that we want to be able to download arbitrary sized blobs for e.g. game or software installers that could be many gigabytes big. Thus, I agree: loading the whole extra-data file into memory to checksum it is not ideal for the purpose it's being used for. That said, I'm just a passer-by trying to do my work and not a Flatpak maintainer :-)

    I have some new features for extra-data I hope to propose as pull-requests here soon, so I will offer to land a PR to fix the "extra-data checksum chunking" in Flatpak at the same time.

  3. changed the title [-][Bug]: glib g_compute_checksum_for_bytes fails on buffers > 8GB (tracking issue)[/-] [+]extra-data mechanism not suitable for multi-gigabyte proprietary apps[/+] on Sep 18, 2026
  4. smcv commented on Sep 18, 2026

    @smcv
    Collaborator

    Practically, the extra-data mechanism is being used to Flatpak package apps for which we don't have distribution rights for the binaries. i.e. commercial software.

    I'm aware of how extra-data works in general. Its original use-case was to install things like the Nvidia proprietary driver, which is considerably smaller. I think everyone involved in Flatpak considers extra-data to be a necessary evil that should be reluctantly used as little as possible, rather than a core feature that gets a lot of time put into it.

    Specifically - what I am working on - we use extra-data to download the 11GB installater .zip for DaVinci Resolve. This is how the software is packaged by Blackmagic and we have no control of that.

    The extra-data mechanism, as currently implemented, isn't suitable for a data blob this large.

    One way to address this would be for the thing that is packaged as a Flatpak app to be just a bootstrapper/downloader, which downloads the actual proprietary thing into its subdirectory of ~/.var/app/. I believe this is how the Discord app works, and it's certainly how Steam works. It does have the disadvantage that the proprietary data can't be shared between users of a multi-user system.

    Another way to address this would be to improve the extra-data mechanism so it downloads the extra-data blob into a file (not into RAM) and checksums it by iterating through the file, maybe a few MiB at a time. Ideally it would also have resumable downloads.

    Or, instead of using extra-data, there could be something that converts a large proprietary blob downloaded out-of-band into a non-redistributable .flatpak bundle that can be installed locally, more like the approach taken by https://github.com/kujeger/flatpak-gog (which outputs Flatpak apps) or https://salsa.debian.org/games-team/game-data-packager (which currently only outputs .deb and similar distro-level packages, but in principle it could output Flatpak apps if someone devoted enough time to it).

  5. night199uk commented on Sep 18, 2026

    @night199uk
    Author

    So we already have a working package for Resolve which works in what sounds like it is your "flatpak-gog style" here:
    https://github.com/night199uk/resolve-flatpak

    i.e. users can create their own .flatpak on demand. But the friction for that is high and users of e.g. media software shouldn't have to understand Flatpak packaging and bundling. This is creating a large barrier of entry.

    My goal would be to get to a state where we can distribute a Resolve installer on Flathub; which would mean either using extra-data or the Discord/Steam style that you mentioned. For e.g. Silverblue and atomic distributions this is more critical; as the alternative is building a complex distrobox which also doesn't really meet the bar for ease of use. I have not investigated the Steam/Discord installers; as extra-data seemed fit for purpose up until this bug. extra-data is what my users were asking me for, thus the route I took.

    So here's the question - is there interest in making extra-data more usable for this purpose or is extra-data seen as a dead end? I intended to land a few PRs over the next few days to make extra-data better, but if it's seen as a dead end I might be wasting my time.

  6. changed the title [-]extra-data mechanism not suitable for multi-gigabyte proprietary apps[/-] [+]extra-data mechanism not suitable for multi-gigabyte downloads[/+] on Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions