Skip to content

[Bug] gh gist edit truncates large gist file and risks data loss #11739

Description

@0xdevalias

Describe the bug

When editing a large gist (secret, so can't share directly) containing a single .md file (~1.6 MB, ~38,072 lines, ~1.7M bytes), using gh gist edit does not load the entire file content. The file is truncated to about 900K bytes (17,000 lines), instead of loading the full content.

This presents a risk: If a user edits and saves, the truncated file will overwrite the gist, resulting in data loss.

Affected version

⇒ gh version
gh version 2.67.0 (2025-02-11)
https://github.com/cli/cli/releases/tag/v2.67.0
  • Platform: macOS

Steps to reproduce the behavior

  1. Have a gist with a large file (e.g., 1.6 MB .md file, ~38,000 lines).
  2. Run gh gist edit <gist-id> (with any editor).
  3. Observe that the file content is cut off (e.g., only 900K of 1.6M bytes loaded).
  4. (Optional) Save the file, which would (presumably) overwrite the gist with a truncated version.

Expected vs actual behavior

Expected:

  • gh gist edit should load and allow editing of the entire file, regardless of size, or should display a clear error message if the file is too large to edit safely.

Actual:

  • The file is silently truncated. No warning or error is given. Risk of accidental data loss if saved (presumably).

Logs

⇒ gh gist clone <gist-id> && cd <gist-id>

⇒ du -sh foo.md
1.6M  foo.md

⇒ cat foo.md | wc -l
38072

⇒ cat foo.md | wc -c
1685200

⇒ EDITOR=cat gh gist edit <gist-id> | wc -l
17000

⇒ EDITOR=cat gh gist edit <gist-id> | wc -c
921600

⇒ EDITOR=cat gh gist edit <gist-id> | wc -c | numfmt --to=iec
900K

Activity

  1. 0xdevalias commented on Sep 15, 2025

    @0xdevalias
    Author

    I haven't deeply verified this, but here is some additional context about the likely root cause and related code-paths; as generated by GitHub Codex (GPT-5) when fed this repo + issue:

    Summary of findings

    • Symptom recap

      • Editing a large single-file gist (~1.6 MB) via gh gist edit results in only ~900 KiB of content being loaded and then, if saved, would overwrite the gist with the truncated content.

    Relevant code paths

    • Where the gist content is fetched and prepared for editing

    • The API fetch used by gh gist edit

      • editRun calls shared.GetGist(...) to obtain gist metadata and file contents, then uses those contents for editing. That function is the likely place where file contents are populated and where truncation must be handled. In editRun:

    Why truncation happens

    • Primary cause: not handling API truncation for large gist files

      • The GitHub Gists REST API returns, for each file, a content string and a truncated flag. For sufficiently large files, the API sets truncated: true and only includes a shortened content preview; the full content must be fetched from files[...].raw_url.
      • In the current flow, gh gist edit uses shared.GetGist to populate Files[].Content and then passes that to the editor unchanged. There is no logic in editRun to detect/resolve truncated gist content before editing. This explains why a large file appears cut and presents a risk of overwriting with the truncated text when saved.
    • Secondary observation: the 900 KiB boundary

      • The measurement of 921,600 bytes (exactly 900 KiB) is a strong indicator that, beyond the API’s own truncation behavior, there may also be an internal limit encountered during editing (e.g., when $EDITOR is a filter like cat). Regardless, the primary correctness issue is that the initial content used for editing is already truncated when fetched.

    Ideal fix

    • Always edit the full file contents

      • In shared.GetGist (or specifically within the gist edit command flow before opening the editor), detect truncated file content and fetch the full text from files[...].raw_url.

        • If you want to minimize network and memory use, do on-demand fetching only for the file selected for editing (not for all files in a multi-file gist).
        • After fetching the raw content, populate gist.Files[filename].Content with the full text so the editor opens with the complete data.
    • Add guardrails

      • If a file reports truncated: true or file.size > len(content), and the raw content fetch fails, show a clear error with guidance (e.g., “File is too large to edit safely. Try gh gist clone and edit locally.”).
      • Do not proceed to editing with truncated content.
    • Tests to prevent regressions

      • Unit test: mock GetGist to return a file with truncated=true and a raw_url; stub the raw fetch to return >1 MB content; verify edit opens full content and the update payload uses the complete text.
      • Unit/integration test: simulate an editor that echoes large input (like cat) and verify no truncation occurs and that the update request body contains the full content.
      • Test for the guardrail path: simulate truncated content with failing raw fetch and assert a clear, actionable error.

    Concrete spots to adjust

    • Ensure shared.GetGist (in the gist shared package) or the gh gist edit flow:

      • Checks for truncated (REST) or analogous GraphQL truncation flags, and then follows files[...].raw_url for the file to be edited before invoking the editor.
    • Review editor invocation in editRun:

      • The code currently uses:

        opts.Edit(editorCommand, filename, gistFile.Content, opts.IO)

        with gistFile.Content as the initial buffer. Ensure that this path uses temp-file editing (so no stdout capture limits apply) before reading the final content back.


    By addressing the fetch-time truncation, gh gist edit will reliably load and save the full contents of large gist files and eliminate the risk of silent data loss.


    And this additional information RE: the related API docs:

    Further API documentation references

    • GitHub Gist REST API

    • GitHub Gist API behavior

      • For large files, the REST API sets truncated: true and provides only a shortened content preview.
      • Full file contents should be fetched via the raw_url property, which delivers the complete file contents in plain text.
  2. babakks commented on Sep 15, 2025

    @babakks
    Member

    Thanks for reporting this bug, @0xdevalias! 🙏

    I can confirm there's a bug in the gist command set. There's this boolean truncated field in the API response that gh does not respect. I think it's probably because the field has been introduced after the gist command set was implemented.

    Notes on the fix

    We need to check all gist retrieval endpoints used and respect the truncated field if it's set to true. The truncated field is available down to GHES 3.14 which is the minimum supported version. So, we're okay introducing it into gh.

    We should be careful not to fetch all contents that are marked as truncated; rather we should only fetch the full content if we actually need the file (like when editing). The URL to fetch the full gist content is available via the raw_url field.

  3. added
    priority-3Affects a small number of users or is largely cosmetic
    gh-gistrelating to the gh gist command
    help wanted candidateIssue may be marked as help-wanted but not yet ready to accept PR
    and removed on Sep 15, 2025
  4. babakks commented on Sep 16, 2025

    @babakks
    Member

    I'll post the acceptance criteria for this change in the next comment.

    As a hint to anyone who's going to work on this, the problem is rooted at fetching Gists from the API. When there's a large file, the API truncates its content and sets the truncated to true for that file. Note that, in such cases, further files in the response will have an empty content (again with truncated: true).

    So, the fix basically involves two things:

    1. Adding truncated and raw_url fields to the response type (See below). Note that since this type is used for both GET and POST HTTP requests, it's necessary to add omitempty to the new fields' JSON tags so that they don't appear in POST requests.
      type GistFile struct {
      Filename string `json:"filename,omitempty"`
      Type string `json:"type,omitempty"`
      Language string `json:"language,omitempty"`
      Content string `json:"content"`
      }
    2. Checking all places where we read file contents from a fetched gist. If the truncated field is set, then we should hit the URL populated in raw_url field to get the full content. Needless to say we shouldn't hit the API if we don't really need the content of the file.
  5. babakks commented on Sep 16, 2025

    @babakks
    Member

    Tip

    On Linux/Unix, a large file can be created with this:

    head /dev/urandom -c 2MiB | base64 > large-file.txt 
    

    Acceptance Criteria

    Affected experiences

    Given I have a gist with ID <GIST-ID> that has a large file (>1MB) named <LARGE-FILE>
    When I run gh gist view <GIST-ID> --raw --filename <LARGE-FILE>
    Then I see the entire content of the file (not truncated)

    Given I have a gist with ID <GIST-ID> that has a large file (>1MB) named <LARGE-FILE>
    When I run gh gist edit <GIST-ID> and select the <LARGE-FILE> in the prompt
    Then I see the entire gist file content in the editor

    Given I have a gist with ID <GIST-ID> that has a large file (>1MB) named <LARGE-FILE>
    When I run gh gist edit <GIST-ID>, select the <LARGE-FILE> in the prompt, and add a few chars to the end of the file
    Then the gist file is updated successfully (this can be verified by using gist clone <GIST-ID>)

    Unaffected experiences

    Given I have a gist with ID <GIST-ID>, and a few large local files named <LARGE-FILE-1>, <LARGE-FILE-2>, and <LARGE-FILE-3>
    When I run

    gh gist edit <GIST-ID> --add <LARGE-FILE-1>
    gh gist edit <GIST-ID> --add <LARGE-FILE-2>
    gh gist edit <GIST-ID> --add <LARGE-FILE-3>
    

    Then I see all files are correctly added to the gist (this can be verified by using gist clone <GIST-ID>)

    Unaffected experiences

    Given I have a few large local files named <LARGE-FILE-1>, <LARGE-FILE-2>, and <LARGE-FILE-3>
    When I run gh gist create <LARGE-FILE-1> <LARGE-FILE-2> <LARGE-FILE-3>
    Then I see all files are correctly added to the new gist (this can be verified by using gist clone <GIST-ID>)

  6. added and removed
    help wanted candidateIssue may be marked as help-wanted but not yet ready to accept PR
    on Sep 16, 2025
  7. babakks commented on Sep 16, 2025

    @babakks
    Member

    @0xdevalias, since this is a niche case with gists, I put down the A/C for this issue so that external contributors can help with it.

    That said, I'm still responsible to provide you with a workaround, as of which I'd recommend you treating gists as ordinary git repositories. For example, to edit a gist file you can do it like this:

    gh gist clone <GIST-ID> my-gist
    cd my-gist
    echo foo >> existing-file
    git add existing-file
    git commit -m 'updated existing file'
    git push

    Of course, you can always use gh api gists commands to interact with the API directly, but I wouldn't recommend that due to the complexity around putting together request body data and handling intricate details.

    Please let me know if the git approach helps with your situation. 🙏

  8. 0xdevalias commented on Oct 3, 2025

    @0xdevalias
    Author

    I put down the A/C for this issue so that external contributors can help with it

    @babakks Thanks :) Looks like there is already a PR for it:


    That said, I'm still responsible to provide you with a workaround, as of which I'd recommend you treating gists as ordinary git repositories.

    Please let me know if the git approach helps with your situation.

    @babakks 👌🏻 This was the workaround I identified as well, and it worked well enough for me to refactor the impacted gist to separate it into multiple files (rather than trying to embed huge JSON responses inside markdown code blocks in a single file)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workinggh-gistrelating to the gh gist commandhelp wantedContributions welcomepriority-3Affects a small number of users or is largely cosmetic

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions