Skip to content

[Bug]: MCP Streamable HTTP client never DELETEs retired sessions — 287 sessions / 146 sockets accumulate against one server #162

Description

@inercia

What happened?

Auggie's Streamable HTTP MCP client never sends DELETE to terminate an MCP session it has stopped using. Each new session mints a fresh Mcp-Session-Id via initialize, and the previous one is abandoned rather than terminated: its TCP connection and standalone SSE GET keepalive stream stay open indefinitely, so the server keeps the session (and everything it pins) alive forever.

Over a normal working day a single auggie process accumulates well over a hundred simultaneously-open MCP sessions against one server.

Measured against a local MCP server (Go, modelcontextprotocol/go-sdk v1.4.0, Streamable HTTP on 127.0.0.1:5757/mcp), with server-side request logging:

Metric (single server, ~29-minute log window) Value
Distinct Mcp-Session-Ids seen on SSE GET 287
Distinct Mcp-Session-Ids that ever sent a POST (i.e. actually in use) 14
DELETE requests received from any client, ever 0
Established TCP connections from the busiest single auggie pid 146

So ~95% of the live sessions are pure keepalive ghosts: they were initialized, used briefly or not at all, then abandoned while still holding an open SSE stream.

lsof -nP -iTCP:5757 -sTCP:ESTABLISHED, grouped by client pid:

146 node 50364
 60 node 61552
 34 node 90807
 19 node 54851
 11 node 67619

Each ghost session costs the server a goroutine set, a file descriptor and an open SSE stream. On our side ~89% of the server process's goroutines were attributable to these abandoned sessions. It is also self-harming for Auggie: 146 sockets to one server saturates the client's own connection pool, and we correlated this with Auggie's own MCP initialization timed out after Ns gate firing (budgets escalating 11s → 634s) while the server's handler was answering in under 1ms throughout.

The MCP spec (2025-03-26, §Session Management) is explicit here:

Clients that no longer need a particular session ... SHOULD send an HTTP DELETE to the MCP endpoint with the Mcp-Session-Id header, to explicitly terminate the session.

What did you expect to happen?

When Auggie retires an MCP transport — the ACP/agent session it belonged to ended, the client reconnected, the server list was reloaded, or the process is shutting down — it should send DELETE <endpoint> with the Mcp-Session-Id header (and close the SSE GET) before dropping the transport, so the server can release the session immediately.

Steady state should be roughly one live MCP session per active agent session, not an ever-growing pile.

Steps to reproduce

  1. Register any Streamable HTTP MCP server globally (~/.augment/settings.json), e.g. http://127.0.0.1:5757/mcp.
  2. Enable request logging on that server (method, path, Mcp-Session-Id).
  3. Run Auggie normally for an hour, starting/ending several agent sessions (or driving it via ACP so sessions are created and retired repeatedly).
  4. Observe on the server:
    • the number of distinct Mcp-Session-Ids climbs monotonically (~8 new sessions/hour/process in our case);
    • the great majority of them only ever issue the standalone GET keepalive, never a POST;
    • not a single DELETE is ever received;
    • lsof -nP -iTCP:<port> -sTCP:ESTABLISHED shows the connection count from the auggie pid growing without bound.

Server-side termination is well tolerated, which confirms the sessions really are abandoned. We deleted two sessions by hand with curl -X DELETE -H 'Mcp-Session-Id: <id>' <endpoint> (both returned 204):

  • an idle ghost (SSE only, zero POSTs ever): the client retried the GET twice, got 404 both times, then tore the transport down cleanly — no error storm, sibling sessions unaffected;
  • a live session that was actively carrying tools/call traffic: subsequent tool calls continued with zero user-visible disruption, transparently failing over to another pooled session the same process already held.

Auggie version

0.34.0 (commit 81042879), installed via npm/Homebrew (@augmentcode/auggie)

Request ID

n/a — this is a client-side transport-lifecycle issue observed from the MCP server side, not a model request failure.

Environment details

Environment
  • OS: macOS (Apple Silicon)
  • Shell: zsh
  • Tool/CLI version: auggie 0.34.0 (commit 81042879)
  • MCP server under test: Go, github.com/modelcontextprotocol/go-sdk v1.4.0, Streamable HTTP transport, 127.0.0.1:5757/mcp
  • Transport: Streamable HTTP (MCP spec 2025-03-26), JSONResponse: true

Anything else we need to know?

Relationship to #149. These are two halves of the same lifecycle gap and a fix for one without the other is incomplete:

Concretely: a server operator's only defence against the pile-up is a SessionTimeout, and turning that on is exactly what exposes #149. Fixing both — DELETE on retire here, and 404 → re-initialize in #149 — closes the loop.

Workaround for server authors (what we did): set an idle-session timeout on the Streamable HTTP handler. With the Go SDK, mcp.StreamableHTTPOptions{SessionTimeout: 30 * time.Minute}. Note the SDK's idle timer keys off POST activity only — the periodic SSE GET recycle does not refresh it — which is what makes this effective against exactly the ghost population described above. Server-side reaping is a mitigation, not a fix: it cannot reclaim a session any earlier than the timeout, and it depends on clients handling the resulting 404 correctly (#149).

Happy to supply the raw server logs or re-run the measurement against a build with a fix.

Activity

  1. inercia commented on Aug 7, 2026

    @inercia
    Author

    Follow-up: the SDK already implements termination — it's just never called

    I dug into where the omission actually lives, since that determines whether this is an auggie fix or an upstream SDK fix. It's auggie's client lifecycle.

    1. The bundled TS SDK has a working terminateSession()

    From /opt/homebrew/lib/node_modules/@augmentcode/auggie/augment.mjs (v0.9.0, minified), in the Streamable HTTP client transport:

    async terminateSession() {
      if (this._sessionId) try {
        let e = await this._commonHeaders(),
            n = { ...this._requestInit, method: "DELETE", headers: e, signal: this._abortController?.signal },
            r = await (this._fetch ?? fetch)(this._url, n);
        if (await r.body?.cancel(), !r.ok && r.status !== 405)
          throw new g2(r.status, `Failed to terminate session: ${r.statusText}`);
        this._sessionId = void 0
      } catch (e) { throw this.onerror?.(e), e }
    }

    It sends the spec-conformant DELETE with the mcp-session-id header (set by _commonHeaders()), and correctly tolerates 405 for servers that don't support termination. Nothing wrong with it.

    2. Nothing ever calls it

    Searching the whole 12,988,515-byte bundle:

    Symbol Occurrences
    terminateSession 1 (the definition above)
    method:"DELETE" 1 (inside that same definition)

    One occurrence means definition only, zero call sites. The remaining 22 "DELETE" string literals in the bundle are all unrelated (OpenTelemetry HTTP-method enums, undici's method table, llhttp constants).

    3. close() deliberately does not terminate

    The transport's own teardown, immediately above terminateSession in the same class:

    async close() {
      this._reconnectionTimeout && (clearTimeout(this._reconnectionTimeout), this._reconnectionTimeout = void 0),
      this._abortController?.abort(),
      this.onclose?.()
    }

    It aborts the in-flight SSE GET and fires onclose — no DELETE. This matches the upstream TS SDK, where termination is intentionally opt-in (a client may want to reconnect to the same session later). So the SDK is behaving as designed; the caller is expected to decide. Auggie never does, so retired transports leak server-side forever.

    4. Same-transport control: a Go client on the identical server does send DELETE

    While verifying the workaround I caught our own MCP-discovery client (go-sdk v1.4.0) hitting the same :5757 endpoint. Full request history of one of its sessions:

    17:38:28.616 GET    ""
    17:38:28.616 POST   notifications/initialized
    17:38:28.616 POST   tools/list
    17:38:28.620 DELETE ""
    

    Seven such clean DELETEs in the same window, versus zero from auggie. Different SDK language, but it confirms the server accepts and correctly handles DELETE — the missing call is entirely on the client side.

    Suggested fix

    Call terminateSession() before/inside close() on the paths where auggie retires a transport for good (session replacement, agent shutdown, MCP server reconfiguration). Guarding with a try/catch is enough — the SDK already treats 405 as success, and a failed DELETE shouldn't block teardown.

    Workaround verification (for other server authors landing here)

    The server-side SessionTimeout mitigation is now confirmed in production. After enabling a 30-minute idle timeout and restarting:

    Metric Before After (36 min uptime)
    Distinct session IDs on SSE GET 287 63
    Distinct session IDs that ever POSTed 14 63
    Ghost sessions (GET-only, never used) 273 0
    Peak TCP conns from busiest auggie pid 146 16
    Server goroutines ~960 ~207

    The reaper fired exactly at the 30-minute boundary (34 sessions expired, 404s clustered at 18:08–18:11, then stopped). Notably it did not disrupt anything: zero errors, POST traffic continued uninterrupted. That's because auggie's oversized session pool absorbs the 404 by failing over to a sibling session — which is only possible because of this bug.

    Which is the reason #149 matters: if this issue is fixed without also fixing the 404 → re-initialize path, that absorbing pool disappears and idle-timeout servers will start producing genuinely wedged tool calls. Worth fixing the two together.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions