Repository navigation
[Bug]: MCP Streamable HTTP client never DELETEs retired sessions — 287 sessions / 146 sockets accumulate against one server #162
Description
Activity
Follow-up: the SDK already implements termination — it's just never called
I dug into where the omission actually lives, since that determines whether this is an auggie fix or an upstream SDK fix. It's auggie's client lifecycle.
1. The bundled TS SDK has a working
terminateSession()From
/opt/homebrew/lib/node_modules/@augmentcode/auggie/augment.mjs(v0.9.0, minified), in the Streamable HTTP client transport:async terminateSession() { if (this._sessionId) try { let e = await this._commonHeaders(), n = { ...this._requestInit, method: "DELETE", headers: e, signal: this._abortController?.signal }, r = await (this._fetch ?? fetch)(this._url, n); if (await r.body?.cancel(), !r.ok && r.status !== 405) throw new g2(r.status, `Failed to terminate session: ${r.statusText}`); this._sessionId = void 0 } catch (e) { throw this.onerror?.(e), e } }
It sends the spec-conformant
DELETEwith themcp-session-idheader (set by_commonHeaders()), and correctly tolerates405for servers that don't support termination. Nothing wrong with it.2. Nothing ever calls it
Searching the whole 12,988,515-byte bundle:
Symbol Occurrences terminateSession1 (the definition above) method:"DELETE"1 (inside that same definition) One occurrence means definition only, zero call sites. The remaining 22
"DELETE"string literals in the bundle are all unrelated (OpenTelemetry HTTP-method enums, undici's method table, llhttp constants).3.
close()deliberately does not terminateThe transport's own teardown, immediately above
terminateSessionin the same class:async close() { this._reconnectionTimeout && (clearTimeout(this._reconnectionTimeout), this._reconnectionTimeout = void 0), this._abortController?.abort(), this.onclose?.() }
It aborts the in-flight SSE
GETand firesonclose— noDELETE. This matches the upstream TS SDK, where termination is intentionally opt-in (a client may want to reconnect to the same session later). So the SDK is behaving as designed; the caller is expected to decide. Auggie never does, so retired transports leak server-side forever.4. Same-transport control: a Go client on the identical server does send DELETE
While verifying the workaround I caught our own MCP-discovery client (go-sdk v1.4.0) hitting the same
:5757endpoint. Full request history of one of its sessions:17:38:28.616 GET "" 17:38:28.616 POST notifications/initialized 17:38:28.616 POST tools/list 17:38:28.620 DELETE ""Seven such clean
DELETEs in the same window, versus zero from auggie. Different SDK language, but it confirms the server accepts and correctly handlesDELETE— the missing call is entirely on the client side.Suggested fix
Call
terminateSession()before/insideclose()on the paths where auggie retires a transport for good (session replacement, agent shutdown, MCP server reconfiguration). Guarding with a try/catch is enough — the SDK already treats405as success, and a failedDELETEshouldn't block teardown.Workaround verification (for other server authors landing here)
The server-side
SessionTimeoutmitigation is now confirmed in production. After enabling a 30-minute idle timeout and restarting:Metric Before After (36 min uptime) Distinct session IDs on SSE GET287 63 Distinct session IDs that ever POSTed14 63 Ghost sessions (GET-only, never used) 273 0 Peak TCP conns from busiest auggie pid 146 16 Server goroutines ~960 ~207 The reaper fired exactly at the 30-minute boundary (34 sessions expired,
404s clustered at 18:08–18:11, then stopped). Notably it did not disrupt anything: zero errors, POST traffic continued uninterrupted. That's because auggie's oversized session pool absorbs the404by failing over to a sibling session — which is only possible because of this bug.Which is the reason #149 matters: if this issue is fixed without also fixing the
404→ re-initialize path, that absorbing pool disappears and idle-timeout servers will start producing genuinely wedged tool calls. Worth fixing the two together.
What happened?
Auggie's Streamable HTTP MCP client never sends
DELETEto terminate an MCP session it has stopped using. Each new session mints a freshMcp-Session-Idviainitialize, and the previous one is abandoned rather than terminated: its TCP connection and standalone SSEGETkeepalive stream stay open indefinitely, so the server keeps the session (and everything it pins) alive forever.Over a normal working day a single
auggieprocess accumulates well over a hundred simultaneously-open MCP sessions against one server.Measured against a local MCP server (Go,
modelcontextprotocol/go-sdkv1.4.0, Streamable HTTP on127.0.0.1:5757/mcp), with server-side request logging:Mcp-Session-Ids seen on SSEGETMcp-Session-Ids that ever sent aPOST(i.e. actually in use)DELETErequests received from any client, everauggiepidSo ~95% of the live sessions are pure keepalive ghosts: they were
initialized, used briefly or not at all, then abandoned while still holding an open SSE stream.lsof -nP -iTCP:5757 -sTCP:ESTABLISHED, grouped by client pid:Each ghost session costs the server a goroutine set, a file descriptor and an open SSE stream. On our side ~89% of the server process's goroutines were attributable to these abandoned sessions. It is also self-harming for Auggie: 146 sockets to one server saturates the client's own connection pool, and we correlated this with Auggie's own
MCP initialization timed out after Nsgate firing (budgets escalating 11s → 634s) while the server's handler was answering in under 1ms throughout.The MCP spec (2025-03-26, §Session Management) is explicit here:
What did you expect to happen?
When Auggie retires an MCP transport — the ACP/agent session it belonged to ended, the client reconnected, the server list was reloaded, or the process is shutting down — it should send
DELETE <endpoint>with theMcp-Session-Idheader (and close the SSEGET) before dropping the transport, so the server can release the session immediately.Steady state should be roughly one live MCP session per active agent session, not an ever-growing pile.
Steps to reproduce
~/.augment/settings.json), e.g.http://127.0.0.1:5757/mcp.Mcp-Session-Id).Mcp-Session-Ids climbs monotonically (~8 new sessions/hour/process in our case);GETkeepalive, never aPOST;DELETEis ever received;lsof -nP -iTCP:<port> -sTCP:ESTABLISHEDshows the connection count from theauggiepid growing without bound.Server-side termination is well tolerated, which confirms the sessions really are abandoned. We deleted two sessions by hand with
curl -X DELETE -H 'Mcp-Session-Id: <id>' <endpoint>(both returned204):POSTs ever): the client retried theGETtwice, got404both times, then tore the transport down cleanly — no error storm, sibling sessions unaffected;tools/calltraffic: subsequent tool calls continued with zero user-visible disruption, transparently failing over to another pooled session the same process already held.Auggie version
0.34.0 (commit 81042879), installed via npm/Homebrew (
@augmentcode/auggie)Request ID
n/a — this is a client-side transport-lifecycle issue observed from the MCP server side, not a model request failure.
Environment details
Environment
github.com/modelcontextprotocol/go-sdkv1.4.0, Streamable HTTP transport,127.0.0.1:5757/mcpJSONResponse: trueAnything else we need to know?
Relationship to #149. These are two halves of the same lifecycle gap and a fix for one without the other is incomplete:
404and wedges on the dead session ID.Concretely: a server operator's only defence against the pile-up is a
SessionTimeout, and turning that on is exactly what exposes #149. Fixing both —DELETEon retire here, and404→ re-initialize in #149 — closes the loop.Workaround for server authors (what we did): set an idle-session timeout on the Streamable HTTP handler. With the Go SDK,
mcp.StreamableHTTPOptions{SessionTimeout: 30 * time.Minute}. Note the SDK's idle timer keys offPOSTactivity only — the periodic SSEGETrecycle does not refresh it — which is what makes this effective against exactly the ghost population described above. Server-side reaping is a mitigation, not a fix: it cannot reclaim a session any earlier than the timeout, and it depends on clients handling the resulting404correctly (#149).Happy to supply the raw server logs or re-run the measurement against a build with a fix.