You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
A refused session delete leaks an instance reference: the session stays pinned in that process and every later delete there is refused #1765
EpisodicMemoryManager.delete_episodic_session takes a reference on the cached instance before it checks whether the instance is in use, and does not release it when it refuses (packages/server/src/memmachine_server/episodic_memory/episodic_memory_manager.py, as of 82b6c6f58 main):
ref_count=awaitself._instance_cache.get_ref_count(session_key) # line 301instance=awaitself._instance_cache.get(session_key) # line 302, takes a referenceifinstanceandref_count>0:
raiseSessionInUseError(session_key, ref_count) # line 304, reference not released
MemoryInstanceCache.get increments the node's ref_count (episodic_memory/instance_lru_cache.py:132, :139). Every refused delete therefore leaves the count one higher than the number of requests holding the instance, so it never returns to zero. close_session reads the count and raises before it calls get (episodic_memory_manager.py:357-362) and does not leak.
Reproduced by calling the manager directly with one cached instance held by one request, at instance_cache_size 0 and at 100, with the same result:
Step
Result
A request holds the instance
ref_count 1
delete_episodic_session
raises SessionInUseError, ref_count 2
The request releases its reference
ref_count 1, instance still cached, not closed
delete_episodic_session again, nothing in flight
raises SessionInUseError
Why it matters
One delete that overlaps a request on the same session leaves that process in a state only a restart clears.
The instance is pinned. With instance_cache_size: 0 an instance is dropped only when its count reaches zero (instance_lru_cache.py:208). With a cache, LRU eviction and the lifetime check both skip a node whose count is positive (:180, :226). The instance is never closed.
A refused delete leaves the reference count as it found it: read the count and raise before calling get, as close_session does.
Notes
Found while reading the instance cache for #1571. The leak is reproduced at the manager; its effect on the open paths is read from the code, not run end to end. The defect goes away when the reference counts and the in-use check are removed (#1755, PR (e)); the fix above is independent of that and of the deletion queue (#1749).
What happened
EpisodicMemoryManager.delete_episodic_sessiontakes a reference on the cached instance before it checks whether the instance is in use, and does not release it when it refuses (packages/server/src/memmachine_server/episodic_memory/episodic_memory_manager.py, as of82b6c6f58main):MemoryInstanceCache.getincrements the node'sref_count(episodic_memory/instance_lru_cache.py:132,:139). Every refused delete therefore leaves the count one higher than the number of requests holding the instance, so it never returns to zero.close_sessionreads the count and raises before it callsget(episodic_memory_manager.py:357-362) and does not leak.Reproduced by calling the manager directly with one cached instance held by one request, at
instance_cache_size0 and at 100, with the same result:ref_count1delete_episodic_sessionSessionInUseError,ref_count2ref_count1, instance still cached, not closeddelete_episodic_sessionagain, nothing in flightSessionInUseErrorWhy it matters
One delete that overlaps a request on the same session leaves that process in a state only a restart clears.
instance_cache_size: 0an instance is dropped only when its count reaches zero (instance_lru_cache.py:208). With a cache, LRU eviction and the lifetime check both skip a node whose count is positive (:180,:226). The instance is never closed.episodic_memory_manager.py:142-156,:253-266), so every later request gets the pinned instance although the row is inDeletedstatus. Without the leak, at the default cache size the instance would be dropped when the overlapping requests finish, and the next open would raiseSessionDeletedError. With it, the stuck project in Deleting a project that is in use returns 204 and then never completes: the deletion worker gives up on SessionInUseError, leaves a partial deletion, and only a restart retries #1577 keeps answering after its in-flight requests have finished.Expected
A refused delete leaves the reference count as it found it: read the count and raise before calling
get, asclose_sessiondoes.Notes
Found while reading the instance cache for #1571. The leak is reproduced at the manager; its effect on the open paths is read from the code, not run end to end. The defect goes away when the reference counts and the in-use check are removed (#1755, PR (e)); the fix above is independent of that and of the deletion queue (#1749).
🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.