Describe the bug
Three independent defects in the short-term memory lifecycle, all present on main at d57f5cb.
1. One cancelled caller cancels summarization for every caller. ShortTermMemoryConsolidator.wait_until_done() awaits the shared worker task directly, so the awaiting task's cancellation propagates into the worker. Every other caller waiting on that worker is cancelled with it, and the episodes the worker had already drained are lost with no retry, because _run_summary_loop clears _pending_episodes before the summary completes.
2. Reset keeps the summary of removed content. ShortTermMemory._do_reset() clears the episode deque and the counters but leaves the consolidator's summary in place. After clear_memory() the instance stays open, so an LLM summary of the cleared content remains queryable from it. delete_session_episodes() clears through clear_memory(), so the same holds for deleted episodes. close() leaves the summary behind as well, although a closed instance refuses reads. Separately, the copy the consolidator persists through save_short_term_memory() is never cleared, so after clear_memory() any new instance for the session restores the summary from it.
3. A cancelled close leaves the instance open. ShortTermMemory.close() sets _closed only after await self._do_reset() returns. If the close is cancelled — a shutdown, an abandoned request — the instance stays open and keeps accepting writes after its owner has been dropped.
Steps to reproduce
These are reproducible at the class level; the steps below are what I actually exercised.
- Shared summarization. Construct a
ShortTermMemory whose model blocks inside summarization, add enough episodes to trigger an eviction, then start two concurrent readers of the summary and cancel one of them. Observed: the surviving reader dies with asyncio.exceptions.CancelledError, raised at the await task in wait_until_done(). Two consequences follow from the code rather than from that run: the worker itself is cancelled, because awaiting a task unshielded propagates the awaiter's cancellation into it, and the batch it had already drained is lost, because _run_summary_loop copies and clears _pending_episodes before the summary completes.
- Summary survives reset. Add episodes, let summarization finish, call
clear_memory(), then read the episodes and the summary. Observed: the episodes come back empty and the summary is still the pre-clear value. Then create a new instance for the same session. Observed, with the test suite's in-memory stand-in for the session data manager: it restores the pre-clear summary from the persisted copy. delete_session_episodes() reaches the same _do_reset(), so the same holds for a session whose episodes were deleted. I have not exercised that end to end through the REST layer; the link is by code path, not by a request I ran.
- Cancelled close. Call
close() and cancel it while _do_reset() is awaiting the consolidator. Observed: a following add_episodes() does not raise ShortTermMemoryClosedError — the instance is still open and accepts the write.
Expected behavior
- An abandoned caller cancels only itself. The shared worker and the other waiters keep going, and no drained batch is silently dropped.
- After
clear_memory() or deleting a session's episodes, no summary derived from the removed content remains readable, from that instance or from a new one for the session, and close() does not leave it held in memory.
close() leaves the instance closed and cleared whether or not it was cancelled.
Environment
- OS: Linux
- MemMachine version:
main at d57f5cb (the package version is derived from VCS, so there is no release number to quote)
- Development language version: Python 3.13.7
Additional context
Related to #1514, fixed by #1517: a deadlock from acquiring the non-reentrant read lock while already holding it. That fix removed the live nested acquire, but the same area still carries the hazard. ShortTermMemory.get_summary() acquires self._lock itself, so a future call to it from inside a locked section deadlocks again whenever a writer is queued, and several locked sections exist today. Two tests in test_rw_locks.py also state the opposite contract — one comment reads "does not block when acquiring read lock again", and a test asserts the nesting is deadlock-free without qualification. Both pass only because no writer is queued; with a writer waiting, the same nesting hangs, so as written they present the nesting that deadlocked in #1514 as a guarantee.
Describe the bug
Three independent defects in the short-term memory lifecycle, all present on
mainatd57f5cb.1. One cancelled caller cancels summarization for every caller.
ShortTermMemoryConsolidator.wait_until_done()awaits the shared worker task directly, so the awaiting task's cancellation propagates into the worker. Every other caller waiting on that worker is cancelled with it, and the episodes the worker had already drained are lost with no retry, because_run_summary_loopclears_pending_episodesbefore the summary completes.2. Reset keeps the summary of removed content.
ShortTermMemory._do_reset()clears the episode deque and the counters but leaves the consolidator's summary in place. Afterclear_memory()the instance stays open, so an LLM summary of the cleared content remains queryable from it.delete_session_episodes()clears throughclear_memory(), so the same holds for deleted episodes.close()leaves the summary behind as well, although a closed instance refuses reads. Separately, the copy the consolidator persists throughsave_short_term_memory()is never cleared, so afterclear_memory()any new instance for the session restores the summary from it.3. A cancelled close leaves the instance open.
ShortTermMemory.close()sets_closedonly afterawait self._do_reset()returns. If the close is cancelled — a shutdown, an abandoned request — the instance stays open and keeps accepting writes after its owner has been dropped.Steps to reproduce
These are reproducible at the class level; the steps below are what I actually exercised.
ShortTermMemorywhose model blocks inside summarization, add enough episodes to trigger an eviction, then start two concurrent readers of the summary and cancel one of them. Observed: the surviving reader dies withasyncio.exceptions.CancelledError, raised at theawait taskinwait_until_done(). Two consequences follow from the code rather than from that run: the worker itself is cancelled, because awaiting a task unshielded propagates the awaiter's cancellation into it, and the batch it had already drained is lost, because_run_summary_loopcopies and clears_pending_episodesbefore the summary completes.clear_memory(), then read the episodes and the summary. Observed: the episodes come back empty and the summary is still the pre-clear value. Then create a new instance for the same session. Observed, with the test suite's in-memory stand-in for the session data manager: it restores the pre-clear summary from the persisted copy.delete_session_episodes()reaches the same_do_reset(), so the same holds for a session whose episodes were deleted. I have not exercised that end to end through the REST layer; the link is by code path, not by a request I ran.close()and cancel it while_do_reset()is awaiting the consolidator. Observed: a followingadd_episodes()does not raiseShortTermMemoryClosedError— the instance is still open and accepts the write.Expected behavior
clear_memory()or deleting a session's episodes, no summary derived from the removed content remains readable, from that instance or from a new one for the session, andclose()does not leave it held in memory.close()leaves the instance closed and cleared whether or not it was cancelled.Environment
mainatd57f5cb(the package version is derived from VCS, so there is no release number to quote)Additional context
Related to #1514, fixed by #1517: a deadlock from acquiring the non-reentrant read lock while already holding it. That fix removed the live nested acquire, but the same area still carries the hazard.
ShortTermMemory.get_summary()acquiresself._lockitself, so a future call to it from inside a locked section deadlocks again whenever a writer is queued, and several locked sections exist today. Two tests intest_rw_locks.pyalso state the opposite contract — one comment reads "does not block when acquiring read lock again", and a test asserts the nesting is deadlock-free without qualification. Both pass only because no writer is queued; with a writer waiting, the same nesting hangs, so as written they present the nesting that deadlocked in #1514 as a guarantee.