Skip to content

[Bug]: Three short-term memory lifecycle defects around cancellation and reset #1719

Description

@xiongzubiao

Describe the bug

Three independent defects in the short-term memory lifecycle, all present on main at d57f5cb.

1. One cancelled caller cancels summarization for every caller. ShortTermMemoryConsolidator.wait_until_done() awaits the shared worker task directly, so the awaiting task's cancellation propagates into the worker. Every other caller waiting on that worker is cancelled with it, and the episodes the worker had already drained are lost with no retry, because _run_summary_loop clears _pending_episodes before the summary completes.

2. Reset keeps the summary of removed content. ShortTermMemory._do_reset() clears the episode deque and the counters but leaves the consolidator's summary in place. After clear_memory() the instance stays open, so an LLM summary of the cleared content remains queryable from it. delete_session_episodes() clears through clear_memory(), so the same holds for deleted episodes. close() leaves the summary behind as well, although a closed instance refuses reads. Separately, the copy the consolidator persists through save_short_term_memory() is never cleared, so after clear_memory() any new instance for the session restores the summary from it.

3. A cancelled close leaves the instance open. ShortTermMemory.close() sets _closed only after await self._do_reset() returns. If the close is cancelled — a shutdown, an abandoned request — the instance stays open and keeps accepting writes after its owner has been dropped.

Steps to reproduce

These are reproducible at the class level; the steps below are what I actually exercised.

  1. Shared summarization. Construct a ShortTermMemory whose model blocks inside summarization, add enough episodes to trigger an eviction, then start two concurrent readers of the summary and cancel one of them. Observed: the surviving reader dies with asyncio.exceptions.CancelledError, raised at the await task in wait_until_done(). Two consequences follow from the code rather than from that run: the worker itself is cancelled, because awaiting a task unshielded propagates the awaiter's cancellation into it, and the batch it had already drained is lost, because _run_summary_loop copies and clears _pending_episodes before the summary completes.
  2. Summary survives reset. Add episodes, let summarization finish, call clear_memory(), then read the episodes and the summary. Observed: the episodes come back empty and the summary is still the pre-clear value. Then create a new instance for the same session. Observed, with the test suite's in-memory stand-in for the session data manager: it restores the pre-clear summary from the persisted copy. delete_session_episodes() reaches the same _do_reset(), so the same holds for a session whose episodes were deleted. I have not exercised that end to end through the REST layer; the link is by code path, not by a request I ran.
  3. Cancelled close. Call close() and cancel it while _do_reset() is awaiting the consolidator. Observed: a following add_episodes() does not raise ShortTermMemoryClosedError — the instance is still open and accepts the write.

Expected behavior

  1. An abandoned caller cancels only itself. The shared worker and the other waiters keep going, and no drained batch is silently dropped.
  2. After clear_memory() or deleting a session's episodes, no summary derived from the removed content remains readable, from that instance or from a new one for the session, and close() does not leave it held in memory.
  3. close() leaves the instance closed and cleared whether or not it was cancelled.

Environment

  • OS: Linux
  • MemMachine version: main at d57f5cb (the package version is derived from VCS, so there is no release number to quote)
  • Development language version: Python 3.13.7

Additional context

Related to #1514, fixed by #1517: a deadlock from acquiring the non-reentrant read lock while already holding it. That fix removed the live nested acquire, but the same area still carries the hazard. ShortTermMemory.get_summary() acquires self._lock itself, so a future call to it from inside a locked section deadlocks again whenever a writer is queued, and several locked sections exist today. Two tests in test_rw_locks.py also state the opposite contract — one comment reads "does not block when acquiring read lock again", and a test asserts the nesting is deadlock-free without qualification. Both pass only because no writer is queued; with a writer waiting, the same nesting hangs, so as written they present the nesting that deadlocked in #1514 as a guarantee.

Activity

  1. changed the title [-][Bug]: Short-term memory lifecycle: one cancelled caller cancels shared summarization, reset keeps the summary of removed content, and a cancelled close leaves the instance open[/-] [+][Bug]: Three short-term memory lifecycle defects around cancellation and reset[/+] on Sep 29, 2026
  2. self-assigned this
    on Sep 29, 2026
  3. added theissue type on Sep 30, 2026
  4. added
    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processes
    on Oct 2, 2026
  5. edwinyyyu commented on Oct 3, 2026

    @edwinyyyu
    Contributor

    Where short-term memory stands in multi-replica deployments is decided in #1757 (declared process-scoped) and #1748; if the project removes short-term memory instead, this closes with it. Until then the three defects stand as filed.


    🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processes

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions