Describe the bug
Summary
A search and an add arriving concurrently on the same session, while a short-term-memory summarization is in flight, deadlock each other permanently. The session is then unusable for the lifetime of the process: every later request on it blocks, including close(), and POST /api/v2/projects/delete returns 204 without deleting anything.
Other sessions are unaffected and /api/v2/health continues to return 200, so the server looks healthy from outside while individual conversations die.
Reproduced on main at 2d28c1c with stock settings (message_capacity left at its default of 64000).
Root cause
AsyncRWLock is writer-preferring and not reentrant: a waiting writer holds _read_gate while it waits for _writer_lock (common/rw_locks.py:80-89). That is what prevents writer starvation, and it is correct on its own.
ShortTermMemory.get_short_term_memory_context acquires the read lock at short_term_memory.py:411, then re-acquires the same lock at :432:
async with self._lock.read_lock(): # :411 first acquire
...
await self._consolidator.wait_until_done() # :414 suspends for the whole LLM call
...
return list(episodes), await self.get_summary() # :432 -> get_summary -> :379
# -> _wait_for_summary_to_finish -> :217
# -> self._lock.read_lock() SECOND acquire
Nested read acquisition is harmless until a writer arrives between the two:
| Step |
Search task |
Add task |
| 1 |
acquires read lock — _readers=1, _writer_lock held by the reader cohort |
|
| 2 |
suspends at :414 awaiting the summarization task |
|
| 3 |
|
add_episodes (:199) wants the write lock: takes _read_gate, blocks on _writer_lock |
| 4 |
LLM returns, resumes, re-acquires the read lock → blocks on _read_gate |
|
| 5 |
blocked forever |
blocked forever |
Steps to reproduce
Reproduction
Two failing tests, no server required. Save as test_deadlock_repro.py and run uv run --all-extras pytest test_deadlock_repro.py -q.
import asyncio
import datetime
import pytest
from memmachine_server.common.episode_store.episode_model import Episode
from memmachine_server.common.language_model.language_model import LanguageModel
from memmachine_server.common.rw_locks import AsyncRWLock
from memmachine_server.episodic_memory.short_term_memory.short_term_memory import (
ShortTermMemory,
ShortTermMemoryParams,
)
class SlowLLM(LanguageModel):
"""Stands in for any real provider: a summary call that takes a second."""
async def generate_response(
self, system_prompt=None, user_prompt=None, tools=None,
tool_choice=None, max_attempts=1,
):
await asyncio.sleep(1.0)
return ("summary text", None)
async def generate_parsed_response(self, *args, **kwargs):
raise NotImplementedError
async def generate_response_with_token_usage(self, *args, **kwargs):
raise NotImplementedError
def _episode(i: int) -> Episode:
return Episode(
uid=f"e{i}",
content=f"message {i} " + "x" * 40,
session_key="s1",
created_at=datetime.datetime.now(datetime.UTC),
producer_id="user",
producer_role="user",
sequence_num=i,
)
@pytest.mark.asyncio
async def test_concurrent_search_and_add_deadlock():
"""A search and an add on the same session, while a summary is in flight."""
memory = ShortTermMemory(
ShortTermMemoryParams(
session_key="s1",
llm_model=SlowLLM(),
summary_prompt_system="system",
summary_prompt_user="{episodes} {summary} {max_length}",
message_capacity=100,
)
)
# Overflow the budget so the background summarizer starts.
await memory.add_episodes([_episode(1), _episode(2), _episode(3)])
async def search():
return await memory.get_short_term_memory_context("query", limit=10)
async def add():
await asyncio.sleep(0.3) # arrive while the summary call is in flight
return await memory.add_episodes([_episode(4)])
# Hangs forever on main. Also leaves the session permanently unusable:
# every later call, including close(), blocks on the same gate.
await asyncio.wait_for(asyncio.gather(search(), add()), timeout=10)
@pytest.mark.asyncio
async def test_nested_read_lock_with_waiting_writer():
"""The underlying lock behaviour, isolated from ShortTermMemory."""
lock = AsyncRWLock()
async def reader():
async with lock.read_lock():
await asyncio.sleep(0.1) # a writer arrives during this await
async with lock.read_lock(): # blocks forever
pass
async def writer():
await asyncio.sleep(0.05)
async with lock.write_lock():
pass
await asyncio.wait_for(asyncio.gather(reader(), writer()), timeout=3)
Both fail on 2d28c1c:
FAILED test_deadlock_repro.py::test_concurrent_search_and_add_deadlock - TimeoutError
FAILED test_deadlock_repro.py::test_nested_read_lock_with_waiting_writer - TimeoutError
2 failed in 16.83s
Over HTTP, default settings
With message_capacity at its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue one POST /api/v2/memories/search and one POST /api/v2/memories 0.4 s apart.
[ 0.03s] create: 201
[ 0.92s] filled the short-term budget; summarizer in flight
[ 20.92s] search: *** no response in 20s ***
[ 21.33s] add (concurrent): *** no response in 20s ***
[ 21.33s] health: 200 healthy
kill -USR1 <pid> — using the existing server/diagnostics.py — shows both tasks parked:
Task 'Task-1077' [pending]
episodic_memory/episodic_memory.py:325 in _query_short_term_memory
short_term_memory.py:432 return list(episodes), await self.get_summary()
short_term_memory.py:379 await self._wait_for_summary_to_finish()
short_term_memory.py:217 async with self._lock.read_lock():
Task 'Task-1091' [pending]
short_term_memory.py:199 async with self._lock.write_lock():
Lock state at that moment: _readers=1, _read_gate.locked()=True, _writer_lock.locked()=True.
Expected behavior
Reproduction
Two failing tests, no server required. Save as test_deadlock_repro.py and run uv run --all-extras pytest test_deadlock_repro.py -q.
import asyncio
import datetime
import pytest
from memmachine_server.common.episode_store.episode_model import Episode
from memmachine_server.common.language_model.language_model import LanguageModel
from memmachine_server.common.rw_locks import AsyncRWLock
from memmachine_server.episodic_memory.short_term_memory.short_term_memory import (
ShortTermMemory,
ShortTermMemoryParams,
)
class SlowLLM(LanguageModel):
"""Stands in for any real provider: a summary call that takes a second."""
async def generate_response(
self, system_prompt=None, user_prompt=None, tools=None,
tool_choice=None, max_attempts=1,
):
await asyncio.sleep(1.0)
return ("summary text", None)
async def generate_parsed_response(self, *args, **kwargs):
raise NotImplementedError
async def generate_response_with_token_usage(self, *args, **kwargs):
raise NotImplementedError
def _episode(i: int) -> Episode:
return Episode(
uid=f"e{i}",
content=f"message {i} " + "x" * 40,
session_key="s1",
created_at=datetime.datetime.now(datetime.UTC),
producer_id="user",
producer_role="user",
sequence_num=i,
)
@pytest.mark.asyncio
async def test_concurrent_search_and_add_deadlock():
"""A search and an add on the same session, while a summary is in flight."""
memory = ShortTermMemory(
ShortTermMemoryParams(
session_key="s1",
llm_model=SlowLLM(),
summary_prompt_system="system",
summary_prompt_user="{episodes} {summary} {max_length}",
message_capacity=100,
)
)
# Overflow the budget so the background summarizer starts.
await memory.add_episodes([_episode(1), _episode(2), _episode(3)])
async def search():
return await memory.get_short_term_memory_context("query", limit=10)
async def add():
await asyncio.sleep(0.3) # arrive while the summary call is in flight
return await memory.add_episodes([_episode(4)])
# Hangs forever on main. Also leaves the session permanently unusable:
# every later call, including close(), blocks on the same gate.
await asyncio.wait_for(asyncio.gather(search(), add()), timeout=10)
@pytest.mark.asyncio
async def test_nested_read_lock_with_waiting_writer():
"""The underlying lock behaviour, isolated from ShortTermMemory."""
lock = AsyncRWLock()
async def reader():
async with lock.read_lock():
await asyncio.sleep(0.1) # a writer arrives during this await
async with lock.read_lock(): # blocks forever
pass
async def writer():
await asyncio.sleep(0.05)
async with lock.write_lock():
pass
await asyncio.wait_for(asyncio.gather(reader(), writer()), timeout=3)
Both fail on 2d28c1c:
FAILED test_deadlock_repro.py::test_concurrent_search_and_add_deadlock - TimeoutError
FAILED test_deadlock_repro.py::test_nested_read_lock_with_waiting_writer - TimeoutError
2 failed in 16.83s
Over HTTP, default settings
With message_capacity at its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue one POST /api/v2/memories/search and one POST /api/v2/memories 0.4 s apart.
[ 0.03s] create: 201
[ 0.92s] filled the short-term budget; summarizer in flight
[ 20.92s] search: *** no response in 20s ***
[ 21.33s] add (concurrent): *** no response in 20s ***
[ 21.33s] health: 200 healthy
kill -USR1 <pid> — using the existing server/diagnostics.py — shows both tasks parked:
Task 'Task-1077' [pending]
episodic_memory/episodic_memory.py:325 in _query_short_term_memory
short_term_memory.py:432 return list(episodes), await self.get_summary()
short_term_memory.py:379 await self._wait_for_summary_to_finish()
short_term_memory.py:217 async with self._lock.read_lock():
Task 'Task-1091' [pending]
short_term_memory.py:199 async with self._lock.write_lock():
Lock state at that moment: _readers=1, _read_gate.locked()=True, _writer_lock.locked()=True.
Environment
MemMachine 0.3.10.dev20+g2d28c1c1e, main at 2d28c1c
- Python 3.12.13,
uv 0.11.26
- macOS 26.6.1 (arm64)
- Config: SQLite episode/segment/vector stores, local
sentence-transformer embedder, openai-chat-completions provider pointed at a local stub. The deadlock is in the locking layer and is independent of the storage backends.
Suggested fix
Suggested fix
The second acquisition buys nothing: the read lock is already held, and wait_until_done() has already been awaited at :414. No new summarization can start in between, because starting one requires the write lock.
--- a/packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py
+++ b/packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py
@@ -429,7 +429,7 @@ class ShortTermMemory:
break
episodes.appendleft(e)
length += msg_len
- return list(episodes), await self.get_summary()
+ return list(episodes), await self._consolidator.summary
@staticmethod
def _compute_episode_length(episode: Episode) -> int:
With this applied:
test_concurrent_search_and_add_deadlock passes.
uv run --all-extras pytest packages/server/server_tests → 1868 passed, 3 skipped, unchanged from before the patch.
test_nested_read_lock_with_waiting_writer still fails, by design — the one-line change removes the only caller that re-enters the lock, not the underlying property. Worth deciding separately whether AsyncRWLock should be made reentrant or simply documented as not reentrant, since the next caller to nest a read lock will hit the same trap. There is precedent: fe2790d ("Fix deadlock due to non-reentrant lock pool", #1312).
Describe the bug
Summary
A search and an add arriving concurrently on the same session, while a short-term-memory summarization is in flight, deadlock each other permanently. The session is then unusable for the lifetime of the process: every later request on it blocks, including
close(), andPOST /api/v2/projects/deletereturns204without deleting anything.Other sessions are unaffected and
/api/v2/healthcontinues to return200, so the server looks healthy from outside while individual conversations die.Reproduced on
mainat2d28c1cwith stock settings (message_capacityleft at its default of 64000).Root cause
AsyncRWLockis writer-preferring and not reentrant: a waiting writer holds_read_gatewhile it waits for_writer_lock(common/rw_locks.py:80-89). That is what prevents writer starvation, and it is correct on its own.ShortTermMemory.get_short_term_memory_contextacquires the read lock atshort_term_memory.py:411, then re-acquires the same lock at:432:Nested read acquisition is harmless until a writer arrives between the two:
Steps to reproduce
Reproduction
Two failing tests, no server required. Save as
test_deadlock_repro.pyand runuv run --all-extras pytest test_deadlock_repro.py -q.Both fail on
2d28c1c:Over HTTP, default settings
With
message_capacityat its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue onePOST /api/v2/memories/searchand onePOST /api/v2/memories0.4 s apart.kill -USR1 <pid>— using the existingserver/diagnostics.py— shows both tasks parked:Lock state at that moment:
_readers=1,_read_gate.locked()=True,_writer_lock.locked()=True.Expected behavior
Reproduction
Two failing tests, no server required. Save as
test_deadlock_repro.pyand runuv run --all-extras pytest test_deadlock_repro.py -q.Both fail on
2d28c1c:Over HTTP, default settings
With
message_capacityat its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue onePOST /api/v2/memories/searchand onePOST /api/v2/memories0.4 s apart.kill -USR1 <pid>— using the existingserver/diagnostics.py— shows both tasks parked:Lock state at that moment:
_readers=1,_read_gate.locked()=True,_writer_lock.locked()=True.Environment
MemMachine
0.3.10.dev20+g2d28c1c1e,mainat2d28c1cuv0.11.26sentence-transformerembedder,openai-chat-completionsprovider pointed at a local stub. The deadlock is in the locking layer and is independent of the storage backends.Suggested fix
Suggested fix
The second acquisition buys nothing: the read lock is already held, and
wait_until_done()has already been awaited at:414. No new summarization can start in between, because starting one requires the write lock.With this applied:
test_concurrent_search_and_add_deadlockpasses.uv run --all-extras pytest packages/server/server_tests→1868 passed, 3 skipped, unchanged from before the patch.test_nested_read_lock_with_waiting_writerstill fails, by design — the one-line change removes the only caller that re-enters the lock, not the underlying property. Worth deciding separately whetherAsyncRWLockshould be made reentrant or simply documented as not reentrant, since the next caller to nest a read lock will hit the same trap. There is precedent:fe2790d("Fix deadlock due to non-reentrant lock pool", #1312).