Skip to content

[Bug]: Deadlock: concurrent search + add on the same session permanently wedges it (non-reentrant read lock) #1514

Description

@daproject85

Describe the bug

Summary

A search and an add arriving concurrently on the same session, while a short-term-memory summarization is in flight, deadlock each other permanently. The session is then unusable for the lifetime of the process: every later request on it blocks, including close(), and POST /api/v2/projects/delete returns 204 without deleting anything.

Other sessions are unaffected and /api/v2/health continues to return 200, so the server looks healthy from outside while individual conversations die.

Reproduced on main at 2d28c1c with stock settings (message_capacity left at its default of 64000).

Root cause

AsyncRWLock is writer-preferring and not reentrant: a waiting writer holds _read_gate while it waits for _writer_lock (common/rw_locks.py:80-89). That is what prevents writer starvation, and it is correct on its own.

ShortTermMemory.get_short_term_memory_context acquires the read lock at short_term_memory.py:411, then re-acquires the same lock at :432:

async with self._lock.read_lock():                    # :411  first acquire
    ...
    await self._consolidator.wait_until_done()        # :414  suspends for the whole LLM call
    ...
    return list(episodes), await self.get_summary()   # :432  -> get_summary -> :379
                                                      #      -> _wait_for_summary_to_finish -> :217
                                                      #      -> self._lock.read_lock()  SECOND acquire

Nested read acquisition is harmless until a writer arrives between the two:

Step Search task Add task
1 acquires read lock — _readers=1, _writer_lock held by the reader cohort  
2 suspends at :414 awaiting the summarization task  
3   add_episodes (:199) wants the write lock: takes _read_gate, blocks on _writer_lock
4 LLM returns, resumes, re-acquires the read lock → blocks on _read_gate  
5 blocked forever blocked forever

Steps to reproduce

Reproduction

Two failing tests, no server required. Save as test_deadlock_repro.py and run uv run --all-extras pytest test_deadlock_repro.py -q.

import asyncio
import datetime

import pytest

from memmachine_server.common.episode_store.episode_model import Episode
from memmachine_server.common.language_model.language_model import LanguageModel
from memmachine_server.common.rw_locks import AsyncRWLock
from memmachine_server.episodic_memory.short_term_memory.short_term_memory import (
    ShortTermMemory,
    ShortTermMemoryParams,
)


class SlowLLM(LanguageModel):
    """Stands in for any real provider: a summary call that takes a second."""

    async def generate_response(
        self, system_prompt=None, user_prompt=None, tools=None,
        tool_choice=None, max_attempts=1,
    ):
        await asyncio.sleep(1.0)
        return ("summary text", None)

    async def generate_parsed_response(self, *args, **kwargs):
        raise NotImplementedError

    async def generate_response_with_token_usage(self, *args, **kwargs):
        raise NotImplementedError


def _episode(i: int) -> Episode:
    return Episode(
        uid=f"e{i}",
        content=f"message {i} " + "x" * 40,
        session_key="s1",
        created_at=datetime.datetime.now(datetime.UTC),
        producer_id="user",
        producer_role="user",
        sequence_num=i,
    )


@pytest.mark.asyncio
async def test_concurrent_search_and_add_deadlock():
    """A search and an add on the same session, while a summary is in flight."""
    memory = ShortTermMemory(
        ShortTermMemoryParams(
            session_key="s1",
            llm_model=SlowLLM(),
            summary_prompt_system="system",
            summary_prompt_user="{episodes} {summary} {max_length}",
            message_capacity=100,
        )
    )
    # Overflow the budget so the background summarizer starts.
    await memory.add_episodes([_episode(1), _episode(2), _episode(3)])

    async def search():
        return await memory.get_short_term_memory_context("query", limit=10)

    async def add():
        await asyncio.sleep(0.3)  # arrive while the summary call is in flight
        return await memory.add_episodes([_episode(4)])

    # Hangs forever on main. Also leaves the session permanently unusable:
    # every later call, including close(), blocks on the same gate.
    await asyncio.wait_for(asyncio.gather(search(), add()), timeout=10)


@pytest.mark.asyncio
async def test_nested_read_lock_with_waiting_writer():
    """The underlying lock behaviour, isolated from ShortTermMemory."""
    lock = AsyncRWLock()

    async def reader():
        async with lock.read_lock():
            await asyncio.sleep(0.1)      # a writer arrives during this await
            async with lock.read_lock():  # blocks forever
                pass

    async def writer():
        await asyncio.sleep(0.05)
        async with lock.write_lock():
            pass

    await asyncio.wait_for(asyncio.gather(reader(), writer()), timeout=3)

Both fail on 2d28c1c:

FAILED test_deadlock_repro.py::test_concurrent_search_and_add_deadlock - TimeoutError
FAILED test_deadlock_repro.py::test_nested_read_lock_with_waiting_writer - TimeoutError
2 failed in 16.83s

Over HTTP, default settings

With message_capacity at its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue one POST /api/v2/memories/search and one POST /api/v2/memories 0.4 s apart.

[  0.03s] create: 201
[  0.92s] filled the short-term budget; summarizer in flight
[ 20.92s] search:           *** no response in 20s ***
[ 21.33s] add (concurrent): *** no response in 20s ***
[ 21.33s] health: 200 healthy

kill -USR1 <pid> — using the existing server/diagnostics.py — shows both tasks parked:

Task 'Task-1077' [pending]
  episodic_memory/episodic_memory.py:325   in _query_short_term_memory
  short_term_memory.py:432                 return list(episodes), await self.get_summary()
  short_term_memory.py:379                 await self._wait_for_summary_to_finish()
  short_term_memory.py:217                 async with self._lock.read_lock():

Task 'Task-1091' [pending]
  short_term_memory.py:199                 async with self._lock.write_lock():

Lock state at that moment: _readers=1, _read_gate.locked()=True, _writer_lock.locked()=True.

Expected behavior

Reproduction

Two failing tests, no server required. Save as test_deadlock_repro.py and run uv run --all-extras pytest test_deadlock_repro.py -q.

import asyncio
import datetime

import pytest

from memmachine_server.common.episode_store.episode_model import Episode
from memmachine_server.common.language_model.language_model import LanguageModel
from memmachine_server.common.rw_locks import AsyncRWLock
from memmachine_server.episodic_memory.short_term_memory.short_term_memory import (
    ShortTermMemory,
    ShortTermMemoryParams,
)


class SlowLLM(LanguageModel):
    """Stands in for any real provider: a summary call that takes a second."""

    async def generate_response(
        self, system_prompt=None, user_prompt=None, tools=None,
        tool_choice=None, max_attempts=1,
    ):
        await asyncio.sleep(1.0)
        return ("summary text", None)

    async def generate_parsed_response(self, *args, **kwargs):
        raise NotImplementedError

    async def generate_response_with_token_usage(self, *args, **kwargs):
        raise NotImplementedError


def _episode(i: int) -> Episode:
    return Episode(
        uid=f"e{i}",
        content=f"message {i} " + "x" * 40,
        session_key="s1",
        created_at=datetime.datetime.now(datetime.UTC),
        producer_id="user",
        producer_role="user",
        sequence_num=i,
    )


@pytest.mark.asyncio
async def test_concurrent_search_and_add_deadlock():
    """A search and an add on the same session, while a summary is in flight."""
    memory = ShortTermMemory(
        ShortTermMemoryParams(
            session_key="s1",
            llm_model=SlowLLM(),
            summary_prompt_system="system",
            summary_prompt_user="{episodes} {summary} {max_length}",
            message_capacity=100,
        )
    )
    # Overflow the budget so the background summarizer starts.
    await memory.add_episodes([_episode(1), _episode(2), _episode(3)])

    async def search():
        return await memory.get_short_term_memory_context("query", limit=10)

    async def add():
        await asyncio.sleep(0.3)  # arrive while the summary call is in flight
        return await memory.add_episodes([_episode(4)])

    # Hangs forever on main. Also leaves the session permanently unusable:
    # every later call, including close(), blocks on the same gate.
    await asyncio.wait_for(asyncio.gather(search(), add()), timeout=10)


@pytest.mark.asyncio
async def test_nested_read_lock_with_waiting_writer():
    """The underlying lock behaviour, isolated from ShortTermMemory."""
    lock = AsyncRWLock()

    async def reader():
        async with lock.read_lock():
            await asyncio.sleep(0.1)      # a writer arrives during this await
            async with lock.read_lock():  # blocks forever
                pass

    async def writer():
        await asyncio.sleep(0.05)
        async with lock.write_lock():
            pass

    await asyncio.wait_for(asyncio.gather(reader(), writer()), timeout=3)

Both fail on 2d28c1c:

FAILED test_deadlock_repro.py::test_concurrent_search_and_add_deadlock - TimeoutError
FAILED test_deadlock_repro.py::test_nested_read_lock_with_waiting_writer - TimeoutError
2 failed in 16.83s

Over HTTP, default settings

With message_capacity at its default and a provider that takes ~3 s to summarize: create a project, send 48 ordinary ~1,400-character messages to fill the budget, then issue one POST /api/v2/memories/search and one POST /api/v2/memories 0.4 s apart.

[  0.03s] create: 201
[  0.92s] filled the short-term budget; summarizer in flight
[ 20.92s] search:           *** no response in 20s ***
[ 21.33s] add (concurrent): *** no response in 20s ***
[ 21.33s] health: 200 healthy

kill -USR1 <pid> — using the existing server/diagnostics.py — shows both tasks parked:

Task 'Task-1077' [pending]
  episodic_memory/episodic_memory.py:325   in _query_short_term_memory
  short_term_memory.py:432                 return list(episodes), await self.get_summary()
  short_term_memory.py:379                 await self._wait_for_summary_to_finish()
  short_term_memory.py:217                 async with self._lock.read_lock():

Task 'Task-1091' [pending]
  short_term_memory.py:199                 async with self._lock.write_lock():

Lock state at that moment: _readers=1, _read_gate.locked()=True, _writer_lock.locked()=True.

Environment

MemMachine 0.3.10.dev20+g2d28c1c1e, main at 2d28c1c

  • Python 3.12.13, uv 0.11.26
  • macOS 26.6.1 (arm64)
  • Config: SQLite episode/segment/vector stores, local sentence-transformer embedder, openai-chat-completions provider pointed at a local stub. The deadlock is in the locking layer and is independent of the storage backends.

Suggested fix

Suggested fix

The second acquisition buys nothing: the read lock is already held, and wait_until_done() has already been awaited at :414. No new summarization can start in between, because starting one requires the write lock.

--- a/packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py
+++ b/packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py
@@ -429,7 +429,7 @@ class ShortTermMemory:
                     break
                 episodes.appendleft(e)
                 length += msg_len
-            return list(episodes), await self.get_summary()
+            return list(episodes), await self._consolidator.summary
 
     @staticmethod
     def _compute_episode_length(episode: Episode) -> int:

With this applied:

  • test_concurrent_search_and_add_deadlock passes.
  • uv run --all-extras pytest packages/server/server_tests → 1868 passed, 3 skipped, unchanged from before the patch.

test_nested_read_lock_with_waiting_writer still fails, by design — the one-line change removes the only caller that re-enters the lock, not the underlying property. Worth deciding separately whether AsyncRWLock should be made reentrant or simply documented as not reentrant, since the next caller to nest a read lock will hit the same trap. There is precedent: fe2790d ("Fix deadlock due to non-reentrant lock pool", #1312).

Activity

  1. added theissue type on Aug 18, 2026
  2. daproject85 commented on Aug 18, 2026

    @daproject85
    Author

    Adding which releases this affects, since the report above doesn't say.

    Both ingredients — the nested read acquisition and the writer-priority _read_gate — have been there continuously since the change that introduced the nesting, so this isn't main-only.

    git tag --contains aa74179 lists 11 releases: v0.2.6, and v0.3.0 through v0.3.9. Spot-checked against the tags:

    Version Nested acquire Writer-priority gate
    v0.2.6 (2026-02-02) yes — :436 → :205 yes
    v0.3.3 (2026-03-27) yes yes
    v0.3.9 (2026-05-18, latest release) yes — :432 → :217 yes

    At v0.3.9 the line numbers are identical to main, so the suggested fix applies unchanged to the latest published release.

    The nesting arrived with aa74179 / PR #949 ("[GH-909] Improve async summary performance", 2026-01-27), which fixed #909 — a real blocking-write bottleneck found while profiling #837. Its change list includes:

    "Ensure queries wait for consolidation to finish so results are always stable and consistent"

    That requirement is what put wait_until_done() and the second get_summary() call inside the read lock. The suggested fix keeps that guarantee: wait_until_done() is still awaited at :414, and no new consolidation can start while the read lock is held, because starting one needs the write lock. It only removes the redundant re-entry.

    Worth noting #837 is still open, so the performance complaint that motivated #949 was never closed out.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions