What happened
MemMachine.delete_session (main/memmachine.py:582-608) flips the row to Deleted and then put_nowaits the session on self._deletion_queue, an asyncio.Queue owned by the process (:117-119). One worker task per process drains it (:342, started at :406). The queue has no persistence: if the process dies, the job is gone and the row stays Deleted. Recovery is start(), which re-enqueues every Deleted session (:409-413). Every replica does that at boot, so N replicas run N concurrent deletions of the same sessions; the losers fail at the row delete with SessionNotFoundError, which the worker logs (:350-355).
The store-level deletes are idempotent (segment store by contract; Qdrant and Milvus return when the registry entry is gone), so the duplication is wasted work and log noise rather than data loss. But the job itself is neither durable nor visible: nothing records that a deletion is pending, running, or failed.
#1577 covers the worker giving up on SessionInUseError and the partial delete it leaves. #1739 adds in-process retry with backoff and a database-polled wait on create. Both keep the queue in process memory.
Expected
Deletion is a durable job: a row written in the same transaction as the status flip, claimed by any replica with SELECT ... FOR UPDATE SKIP LOCKED, retried with backoff, and reaching a terminal failed state the API can show. Boot does not fan out: a replica picks up only unclaimed jobs.
Notes
Code read at d6068cdbf (main), paths under packages/server/src/memmachine_server/. The segment store purge and the collection registry purge in #1734 already use the claim pattern; the deletion job is the same shape one level up.
🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.
What happened
MemMachine.delete_session(main/memmachine.py:582-608) flips the row toDeletedand thenput_nowaits the session onself._deletion_queue, anasyncio.Queueowned by the process (:117-119). One worker task per process drains it (:342, started at:406). The queue has no persistence: if the process dies, the job is gone and the row staysDeleted. Recovery isstart(), which re-enqueues everyDeletedsession (:409-413). Every replica does that at boot, so N replicas run N concurrent deletions of the same sessions; the losers fail at the row delete withSessionNotFoundError, which the worker logs (:350-355).The store-level deletes are idempotent (segment store by contract; Qdrant and Milvus return when the registry entry is gone), so the duplication is wasted work and log noise rather than data loss. But the job itself is neither durable nor visible: nothing records that a deletion is pending, running, or failed.
#1577 covers the worker giving up on
SessionInUseErrorand the partial delete it leaves. #1739 adds in-process retry with backoff and a database-polled wait on create. Both keep the queue in process memory.Expected
Deletion is a durable job: a row written in the same transaction as the status flip, claimed by any replica with
SELECT ... FOR UPDATE SKIP LOCKED, retried with backoff, and reaching a terminal failed state the API can show. Boot does not fan out: a replica picks up only unclaimed jobs.Notes
Code read at
d6068cdbf(main), paths underpackages/server/src/memmachine_server/. The segment store purge and the collection registry purge in #1734 already use the claim pattern; the deletion job is the same shape one level up.🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.