Skip to content

Session lifecycle is single-process: locks, in-use checks and row-vs-storage creation are not arbitrated across replicas #1655

Description

@edwinyyyu

Status 2026-10-02. Finding 5 is fixed by the collection registry chain #1734 -> #1735 -> #1736 (SQL registry, create by primary key, an incarnation per collection life, fenced handles). Findings 1 to 4 are tracked under #1755, which states the target semantics and a PR split. The disposition below (fix together in #1579) is superseded: #1579 is deferred as a whole, and targeted session-layer changes against #1755's semantics are in scope.

Summary

The session layer (EpisodicMemoryManager and SessionDataManagerSQLImpl) is safe for one server process. Under several processes on different nodes sharing one backend, with no session affinity, every lifecycle operation on a session is a read-then-write guarded by a process-local lock, so two replicas acting on the same session key coordinate on nothing. The storage layer underneath is now mostly cross-process safe (segment store since #1548; vector stores through the chain #1622-#1618, and #1631 in particular), which moves the remaining gaps up to this layer. This issue records the concrete findings so they are fixed together in the tenant lifecycle redesign (#1579, design/tenant_lifecycle.md) rather than patched one PR at a time. It links the issues that already carry parts of it.

Findings are against the vector store chain at its current tip (#1622..#1618 on speedkick); line references are to packages/server/src/memmachine_server/.

1. Session mutual exclusion is per process

episodic_memory_manager.py:91-94: _session_locks is a defaultdict[str, AsyncRWLock] and _close_lock an AsyncRWLock; create_session, open_episodic_memory, close_session and delete_session take the session's lock. These are asyncio locks in one process. Two replicas creating, opening or deleting the same key take two unrelated locks.

2. Create is a read-then-insert with the race outcome untranslated

session_manager/session_data_manager_sql_impl.py, create_new_session_if_not_exist: SELECT the row, then INSERT. session_key is the primary key, so the database does arbitrate, but the loser gets a raw IntegrityError (a 500) rather than the idempotent accept or SessionAlreadyExistsError. Already filed as #1543 (which also covers the lost update in update_session_episodic_config); listed here because #1622 puts more behind it (finding 3).

3. Row-then-storage creation has no repair path (introduced by #1622)

#1622 creates a session's storage (segment-store and vector-store partitions) with the session row instead of on the first request, and the stores' create_partition is strict. The data manager therefore returns whether this call inserted the row, and the manager creates storage only then (episodic_memory_manager.py, _create_session). Contract:

state data manager manager
no row inserts, returns True creates both partitions
row exists, active, configuration/param_data/user_metadata equal returns False accepts the row; storage assumed present
row exists and anything differs, or status Deleted raises SessionAlreadyExistsError propagates

The row commits first, then the partitions are created. A crash or backend failure between the two leaves an active row with no partitions; every later equivalent create returns False and never repairs it, and every request on the session fails with SessionPartitionMissingError (long_term_memory/service_locator.py, _event_params). The failure is loud and the row stays the source of truth, but the only recovery is delete-and-recreate. Row-first was chosen over storage-first so that a crash never leaves partitions no row knows about, which nothing would reclaim. Across replicas, a create on one node interleaved with a delete on another can produce either state, since finding 1 gives them no common lock.

Not patched in #1622 on purpose: a repair on the False path (get-then-create at the session layer) would be one more process-local read-then-write; the redesign's step table and reconciler are the mechanism that makes the partial state recoverable from any process.

4. "In use" and eviction are per process

delete_session (episodic_memory_manager.py:293-300) refuses with SessionInUseError when this process's instance cache holds the session with a positive ref count, and max_life_time eviction runs per cache. A delete on replica A proceeds while replica B holds the session open. The stores now fence this below the session layer -- a handle bound to a deleted partition raises on every operation (segment store: SegmentStorePartitionHandleStaleError, #1548; vector store: VectorStorePartitionHandleStaleError, #1631) -- so a stale handle cannot write into a successor, but the session layer does not know, and the cached EpisodicMemory on replica B keeps raising until eviction. #1571 covers the eviction and the status mapping for the segment store's error; the vector store's error now needs the same handling.

5. Partition creation on Qdrant and Milvus is not atomic across processes

After #1631 the data path, deletion and reclamation are cross-process safe on Qdrant and Milvus (every operation fences on the registry, delete is a registry write, purge is idempotent from any process). Creation is not: _client_partition_locks is per process (qdrant_vector_store.py, its own comment says it serializes callers within a process and not across workers), and create_partition is read-registry-then-write with no compare-and-set.

  • Qdrant: two replicas creating the same key both succeed; the registry point is an upsert, so the last registration wins and the other incarnation is orphaned (nothing was written under it; a handle bound to it raises from then on). Data-safe, but two creates instead of one create and one AlreadyExists.
  • Milvus: the registry insert does not enforce primary-key uniqueness (Milvus dedups on upsert only), so two concurrent creates leave two live entries under one key; _live_entry reads the first and the other replica's handle is stale at once. Data-safe, wrong contract.

The SQLite stores are unaffected (unique constraint plus an in-transaction queue check), and are single-node by construction; the engine-backed SQLite store is additionally single-process per partition (in-process search engine, index file and write lock), which is what the VectorStore docstring's "at most one process" sentence (vector_store.py:172) still describes. #1524/#1525 carry the cross-process registry that fixes the Qdrant/Milvus create race.

What is already cross-process safe

Proposed disposition

Fix findings 1-4 together in the tenant lifecycle redesign (#1579): the session row becomes the single, database-arbitrated creation point (finding 2 folds into #1543's translation of the constraint violation), the step table records row-created / storage-created so a reconciler on any replica finishes or reverts a partial creation (finding 3), and liveness moves from process-local ref counts to the stores' fencing plus handle eviction (finding 4, with #1571). Finding 5 is #1524/#1525's cross-process registry; once the session row is the creation point, a store-level create race can only occur between two owners of the same row, which the row prevents.

Related: #1574 (tracking), #1543, #1524, #1525, #1571, #1579.

Investigated and written by Claude (Claude Code), filed from the account of the user who commissioned the investigation.

Activity

  1. added theissue type on Sep 16, 2026
  2. edwinyyyu commented on Sep 16, 2026

    @edwinyyyu
    ContributorAuthor

    Finding 5 (Qdrant and Milvus partition creation not atomic across processes) is fixed by #1656, slice 11 of the vector store chain: both stores keep their partition registry and purge queue in the deployment's relational database (SqlPartitionRegistry, registry_database on the conf), so creation is a primary-key insert (a racing creator on any process gets AlreadyExists), deletion one transaction, purge a row lock (FOR UPDATE SKIP LOCKED on PostgreSQL), and the per-process lock tables are gone. The VectorStore contract's one-process clause is replaced by a per-store statement (the engine-backed SQLite store keeps a partition's index in one process; sqlite-vec is shared by one node's processes). Findings 1-4 (the session layer) stand.

  3. edwinyyyu commented on Sep 16, 2026

    @edwinyyyu
    ContributorAuthor

    Correction to the comment above: #1656 was folded into #1631 (slice 10), which now carries incarnations and the SQL-arbitrated registry as one change; the fix for finding 5 is the same, its home is #1631.

  4. added
    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)
    on Sep 29, 2026
  5. edwinyyyu commented on Oct 2, 2026

    @edwinyyyu
    ContributorAuthor

    Status 2026-10-02: finding 5 (Qdrant and Milvus partition creation not atomic across processes) is fixed by the collection registry chain #1734 -> #1735 -> #1736 (SQL registry, create by primary key, an incarnation per collection life, fenced handles, purge claimed with SKIP LOCKED). Findings 1 to 4 are tracked under #1755, which states the target semantics for the session layer and a PR split. With #1579 deferred, targeted changes against those semantics are in scope.


    🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

  6. added
    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processes
    recoveryA failure or crash leaves durable state nothing repairs: partial writes, lost jobs, no retry
    on Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processeshorizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)recoveryA failure or crash leaves durable state nothing repairs: partial writes, lost jobs, no retry

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions