Skip to content

Tracking: make the session lifecycle safe across replicas with a registry row and durable jobs #1755

Description

@edwinyyyu

Priority 1 of the horizontal-scaling work (#1574). Read this before any of the issues it tracks.

Why one issue

The session layer's defects are one defect seen from several sides: the session row is not the arbiter of the lifecycle, so each operation re-arbitrates in process memory (locks, ref counts, an in-process queue). Fixing them one lock at a time, a lease here and a poll there, leaves the semantics where they are. This issue states the target semantics once and says which open issue each piece closes.

The storage layer below is done or in review. The segment store (#1661) and the collection registry chain (#1733 to #1736) give every store a registry row per key, strict create by primary key, an incarnation per life, fenced handles, logical delete and a concurrent-safe purge. The session layer needs the same shape one level up, and nothing else.

At a glance

Issue What is wrong Best fix Step
#1543 Create is SELECT then INSERT and leaks IntegrityError; config update is read-modify-write INSERT on the primary key and translate the violation; one-statement or versioned update 1
#1575, #1600 Deletion calls semantic cleanup unconditionally; enabled honored only at startup Refuse the semantic manager in the composition root and route in MemMachine (filter target_memories, skip the delete seam at the call site), per the correction on #1600; #1584 removes the shadowing duplicate first 0, independent
#1749 Deletion queue is in process memory; every replica re-enqueues all Deleted rows at boot Durable delete job, claimed with SKIP LOCKED from any replica 3
#1577 Worker gives up on SessionInUseError, leaves a partial delete, only a restart retries Same job, no in-use check, terminal failed state the API shows 3 and 5
#1765 A refused delete leaks an instance reference; the session stays pinned and every later delete in that process is refused Read the count and raise before taking the reference; moot once the ref counts go (step 5) 0, independent
#1655 Per-process locks and ref counts; row-then-storage gap once #1622 lands Registry states plus jobs; remove the locks (findings 1 to 4; 5 is closed by #1733 to #1736) 2 to 5
#1571 Stale handle has no status mapping; a cached handle on another replica never recovers Map the stale errors at the API and evict the instance 5
#1576 (PR #1624) Add and search create sessions No memory request creates a session 6
#1728 Delete builds the storage in order to delete it Delete by key, create nothing 7
PRs #1622, #1625 Storage created with the row, never on a request Keep; add the provisioning state so a crash between row and storage is repairable 4
PR #1739 Create during a pending delete Keep its wait and 503 semantics; its in-process retry moves into the job with 3

Target semantics

  1. The sessions row is the registry. States provisioning, active, deleting. Create is one INSERT on the primary key; the loser gets SessionAlreadyExistsError or the idempotent accept, never a driver error. Config updates are one statement or carry a version.
  2. Storage is provisioned with the row by a job, not on a request. [session storage 1/2] Create a session's storage with the session, never on a request #1622 creates the partition and collection at create time; the provisioning state plus a job row make a crash between row and storage repairable from any replica. No memory request creates a session.
  3. Delete is a state flip plus a durable job. The job row is written in the same transaction as the flip, claimed by any replica with FOR UPDATE SKIP LOCKED, retried with backoff, and lands in a terminal failed state the API can show. Store deletes are O(1) logical; the row goes when the job completes; a create during deletion answers as Wait for a pending delete before re-creating a session #1739 does. Boot does not fan out.
  4. No per-process arbitration. _session_locks, _close_lock, the in-use check and ref counts go. In-flight writers on another replica are fenced by the stores; the stale-handle errors are mapped at the API and evict the cached instance. Deleting a session creates nothing.
  5. Instances hold no per-session state. Once short-term memory is a process-scoped component (Tracking: declare a concurrency scope per component so single-process components stay usable without blocking multi-replica deployments #1757), an EpisodicMemory is handles plus a segmenter and a deriver, cheap to build per request. The LRU cache, its janitor and the lifetime checker can go, which removes the cross-replica cache coherence gap that Tracking: move multitenancy from sharded (one process per collection, per-tenant tables and shard keys) to non-sharded (shared structures served by any process) #1574 listed as out of scope.
  6. Deletion skips the semantic seam at the call site when semantic memory is disabled, and the composition root refuses to build the semantic manager. No component below reads the flag (semantic_memory.enabled is honored only at startup: search, add and list still build and query the semantic stack whenever a request names the semantic type #1600).

Not part of this

A general lock service (#1722 has no consumer here; the arbitration above is rows and constraints). Request routing or session affinity. The full #1579 redesign, which stays the reference for the vocabulary (registry, job, fence) along with #1734.

PR split

Each is reviewable alone; (c) and (e) are the only two that touch the same files.

Tracked


🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Activity

  1. added
    priority: highIssue is urgent or highly impactful. Needs to be addressed as soon as possible.
    keep-openPrevents the auto-close task from closing this issue.
    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)
    on Oct 2, 2026
  2. added theissue type on Oct 2, 2026
  3. added
    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processes
    recoveryA failure or crash leaves durable state nothing repairs: partial writes, lost jobs, no retry
    and removed
    priority: highIssue is urgent or highly impactful. Needs to be addressed as soon as possible.
    on Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processeshorizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)keep-openPrevents the auto-close task from closing this issue.recoveryA failure or crash leaves durable state nothing repairs: partial writes, lost jobs, no retry

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions