You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Tracking: make the session lifecycle safe across replicas with a registry row and durable jobs #1755
Priority 1 of the horizontal-scaling work (#1574). Read this before any of the issues it tracks.
Why one issue
The session layer's defects are one defect seen from several sides: the session row is not the arbiter of the lifecycle, so each operation re-arbitrates in process memory (locks, ref counts, an in-process queue). Fixing them one lock at a time, a lease here and a poll there, leaves the semantics where they are. This issue states the target semantics once and says which open issue each piece closes.
The storage layer below is done or in review. The segment store (#1661) and the collection registry chain (#1733 to #1736) give every store a registry row per key, strict create by primary key, an incarnation per life, fenced handles, logical delete and a concurrent-safe purge. The session layer needs the same shape one level up, and nothing else.
Deletion calls semantic cleanup unconditionally; enabled honored only at startup
Refuse the semantic manager in the composition root and route in MemMachine (filter target_memories, skip the delete seam at the call site), per the correction on #1600; #1584 removes the shadowing duplicate first
Keep its wait and 503 semantics; its in-process retry moves into the job
with 3
Target semantics
The sessions row is the registry. States provisioning, active, deleting. Create is one INSERT on the primary key; the loser gets SessionAlreadyExistsError or the idempotent accept, never a driver error. Config updates are one statement or carry a version.
Delete is a state flip plus a durable job. The job row is written in the same transaction as the flip, claimed by any replica with FOR UPDATE SKIP LOCKED, retried with backoff, and lands in a terminal failed state the API can show. Store deletes are O(1) logical; the row goes when the job completes; a create during deletion answers as Wait for a pending delete before re-creating a session #1739 does. Boot does not fan out.
No per-process arbitration._session_locks, _close_lock, the in-use check and ref counts go. In-flight writers on another replica are fenced by the stores; the stale-handle errors are mapped at the API and evict the cached instance. Deleting a session creates nothing.
A general lock service (#1722 has no consumer here; the arbitration above is rows and constraints). Request routing or session affinity. The full #1579 redesign, which stays the reference for the vocabulary (registry, job, fence) along with #1734.
PR split
Each is reviewable alone; (c) and (e) are the only two that touch the same files.
Priority 1 of the horizontal-scaling work (#1574). Read this before any of the issues it tracks.
Why one issue
The session layer's defects are one defect seen from several sides: the session row is not the arbiter of the lifecycle, so each operation re-arbitrates in process memory (locks, ref counts, an in-process queue). Fixing them one lock at a time, a lease here and a poll there, leaves the semantics where they are. This issue states the target semantics once and says which open issue each piece closes.
The storage layer below is done or in review. The segment store (#1661) and the collection registry chain (#1733 to #1736) give every store a registry row per key, strict create by primary key, an incarnation per life, fenced handles, logical delete and a concurrent-safe purge. The session layer needs the same shape one level up, and nothing else.
At a glance
IntegrityError; config update is read-modify-writeenabledhonored only at startupMemMachine(filtertarget_memories, skip the delete seam at the call site), per the correction on #1600; #1584 removes the shadowing duplicate firstSKIP LOCKEDfrom any replicaSessionInUseError, leaves a partial delete, only a restart retriesTarget semantics
sessionsrow is the registry. Statesprovisioning,active,deleting. Create is one INSERT on the primary key; the loser getsSessionAlreadyExistsErroror the idempotent accept, never a driver error. Config updates are one statement or carry a version.provisioningstate plus a job row make a crash between row and storage repairable from any replica. No memory request creates a session.FOR UPDATE SKIP LOCKED, retried with backoff, and lands in a terminal failed state the API can show. Store deletes are O(1) logical; the row goes when the job completes; a create during deletion answers as Wait for a pending delete before re-creating a session #1739 does. Boot does not fan out._session_locks,_close_lock, the in-use check and ref counts go. In-flight writers on another replica are fenced by the stores; the stale-handle errors are mapped at the API and evict the cached instance. Deleting a session creates nothing.EpisodicMemoryis handles plus a segmenter and a deriver, cheap to build per request. The LRU cache, its janitor and the lifetime checker can go, which removes the cross-replica cache coherence gap that Tracking: move multitenancy from sharded (one process per collection, per-tenant tables and shard keys) to non-sharded (shared structures served by any process) #1574 listed as out of scope.Not part of this
A general lock service (#1722 has no consumer here; the arbitration above is rows and constraints). Request routing or session affinity. The full #1579 redesign, which stays the reference for the vocabulary (registry, job, fence) along with #1734.
PR split
Each is reviewable alone; (c) and (e) are the only two that touch the same files.
Tracked
🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.