Skip to content

Tracking: declare a concurrency scope per component so single-process components stay usable without blocking multi-replica deployments #1757

Description

@edwinyyyu

Priority 3 of the horizontal-scaling work (#1574), after the session lifecycle (#1755) and the background-work primitive (#1756). It is the smallest change of the three and unblocks keeping several components rather than removing them.

Why

Several components are correct in one process and wrong in several. The project wants to keep declarative memory for existing deployments and does not need it to scale, and does not depend on short-term memory at all. Routing requests by session so each replica owns its sessions does not deliver that: semantic sets are org-level and shared across projects (#1732), a failover loses the per-process state, and a routing layer is one more thing every later component must respect. A declared scope gives the same guarantee at boot, in one place, with no runtime cost.

At a glance

Component Today Scope to declare Issue
PostgreSQL segment store, episode store, session table safe from any replica CLUSTER done (#1661)
Qdrant and Milvus stores on the SQL registry safe from any replica after the chain CLUSTER #1733 to #1736
sqlite-vec store, SQLite SQL stores shared by one host's processes HOST #1733 scope statements
Engine-backed SQLite vector store one process per partition (in-process index) PROCESS #1733 scope statements
Short-term memory deque in process memory PROCESS #1748
Neo4j vector graph store, Nebula vector graph store (declarative memory) per-process caches PROCESS #1753
Neo4j semantic storage per-process set cache PROCESS #1753
Semantic memory as a whole ingestion loop in every replica PROCESS until #1745 lands, then CLUSTER #1745

Single fix

A ConcurrencyScope enum, PROCESS < HOST < CLUSTER, that every component computes from its parameters at construction and composes as the minimum of its parts. The deployment declares its own scope (server.concurrency_scope), and startup refuses any component narrower than the deployment, naming the component. Nothing is checked at runtime. Prior art: #1531 (closed) and design/concurrency_scopes.md on its branch; the scope round of design/server_redesign.md (#1579).

What this buys

Declarative memory and short-term memory stay as they are for single-replica deployments, with no code change beyond the declaration. A multi-replica deployment fails at boot with a clear message instead of serving wrong answers. Each later fix (claimed ingestion, a database-backed short-term memory, a loaded Nebula schema) is a one-word scope change on the component.

Not part of this

Making any PROCESS component scale. Removing declarative memory or short-term memory; removal becomes a separate decision that this issue makes unnecessary.

Tracked


🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Metadata

Metadata

Assignees

No one assigned

    Labels

    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)keep-openPrevents the auto-close task from closing this issue.

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions