Priority 3 of the horizontal-scaling work (#1574), after the session lifecycle (#1755) and the background-work primitive (#1756). It is the smallest change of the three and unblocks keeping several components rather than removing them.
Why
Several components are correct in one process and wrong in several. The project wants to keep declarative memory for existing deployments and does not need it to scale, and does not depend on short-term memory at all. Routing requests by session so each replica owns its sessions does not deliver that: semantic sets are org-level and shared across projects (#1732), a failover loses the per-process state, and a routing layer is one more thing every later component must respect. A declared scope gives the same guarantee at boot, in one place, with no runtime cost.
At a glance
| Component |
Today |
Scope to declare |
Issue |
| PostgreSQL segment store, episode store, session table |
safe from any replica |
CLUSTER |
done (#1661) |
| Qdrant and Milvus stores on the SQL registry |
safe from any replica after the chain |
CLUSTER |
#1733 to #1736 |
| sqlite-vec store, SQLite SQL stores |
shared by one host's processes |
HOST |
#1733 scope statements |
| Engine-backed SQLite vector store |
one process per partition (in-process index) |
PROCESS |
#1733 scope statements |
| Short-term memory |
deque in process memory |
PROCESS |
#1748 |
| Neo4j vector graph store, Nebula vector graph store (declarative memory) |
per-process caches |
PROCESS |
#1753 |
| Neo4j semantic storage |
per-process set cache |
PROCESS |
#1753 |
| Semantic memory as a whole |
ingestion loop in every replica |
PROCESS until #1745 lands, then CLUSTER |
#1745 |
Single fix
A ConcurrencyScope enum, PROCESS < HOST < CLUSTER, that every component computes from its parameters at construction and composes as the minimum of its parts. The deployment declares its own scope (server.concurrency_scope), and startup refuses any component narrower than the deployment, naming the component. Nothing is checked at runtime. Prior art: #1531 (closed) and design/concurrency_scopes.md on its branch; the scope round of design/server_redesign.md (#1579).
What this buys
Declarative memory and short-term memory stay as they are for single-replica deployments, with no code change beyond the declaration. A multi-replica deployment fails at boot with a clear message instead of serving wrong answers. Each later fix (claimed ingestion, a database-backed short-term memory, a loaded Nebula schema) is a one-word scope change on the component.
Not part of this
Making any PROCESS component scale. Removing declarative memory or short-term memory; removal becomes a separate decision that this issue makes unnecessary.
Tracked
🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.
Priority 3 of the horizontal-scaling work (#1574), after the session lifecycle (#1755) and the background-work primitive (#1756). It is the smallest change of the three and unblocks keeping several components rather than removing them.
Why
Several components are correct in one process and wrong in several. The project wants to keep declarative memory for existing deployments and does not need it to scale, and does not depend on short-term memory at all. Routing requests by session so each replica owns its sessions does not deliver that: semantic sets are org-level and shared across projects (#1732), a failover loses the per-process state, and a routing layer is one more thing every later component must respect. A declared scope gives the same guarantee at boot, in one place, with no runtime cost.
At a glance
Single fix
A
ConcurrencyScopeenum,PROCESS < HOST < CLUSTER, that every component computes from its parameters at construction and composes as the minimum of its parts. The deployment declares its own scope (server.concurrency_scope), and startup refuses any component narrower than the deployment, naming the component. Nothing is checked at runtime. Prior art: #1531 (closed) anddesign/concurrency_scopes.mdon its branch; the scope round ofdesign/server_redesign.md(#1579).What this buys
Declarative memory and short-term memory stay as they are for single-replica deployments, with no code change beyond the declaration. A multi-replica deployment fails at boot with a clear message instead of serving wrong answers. Each later fix (claimed ingestion, a database-backed short-term memory, a loaded Nebula schema) is a one-word scope change on the component.
Not part of this
Making any PROCESS component scale. Removing declarative memory or short-term memory; removal becomes a separate decision that this issue makes unnecessary.
Tracked
🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.