Repository navigation
Conversation
…laims purge_partition existed to serve a promise the delete path can no longer make: prompt physical erasure is not something the store can keep on every dialect (on SQLite any writer past the busy timeout fails), and the inline drain that used it, first of the global queue and then scoped to the deleted key, added a second purger, a polling loop while an entry was held, and a slow-purge warning, all to shorten a window that no contract requires to be short. drop_session_partition now deletes the collection and the partition, nulls its handles, and returns; the partition is unreachable at once and its rows are reclaimed by the resource manager's sweeper within its interval. The ABC keeps one purge method, the sweeper, whose contract already says a deployment must run it. The composite queue index, the held-entry pause, the existence read, the slow-purge warning and their tests go with the method; the sweeper's same-tick ordering wording and the queue-row docstring's forensic key stay. A per-tenant reclaim step returns with the tenant lifecycle layer (MemMachine#1579), where single-use keys make it job-like: progress, retry and failure per tenant, which the global sweeper cannot attribute. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
2d315ac to
7e33dc2
Compare
…laims purge_partition existed to serve a promise the delete path can no longer make: prompt physical erasure is not something the store can keep on every dialect (on SQLite any writer past the busy timeout fails), and the inline drain that used it, first of the global queue and then scoped to the deleted key, added a second purger, a polling loop while an entry was held, and a slow-purge warning, all to shorten a window that no contract requires to be short. drop_session_partition now deletes the collection and the partition, nulls its handles, and returns; the partition is unreachable at once and its rows are reclaimed by the resource manager's sweeper within its interval. The ABC keeps one purge method, the sweeper, whose contract already says a deployment must run it. The composite queue index, the held-entry pause, the existence read, the slow-purge warning and their tests go with the method; the sweeper's same-tick ordering wording and the queue-row docstring's forensic key stay. A per-tenant reclaim step returns with the tenant lifecycle layer (MemMachine#1579), where single-use keys make it job-like: progress, retry and failure per tenant, which the global sweeper cannot attribute. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Design proposal for tenant lifecycle management above the stores: a tenant record with step rows and a reconciler, UUID-only store keys minted per tenant lifetime, a single tenant handle, and the resource contracts the segment and vector stores expose to that layer. Rebuilt on speedkick, which carries the segment store overhaul (MemMachine#1548) the document refers to. Tracking: MemMachine#1574. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
7e33dc2 to
32c813f
Compare
The lifecycle design cannot be fixed inside the current API, configuration and wiring, so the document now covers the server: tenants (name + UUID id, renamable, name in one SQL table, UUID keys in every store), a control plane (tenant service, job table, reconciler) separate from the data plane (subsystems serving data operations by tenant id in the path, no routing handle), the subsystem registration contract (schema with mutable/immutable options, provision and delete hooks), event memory with the segment store as system of record and no episode store, store contracts including a SQL ledger for vector collections with a derived reclaim grace period, a declarative configuration document with providers and tenant templates, eager startup by constructor injection, the v1 HTTP API with one error handler, and schema management (Alembic per component, boot modes, container provisioning). Short-term, semantic and declarative memory are not wired in. File renamed from tenant_lifecycle.md. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
… config Every store now rejects operations on a deleted key by itself, with its registry row in SQL: writes pin it FOR SHARE, the logical delete takes it FOR UPDATE and so waits out in-flight writes, reads verify it. The vector store adopts this through a SQL ledger, and dead keys stay as tombstones swept at a bounded rate, so no clock is compared anywhere and no key is forgotten while a record could exist. The segment store no longer fences on behalf of the subsystem. Subsystems own their per-tenant state; only the tenant service reads the tenant table. The reconciler is a role a deployment runs in as many or as few processes as it needs. An event store returns as the system of record, designed new, so derived data can be rebuilt. Configuration is the components' own parameter models with typed references resolved by the loader; nothing receives a catalog. Schema is upgraded only by an operator's command and verified at startup. The SQLite vector stores keep per-collection tables, as their docstring's reason for avoiding vec0 partition keys stands. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
The SQLite stores are not for large deployments; they are fine as long as they work and obey the contracts. State the rule as one about cost at scale in the requirements, the schema principle, and both places that mention the SQLite stores' per-collection tables. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
State the rule as two halves: component schema runs only in the setup command, which serves nothing and cannot race, and serving and reconciler processes verify it; tenant-specific DDL is the only DDL allowed elsewhere, avoided where it would be expensive at scale. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
…name to episodic memory Stores fence on the caller's UUID alone: the registry row keyed by it is the whole fence, and the segment store's incarnation goes; replacement is the subsystem's, by minting a new key per generation and recording it in its own per-tenant row. The fence section now defines the write step as a FOR SHARE row lock released by the database at transaction end, with the two settings that bound a live session holding it. The event store is its own component, the system of record, with positions that subsystems process by; EventMemory is renamed episodic memory, so "event" is the caller's ingestion type. The uuid5-derived segment and derivative ids are withdrawn; idempotency is per event, by forgetting an event's derived rows before reprocessing. A table surveys Qdrant, Milvus, pgvector, Pinecone, S3 Vectors, Weaviate and Chroma at the tier that scales, for per-tenant objects, rejection, listing and reclaim. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
A full rebuild costs what an ingestion costs, so it is one, into a new tenant. The rebuild job, its endpoint, and the per-generation keys it needed go; the tenant id is the key in every store again, and the subsystem's per-tenant row holds only the watermark and the applied configuration. The event store keeps its two uses: repair of partial processing and processing history for a subsystem enabled later. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Chroma's own multi-tenancy write-up warns that metadata filtering slows as users and documents grow, so its row now reads collection per tenant. Its per-tenant object is a Collection handle: get_collection is one round trip resolving the name to the collection's UUID (verified in chromadb/api/fastapi.py), after which every data operation addresses the UUID. The paragraph under the table says what that costs, that the instance cache pays it once per open, and that storing the UUID in the ledger removes the call. Per-collection cost and Chroma Cloud's collection count are marked unverified. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Every store operation takes the key. The registry row that fences it also holds what addresses the tenant on the backend (codec configuration, container, collection UUID), so nothing is opened or closed per tenant, a process holds no per-tenant state, and a configuration update takes effect on the next request. Chroma's collection UUID is recorded in the ledger at creation and operations go to its HTTP API by that UUID; Weaviate's wrapper is built per call. The per-tenant instance cache, its TTL, and MemMachine#1548's partition handle go. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Replace the "verified before an implementation" note for S3 Vectors with the documented figures: 10,000 indexes per bucket, top-K up to 10,000 per query, DeleteVectors by key at 500 per call and no delete by filter, filterable metadata 2 KB per vector, filters evaluated during the search, numeric-only range comparisons. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Adopt MemMachine#1531's ConcurrencyScope: every component computes its scope from its params, a composition's scope is the minimum of its parts', the deployment declares the scope it runs at, and startup refuses any component narrower than that. The horizontal scaling requirement is stated at cluster scope; the SQLite stores declare process or machine and are held to every contract within it. Scope declarations are tabulated, the file lock is named as what gives SQLite-backed stores machine scope, and "large deployment" wording is replaced by scope. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
…as reference Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
From the registry session's review of the draft. Store creates are strict and raise on any row under the key; idempotency is the component's ensure, which knows the key's provenance, and a row in a non-live state is a reused key that raises. A table states what every operation does with a non-live key. Tombstones are kept by default; pruning is an operator's trade gated on a clean sweep, with what it gives up stated. The fence's cost is stated for sizing: a pooled connection held across each remote write, one row read per query. Containers are retired by the schema command once undeclared and unreferenced. maintain runs without exclusion and says why that is safe. MemMachine#1530 is recorded as agreeing. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
|
Chroma facts for the backend table, measured against chromadb 1.5.9 (local Rejects a write to a dead tenant: yes, but not classifiably. Operations route by the collection's UUID, not its name, so a handle obtained before a delete faults afterwards rather than writing into a replacement: Same result on Duplicate create is rejected, but the error is not typed. The same condition surfaces differently by transport, and never as the So a store distinguishing already-exists from any other failure has to match the message. Worth noting in the row, since Creation is genuinely atomic, not merely rejected. Eight concurrent The uniqueness constraint lives in the persisted sysdb rather than client memory — a second independent Metadata values must be scalars. A nested dict raises Not verified, so the doc's caveat should stay: per-collection cost on a single node, and the collection counts Chroma Cloud actually supports. I only exercised correctness and concurrency semantics, not scale. |
Where a backend's tenant is a native object (a Chroma collection, a Weaviate tenant, a SQLite table), create is two steps, so the ledger gains a creating state that ensure resumes, and the row carries the object's address. The Chroma row and the paragraph under the table take the registry session's chromadb 1.5.9 findings from the MemMachine#1579 comment: a stale UUID raises NotFoundError, so the UUID is the fence; duplicate create is rejected but untyped, so already-exists is told by message; concurrent creates yield one winner; metadata values are scalars. The per-collection cost and Cloud collection counts stay marked unverified. schema status and prune's dry run report prunable versus awaiting tombstones. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
Same signature and reason; MemMachine#1530's names are reused and separated by incarnations, so its create raising means exists, while here a row under a never-reused key is a violated invariant and kept tombstones are load-bearing. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
…d-poll, constructor-derived configuration Tombstones move to the tenant service: a deleted tenant's row stays as the one detector of a reused id, refusing a duplicate at mint before any store is touched, and drives a periodic re-sweep through every component's reclaim, so no store keeps tombstones and the stores cannot disagree about a reused key. The reconciler's lease becomes a held row lock, and a paragraph says why locks rather than leases everywhere. Creation, deletion and configuration updates respond 202 and are polled or waited on; no lifecycle request fails for a job the reconciler will retry; the get-or-create flag goes, the 409 carrying the existing tenant. Configuration updates keep the tenant active and reach every process through the component's per-tenant row; an option is immutable exactly when changing it would touch existing data. Configuration is the resources' constructor arguments: plain classes with typed constructors, one flat resources map with unique ids and a kind table as the only registration point, dependencies reflected from annotations, third-party clients as factory kinds; the relation to the earlier resource_initializer proposal is stated. Concurrency scope levels are process, host, cluster. Both SQLite vector stores move to shared tables, and no store creates a table per tenant. All numeric defaults are removed; settings are named. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
…llables A duplicate-name create is 409 tenant_exists and nothing else; a caller wanting the existing tenant looks it up. The sqlite-vec paragraph records that a vec0 partition key prunes the KNN to the tenant's own chunks (two orders of magnitude on 0.1.9 with 400 tenants) and that the per-partition cost is a chunk allocation both layouts paid identically, with chunk_size as the knob. A kind names a callable, class or factory, ours or third-party; Params models are optional grouping; the loader reflects the signature, validates scalars, and instance-checks resolved dependencies against annotations via validate_call, which is stated as exactly what a runtime check can and cannot prove. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
System fields are first-class typed parameters (since, before, producers) stored under a reserved memmachine_ key namespace so stores filter them with the same machinery as user properties. User properties are scalar, bounded, immutable, copied verbatim to derived data; an opaque payload is proposed for unfilterable metadata. Filter indexes are declared once per vector store in configuration and never created dynamically; undeclared keys are filtered in the segment store. A filter is a constructed closed-union tree, a JSON object under a generated schema at the API and in MCP, never a string language. Routing is per backend capability and selectivity: declared keys during the search, undeclared keys by a selectivity probe choosing an allowlist or bounded post-filter. Every count is a maximum; nothing is called top k. Reference: the default branch of edwinyyyu/MemMachine. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5
…k is one statement per direction Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
…der, fix a fresh table's plans Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
…ex when measured Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
…'s new name Records the MemMachine#1659 decisions: the store's write() transaction inside which the event memory upserts its vector records, the event rows that hold an event once with rejection by the database's primary key rather than a caller convention, delete_events and the by-event derivative lookup, the exclusive fence read repair uses, the purge order, the tenant deletion order, the residue analysis and the designs rejected for it (record-state ledger, change data capture, vector outbox), and the rename to EventMemoryStore, config key included. The shared-tables doc keeps its name here; MemMachine#1659 renames it. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Claude Code and Codex as clients of the tenant-scoped v1 API the server redesign specifies, on today's server: search and expand as MCP tools served by the server at /v1/mcp with the tenant in a header, capture as a Stop hook posting events, an installer for the agents' configs. Records what is carried over from the in-process claude_memory design, the tenant, session and source mapping, full ids now with shortening as a client concern later, the PR plan starting with recall, what is deferred and why, and the decisions that are expensive to reverse. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Tenant is roughly one human user with lifecycle-free sessions and a v1 lifecycle of its own that must not step on the v2 API's; projects are user-defined properties; the captured event uuid is the transcript entry's; messages-only embedding is the initial deriver policy and applies going forward; MCP is served by the server over HTTP. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The installer ships as `memmachine agent install` and `memmachine agent disable`, and capture is the third slice, with the installer, since the stack was cut to three. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
|
Status 2026-09-30, recorded here so it is not only in chat: this design is deferred as a whole (too large to ship as one change); the priority is working horizontal scaling first. It stays the reference for vocabulary and target shapes (registry row, incarnation, durable jobs, declared scope). Work proceeds as targeted changes under #1755 (session lifecycle), #1756 (background work), #1757 (concurrency scope) and #1758 (replica-safe defaults), all under #1574. 🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu. |
One document for the Event type as the three event-memory changes on main shape it: identity and immutability, the event timestamp, source, properties, context parts, and block kinds, with the reason for each choice and the alternatives considered; the system-field criterion and the reserved key namespace; a field-by-field mapping of Episode onto Event, since the proposal is for Event to become the server's top-level type; and, scoped to the episodic memory, how an event is segmented, derived, composed, rendered, and stored, ending with the comparison of one type pair per episode against blocks and parts that mix. Co-Authored-By: Claude Fable 5.1 <[email protected]>
Summary
A design proposal, no code:
design/server_redesign.mdand one specification per component underdesign/components/. It started as a tenant lifecycle design and now covers the server, because the lifecycle cannot be fixed inside the current API, configuration and wiring. Againstspeedkick, which carries the segment store overhaul it refers to (#1548, the merged copy of #1545). Tracking: #1574. Line references are tospeedkickat 7752e4c. Settings are named, never given numeric defaults. The premise: this is a clean slate; the first deployment carries no data forward.Decisions the document records
active,deleting,deleted); it exists the moment its row is inserted and holds nothing by itself. Each component's per-tenant resources (the event store's partition, a memory subsystem's stores) are enabled, configured and disabled on their own throughPUT/DELETE /v1/tenants/{id}/components/{name}, each with its own state, jobs and tombstone; a template is the set of sections a create enables at once. Deleting a tenant disables every component and retires the id. The id is a UUID minted per lifetime and is the physical key in every store (no incarnation anywhere); a deleted tenant's row keeps only the id anddeleted_at(no former name, so no personal data outlives the tenant in the registry).clean_at; after a retention of at least a day on the database clock a verification round runs, and only if it finds nothing again is the tombstone removed (finding something clearsclean_atand the rounds start over). The tenant row goes when its last component tombstone does. Every remote client must have a request timeout, which is what makes the retention safe.pendingtorunning), one job per claim, on its own connection; the hook runs with no database lock held on any row. Arunningjob whose claim is older thanreclaim_afteris taken over, which is the one thing time decides in the reconciler and is liveness only: a slow step may run twice, and idempotent hooks plus store fences make the second run a wasted write. Ordering comes from the state machine on the component row, idempotent hooks, strict creates, and the sweep step callingdeletebefore it purges, so a resource a lateprovisioncreated is still unlinked and removed. Eligibility is computed at claim fromlast_run_at,last_outcomeandattempts(backoff capped, reset on success; a replay with more log continues at once);reset_replayupdates existing rows only.session_idandsource_idbesidetimestamp(system fields with server semantics: the one total order is within a session; the source is the filterable identity, its readable name a context part) and aContextthat is a mapping from part kind to one registered Pydantic part. Blocks are a registered family of kinds; a segment is one block, so a block's kind is a system field of the segment.EventandContextare reshaped for one reason: to prohibit dynamic index creation. The event store keeps events and a per-tenant log of additions and deletions with commit-ordered positions; every field round-trips, the timestamp with its ingested offset. Ingest responds 202 once durable and the replay job is pending; processing is thereplayjob per memory subsystem, one consumer each, which advances its watermark only after both derived stores hold the batch, and?wait=polls the watermarks. The job carries no content and no ids: the log holds them. No rebuild of derived data: a reprocessing is a new ingestion into a new tenant. Log compaction is an operator command.provision, which knows the key's provenance. Data consumers hold stateless handles (EventPartition,SegmentPartition,VectorCollection), the store's operations bound to one key with no method taking a key, so a wrong key is unrepresentable past construction.process<host<cluster; Feat: Declare concurrency scope and data-plane contracts on storage ABCs (collection registry stack 4/5) #1531 is the reference): every resource computes its scope from its constructor arguments, the deployment declares the scope it runs at, and startup refuses anything narrower. The reconciler is a role.kind; composition is one Python function; settings are data,ServerSettingsthe root of the file; scoped views wherever a shared resource serves several holders; three scopes told apart by identity (resources, configured objects, call arguments).EpisodicMemoryManageris the resource; it builds oneEpisodicMemoryper request from the tenant's handles, embedder, format and cached segmenter and deriver, and runs the search stages.EpisodicMemory.queryis the vector stage with one limit and one threshold (min_similarity, cosine) and returns hits in descending score, each a context window with the matched segment marked; the manager runs the optional reranking stage over the rendered windows with its own candidate count and threshold, so over-fetching is one limit set above another.SearchOptionsis one model: the tenant'ssearchsection with every field set, the request's overrides with every field optional. Events are not returned by search. Expansion walks the anchor's session in the one total order (session, timestamp, event position, index, offset), counted in segments, the one unit, and returns the two sides without the anchor, which the caller holds, so a neighborhood survives an anchor that fails the filter (Let a filtered-out seed still anchor its segment context #1498); MCP tools expose no expansion counts to the model.agentic_expansionwhere it runs in production: with it on,encodetreats near-duplicate derivatives as a cluster (stored neighbors above a cosine threshold, earlier derivatives of the same batch, itself) and trims a cluster over the target size from the temporal middle, deleting stored derivatives and skipping new ones; segments and events are untouched, so the loss is to search only. A mutable tenant option whose threshold is calibrated per embedder.memmachine_key namespace. User properties are scalar, bounded, immutable, copied verbatim to derived data; filter indexes are declared once per vector store, never dynamically; undeclared keys are filtered in the segment store, whose expression indexes a deployment names in settings and the schema command creates. A filter is a constructed closed-union tree, a JSON object at the API and in MCP. Routing by backend capability: the declared part is evaluated during the search, the undeclared part afterward by the segment store with bounded widening; the store scores nothing by id. Every count is a maximum./v1/tenants,/v1/tenants/{id}/components/{name},/v1/tenants/{id}/events,/v1/tenants/{id}/episodic-memory/*, one error handler, closed error-code set includingcomponent_not_enabledandcomponent_not_active, no tracebacks in responses.include_objectlimited to its tables;memmachine schema upgradeis the only thing that runs DDL that is not tenant-specific, creates the declared property indexes outside its transaction, and provisions containers;servefails only when a database is behind its code, so a previous release may restart after a migration under expand/contract.Component specifications
design/components/: one file per component (tenant service with registry, component enablement, jobs, reconciler and tombstone pass; key registry; event store; segment store; vector store;EpisodicMemory;EpisodicMemoryManager; context; blocks; ingest service; filters and properties; server and settings), each with its API, storage, fencing, settings and the changes required of an existing component; identifiers are typedUUIDthroughout. Race matrices tabulate every concurrent pair on a tenant or component row and on the data path. Every SQL-backed component's schema is given with SQLAlchemy types mapped to PostgreSQL and SQLite, constraints and indexes. Every contract is an ABC; every store is two, the store (lifecycle, and constructing handles) and the handle (data). Types are few and each answers one need (StoredEvent,IngestResult,LogEntry,SearchHit,Tenant,SearchOptions); listings take the last position or name as the cursor. Job kinds are four (provision;delete, the unlink;sweep;replay), all defined by the tenant service;MemorySubsystemextendsTenantComponentwithreplayandwatermark.One naming scheme for values: settings, tenant configuration (per component row, with options), templates, overrides, defaults, partition configuration, request parameters, job arguments; in identifiers
SettingsandConfig.What must be built first
A section names what the first deployment freezes and so must precede the first tenant (the key scheme, including the physical key inside each store, and the tenant tables with tombstones; the event store with positions; stored-record conventions; a registry row per vector key; the public surface; Alembic from the first migration), what must land before churn, what can follow at any time, and the narrowest first deployment.
Relation to issues
Listed in the document's "Relation to open issues": #1548, #1530, #1531, #1571, #1575, #1576, #1577, #1572, #1573, #1564, #1565, #1537, #1563, #1535, #1570, #1542, and #1574 as the tracking issue.
Open questions
Hierarchy (flat proposed; an org-level share if semantic memory returns); event size limits; readable metadata as a
jsonblock kind or a designed field; retention by age, source or session as a job kind; which vector backends the first release implements; selecting the ungrouped stream in a search; log compaction as a scheduled duty.🤖 Generated with Claude Code
https://claude.ai/code/session_01MbYdqGZsuws6Z2WHYfCCR5