Skip to content

[session storage 1/2] Create a session's storage with the session, never on a request - #1622

Draft
edwinyyyu wants to merge 26 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/session-owned-storage-speedkick
Draft

edwinyyyu wants to merge 26 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/session-owned-storage-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

The event backend created a session's vector store collection and segment store partition on the first request that opened the session, so a search or a write for an unknown session created storage as a side effect, and the service locator was the only place that knew both stores' create paths.

The owner is the session. Every path that creates a session row runs through EpisodicMemoryManager._create_session, which inserts the row and, when the row is new, creates the session's partitions in its segment store and its vector store (create_episodic_memory_storage); an equivalent re-create accepts the row and leaves the storage alone, so create_or_validate_session now returns whether it created the row. The request path binds handles with the stores' lookups and raises SessionPartitionMissingError when a partition is absent: a session without its storage is broken, not new. Deleting a session with no open instance deletes its partitions by key, so a session whose storage was never fully created can still be deleted. MemMachine.create_session goes through the manager for the same reason.

The semantic manager owns its one collection and creates it, once, at the storage's first use. With that, nothing calls the stores' open-or-create.

No memory request creates a session on main (#1624): sessions are created only through MemMachine.create_session, which now goes through the manager.

Adaptation for main: the created flag is added to #1677's create_or_validate_session (on speedkick the same hunk sits on #1539's create_new_session_if_not_exist). main still has the manager's open-or-create, which #1624 removes on its own, so its create branch goes through _create_session too, and the storage test covers a session created and reopened through it. The request path's lookups are on the one-collection store of [vector store scale-out 6/6]: get_partition(session_key) on the event backend's store, and the semantic manager creates its one partition strictly (get, create, get) rather than through open-or-create.

Position: on [vector store scale-out 6/6] directly, in parallel with [search results], since nothing here reads a score or a stored vector.

Stack

21 open PRs: three independent PRs, and the vector store tree of short parallel branches. Every PR in the tree has feat/horizontal-scaling as its GitHub base, and the independent PRs have main. The branches are in a fork, and a pull request can target only this repository's branches, so the on column gives the order the PRs build on each other. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main
— #1786 Refuse a filter value of the wrong type for a datetime column, and answer an invalid list filter with 422 (port of #1620) main
— #1792 Refuse property values that some store refuses or alters where an episode enters main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1702 and #1670 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (merged into feat/horizontal-scaling) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs feat/horizontal-scaling
[vector store scale-out 3/6] #1734 (merged into feat/horizontal-scaling) Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge feat/horizontal-scaling
[vector store scale-out 4/6] #1735 (merged into feat/horizontal-scaling) Move the Qdrant store onto the collection registry feat/horizontal-scaling
[vector store scale-out 5/6] #1736 (merged into feat/horizontal-scaling) Move the Milvus store onto the collection registry, against a Milvus server feat/horizontal-scaling
— #1775 (merged into feat/horizontal-scaling) Accept attempts exhausted in the lifecycle churn contract, and say which Qdrant operations filter on the incarnation feat/horizontal-scaling
— #1779 (merged into feat/horizontal-scaling) Classify Qdrant errors by status code alone feat/horizontal-scaling
— #1813 (merged into feat/horizontal-scaling) Create Qdrant collections with strict mode off feat/horizontal-scaling
— #1788 Refuse a repeated record UUID or a non-finite property value at upsert, and state the datetime property contract feat/horizontal-scaling
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1702 Keep undeclared properties out of the vector store feat/horizontal-scaling
[user properties 2/2] #1670 Remove per-project filterable properties (port of #1606) #1702
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1670
[session storage 1/2] #1622 (this PR) Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1627
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1628
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1663
[sqlite store fixes 1/7] #1460 Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[milvus options] #1741 Let a deployment tune a Milvus collection's vector index and its searches #1618

This PR is its one commit, 809d796bd, stacked on #1627. #1625 is stacked on it.

Verification

At this PR's head 809d796bd, on 2026-10-09, on feat/horizontal-scaling with #1813 merged: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean.

Before the rebase onto #1813's merge: At this PR's head 9087947ff, on 2026-10-09: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean. The full suites were not run at this head: the server suite without integration tests last passed at 20a2a2236, on 2026-10-08, 2175 tests, and the integration tests of the vector stores, the resource manager, episodic memory, and semantic storage at 20a2a2236, on 2026-10-08, 687 tests, against PostgreSQL 16, Neo4j, Qdrant 1.19.1, and Milvus 2.6.24 in containers. This head differs from those by #1779's change to the Qdrant store and the rebase onto the merged #1736.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

@edwinyyyu
edwinyyyu force-pushed the feat/session-owned-storage-speedkick branch 5 times, most recently from 4266b09 to 5ce230e Compare September 14, 2026 23:16
@edwinyyyu edwinyyyu changed the title [vector store 3/13] Create a session's storage with the session, never on a request (speedkick) [vector store 3/12] Create a session's storage with the session, never on a request (speedkick) Sep 14, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/session-owned-storage-speedkick branch 2 times, most recently from 4266b09 to 0a7b52a Compare September 15, 2026 17:25
@edwinyyyu edwinyyyu changed the title [vector store 3/12] Create a session's storage with the session, never on a request (speedkick) [vector store 3/13] Create a session's storage with the session, never on a request (speedkick) Sep 15, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/session-owned-storage-speedkick branch 5 times, most recently from 40985ea to 90175ca Compare September 15, 2026 19:16
edwinyyyu and others added 26 commits October 9, 2026 17:04
The event backend wrote every property of an event into its vector record
and mapped the caller's whole filter onto the vector store, so a user key
that a deployment never declared was both stored and filtered there: on
Qdrant and Milvus as unindexed payload a filtered query scans for. The
segment store already holds every property and already receives the whole
filter for the context windows, so the vector side only duplicated work the
segment store does anyway.

The vector record now carries the keys the collection declares:
EventMemory's reserved timestamp, the `_`-prefixed system properties an
adapter stamps on the event, and the keys a project's `properties_schema`
declares. The vector store is queried with the conjuncts of the filter
that name only such fields; a conjunct is dropped whole when any field
under it is undeclared, so dropping only ever widens the vector search,
and the segment store narrows it back on the windows. An undeclared key
never reaches the vector store, so a tenant's undeclared properties cannot
shape what it stores or scans.

`filter_fields` joins the filter parser: every field name a tree addresses.

Rebased onto MemMachine#1631, where the vector record no longer carries the segment
uuid (the segment store maps a derivative to its segment).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ck) (MemMachine#1606)

* Regenerate the OpenAPI document under the locked FastAPI

`docs/openapi.json` predates the FastAPI release in `uv.lock`
(0.141.1), whose `ValidationError` component carries `input` and `ctx`;
regenerating the document with `docs/tools/generate_openapi.py` adds the
two fields and changes nothing else. Separate from the API changes above
it so their diffs of this file show only what they change.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

* Remove per-project filterable properties

A project could declare `properties_schema`, a set of caller property keys
with types, on its long-term memory configuration; the event backend merged
it into the vector store collection's indexed schema and rejected filters on
any other `m.<key>`. That let a tenant create database resources (indexes,
columns) by naming them in a request, which is what forced per-collection
native resources named by a hash of their schema on the backends that limit
them.

The option is removed from the server configuration, the project API and
the memory-configuration API, the Python SDK, the sample configurations,
the configuration docs and the OpenAPI document. A filter may name any
`m.<key>`; the stores evaluate it on the properties they hold. What a store
indexes is decided by the deployment, not per project.

A breaking API change on `speedkick`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Rebased onto MemMachine#1631: the per-project schema also leaves MemMachine#1631's service
locator, which creates the session's collection in a retry loop, and the
commented option goes from the event sample configuration MemMachine#1698 added.

---------

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
(cherry picked from commit a8322a7)
…res' lookup to get_partition

A vector store's logical collection becomes a partition, the segment
store's word for the same thing, and both stores' lookup is get_partition,
answering None like a Python get. Identifiers only, produced by the script
below; the (namespace, name) identity, the per-partition config and every
docstring are as they were, and the next change gives them their meaning.
The native clients' create_collection and delete_collection keep their
names.

The vocabulary also reaches what sits under the rename: the registry
package, its modules, classes and tables, the `collection_registry` option
of a Qdrant or Milvus backend and the store parameter and attribute that
carry it (`partition_registry`), the stale-handle and pending errors, the
registry-backed base and its handle lookup, the purge method
(`purge_deleted_partitions`, the segment store's name for it) and
open-or-create.

```sh
set -e
cd "$(git rev-parse --show-toplevel)"
git mv packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_collection.py \
       packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_partition.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_lifecycle_contract.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_lifecycle_contract.py
git mv packages/server/src/memmachine_server/common/vector_store/collection_registry \
       packages/server/src/memmachine_server/common/vector_store/partition_registry
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/partition_registry.py
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_partition_registry.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_registry \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry
git mv packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_collection_registry.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_partition_registry.py
git ls-files -z 'packages/server/*.py' 'docs/*.mdx' 'sample_configs/*.sample' 'sample_configs/*.yml' 'deployments/helm/templates/*.yaml' | xargs -0 perl -0pi -e '
  s/VectorStoreCollection(?!Config)/VectorStorePartition/g;
  s/in_memory_vector_store_collection/in_memory_vector_store_partition/g;
  s/collection_lifecycle_contract/partition_lifecycle_contract/g;
  s/CollectionLifecycleContract/PartitionLifecycleContract/g;
  s/collection_registry/partition_registry/g;
  s/SQLAlchemyCollectionRegistry/SQLAlchemyPartitionRegistry/g;
  s/CollectionRegistry/PartitionRegistry/g;
  s/RegisteredCollection/RegisteredPartition/g;
  s/get_registered_collection/get_registered_partition/g;
  s/_build_collection_handle/_build_partition_handle/g;
  s/\bCollectionT\b/PartitionT/g;
  s/register_collection\b/register_partition/g;
  s/\b_Collection\b/_Partition/g;
  s/CollectionRow/PartitionRow/g;
  s/vector_store_collection(?!_schema|_namespace)/vector_store_partition/g;
  s/open_or_create_collection/open_or_create_partition/g;
  s/open_collection/get_partition/g;
  s/_purge_deleted_collections_forever/_purge_deleted_vector_store_partitions_forever/g;
  s/purge_deleted_collections/purge_deleted_partitions/g;
  s/def create_collection\(/def create_partition(/g;
  s/def delete_collection\(/def delete_partition(/g;
  s/\.create_collection\((\s*namespace=)/.create_partition($1/g;
  s/\.delete_collection\((\s*namespace=)/.delete_partition($1/g;
  s/\.create_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.create_partition/g;
  s/\.delete_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.delete_partition/g;
  s/"create_collection"/"create_partition"/g;
  s/"delete_collection"/"delete_partition"/g;
  s/"open_or_create_collection"/"open_or_create_partition"/g;
  s/only delete_collection is invoked/only delete_partition is invoked/g;
  s/test_delete_collection_/test_delete_partition_/g;
  s/open_partition/get_partition/g;
  s/(vector_store_partition \(VectorStorePartition\):\n\s+)Vector store collection\./$1Vector store partition./g;
'
uv run ruff check --fix --quiet packages/server
uv run ruff format --quiet packages/server
```

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…uilt by the composition root

A vector store was a factory of logical collections, each identified by
a (namespace, name) pair and created with its own dimensions, metric and
schema; Qdrant and Milvus shared one native collection among logical
collections of equal configuration, under a name derived from a hash of
that configuration, and the registry mapped (namespace, name) to it.

A store is now one collection: `VectorStore(vector_store_name,
vector_dimensions, similarity_metric, indexed_properties)` names its one
native collection (or its tables and index files) at construction, every
partition of it shares the store's dimensions, metric and schema, and
`provision()` creates the store's durable resources idempotently, before
`startup`. `create_partition(key)`, `open_or_create_partition(key)`,
`get_partition(key)` and `delete_partition(key)` take a string key; a
partition is the records carrying its incarnation in the store's native
collection (Qdrant, Milvus) or a pair of tables (the SQLite stores), and
the registry records what each partition was created under, so a store
built with other dimensions, another metric or another schema raises
`VectorStorePartitionSchemaMismatchError` instead of reading columns and
vectors that are not there. "Collection" names only the native Qdrant or
Milvus collection. Vector store names match `[a-z0-9_]+` and are at most
32 bytes, the rule partition and property keys follow, so every name
works on every backend: the Qdrant store names its collection by the
vector store name, and the Milvus store by `sys_` followed by it, since
Milvus requires a leading letter or underscore and names its own
internals with a leading underscore. The hash-derived native names go,
and with them the per-partition config.

The partition registry is keyed by partition key within one store: its
tables, `partition_registry_pt` and `partition_registry_gc`, are shared
by every store on a database and keyed by vector store name, so the
registries of several stores share one database and nothing else.
`mark_live` and `unregister_incarnation` take the incarnation alone. On
the registry-backed base, the storage a store's partitions share is
prepared by `provision()` (`_prepare_storage()`), and the storage a
partition keeps of its own by `_prepare_partition_storage(key,
incarnation)`, between its registration and its mark; the Qdrant and
Milvus stores keep nothing per partition. A purge round works on one
incarnation in the store's own collection.

`DatabaseManager.get_vector_store(backend, vector_store_name=,
vector_dimensions=, similarity_metric=, indexed_properties=)` builds and
caches one store per (backend, vector store name), provisioning its
registry and then the store; asking for a name again with other
dimensions, another metric or other keys is a
`VectorStoreConfigurationError`. The event backend's store is named by a
UUIDv5 of its embedder id in the UUIDv5 of its backend's key in a namespace
fixed for the event backend, written as 32 hexadecimal digits, so neither
the key nor the id is constrained; stores on different backends never share
a name, so their registries stay apart when the backends share a registry
database, since a store name identifies one store among all the stores
whose registries share it.
Semantic memory's one store is named `semantic_memory` whatever its
embedder and holds every org's features in one partition,
`semantic_memory`, as main's one collection does.
The event backend opens a session's partition with
`open_or_create_partition`, which waits on one another worker is still
creating.

The SQLite stores change shape only: their registry tables become
`vector_store_sqlite_pt` and `vector_store_sqlite_vec_pt`, keyed by
vector store name and partition key, the pending-operation log is keyed
the same way, and per-partition table names embed the vector store name.
No migration is provided; existing SQLite vector data is orphaned.

Rebuilt as one change on MemMachine#1631: this PR's earlier history carried copies
of MemMachine#1631's commits as of `38684b385`, its own commits on them, and
mirrors of MemMachine#1631's later commits in its terms. Its code is that history's
final tree, which was verified, merged with main and without MemMachine#1624's
commits. The design documents are still MemMachine#1631's and describe collections.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…up, not a separate provision

A vector store and its partition registry each had a provision() that
created their durable resources, run by the composition root before
startup(). The split between provisioning a store and starting it is
MemMachine#1570's to make for every store at once, so here startup() does both
again, as on main: a store's startup prepares the storage its
partitions share (the native collection and its indexes, or the SQLite
tables), and the registry's startup creates its tables, as MemMachine#1631's does.
The database manager starts the registry, then the store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…er, live from resolve

The registry is addressed by partition key only. `register(key, schema)`
answers a PendingRegistration, whose `mark_live()` answers the same life
as a LiveRegistration or raises VectorStorePartitionDeletedError when the
partition was deleted meanwhile, and whose `unregister()` abandons that
life alone. `resolve(key)` answers the LiveRegistration of the live
partition, None when there is none, and raises
VectorStorePartitionPendingError, which now carries the pending
partition's schema, when it is pending. A LiveRegistration's
`require_current()` raises VectorStorePartitionHandleStaleError once its
partition is deleted. `get`, `mark_live(incarnation)`,
`unregister_incarnation` and RegisteredPartition are gone, so deletion is
by key or by a pending registration, never through a live one. The
SQLAlchemy registry supplies frozen-dataclass registrations holding the
engine and the vector store name, and a module function runs the
unregistration transaction.

The base handle takes its live registration, fences on
`require_current()`, and `_partition_handle` builds a handle from one.
`create_partition` lets VectorStorePartitionDeletedError propagate, and
the VectorStore interface names it; open-or-create catches it and creates
again, refuses a pending partition of another schema at once from the
pending error's schema, and re-raises the last pending error. A
`get_partition` of a pending partition of another schema still reports the
mismatch first. The Qdrant and Milvus handles take the registration in
place of the key, the incarnation and the registry lookup.

Tests follow: the registry's tests answer registrations and gain one for
`require_current`; the base's tests patch the pending registration type's
`unregister` and expect the deleted error from a creation a deletion
undid; the lifecycle contract patches the live registration type's
`require_current`, registers its racing winners through pending
registrations, and counts the deleted error among churn's outcomes; the
Qdrant and Milvus tests that build a handle on a mocked client give it a
registration that stays current.

The same change as MemMachine#1734's 735c649, f88aaa8, e25be60 and
b2bb2d8 and the handle halves of MemMachine#1735's 9096edd and MemMachine#1736's
5582cf4, for this PR's partition registry. The event-backend locator
change has no counterpart: the locator here opens its partition with
open-or-create, which creates again itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nd their partitions

The design documents came from MemMachine#1733-MemMachine#1736, where a store holds logical
collections addressed by namespace and name, each with its own
configuration. Here a store is one collection, with its dimensions, metric
and declared schema fixed at construction, and a partition is one tenant's
records in it, addressed by key. The documents say so:

- the collection registry document becomes the partition registry
  document: a registry belongs to one store and is addressed by partition
  key; its tables are `partition_registry_pt` and `partition_registry_gc`,
  and a tombstone needs no location, since the incarnation alone finds a
  dead partition's records in the store's native collection; a partition
  created under another schema is refused; `startup` creates the tables, as
  every store's startup creates its durable resources; a decision records
  that a store is one collection;
- the Qdrant and Milvus documents lay out one native collection per store,
  named by the vector store name (`sys_` and the name on Milvus), created at
  startup, where a partition's creation makes nothing in the backend;
- the overview, isolation, consistency and purge documents speak of
  partitions, `purge_deleted_partitions` and `settle(partition)`, and the
  overview describes the store and its partitions.

The measurements and their conditions are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ain open-or-create's give-up

Two review changes to MemMachine#1734's registry and base store, for this PR's
partitions:

- `Reservation.cancel()` acts only while the partition is pending, as
  `confirm()` does. A creator whose confirmation committed but whose
  answer was lost, and which then cancels, leaves the live partition
  alone: only a deletion by key ends it. The registry contract, the
  SQLAlchemy registry and the registry design document say so, and a
  test cancels after a confirmation.
- Open-or-create's `VectorStoreAttemptsExhaustedError` is raised from the
  race it last lost, the last `VectorStorePartitionAlreadyExistsError` or
  `VectorStorePartitionDeletedError` it caught. A partition that stays
  pending still raises the pending error itself. A base-store test checks
  the cause.

The same changes as MemMachine#1734's 9daf7e7 and 93a85da, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…r is cancelled

A creation cancelled its reservation only when preparing the partition's
storage raised. A confirmation that raised, or a creation cancelled while
it confirmed, left the partition pending until someone deleted it.

The base store's creation step now prepares the storage and confirms the
reservation together, and cancels the reservation if either raises or the
creation is cancelled, shielded as before; the cancel's task reports its
own failure. The cancel acts only on a pending partition, so a
confirmation that committed before its failure was observed stands.
create_partition and open-or-create both go through it; open-or-create
still takes a confirmation's VectorStorePartitionDeletedError as a race to
create again. Base-store tests cover a failed confirmation, a cancelled
one, and one that committed before failing. The registry design document
says so, and the purge document says which writers wait on SQLite's lock
during a purge round.

The same changes as MemMachine#1734's f4c585e and 4dde7c8, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The query contract said nothing of a match scoring exactly the threshold,
and Qdrant dropped it: the server compares its threshold in single precision
and keeps only scores strictly better than it.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the server the adjacent single-precision value on
the worse side, and applies the caller's threshold exactly itself; a
threshold beyond single precision is not sent. Both SQLite stores already
keep the match, and each now tests it on every metric it supports.

The same changes as MemMachine#1735's ccba117 and 6bbc131, for this PR's stores.
MemMachine#1736's cc29d27, Milvus's test of the same, is ported with the store
tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…cycle contract's drain

MemMachine#1735's and MemMachine#1736's latest tests, for this PR's partitions:

- Every Qdrant store test runs against a Qdrant server, over REST and gRPC:
  local mode ignores payload indexes and raises its own exceptions where a
  server answers not-found or already-exists, so it is no longer a fixture.
  The tests on a mocked client stay in the default suite.
- Both stores gain tests that partitions and stores of different names keep
  their records apart, that purge rounds reclaim a write landing after a
  round and drain two stores' tombstones alone, that ranking, scores and the
  threshold follow every similarity metric, that a point or entity the
  server refuses on its own raises, that concurrent creations and startups
  agree, that two stores churning one registry keep every partition exact,
  and that seeded operation sequences, and on Milvus random filters, agree
  with a model. The Milvus tests also pin its read consistency and its
  settings with values other than their defaults.
- The lifecycle contract's drain fails after a bounded number of rounds, it
  settles before checking that a new life is empty, and it checks a stale
  upsert by what the partition holds rather than by its registry reads.

The same changes as MemMachine#1735's 30c30f4, e9e71b5, 6c2362b, 4d3241c,
836db21, 252bcd2, 5e8c049, ef8d3ab, 276a034, 4fb5bd4 and
2aabd3d, and MemMachine#1736's 509c235, 7f1eb3b, 6e312e0, 6d1ed1e,
25e7948, 6f36877, d43e9e6, 6dfbf59 and cc29d27, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The base store's purge contract says a round that finds the incarnation's
storage missing returns False, but neither store's round checked: once its
native collection was dropped from outside, every round on it raised,
counted against the tombstone, and dead-lettered it after ten.

A Qdrant round that the server answers not found, over REST or gRPC, and a
Milvus round that finds no native collection, now return False: the
collection is gone with everything in it, so the tombstone is retired. Each
store has a test that drops its native collection and drains the purge.

The same handling as MemMachine#1735's and MemMachine#1736's stores, whose rounds already
checked for a missing native collection; MemMachine#1736's 2575af0 tests it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A round that never ended was counted by the next claim, which took a
CASE-shaped update that both recorded the lost round and ended its claim;
the count lived apart from the claims that made it.

Each claim now counts its attempt: the claim is a plain update that
increments `consecutive_attempts` (renamed from `consecutive_failed_rounds`),
stamps `claimed_at`, and bumps the generation, and `_TombstoneClaim.attempt`
carries the count. A tombstone is claimable when it has no attempts, when no
claim is open and the backoff has passed since its last failure, or when an
open claim has outlived its lease plus the backoff. A raised round ends its
claim and stamps `last_failed_at`; a cancelled round ends its claim and takes
its attempt back; a round that found records resets the attempts. After
`_MAX_PURGE_ATTEMPTS` attempts a tombstone is dead-lettered, and a last
attempt that raises is reported. A retry logs which attempt it is, and a
raised round's error names its incarnation and attempt. The registry tests,
the purge design document, and the registry design document follow.

The same change as MemMachine#1734's 1732038, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
`consecutive_attempts` counted the purge rounds claimed since a round last
found records, but its name said only that they were in a row. It is
`attempts_without_progress`, and `_MAX_PURGE_ATTEMPTS` is
`_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The column's comment, the claim's
attempt, the backoff parameter's description, the dead-letter log, the
class docstring, the tests, the purge design document, and the registry
design document's table follow; nothing else changes.

The same change as MemMachine#1734's 96a8b5b, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…urrent claim

The measured cost of the claim came from earlier forms of it, which held a
transaction across the round. The section now gives the figures measured on
2026-10-06 against the claim this registry ports, naming the MemMachine#1734 commits
they were taken at: the backoff scan with 1k, 10k, and 100k tombstones
backing off, and interactive throughput and liveness latency beside two
sweepers on PostgreSQL and on SQLite.

The same change as MemMachine#1734's bab379b, for this PR's purge document.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's review round, for this PR's Milvus store:

- Filter strings reach Milvus as UTF-8, so a character outside the Basic
  Multilingual Plane parses, and an undeclared property's condition requires
  the stored type tag to be one the value compares with.
- Startup's already-exists guard goes, with its mocked-client test: Milvus
  answers a create of an existing collection with the same schema with
  success.
- The store no longer re-sorts search results, which Milvus returns best
  first.
- The mocked-client purge test disposes its registry engine when it fails.

The same changes as MemMachine#1736's 1a7fffa, d6fc219, 9fd7f76, 6fcd8c5, and
c264b3d, for this PR's store. MemMachine#1736's 28e8c50 and b3ef4f1 have no
counterpart here: startup prepares the store's one native collection before
any purge round runs, and the store refuses an unsupported metric at
construction, before anything is reserved.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The declared path counted an int and a float as comparable either way, so
a float filter value on a declared int property reached Milvus as a float
literal against an INT64 field, which Milvus refuses to parse ("cannot
cast value to Int64", code 1100). A float now compares only with a float
property, and matches no int one, as a value of another type matches
nothing; an int still compares with a float property by value.

The new test filters a declared float property with an int and a declared
int property with a float; it fails with the float counted as comparable
with an int property.

The same change as MemMachine#1736's 31fc7c4, for this PR's store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…y name

MemMachine#1736's baabf48 creates Milvus native collections at Bounded by name,
and this PR's base carries that into startup's create, so the read level
the purge and the tombstone retention rely on no longer comes from
pymilvus's default. The consistency test now spies create_collection and
checks the level it names, as MemMachine#1736's test does; it fails with the level
dropped from the create. The design document and the store's docstring
say the store creates its one native collection at Bounded.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's f56686f halves a Milvus upsert refused with RESOURCE_EXHAUSTED,
and this PR's base carries the halving into the partition handle's
`_upsert`. Its tests come here in partition terms: the integration test
upserts 1,200 records with a 60,000-character property, about 72 MB, and
fails with RESOURCE_EXHAUSTED without the halving; the mocked-client tests
pin the halving, a single refused entity raising, and a timeout or another
refusal sent once, on a handle `_partition_on` builds, which the mocked
delete test now uses too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nt alone

MemMachine#1736's 49b9ec2 drops the offset field beside a declared Milvus
datetime, keeping its UTC instant alone, and this PR's base carries that
into the store. The tests follow in partition terms: the native
collection's fields are exactly the fixed ones and one per declared
property; a declared datetime is stored as its instant in UTC; and a store
whose declared datetime has the longest property key, 32 bytes, writes its
one field, reads it back, and matches the record by its instant, the test
MemMachine#1736's b015229 adds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The test of MemMachine#1813, merged into this PR's base, in partition terms: the
server defaults new collections to strict mode, the store's collection
reads back with it off, and a filter on a partition's unindexed property
is served.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…quest

The event backend created a session's vector store collection and segment
store partition on the first request that opened the session, so a search
or a write for an unknown session created storage as a side effect, and
the service locator was the only place that knew both stores' create paths.

The owner is the session. Every path that creates a session row runs
through EpisodicMemoryManager._create_session, which inserts the row and,
when the row is new, creates the session's partitions in its segment store
and its vector store (create_episodic_memory_storage); an equivalent
re-create accepts the row and leaves the storage as it is. The request
path binds handles with the stores' lookups and raises
SessionPartitionMissingError when a partition is absent: a session without
its storage is broken, not new. Deleting a session with no open instance
deletes its partitions by key, so a session whose storage was never fully
created can still be deleted. MemMachine.create_session goes through the
manager for the same reason.

The semantic manager owns its one collection and creates it, once, at the
storage's first use. With that, nothing calls the stores' open-or-create.

The API is unchanged: the manager's open-or-create still creates a session
a memory request names, now through the same path.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant