Repository navigation
[vector store alt 8/9] Mint an incarnation per collection life on Qdrant; delete logically, reclaim by purge (speedkick) - #1643
Closed
edwinyyyu wants to merge 10 commits into
Conversation
Reverts the second commit of MemMachine#1606 (a8322a7), "Remove per-project filterable properties", and keeps its first, the OpenAPI regeneration under the locked FastAPI 0.141.1: docs/openapi.json is regenerated here with the same tool, so it differs from the pre-MemMachine#1606 document only by the ValidationError fields that regeneration added. A project declares `properties_schema`, a set of caller property keys with types, on its long-term memory configuration; the event backend merges it into the vector store collection's indexed schema and rejects filters on any other `m.<key>`. Under this design the schema is fixed for the life of the project, a schema change is a delete and a create, and the vector store maps every distinct schema to one physical collection that all projects declaring it share. What that costs, per backend, is written where a tenant will see it: a warning in the configuration docs and the field descriptions on the server configuration, the project API, the memory-configuration API and the Python SDK. Alternative to the one-collection chain (MemMachine#1622..MemMachine#1616): the two exclude each other. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…quest The event backend created a session's vector store collection and segment store partition on the first request that opened the session, so a search or a write for an unknown session created storage as a side effect, and the service locator was the only place that knew both stores' create paths. The owner is the session. Every path that creates a session row runs through EpisodicMemoryManager._create_session, which inserts the row and, when the row is new, creates the session's partitions in its segment store and its vector store (create_episodic_memory_storage); an equivalent re-create accepts the row and leaves the storage as it is. The request path binds handles with the stores' lookups and raises SessionPartitionMissingError when a partition is absent: a session without its storage is broken, not new. Deleting a session with no open instance deletes its partitions by key, so a session whose storage was never fully created can still be deleted. MemMachine.create_session goes through the manager for the same reason. The session's vector store collection is created with the system keys and the project's own `properties_schema`, resolved and merged as the request path did before. The semantic manager owns its one collection and creates it, once, at the storage's first use. With that, nothing calls the stores' open-or-create. The API is unchanged: the manager's open-or-create still creates a session a memory request names, now through the same path. Ported from MemMachine#1622 (4266b09) onto the chain that keeps per-project filterable properties. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The simple chatbot example, the TypeScript REST demo and the Dify plugin's add-memory tool wrote to a project without creating it, relying on the write to create it. Each now creates its project before its first memory request and accepts 409 as the project already existing. No behavior changes for them; they stop depending on a write creating a project. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Adding memories to, or searching, a project that did not exist created it, with the server's default configuration, without the caller's knowledge. Now only the create-project request creates a project: a write or a search opens the session and answers 404 for an unknown project, as the search endpoint already promised; the manager's open-or-create goes. Two callers depended on the implicit creation. `org_id` and `project_id` default to `universal`, so the API promises the project `universal/universal`; the server creates it, once, at startup, and leaves one that already exists as it is. The MCP add tool names its own project and has no create-project counterpart, so it creates the project it writes to, once, and says so. The API doc strings and the OpenAPI document say which requests create a project. A breaking API change on `speedkick`. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Nothing calls them since a session's storage is created with the session: `open_or_create_collection` and `close_collection` leave the vector store interface and its four backends, `open_or_create_partition` and `close_partition` leave the segment store interface and its implementation, and the two config-mismatch errors that only open-or-create raised go with them. A store creates on `create_*`, strictly, and looks up on `open_*`, answering None; create-if-absent is the owner's, where the key's provenance is known. Source changes are deletions only. The tests that exercised open-or-create as a fixture use a test-side create-if-absent instead, and the tests of its own semantics go. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A lookup that answers None for an absent collection or partition is a Python get, not an open. Identifiers only, produced by the script below, plus the two docstring lines that said "opened"; every contract is as it was. ```sh set -e cd "$(git rev-parse --show-toplevel)" git ls-files -z 'packages/server/*.py' | xargs -0 perl -0pi -e ' s/\bopen_collection\b/get_collection/g; s/\bopen_partition\b/get_partition/g; ' uv run ruff check --fix --quiet packages/server uv run ruff format --quiet packages/server ``` The `open -> get` half of MemMachine#1626 (e9783d9); the `collection -> partition` half is not carried, because its premise is a store that is one collection. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`QdrantConf.is_distributed` put every logical collection in its own shard of the physical collection so that deleting one was a shard drop: an O(1) deletion in a cluster, at the price of a shard, with its own segments and graphs, per logical collection. The lifecycle that follows this change makes a deletion a registry write, visible at once, and reclaims the points afterward by a filter on the incarnation; the shard has nothing left to buy. The option goes from the configuration, the parameters, the database manager and the store, along with the shard-key selectors on every operation and the tests that ran a one-node cluster for it. A breaking configuration change on `speedkick`; the option was never documented. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A caller may filter on any property, declared in the collection's schema or not; an undeclared key is evaluated against the payload of the tenant's own candidate points, without an index. Qdrant Cloud enables strict mode on new collections by default, and strict mode rejects a filter on an unindexed payload key, so on such a deployment every undeclared filter failed. The store now creates its collections, the physical ones that hold points and the per-namespace registries, with `strict_mode_config` disabled. Collections created by an earlier version keep the setting they have. The chain that indexes every filterable key keeps strict mode on; this one, which lets a project declare what is indexed, turns it off. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…reclaim by purge A logical collection is identified to callers by its (namespace, name) pair and, inside the Qdrant store, by an incarnation the store mints when the collection is created. Points carry the incarnation, never the name, and a point's id is derived from the incarnation and the record uuid, so a collection deleted and re-created under the same pair starts empty, its predecessor's points are never adopted by or reclaimed out from under the successor, and neither a stale handle nor a co-tenant of the shared physical collection can overwrite a point a live incarnation holds under the same record uuid. Handles are bound to one incarnation: once it is deleted, every operation of the handle raises VectorStoreCollectionHandleStaleError. delete_collection becomes two registry writes: the collection's live entry goes, conditionally on the incarnation the delete read, so a creation that raced in from another process keeps its entry, and an entry in the store's purge queue follows, stamped with the deletion time and the physical collection the incarnation lived in. The queue is one collection per store, `memmachine__purge_queue`, so one sweep serves every namespace. The new purge_deleted_collections reclaims one dead incarnation per call, oldest first, by a filter on the incarnation; safe to repeat and to run from several processes. Qdrant has no transactions: a write that passed its fence before the deletion lands under the dead incarnation and is reclaimed with it, and a crash between the two registry writes leaves the points unreachable and unreclaimed, a leak, never a collection the purge takes from under a live one. The VectorStore contract now says what every implementation guarantees (deletion makes the collection unreachable at once and is idempotent; a collection created under a deleted pair starts empty) and what it leaves to the implementation (whether reclamation is immediate or deferred to purge_deleted_collections, and what a handle held across a deletion does). The SQLite stores and the Milvus store keep their immediate reclamation, answer False from purge_deleted_collections, and state their handle behavior in their docstrings. Existing Qdrant data is not migrated: the payload key, the point ids, the registry entries and the purge queue are new. collection_lifecycle_contract.py holds the tests a conforming store satisfies, mixed into the Qdrant module and run in local, REST and gRPC modes: stale handles, empty re-creation, idempotent deletion, purge reclaiming what deletion deferred and leaving every other collection alone -- of the same configuration, of another, in another namespace, and one re-created under another configuration while its predecessor awaits purge -- and the same record uuid held by two collections of one physical collection. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The first time a vector store is handed out, the resource manager starts the same purge loop it runs for segment stores, one per store name; close() cancels both sets before the stores shut down. A store that reclaims in delete_collection answers False and costs one call per tick. Mechanical churn in the same change: the loop takes the store's bound purge and a label for its failure log line, and its interval and pause constants lose their SEGMENT_STORE_ prefix. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
This was referenced Sep 15, 2026
Contributor
Author
|
Reopen if stack chosen. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Alternative to the
[vector store N/13]chain (#1597, #1622..#1616): the two exclude each other. This chain keeps per-projectproperties_schema(user-defined metadata indexes) and conforms Qdrant to the incarnation lifecycle; SQLite and Milvus follow later. Every PR is a draft.Purpose of the change
A logical collection is identified to callers by its (namespace, name) pair and, inside the Qdrant store, by an incarnation minted when it is created. Points carry the incarnation, never the name, and a point's id is
uuid5(incarnation, record uuid), so a collection deleted and re-created under the same pair starts empty, its predecessor's points are never adopted by or reclaimed out from under the successor, and neither a stale handle nor a co-tenant of the shared physical collection can overwrite a point a live incarnation holds under the same record uuid (with bare record-uuid ids, as onspeedkickand in #1631, they could). Handles are bound to one incarnation: once it is deleted, every operation of the handle raisesVectorStoreCollectionHandleStaleError.delete_collectionbecomes two registry writes: the live entry goes, conditionally on the incarnation the delete read, and an entry in the store's purge queue (memmachine__purge_queue, one per store, so one sweep serves every namespace) follows, stamped with the deletion time and the physical collection the incarnation lived in.purge_deleted_collectionsreclaims one dead incarnation per call, oldest first, by a filter on the incarnation; the resource manager runs it, one loop per store, as it runs the segment stores' purge. Qdrant has no transactions: a write that passed its fence before the deletion lands under the dead incarnation and is reclaimed with it; a crash between the two registry writes leaves the points unreachable and unreclaimed, a leak, never a collection the purge takes from under a live one.The
VectorStorecontract states what every implementation guarantees (deletion makes the collection unreachable at once and is idempotent; a collection created under a deleted pair starts empty) and what it leaves to the implementation (immediate or deferred reclamation; what a handle held across a deletion does). Qdrant is the implementation that conforms here; the SQLite stores and the Milvus store keep their immediate reclamation, answer False frompurge_deleted_collections, and state their handle behavior in their docstrings. Existing Qdrant data is not migrated.collection_lifecycle_contract.pyholds the tests a conforming store satisfies, run against Qdrant in local, REST and gRPC modes: stale handles, empty re-creation, idempotent deletion, purge reclaiming what deletion deferred and leaving every other collection alone (same configuration, another configuration, another namespace, and one re-created under another configuration while its predecessor awaits purge), and the same record uuid held by two collections of one physical collection.Tenant-created physical resources: limits and cross-tenant effects
Every project (tenant) declares its own
properties_schema, fixed for the life of the project (to change it, delete the project and create it again; its memories start empty). Filters may name anym.<key>: a declared key is indexed where the backend supports it, an undeclared key is evaluated without an index. On Qdrant and Milvus the vector store maps every distinct(namespace, vector dimensions, properties_schema)to one physical collection that all projects declaring that schema share, created the first time a project declares it and never dropped by this code. One tenant's choices therefore change what other tenants get. Verified 2026-09-15 against the vendor docs and, for the Qdrant planner, Qdrant'sdevsource (read_view/filtering.rs,read_view/dispatch.rs,sample_estimation.rs).max_payload_index_count100, rejection of unindexed filters) do not apply; collections created by an earlier MemMachine keep their own setting.is_tenantkeyword index, which is the planner's primary clause, so the candidates are the project's own points, never the physical collection. An undeclared condition estimates as unknown (min 0), so the planner brute-forces the project's posting list with one payload read per point when the project's count in a segment is belowfull_scan_threshold(10,000 KB / vector bytes: ~1,700 points at 1536 dims), and otherwise samples 1,000 points to choose between that brute force and HNSW traversal with a payload read per visited node, where the missingpayload_mlinks for the undeclared value fragment traversal under a restrictive condition (fewer results). Payload is in thecoldtier (disk) by default. The cost lands on the querying project; other tenants feel it only as load on the shared node.$meta, 65,536 bytes per row, all keys together) and again into thepropertiesJSON field, and the filter, tenant term included, is evaluated per row over the project's hash partition (1 of 16 by default), which holds every tenant hashing there. Milvus allows 64 fields and 1,024 partitions per collection, VARCHAR <= 65,535.sqlite_vectakes thelimitnearest neighbors and applies the filter afterwards (k = min(limit, 4096)), so a selective filter returns fewer results thanlimit, declared or not; the search-engine store evaluates the filter per visited candidate with one SQL statement each and indexes no key.Stack
Slice 8 of 9, every PR targeting
speedkick; merge bottom-up.This PR's own change is its commits
ff8bcd62aand94cd5dfb5; the rest of its diff is the slices under it, and drops out as they merge. Stacked on slice 7; slice 9 is stacked on it.Verification
ruff check,ruff format --check,ty check,pytest packages/server/server_tests(integration deselected): server 1931 passed at ff8bcd6 and 1932 passed at 94cd5df, 3 skipped; Qdrant integration lane (REST + gRPC) 140 passed; run at this slice's own commit on 2026-09-15.🤖 Generated with Claude Code