Skip to content

[vector store alt 9/9] Bound every request to a remote vector store by a configured timeout (speedkick) - #1644

Closed
edwinyyyu wants to merge 11 commits into
MemMachine:speedkickfrom
edwinyyyu:alt/vector-store-request-timeout-speedkick
Closed

edwinyyyu wants to merge 11 commits into
MemMachine:speedkickfrom
edwinyyyu:alt/vector-store-request-timeout-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Alternative to the [vector store N/13] chain (#1597, #1622..#1616): the two exclude each other. This chain keeps per-project properties_schema (user-defined metadata indexes) and conforms Qdrant to the incarnation lifecycle; SQLite and Milvus follow later. Every PR is a draft.

Purpose of the change

Port of #1630: request_timeout on QdrantConf and MilvusConf, in seconds, passed to the client; required, with no default; the configuration wizard supplies 30 seconds. Sample configurations and the configuration docs show it. A breaking configuration change on speedkick.

Stack

Slice 9 of 9, every PR targeting speedkick; merge bottom-up.

# PR change
1 #1636 Restore per-project filterable properties
2 #1637 Create a session's storage when the session is created, never on a request
3 #1638 Make the scripts and examples that write to a project create it first
4 #1639 Make no memory request create a project
5 #1640 Remove open-or-create and close from both stores
6 #1641 Rename both stores' lookup to get_collection and get_partition
7 #1642 Remove custom sharding from the Qdrant store, and disable strict mode on its collections
8 #1643 Mint an incarnation per collection life on Qdrant; delete logically, reclaim by purge
9 #1644 (this PR) Bound every request to a remote vector store by a configured timeout

This PR's own change is its commit d1bfc5348; the rest of its diff is the slices under it, and drops out as they merge. Stacked on slice 8.

Verification

ruff check, ruff format --check, ty check, pytest packages/server/server_tests (integration deselected): server 1932 passed, 3 skipped; run at this slice's own commit on 2026-09-15.

🤖 Generated with Claude Code

edwinyyyu and others added 11 commits September 15, 2026 12:00
Reverts the second commit of MemMachine#1606 (a8322a7), "Remove per-project
filterable properties", and keeps its first, the OpenAPI regeneration
under the locked FastAPI 0.141.1: docs/openapi.json is regenerated here
with the same tool, so it differs from the pre-MemMachine#1606 document only by the
ValidationError fields that regeneration added.

A project declares `properties_schema`, a set of caller property keys with
types, on its long-term memory configuration; the event backend merges it
into the vector store collection's indexed schema and rejects filters on
any other `m.<key>`. Under this design the schema is fixed for the life of
the project, a schema change is a delete and a create, and the vector
store maps every distinct schema to one physical collection that all
projects declaring it share. What that costs, per backend, is written
where a tenant will see it: a warning in the configuration docs and the
field descriptions on the server configuration, the project API, the
memory-configuration API and the Python SDK.

Alternative to the one-collection chain (MemMachine#1622..MemMachine#1616): the two exclude
each other.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…quest

The event backend created a session's vector store collection and segment
store partition on the first request that opened the session, so a search
or a write for an unknown session created storage as a side effect, and
the service locator was the only place that knew both stores' create paths.

The owner is the session. Every path that creates a session row runs
through EpisodicMemoryManager._create_session, which inserts the row and,
when the row is new, creates the session's partitions in its segment store
and its vector store (create_episodic_memory_storage); an equivalent
re-create accepts the row and leaves the storage as it is. The request
path binds handles with the stores' lookups and raises
SessionPartitionMissingError when a partition is absent: a session without
its storage is broken, not new. Deleting a session with no open instance
deletes its partitions by key, so a session whose storage was never fully
created can still be deleted. MemMachine.create_session goes through the
manager for the same reason.

The session's vector store collection is created with the system keys and
the project's own `properties_schema`, resolved and merged as the request
path did before.

The semantic manager owns its one collection and creates it, once, at the
storage's first use. With that, nothing calls the stores' open-or-create.

The API is unchanged: the manager's open-or-create still creates a session
a memory request names, now through the same path.

Ported from MemMachine#1622 (4266b09) onto the chain that keeps per-project
filterable properties.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The simple chatbot example, the TypeScript REST demo and the Dify plugin's
add-memory tool wrote to a project without creating it, relying on the
write to create it. Each now creates its project before its first memory
request and accepts 409 as the project already existing. No behavior
changes for them; they stop depending on a write creating a project.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Adding memories to, or searching, a project that did not exist created it,
with the server's default configuration, without the caller's knowledge.
Now only the create-project request creates a project: a write or a
search opens the session and answers 404 for an unknown project, as the
search endpoint already promised; the manager's open-or-create goes.

Two callers depended on the implicit creation. `org_id` and `project_id`
default to `universal`, so the API promises the project
`universal/universal`; the server creates it, once, at startup, and leaves
one that already exists as it is. The MCP add tool names its own project
and has no create-project counterpart, so it creates the project it writes
to, once, and says so. The API doc strings and the OpenAPI document say
which requests create a project.

A breaking API change on `speedkick`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Nothing calls them since a session's storage is created with the session:
`open_or_create_collection` and `close_collection` leave the vector store
interface and its four backends, `open_or_create_partition` and
`close_partition` leave the segment store interface and its
implementation, and the two config-mismatch errors that only open-or-create
raised go with them. A store creates on `create_*`, strictly, and looks up
on `open_*`, answering None; create-if-absent is the owner's, where the
key's provenance is known.

Source changes are deletions only. The tests that exercised
open-or-create as a fixture use a test-side create-if-absent instead, and
the tests of its own semantics go.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A lookup that answers None for an absent collection or partition is a
Python get, not an open. Identifiers only, produced by the script below,
plus the two docstring lines that said "opened"; every contract is as it
was.

```sh
set -e
cd "$(git rev-parse --show-toplevel)"
git ls-files -z 'packages/server/*.py' | xargs -0 perl -0pi -e '
  s/\bopen_collection\b/get_collection/g;
  s/\bopen_partition\b/get_partition/g;
'
uv run ruff check --fix --quiet packages/server
uv run ruff format --quiet packages/server
```

The `open -> get` half of MemMachine#1626 (e9783d9); the `collection -> partition`
half is not carried, because its premise is a store that is one
collection.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
`QdrantConf.is_distributed` put every logical collection in its own shard
of the physical collection so that deleting one was a shard drop: an O(1)
deletion in a cluster, at the price of a shard, with its own segments and
graphs, per logical collection.

The lifecycle that follows this change makes a deletion a registry write,
visible at once, and reclaims the points afterward by a filter on the
incarnation; the shard has nothing left to buy. The option goes from the
configuration, the parameters, the database manager and the store, along
with the shard-key selectors on every operation and the tests that ran a
one-node cluster for it. A breaking configuration change on `speedkick`;
the option was never documented.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
A caller may filter on any property, declared in the collection's schema
or not; an undeclared key is evaluated against the payload of the
tenant's own candidate points, without an index. Qdrant Cloud enables
strict mode on new collections by default, and strict mode rejects a
filter on an unindexed payload key, so on such a deployment every
undeclared filter failed.

The store now creates its collections, the physical ones that hold points
and the per-namespace registries, with `strict_mode_config` disabled.
Collections created by an earlier version keep the setting they have.
The chain that indexes every filterable key keeps strict mode on; this
one, which lets a project declare what is indexed, turns it off.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
…reclaim by purge

A logical collection is identified to callers by its (namespace, name) pair
and, inside the Qdrant store, by an incarnation the store mints when the
collection is created. Points carry the incarnation, never the name, and a
point's id is derived from the incarnation and the record uuid, so a
collection deleted and re-created under the same pair starts empty, its
predecessor's points are never adopted by or reclaimed out from under the
successor, and neither a stale handle nor a co-tenant of the shared
physical collection can overwrite a point a live incarnation holds under
the same record uuid. Handles are bound to one incarnation: once it is
deleted, every operation of the handle raises
VectorStoreCollectionHandleStaleError.

delete_collection becomes two registry writes: the collection's live entry
goes, conditionally on the incarnation the delete read, so a creation that
raced in from another process keeps its entry, and an entry in the store's
purge queue follows, stamped with the deletion time and the physical
collection the incarnation lived in. The queue is one collection per
store, `memmachine__purge_queue`, so one sweep serves every namespace. The
new purge_deleted_collections reclaims one dead incarnation per call,
oldest first, by a filter on the incarnation; safe to repeat and to run
from several processes. Qdrant has no transactions: a write that passed
its fence before the deletion lands under the dead incarnation and is
reclaimed with it, and a crash between the two registry writes leaves the
points unreachable and unreclaimed, a leak, never a collection the purge
takes from under a live one.

The VectorStore contract now says what every implementation guarantees
(deletion makes the collection unreachable at once and is idempotent; a
collection created under a deleted pair starts empty) and what it leaves
to the implementation (whether reclamation is immediate or deferred to
purge_deleted_collections, and what a handle held across a deletion does).
The SQLite stores and the Milvus store keep their immediate reclamation,
answer False from purge_deleted_collections, and state their handle
behavior in their docstrings.

Existing Qdrant data is not migrated: the payload key, the point ids, the
registry entries and the purge queue are new.

collection_lifecycle_contract.py holds the tests a conforming store
satisfies, mixed into the Qdrant module and run in local, REST and gRPC
modes: stale handles, empty re-creation, idempotent deletion, purge
reclaiming what deletion deferred and leaving every other collection alone
-- of the same configuration, of another, in another namespace, and one
re-created under another configuration while its predecessor awaits purge
-- and the same record uuid held by two collections of one physical
collection.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The first time a vector store is handed out, the resource manager starts
the same purge loop it runs for segment stores, one per store name;
close() cancels both sets before the stores shut down. A store that
reclaims in delete_collection answers False and costs one call per tick.
Mechanical churn in the same change: the loop takes the store's bound
purge and a label for its failure log line, and its interval and pause
constants lose their SEGMENT_STORE_ prefix.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
The Qdrant and Milvus clients were built without a timeout, so a remote
write could hang a request indefinitely. `request_timeout` on QdrantConf
and MilvusConf, in seconds, is passed to the client; it is required, with
no default, so a deployment states how long it is willing to wait, and the
configuration wizard supplies 30 seconds as the starting point. The sample
configurations and the configuration docs show the option.

A breaking configuration change on `speedkick`. Ported from MemMachine#1630
(d87b9ba) onto the chain that keeps per-project filterable properties.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Reopen if stack chosen.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant