Skip to content

[vector store scale-out 4/6] Move the Qdrant store onto the collection registry - #1735

Merged
edwinyyyu merged 24 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-qdrant-registry-speedkick
Oct 6, 2026
Merged

edwinyyyu merged 24 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-qdrant-registry-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Summary

The Qdrant store moves onto the collection registry, so any number of server processes sharing its registry may serve its collections.

  • Catalog: the collection registry in the relational database QdrantConf.collection_registry names (required). The __registry collections, registry_replication_factor and the process-local locks go. The database manager starts the registry before it opens the client.

  • Incarnations:

    • Every point carries its collection's incarnation in sys-incarnation, under the tenant index the name had, and every search filters on it.
    • A point's id is uuid5(incarnation, record UUID), with the record UUID kept in the payload, so collections sharing a native collection never share a point id.
  • Purge:

    • A round deletes an incarnation's points with one filter-delete, about 1.3 s per 1M points on Qdrant 1.19.1. It stalls writes and scrolls on the shard, but not searches.
    • Deleting in batches collapsed once a vacuum rebuilt the segments mid-purge.
    • QdrantConf.tombstone_retention_seconds (a day by default) is the retention. The configuration refuses one below 10 x request_timeout_seconds + 300 s.
  • Requests: every request is bounded by QdrantConf.request_timeout_seconds (default 30).

    • An upsert refused with a 400 or a 413 is halved until it fits. Any other error, a timeout included, raises at once.
    • Upserts and deletes wait for Qdrant to apply them. Not waiting saved about 1 ms at the median and much of the write tail, but changed neither searches, CPU nor what Qdrant sustains; the measurements are in the design document.
  • Score threshold: a match scoring exactly score_threshold is returned, as the contract now states and the SQLite and Milvus stores already do; both SQLite stores now have the boundary test too. Qdrant keeps only scores strictly better than its threshold, so the store sends it the adjacent single-precision value on the worse side and checks the caller's threshold on the returned scores. Sending the threshold, rather than only checking it in Python, keeps Qdrant from returning matches it cuts: up to 0.9 ms (27%) per query at limit 100 over gRPC.

  • Registry engines: the collection registry refuses an engine whose pool shares one connection (StaticPool, which aiosqlite uses for :memory:) or an in-memory SQLite database, as the segment store and the SQLite vector stores do. It lands here, after the wiring tests move to file databases.

  • Tests: every Qdrant test runs against a real Qdrant over REST and gRPC; the Python local mode, which behaves differently, is gone. The collection lifecycle contract (collection_lifecycle_contract.py) is mixed into the Qdrant tests. It covers:

    • stale handles and empty re-creation;
    • lost creation races and failed preparations;
    • idempotent deletion, the purge, and a write landing under a dead incarnation;
    • the registry lookups each operation makes;
    • identifiers with a trailing newline.

    The Qdrant suite adds filtered-query and namespace isolation; purge rounds on never-created storage and on a write landing after a round; a point refused on its own; ranking, scores and the threshold boundary for every metric; a seeded operation sequence against a model; and churn across two stores sharing one registry, with purgers running. The contract's drain is bounded, and a stale upsert is checked by what it stores.

    The integration tests run Qdrant 1.19.1, which the store's design was measured against.

Breaking, with no migration:

  • Existing Qdrant data is not carried over: drop the Qdrant collections before upgrading. The native collections keep their names, so points left in one stay there, invisible and never purged, and dropping the collection after the upgrade drops the new data too.
  • Every Qdrant configuration must name a collection_registry. The samples, the Helm chart (db_postgres), the configuration wizard and the configuration docs name one.

Commits

  1. Test the Qdrant store against Qdrant 1.19.1.
  2. Bound every Qdrant request by a configured timeout.
  3. Move the Qdrant store onto the collection registry.
  4. Document how the Qdrant store meets the shared contracts.
  5. Build Qdrant handles from live registrations.
  6. Use the registry's reservations and registrations in the Qdrant store and the contract.
  7. Drop the lifecycle contract's stale check of a query with limit 0.
  8. Say to drop the Qdrant collections before upgrading, and list the registry.
  9. Run every Qdrant store test against a real Qdrant server.
  10. Give the wiring tests' collection registry a file database.
  11. Refuse engines that share one connection or hold SQLite in memory.
  12. Pin the Qdrant client's timeout wiring with a non-default value.
  13. Bound the lifecycle contract's drain and settle before empty checks.
  14. Check a stale upsert by what it stores, not by its registry reads.
  15. Test that filtered queries and namespaces keep Qdrant collections apart.
  16. Test Qdrant purge rounds on missing storage and on a write after a round.
  17. Test that a Qdrant upsert raises when a point is refused on its own.
  18. Test ranking, scores and threshold direction for every similarity metric.
  19. Test a seeded Qdrant operation sequence against a model.
  20. Test concurrent churn across two Qdrant stores sharing one registry.
  21. Make the Qdrant creation race tests race, and drop a dead index check.
  22. Keep a Qdrant match that scores exactly the score threshold.
  23. Test that the SQLite stores keep a match scoring exactly the threshold.
  24. Use serial commas in the Qdrant document, and shorten the threshold comment.

Stack

19 open PRs: one independent PR, and the vector store tree of short parallel branches. #1735's and #1736's GitHub base is feat/horizontal-scaling and the others' is main, since the branches are in a fork and a pull request can target only this repository's branches; the on column gives the order the PRs build on each other instead. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1670 and #1702 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (merged into feat/horizontal-scaling) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs feat/horizontal-scaling
[vector store scale-out 3/6] #1734 (merged into feat/horizontal-scaling) Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge feat/horizontal-scaling
[vector store scale-out 4/6] #1735 (this PR) Move the Qdrant store onto the collection registry feat/horizontal-scaling
[vector store scale-out 5/6] #1736 Move the Milvus store onto the collection registry, against a Milvus server #1735
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1670 Remove per-project filterable properties (port of #1606) #1736
[user properties 2/2] #1702 Keep user properties out of the vector store #1670
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1702
[session storage 1/2] #1622 Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1627
[sqlite store fixes 1/7] #1460 Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1663
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1628

This PR is its 24 commits, 01588deb8, 26cc107e0, 33740e0f1, 30c34777f, 9096edddf, 87edbad38, e63300bc3, a40bc5013, 30c30f4ea, 3d7761ead, 81b1bf51b, e9e71b512, 6c2362b49, 4d3241c55, 836db2115, 252bcd252, 5e8c0490f, ef8d3aba6, 276a034d1, 4fb5bd421, 2aabd3d5c, ccba117fe, 6bbc1311b, 106e1067a, directly on feat/horizontal-scaling. #1736 is stacked on it. Split from #1631 (closed).

Verification

Rebased onto feat/horizontal-scaling after #1734 merged there as 5a846570a: every commit's tree is unchanged, so the results below hold.

Commits 9-24, rebased with the stack onto feat/horizontal-scaling: at the head, ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2076); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (586 passed, 6 skipped). Each new or changed test failed under a named mutation of the code, listed in its commit message. Against Qdrant 1.19.1: a scroll or filter-delete on a missing collection raises 404 over REST and NOT_FOUND over gRPC, and creating an existing collection raises 409 or ALREADY_EXISTS, even when creates race; scores and threshold direction match the contract on all four metrics.

After the fourth review round of #1733 and #1734: the commits are rebased onto #1734's head 70e3bb82b, patch-identical to the ones verified at 0c4f87037, with commit 3's conflict in qdrant_vector_store.py resolving to the commit's own file. At the head: ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2135); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (534 passed, 6 skipped).

After the third review round of #1734: commits 1-7 are patch-identical to the ones verified below, rebased onto #1734's head 8d26e27f6, except commit 3, whose conflict in qdrant_vector_store.py resolves to the commit's own rewritten file. Commit 8 adds documentation only. At commit 8, as verified before #1734's last test-only change: ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2128); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (534 passed, 6 skipped).

After the second review round of #1733 and #1734: commits 1-6 are patch-identical to the ones verified below, rebased onto #1734's head 34d197eb1, except commit 3. That commit also carries #1733's new Qdrant test of a limit that is not positive, placed in TestFilters beside the score-threshold test, where the rewritten test file keeps it. At commits 3-6, the lifecycle contract expects a stale handle's query with limit 0 to raise the stale error, but #1734's handle now refuses the limit first, with a ValueError; commit 7 drops that check. At commit 7: ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2125); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (534 passed, 6 skipped).

Rebased onto #1734's registrations. Commits 1-4 are patch-identical to the ones verified below. Commits 3 and 4 predate the collection registry's registration API, so on this base they do not type-check until commit 5, which adapts the Qdrant handle, the collection lifecycle contract and the Qdrant tests. At commit 5: ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2111); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (529 passed, 6 skipped).

After #1734's rename: commits 1-5 are patch-identical to the ones above, rebased onto #1734's twelfth commit, and do not type-check there until commit 6, which renames the Qdrant handle's and the lifecycle contract's uses. At commit 6: ruff check, ruff format --check and ty check are clean; the full server suite without integration tests passes (2111); in test containers, the integration tests pass against PostgreSQL and Qdrant 1.19.1 (529 passed, 6 skipped).

Before the rebase:

At each commit:

  • ruff check, ruff format --check and ty check are clean, run as CI runs them.
  • Commit 1 changes only the test container's image; the integration run at the head covers it.
  • The full server suite without integration tests passes at commit 2 (2083 tests) and at commit 3 (2109 tests).
  • Commit 4 adds the design document only.

At the head, in test containers: the integration tests pass against PostgreSQL and Qdrant 1.19.1 (528 passed, 6 skipped). They cover the vector store (the Qdrant store with the collection lifecycle contract, over local, REST and gRPC clients), the vector-store semantic storage, event memory, long-term memory and the resource manager. The 6 skipped are three Neo4j-specific semantic storage cases, which skip on the other backends, and three long-term memory tests, for want of a NebulaGraph server.

🤖 Generated with Claude Code

This was referenced Oct 1, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/vector-store-qdrant-registry-speedkick branch from 093c0dd to db6d46c Compare October 1, 2026 22:30
@edwinyyyu edwinyyyu added the horizontal scaling Wrong or unsafe when more than one server process serves the same backends (replicas or workers) label Oct 1, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/vector-store-qdrant-registry-speedkick branch 2 times, most recently from dcc108b to d823255 Compare October 2, 2026 03:19
@edwinyyyu
edwinyyyu marked this pull request as ready for review October 2, 2026 17:08

@marvinyu-memverge marvinyu-memverge left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Edwin. Reviewed at d823255 (the 7 commits on top of #1734). The subclass meets the registry's contract cleanly, and the uuid5 point ids are a nice way to keep incarnations sharing a native collection apart. One ask, one optional nit.

1. "Existing Qdrant data is orphaned" understates what an operator will find after upgrading.

The native collection name is unchanged since v0.3.9 ({namespace}__{sha256 of the config}, and the config serializes the same way), and _prepare_storage accepts a collection that already exists. So wherever a collection's config is unchanged, the new incarnation writes into the same native collection that holds the old points. The old points are invisible (every search filters on sys-incarnation) and never purged (no tombstone names them). They also can't be reclaimed by dropping the old collections after the upgrade, because that drops the new data with them.

For the event backend this means each pre-upgrade session gets a new, empty collection, and its episodes stop being recalled while SQL still holds them.

Could the breaking note say how to clean up? Either drop the Qdrant collections before upgrading, or afterwards delete the points that have no sys-incarnation.

Nit (optional): the config summary in deployments/helm/README.md (event_vector_store: { provider: qdrant, config: { host, port, grpc_port, prefer_grpc, https, api_key ... } }) doesn't list collection_registry. The configmap template does.

Verified:

  • The incarnation filter is ANDed with the caller's filter on the only search path, and a negated filter stays inside it.
  • Point ids are uuid5(incarnation, record_uuid), so they can't collide with each other across incarnations or with the old uuid4 ids.
  • A purge round deletes only the dead incarnation's points, and treats a missing native collection as nothing left.
  • Upsert, delete and the purge delete all pass wait=True, which matches "durable once it returns".
  • An upsert halves only on a 400 or 413, and a timeout now raises instead of halving.
  • Every config surface names a registry (the Helm configmap, the samples, the wizard, the docs), and a missing one fails at config parse.

edwinyyyu and others added 10 commits October 6, 2026 14:09
Two collections of one namespace and configuration share a native
collection, and the existing isolation tests either queried without a
filter or used a filter only the querying collection's records met, so
a search whose property filter displaced the incarnation filter, or sat
beside it under `should`, passed them all.

- A filtered query returns only its own collection's records: both
  collections hold records meeting every filter (a comparison, In, Or,
  Not(IsNull) on an undeclared property, And), and each query must
  return exactly its own. Replacing the incarnation filter with the
  property filter, or putting both under `should`, fails it.
- Collections of two namespaces keep their records in separate storage,
  as VectorStore states: an upsert into one namespace's collection adds
  to that namespace's storage and leaves the other's count unchanged,
  under one name and configuration. Building the native collection's
  name without the namespace fails it.

The lifecycle contract's count_stored hook now uses the module's
_count_stored helper, which the namespace test also reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
- A tombstone whose storage was never made is retired: a creation
  whose storage preparation fails before touching Qdrant cancels its
  reservation, queuing a tombstone for a native collection that does
  not exist. Its round must find nothing and retire it, raising
  nothing: the first purge call runs a round and the next finds none
  due. On Qdrant 1.19.1 the round's scroll of the missing collection
  raises UnexpectedResponse 404 over REST and AioRpcError NOT_FOUND
  over gRPC (a filter-delete raises the same). Removing the not-found
  handling around the scroll fails it on both transports; classifying
  only REST's 400, or gRPC's INVALID_ARGUMENT, as not found fails it
  on that transport.
- A write landing after a purge round is reclaimed by the next: after
  a round deletes a deleted collection's records, a write through the
  stale handle's backend path, past its liveness checks, lands under
  the dead incarnation, as a write in flight across the deletion does.
  Draining must leave nothing under the incarnation. A round that
  reports nothing found after deleting retires the tombstone on its
  first round and leaves the late write in place, failing it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
An upsert refused with a 400 or 413 is halved until its halves fit or a
single point is refused, and the store's docstring says the single
point's refusal raises. The halving test refused only batches over a
size, so every point was eventually accepted. The new test refuses one
point whenever it is sent: the upsert must raise the refusal, with its
status, rather than return as if the point were accepted. Returning,
instead of raising, once halving reaches a single point fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The suite queried cosine collections only, so a metric mapped to the
wrong Qdrant distance, or a threshold kept on the wrong side for a
distance, went unseen. For each of COSINE, DOT, EUCLIDEAN and
MANHATTAN, on a real Qdrant over REST and gRPC, five vectors that the
four metrics rank in four different orders, with no ties, are queried
against one vector:

- matches come best first, highest for a similarity and lowest for a
  distance;
- each score is what QueryMatch defines: cosine similarity, dot
  product, Euclidean distance or Manhattan distance;
- a threshold halfway between the second and third best scores keeps
  exactly the best two: scores above it for a similarity, below it for
  a distance.

Mapping EUCLIDEAN to Qdrant's Manhattan distance or back, COSINE to
dot product or back, dropping the threshold, negating it for distances,
or applying it as a similarity for every metric fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Two collections of one namespace and configuration, so sharing a native
collection, take 200 seeded steps over a pool of 12 sorted uuid4s:
upserts and re-upserts that change or drop properties (one of them
undeclared), deletes of a collection's own UUIDs, the other
collection's and absent ones, filtered queries (comparisons, In, IsNull,
Not, And, Or) with a limit of the whole pool, deletion and re-creation
under the same name, and purge drains. After every step, each
collection's stored record UUIDs, read past the store, and an
unfiltered query's matches (each record once, scored by its latest
vector, best first) agree with a model of each collection's records; a
query step's filtered matches agree with the model's filter; a drain
leaves exactly the model's records in the native collection.

Merging an upsert into the stored point's payload, so that a dropped or
changed property keeps its old value, fails it; so does deriving point
ids from the record UUID alone, searching without the incarnation
filter, and a purge round that never deletes what it finds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Two stores, one on a REST and one on a gRPC client, each with its own
registry object and engine on one registry database (a SQLite file, and
PostgreSQL), serve eight seeded workers. Each worker owns four record
UUIDs, so its records in each collection life follow from its own
operations, and takes 50 steps over three collection names: upserts and
deletes of its records, queries filtered on its own records, and
deletion of a collection, half the time followed by its creation. A
purger per store drains whenever a collection is deleted.

- Every operation returns or raises a documented domain error (stale
  handle, already exists, pending, deleted, attempts exhausted); any
  backend or database error fails the test, and a deadline turns a
  deadlock into a failure.
- A query whose collection was live throughout it, which a delete of
  nothing confirms afterwards, returns exactly its worker's records;
  any query returns none of the worker's other lives' records.
- After quiescing and draining, each live collection's stored records
  and an unfiltered query equal what its workers wrote and kept, and a
  deleted life keeps no record but one whose upsert raised as stale: a
  write in flight across a deletion can land after a round found the
  incarnation empty, the race the tombstone retention closes, and the
  test runs with no retention.

Searching without the incarnation filter, deriving point ids from the
record UUID alone, misclassifying REST's 409 on an existing native
collection, a purge round that never deletes, and a purge round that
never returns each fail it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The two-workers test checked that the native collection ended up
indexed, but only the reservation's winner prepares storage, on a
native collection no one else touches, so that check could not fail
the way its docstring said; it is gone, and the docstring says what the
test does. Its creators now reserve the name together, behind a
barrier, so the loser always loses at the reservation rather than
finding the winner already pending or live. It also writes a record
through one handle and reads it through the other, so agreeing on one
collection is shown by state. An open-or-create that raises on losing
the reservation, and a strict create that treats a taken name as
created, each fail it.

A new test races what the old one meant to: two workers create two
collections of one namespace and configuration at once, both
preparations held at a barrier until both start, so both create the
shared native collection together and one finds it already there. Both
creations must succeed and each collection must hold its own records,
over REST and over gRPC. Checking whether the native collection exists
before creating it, without guarding the create, fails it on both
transports while every sequential test passes; so does misclassifying
the transport's already-exists answer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Qdrant rounds a query's score_threshold to single precision and keeps only
scores strictly better than it, on every metric and transport, so a match
scoring exactly the threshold was dropped, even a perfect cosine match under
a threshold of 1.0. The SQLite stores and Milvus keep it, and the contract
did not say which.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the adjacent single-precision value on the worse side
(numpy's nextafter in float32), so Qdrant keeps a score equal to the
threshold, and checks the caller's threshold on the returned scores, so a
score Qdrant's rounding let through but that falls short of the threshold as
returned is dropped. A threshold beyond single precision is not sent; the
store's check applies it.

Sending the threshold, rather than only checking it in Python, keeps Qdrant
from returning matches the threshold cuts. Measured against Qdrant 1.19.1
(4 collections of 5,000 768-dimension points sharing one native collection,
300 queries per cell with the two variants alternated, a threshold passing a
quarter of the limit; AC power), p50 per query, threshold sent vs checked in
Python only: REST 2.89 vs 3.04 ms at limit 10, 2.85 vs 2.96 at 40, 3.13 vs
3.42 at 100; gRPC 2.86 vs 2.93 ms at 10, 3.36 vs 3.80 at 40, 3.33 vs 4.23 at
100. Both returned identical matches.

The new test, on every metric and on REST and gRPC, sets the threshold to a
match's reported score (the match is kept) and one step better (it is
dropped). It fails with the threshold sent unchanged, stepped the wrong way,
without the store's check, or with that check strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The contract says a match scoring exactly the score threshold is returned.
Both SQLite stores already keep it; each now has the test the Qdrant store
has, on every metric it supports (cosine, dot and Euclidean on the USearch
store; cosine and Euclidean on sqlite-vec): a threshold equal to a match's
reported score keeps the match, and one a step better drops it. Each fails
with its store's threshold check made strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@edwinyyyu
edwinyyyu force-pushed the feat/vector-store-qdrant-registry-speedkick branch from f9cae33 to 106e106 Compare October 6, 2026 21:10
@edwinyyyu
edwinyyyu merged commit f4cb7bd into MemMachine:feat/horizontal-scaling Oct 6, 2026
41 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…n registry (MemMachine#1735)

* Test the Qdrant store against Qdrant 1.19.1

The integration tests ran Qdrant 1.17.0. The Qdrant store's design,
which follows in this stack, was measured against 1.19.1: its one-shot
purge of an incarnation by filter, and the per-tenant index layout. 1.17.0
predates the filter-resolution fence that 1.19.0 added to filter deletes
(qdrant#9678), whose extra cost 1.19.1 no longer shows, so the tests
could not see the behavior the store is tuned for.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound every Qdrant request by a configured timeout

The Qdrant client was built with qdrant-client's default timeout, so how
long a write to Qdrant can be in flight was nothing the configuration
stated. The purge that follows in this stack waits out a retention longer
than any write can be in flight, and the request timeout is the part of
that time the store controls.

`QdrantConf.request_timeout_seconds`, a positive whole number of seconds
defaulting to 30, is passed to the client. The sample configurations and
the configuration docs show it. The tests' Qdrant clients are built by
one fixture with the configuration's default, so they run with the
timeout a configured store has.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Move the Qdrant store onto the collection registry

The Qdrant store kept its catalog in a `__registry` collection per
namespace and serialized creating and deleting a collection with locks in
one process. A collection's name was the tenant discriminator on its
points, so a handle kept writing into a collection deleted and created
again under its name (MemMachine#1563), and a write in flight during a deletion
outlived it.

`QdrantVectorStore` is now a `RegistryBackedVectorStore`, and any number
of processes sharing its registry may serve its collections:
- Its catalog is the collection registry in the relational database
  `QdrantConf.collection_registry` names. The `__registry` collections,
  `registry_replication_factor` and the process-local locks go. The
  database manager builds and starts the registry before it opens the
  client, so a registry database it cannot resolve leaves no client open.
- Every point carries its collection's incarnation in `sys-incarnation`,
  in place of the name, under the same tenant index, and every search
  filters on it.
- A point's id is `uuid5(incarnation, record UUID)`, and the record UUID
  is kept in the payload as `sys-record_uuid`, which a search returns.
  The logical collections sharing a native collection share its id space:
  with the record UUID as the id, an upsert of a UUID another collection
  held replaced that collection's point.
- Preparing a collection's storage creates its native collection and
  payload indexes, each under its own already-exists guard, so a creation
  that failed part way is completed by the next.
- A purge round looks for one point under the incarnation and, finding
  one, deletes the incarnation's points with one filter-delete.
  `QdrantConf.tombstone_retention_seconds`, a day by default, is the
  retention, and the configuration refuses one below 10 x
  `request_timeout_seconds` + 300 seconds.
- An upsert is halved only when Qdrant or a proxy refuses it as sent,
  with a 400 or a 413. Any other error raises at once: a timed-out upsert
  may still be applied, and sending it again adds load to a server
  already too slow.

The collection lifecycle contract (`collection_lifecycle_contract.py`),
which a store's tests mix in with hooks that read the backend directly,
runs on Qdrant: stale handles, a collection created again starting empty,
creation races and their outcomes, failed preparations, the purge, a
write landing under a dead incarnation, and the registry lookups each
operation makes. The store's own tests check what Qdrant holds by
scrolling the incarnation past the store.

Breaking: existing Qdrant data is orphaned, since its points carry names
and its catalog is in the `__registry` collections; no migration is
included. Every Qdrant store needs a relational database for its
registry, which the samples, the Helm chart, the configuration wizard
and the configuration docs name.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Document how the Qdrant store meets the shared contracts

`design/qdrant_vector_store.md` records the Qdrant store's layout, its
derived point ids and why they are one-way, filtered-search correctness,
the purge by filter, and its consistency on one node and replicated,
with the measurements behind each choice.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Build Qdrant handles from live registrations

The collection registry now answers registrations: a PendingRegistration
from register, which its creator marks live, and a LiveRegistration from
resolve and mark_live, whose require_current is a handle's fence. The
Qdrant handle takes its live registration in place of the namespace,
name, incarnation, configuration and registry lookup, and
_build_collection_handle builds it from one.

The collection lifecycle contract follows:
- the tests that fail a check or count checks patch the registration
  type's require_current, where they replaced the handle's lookup;
- the racing winners register and mark live through a pending
  registration, and the winner is found with resolve;
- churn counts VectorStoreCollectionDeletedError, which a creation
  undone by a concurrent deletion now raises, among the domain's
  outcomes.
The Qdrant tests that build a handle on a mocked client give it a live
registration that stays current.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use the registry's reservations and registrations in the Qdrant store and the contract

MemMachine#1734 names the registry's handles for their holders: `reserve` answers
a Reservation, whose `confirm` answers a Registration, and whose `cancel`
gives the name back. The Qdrant handle takes a Registration, and the
collection lifecycle contract's racing winners reserve, then confirm.
Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the lifecycle contract's stale check of a query with limit 0

A limit at or below zero is now refused as invalid input, checked
before the handle's liveness, so it no longer stands for a query with
nothing to send to the backend; the query with no vectors still does.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say to drop the Qdrant collections before upgrading, and list the registry

The user docs did not say what happens to existing Qdrant data. The
store keeps its native collections' names, so points left in one stay,
invisible and never purged, and dropping the collection after the
upgrade drops the new data too; databases.mdx now says to drop the
collections before upgrading. The Helm README's configuration summary
also lists collection_registry, which the configmap template sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Run every Qdrant store test against a real Qdrant server

Qdrant's local mode answers differently from a server: it ignores
payload indexes and raises its own exceptions where a server answers
404 or 409 over REST or NOT_FOUND or ALREADY_EXISTS over gRPC. The
store fixture's client is now a testcontainers Qdrant over REST or
gRPC, marked integration, so every store test exercises what a
deployment runs. The tests that build a handle on a mocked client stay
in the default suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Give the wiring tests' collection registry a file database

The Qdrant wiring tests gave the collection registry an in-memory
SQLite database, which aiosqlite serves from one shared connection;
the segment store and the SQLite vector stores refuse such an engine,
and the registry is to do the same. Each test now puts the registry in
a SQLite file under its own tmp_path. What each test asserts is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse engines that share one connection or hold SQLite in memory

The registry's arbitration rests on each operation running in a transaction
of its own: a reservation's insert and queue check, a deletion's row and
tombstone, a purge claim. On an engine whose pool is StaticPool, every
session shares one connection, so concurrent operations would run inside one
another's transactions. On in-memory SQLite each connection gets a separate
database, so the registry's state would not be shared even within one
process. A SQLite configuration whose path is ":memory:" produces the former
(aiosqlite uses StaticPool for it). The segment store and both SQLite vector
stores already refuse both; the registry now does too, with the same
messages.

The dialect and SQLite-version refusal tests now use file engines, so each
checks only its own refusal. The new tests fail with either check removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Pin the Qdrant client's timeout wiring with a non-default value

The wiring test asserted that the client got a 30 s timeout, which is
also QdrantConf's default, so a manager that ignored
request_timeout_seconds and passed 30 passed it. It now configures 7 s
and asserts the client is built with 7; hardcoding 30, or dropping the
timeout from the client's arguments, fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound the lifecycle contract's drain and settle before empty checks

The contract drained purges with an unbounded while loop, so a store
whose purge round kept reporting records hung the suite. The drain now
fails the test once 1,000 rounds have each found something, far more
than the few deleted collections a contract test leaves; a Qdrant
round that always reports records now fails the purge tests instead of
hanging them.

The two tests that check that a collection created again under a
deleted one's name starts empty now call the store's settle hook first.
On a store whose reads lag its writes, a query could miss the old
life's record for that reason alone, and the check would pass even if
the new life could reach it; after settling it fails only when the new
life cannot. Dropping the incarnation filter from Qdrant's search
fails both.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check a stale upsert by what it stores, not by its registry reads

The contract counted require_current calls: two per upsert, one per
query or delete. That pinned how a handle checks, not what the check
guarantees, and a store that checked differently but kept the
guarantee failed it. The test that replaces it pins the state: an
upsert through a handle whose collection is already deleted raises
StaleError, and the dead life's records, read past the store, are
exactly those it held before. Removing the handle's liveness check
before its write lets the record land under the dead incarnation and
fails it.

The rest of what the counts stood for is pinned by state elsewhere in
the contract: a stale query and a stale delete raise, and an upsert
whose collection is deleted between its check and its write raises.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that filtered queries and namespaces keep Qdrant collections apart

Two collections of one namespace and configuration share a native
collection, and the existing isolation tests either queried without a
filter or used a filter only the querying collection's records met, so
a search whose property filter displaced the incarnation filter, or sat
beside it under `should`, passed them all.

- A filtered query returns only its own collection's records: both
  collections hold records meeting every filter (a comparison, In, Or,
  Not(IsNull) on an undeclared property, And), and each query must
  return exactly its own. Replacing the incarnation filter with the
  property filter, or putting both under `should`, fails it.
- Collections of two namespaces keep their records in separate storage,
  as VectorStore states: an upsert into one namespace's collection adds
  to that namespace's storage and leaves the other's count unchanged,
  under one name and configuration. Building the native collection's
  name without the namespace fails it.

The lifecycle contract's count_stored hook now uses the module's
_count_stored helper, which the namespace test also reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test Qdrant purge rounds on missing storage and on a write after a round

- A tombstone whose storage was never made is retired: a creation
  whose storage preparation fails before touching Qdrant cancels its
  reservation, queuing a tombstone for a native collection that does
  not exist. Its round must find nothing and retire it, raising
  nothing: the first purge call runs a round and the next finds none
  due. On Qdrant 1.19.1 the round's scroll of the missing collection
  raises UnexpectedResponse 404 over REST and AioRpcError NOT_FOUND
  over gRPC (a filter-delete raises the same). Removing the not-found
  handling around the scroll fails it on both transports; classifying
  only REST's 400, or gRPC's INVALID_ARGUMENT, as not found fails it
  on that transport.
- A write landing after a purge round is reclaimed by the next: after
  a round deletes a deleted collection's records, a write through the
  stale handle's backend path, past its liveness checks, lands under
  the dead incarnation, as a write in flight across the deletion does.
  Draining must leave nothing under the incarnation. A round that
  reports nothing found after deleting retires the tombstone on its
  first round and leaves the late write in place, failing it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that a Qdrant upsert raises when a point is refused on its own

An upsert refused with a 400 or 413 is halved until its halves fit or a
single point is refused, and the store's docstring says the single
point's refusal raises. The halving test refused only batches over a
size, so every point was eventually accepted. The new test refuses one
point whenever it is sent: the upsert must raise the refusal, with its
status, rather than return as if the point were accepted. Returning,
instead of raising, once halving reaches a single point fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test ranking, scores and threshold direction for every similarity metric

The suite queried cosine collections only, so a metric mapped to the
wrong Qdrant distance, or a threshold kept on the wrong side for a
distance, went unseen. For each of COSINE, DOT, EUCLIDEAN and
MANHATTAN, on a real Qdrant over REST and gRPC, five vectors that the
four metrics rank in four different orders, with no ties, are queried
against one vector:

- matches come best first, highest for a similarity and lowest for a
  distance;
- each score is what QueryMatch defines: cosine similarity, dot
  product, Euclidean distance or Manhattan distance;
- a threshold halfway between the second and third best scores keeps
  exactly the best two: scores above it for a similarity, below it for
  a distance.

Mapping EUCLIDEAN to Qdrant's Manhattan distance or back, COSINE to
dot product or back, dropping the threshold, negating it for distances,
or applying it as a similarity for every metric fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test a seeded Qdrant operation sequence against a model

Two collections of one namespace and configuration, so sharing a native
collection, take 200 seeded steps over a pool of 12 sorted uuid4s:
upserts and re-upserts that change or drop properties (one of them
undeclared), deletes of a collection's own UUIDs, the other
collection's and absent ones, filtered queries (comparisons, In, IsNull,
Not, And, Or) with a limit of the whole pool, deletion and re-creation
under the same name, and purge drains. After every step, each
collection's stored record UUIDs, read past the store, and an
unfiltered query's matches (each record once, scored by its latest
vector, best first) agree with a model of each collection's records; a
query step's filtered matches agree with the model's filter; a drain
leaves exactly the model's records in the native collection.

Merging an upsert into the stored point's payload, so that a dropped or
changed property keeps its old value, fails it; so does deriving point
ids from the record UUID alone, searching without the incarnation
filter, and a purge round that never deletes what it finds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test concurrent churn across two Qdrant stores sharing one registry

Two stores, one on a REST and one on a gRPC client, each with its own
registry object and engine on one registry database (a SQLite file, and
PostgreSQL), serve eight seeded workers. Each worker owns four record
UUIDs, so its records in each collection life follow from its own
operations, and takes 50 steps over three collection names: upserts and
deletes of its records, queries filtered on its own records, and
deletion of a collection, half the time followed by its creation. A
purger per store drains whenever a collection is deleted.

- Every operation returns or raises a documented domain error (stale
  handle, already exists, pending, deleted, attempts exhausted); any
  backend or database error fails the test, and a deadline turns a
  deadlock into a failure.
- A query whose collection was live throughout it, which a delete of
  nothing confirms afterwards, returns exactly its worker's records;
  any query returns none of the worker's other lives' records.
- After quiescing and draining, each live collection's stored records
  and an unfiltered query equal what its workers wrote and kept, and a
  deleted life keeps no record but one whose upsert raised as stale: a
  write in flight across a deletion can land after a round found the
  incarnation empty, the race the tombstone retention closes, and the
  test runs with no retention.

Searching without the incarnation filter, deriving point ids from the
record UUID alone, misclassifying REST's 409 on an existing native
collection, a purge round that never deletes, and a purge round that
never returns each fail it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Make the Qdrant creation race tests race, and drop a dead index check

The two-workers test checked that the native collection ended up
indexed, but only the reservation's winner prepares storage, on a
native collection no one else touches, so that check could not fail
the way its docstring said; it is gone, and the docstring says what the
test does. Its creators now reserve the name together, behind a
barrier, so the loser always loses at the reservation rather than
finding the winner already pending or live. It also writes a record
through one handle and reads it through the other, so agreeing on one
collection is shown by state. An open-or-create that raises on losing
the reservation, and a strict create that treats a taken name as
created, each fail it.

A new test races what the old one meant to: two workers create two
collections of one namespace and configuration at once, both
preparations held at a barrier until both start, so both create the
shared native collection together and one finds it already there. Both
creations must succeed and each collection must hold its own records,
over REST and over gRPC. Checking whether the native collection exists
before creating it, without guarding the create, fails it on both
transports while every sequential test passes; so does misclassifying
the transport's already-exists answer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Keep a Qdrant match that scores exactly the score threshold

Qdrant rounds a query's score_threshold to single precision and keeps only
scores strictly better than it, on every metric and transport, so a match
scoring exactly the threshold was dropped, even a perfect cosine match under
a threshold of 1.0. The SQLite stores and Milvus keep it, and the contract
did not say which.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the adjacent single-precision value on the worse side
(numpy's nextafter in float32), so Qdrant keeps a score equal to the
threshold, and checks the caller's threshold on the returned scores, so a
score Qdrant's rounding let through but that falls short of the threshold as
returned is dropped. A threshold beyond single precision is not sent; the
store's check applies it.

Sending the threshold, rather than only checking it in Python, keeps Qdrant
from returning matches the threshold cuts. Measured against Qdrant 1.19.1
(4 collections of 5,000 768-dimension points sharing one native collection,
300 queries per cell with the two variants alternated, a threshold passing a
quarter of the limit; AC power), p50 per query, threshold sent vs checked in
Python only: REST 2.89 vs 3.04 ms at limit 10, 2.85 vs 2.96 at 40, 3.13 vs
3.42 at 100; gRPC 2.86 vs 2.93 ms at 10, 3.36 vs 3.80 at 40, 3.33 vs 4.23 at
100. Both returned identical matches.

The new test, on every metric and on REST and gRPC, sets the threshold to a
match's reported score (the match is kept) and one step better (it is
dropped). It fails with the threshold sent unchanged, stepped the wrong way,
without the store's check, or with that check strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that the SQLite stores keep a match scoring exactly the threshold

The contract says a match scoring exactly the score threshold is returned.
Both SQLite stores already keep it; each now has the test the Qdrant store
has, on every metric it supports (cosine, dot and Euclidean on the USearch
store; cosine and Euclidean on sqlite-vec): a threshold equal to a match's
reported score keeps the match, and one a step better drops it. Each fails
with its store's threshold check made strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use serial commas in the Qdrant document, and shorten the threshold comment

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…er, live from resolve

The registry is addressed by partition key only. `register(key, schema)`
answers a PendingRegistration, whose `mark_live()` answers the same life
as a LiveRegistration or raises VectorStorePartitionDeletedError when the
partition was deleted meanwhile, and whose `unregister()` abandons that
life alone. `resolve(key)` answers the LiveRegistration of the live
partition, None when there is none, and raises
VectorStorePartitionPendingError, which now carries the pending
partition's schema, when it is pending. A LiveRegistration's
`require_current()` raises VectorStorePartitionHandleStaleError once its
partition is deleted. `get`, `mark_live(incarnation)`,
`unregister_incarnation` and RegisteredPartition are gone, so deletion is
by key or by a pending registration, never through a live one. The
SQLAlchemy registry supplies frozen-dataclass registrations holding the
engine and the vector store name, and a module function runs the
unregistration transaction.

The base handle takes its live registration, fences on
`require_current()`, and `_partition_handle` builds a handle from one.
`create_partition` lets VectorStorePartitionDeletedError propagate, and
the VectorStore interface names it; open-or-create catches it and creates
again, refuses a pending partition of another schema at once from the
pending error's schema, and re-raises the last pending error. A
`get_partition` of a pending partition of another schema still reports the
mismatch first. The Qdrant and Milvus handles take the registration in
place of the key, the incarnation and the registry lookup.

Tests follow: the registry's tests answer registrations and gain one for
`require_current`; the base's tests patch the pending registration type's
`unregister` and expect the deleted error from a creation a deletion
undid; the lifecycle contract patches the live registration type's
`require_current`, registers its racing winners through pending
registrations, and counts the deleted error among churn's outcomes; the
Qdrant and Milvus tests that build a handle on a mocked client give it a
registration that stays current.

The same change as MemMachine#1734's 735c649, f88aaa8, e25be60 and
b2bb2d8 and the handle halves of MemMachine#1735's 9096edd and MemMachine#1736's
5582cf4, for this PR's partition registry. The event-backend locator
change has no counterpart: the locator here opens its partition with
open-or-create, which creates again itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
The query contract said nothing of a match scoring exactly the threshold,
and Qdrant dropped it: the server compares its threshold in single precision
and keeps only scores strictly better than it.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the server the adjacent single-precision value on
the worse side, and applies the caller's threshold exactly itself; a
threshold beyond single precision is not sent. Both SQLite stores already
keep the match, and each now tests it on every metric it supports.

The same changes as MemMachine#1735's ccba117 and 6bbc131, for this PR's stores.
MemMachine#1736's cc29d27, Milvus's test of the same, is ported with the store
tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…cycle contract's drain

MemMachine#1735's and MemMachine#1736's latest tests, for this PR's partitions:

- Every Qdrant store test runs against a Qdrant server, over REST and gRPC:
  local mode ignores payload indexes and raises its own exceptions where a
  server answers not-found or already-exists, so it is no longer a fixture.
  The tests on a mocked client stay in the default suite.
- Both stores gain tests that partitions and stores of different names keep
  their records apart, that purge rounds reclaim a write landing after a
  round and drain two stores' tombstones alone, that ranking, scores and the
  threshold follow every similarity metric, that a point or entity the
  server refuses on its own raises, that concurrent creations and startups
  agree, that two stores churning one registry keep every partition exact,
  and that seeded operation sequences, and on Milvus random filters, agree
  with a model. The Milvus tests also pin its read consistency and its
  settings with values other than their defaults.
- The lifecycle contract's drain fails after a bounded number of rounds, it
  settles before checking that a new life is empty, and it checks a stale
  upsert by what the partition holds rather than by its registry reads.

The same changes as MemMachine#1735's 30c30f4, e9e71b5, 6c2362b, 4d3241c,
836db21, 252bcd2, 5e8c049, ef8d3ab, 276a034, 4fb5bd4 and
2aabd3d, and MemMachine#1736's 509c235, 7f1eb3b, 6e312e0, 6d1ed1e,
25e7948, 6f36877, d43e9e6, 6dfbf59 and cc29d27, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
The base store's purge contract says a round that finds the incarnation's
storage missing returns False, but neither store's round checked: once its
native collection was dropped from outside, every round on it raised,
counted against the tombstone, and dead-lettered it after ten.

A Qdrant round that the server answers not found, over REST or gRPC, and a
Milvus round that finds no native collection, now return False: the
collection is gone with everything in it, so the tombstone is retired. Each
store has a test that drops its native collection and drains the purge.

The same handling as MemMachine#1735's and MemMachine#1736's stores, whose rounds already
checked for a missing native collection; MemMachine#1736's 2575af0 tests it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

horizontal scaling Wrong or unsafe when more than one server process serves the same backends (replicas or workers)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants