Skip to content

[vector store scale-out 3/6] Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge - #1734

Merged
edwinyyyu merged 38 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-collection-registry-speedkick
Oct 6, 2026
Merged

edwinyyyu merged 38 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-collection-registry-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Summary

This PR adds what lets any number of server processes share a vector store backend: a SQL registry that arbitrates collections, an incarnation per collection life, and a purge of deleted collections' storage. It also adds the base class that the Qdrant and Milvus stores move onto in the next two PRs.

Today the Qdrant and Milvus stores keep their catalogs inside the backend, which has no transactions or unique constraints. Two processes creating one collection could both succeed. A handle in one process kept writing into a collection another process had deleted and created again (#1563). A deletion raced the writes in flight. The VectorStore contract answered this by requiring that one process manage a collection, which no consumer arranged.

  • The contract states a collection's lifetime:

    • A collection created again under a deleted one's name starts empty.
    • A handle to a deleted collection may raise VectorStoreCollectionHandleStaleError.
    • Opening a collection still being created raises VectorStoreCollectionPendingError.
    • delete_collection makes the collection unreachable. A store that reclaims storage later does so in the new purge_deleted_collections.
    • Each store states which processes may share its collections.

    Every store still reclaims at deletion here, so each one's purge returns False.

  • VectorStoreCollectionRegistry, on PostgreSQL or SQLite through SQLAlchemyVectorStoreCollectionRegistry, holds each collection's name, configuration and incarnation: a UUID minted for each life of the collection.

    • Its primary key decides creation races across processes.
    • A collection is registered pending, its storage is prepared, and it is marked live.
    • Deleting it is one transaction that queues a tombstone.
    • A tombstone comes due after a retention longer than any write can be in flight. run_purge_round then claims the oldest due tombstone, calls the store's round with its namespace, configuration and incarnation, and records what the round returns, with backoff on failure and dead-lettering after 10 attempts without progress, a round that finds and deletes records resetting the count.
    • The claim is a lease: one committed write counts the attempt, stamps the tombstone's claimed_at, and bumps its claim_generation; the round runs with no transaction open, and the write that ends the claim is conditioned on the generation. So no lock is held across the round's remote calls (on SQLite, whose write lock covers the whole file, every store sharing the database stays writable), and purgers on every process split a backlog. The lease saves repeated rounds; correctness rests on rounds being safe to repeat.
    • A claim holds until its round ends or purge_lease_seconds (300 by default) has passed, on the database clock. Each claim counts an attempt, so a tombstone whose rounds keep killing their purger dead-letters too; it is claimed again after its lease and the backoff. A cancelled round takes its attempt back.
    • The registry is addressed by name: reserve, resolve, unregister and run_purge_round. What acts on one life of a collection is a handle bound to its incarnation, named for its holder. reserve answers a Reservation, which the creator confirms once the collection's storage is prepared, or cancels while the collection is pending, as a creation does when preparing or confirming raises or is cancelled; once it is confirmed, only a deletion by name ends the collection. resolve and confirm answer a Registration, whose require_current is a handle's liveness fence. A handle's fields never change, and no caller passes an incarnation back to the registry. The design document's "Reservations and registrations" section explains why, including why the names are roles rather than states.
    • The logic follows the segment store's registry (Overhaul segment store: shared tables with incarnation-scoped tenant keys (port of #1548) #1661) wherever it can.
  • RegistryBackedVectorStore makes every registry call a store on the registry makes: create, open-or-create, open, delete and the purge round, each under the store's operation tracker. Its handle's upsert, query and delete check their inputs (a query's limit included), and check the handle's liveness around the backend call. An open-or-create that gives up raises from the race it last lost. A subclass supplies only the storage preparation, the backend calls and a purge round.

  • A sweeper per vector store, started by the resource manager, calls the purge.

  • The event backend's service locator opens the collection a racing creator won, or creates it when the winner's creation was undone. When it gives up, its error carries the last error it retried, so a pending collection's registration time reaches the log.

  • The segment store's insert reports a rejected registration without deciding to retry; its callers, which count the attempts, decide.

  • Every store checks namespaces and names with one helper, require_identifiers, whose message names the rule an identifier breaks.

Reviewing: start with design/vector_store_horizontal_scaling.md, which links the shared documents for the registry, the purge, consistency and isolation. They describe the design the Qdrant and Milvus PRs complete. Their rows for those stores, and their links to the per-store documents, describe those PRs, which add the per-store documents.

Commits

  1. State a collection's lifetime in the vector store contract.
  2. Add a SQL collection registry for vector stores whose backend cannot arbitrate one.
  3. Add RegistryBackedVectorStore, a base for stores whose collections a registry arbitrates.
  4. Run each vector store's purge from the resource manager.
  5. Open the event backend's collection a racing creator won.
  6. Let the segment store's callers decide to mint again, not the insert.
  7. Document the vector store's horizontal scaling design.
  8. Put "pending" before the noun in the registry's docstrings.
  9. State what mark_live does with an incarnation that is not pending.
  10. Raise from mark_live when the collection is no longer pending, not a bool.
  11. Answer registrations from the collection registry: pending from register, live from resolve.
  12. Name the registry's handles for their holders: a reservation and a registration.
  13. Leave a confirmed collection alone when its reservation is cancelled.
  14. Refuse a query limit that is not positive in registry-backed handles.
  15. State how long the purge claim holds its transaction.
  16. Chain the locator's give-up error to the last error it retried.
  17. Never shift the purge backoff by a negative count.
  18. Chain open-or-create's give-up error to the race it last lost.
  19. Drop the limit of 0 from the registry doc's no-op calls.
  20. Report a dead-lettered tombstone at or past the bound.
  21. Run a purge round inside the registry instead of handing out a claim.
  22. Claim a purge round on SQLite with a write, as the segment store does.
  23. Report a failed reservation cancel from the cancel's own task.
  24. Check identifiers with the shared helper in every store.
  25. Track open_collection like the other lifecycle calls.
  26. Say to drop the native collections before upgrading.
  27. Cancel the reservation when confirming a creation fails or is cancelled.
  28. Say which writers wait on SQLite's lock during a purge round.
  29. Wrap the registry doc's creation-cleanup paragraph.
  30. Lease the registry's purge claims instead of holding a transaction across the round.
  31. Name the purge backoff's first delay base_purge_retry_backoff_seconds.
  32. Count a purge round that never ended as failed, and end a cancelled round's claim.
  33. Test the registry's wiring, atomicity, isolation and behavior under concurrency.
  34. Name the purge cutoffs and the consecutive count, and shorten comments.
  35. Split the purge round into named steps, and return a typed tombstone claim.
  36. Count each purge attempt in its claim.
  37. Name the purge count attempts_without_progress, and tighten the purge docstrings.
  38. Replace the purge claim's measured costs with figures from the current claim.

Stack

20 open PRs: one independent PR, and the vector store tree of short parallel branches. #1734–#1736's GitHub base is feat/horizontal-scaling and the others' is main, since the branches are in a fork and a pull request can target only this repository's branches; the on column gives the order the PRs build on each other instead. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1670 and #1702 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (merged into feat/horizontal-scaling) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs feat/horizontal-scaling
[vector store scale-out 3/6] #1734 (this PR) Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge feat/horizontal-scaling
[vector store scale-out 4/6] #1735 Move the Qdrant store onto the collection registry #1734
[vector store scale-out 5/6] #1736 Move the Milvus store onto the collection registry, against a Milvus server #1735
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1670 Remove per-project filterable properties (port of #1606) #1736
[user properties 2/2] #1702 Keep user properties out of the vector store #1670
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1702
[session storage 1/2] #1622 Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1627
[sqlite store fixes 1/7] #1460 Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1663
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1628

This PR is its 38 commits, 4fef4627c, 240fe4782, 5d2239811, 12519520f, 9b991faa9, 2d8fb3b68, 40d5aeada, 735c6495e, f88aaa857, e25be6000, b2bb2d82f, ffa954c7c, 9daf7e7f8, fe2debc0b, 00a7436f7, 3444fade2, 819fb11de, 93a85da87, d0417852b, cd13045c1, 35b57b78e, 0a68f3794, d05556b92, 118d23b16, 5225933db, 783967e6f, f4c585e84, 4dde7c85a, f8d091be2, f1e9de89a, 054b07954, a4e2b24de, 900a9269c, 2a3d87cec, 503687c35, 1732038de, 96a8b5bf6, bab379b0b, directly on feat/horizontal-scaling. #1735 is stacked on it. Split from #1631 (closed).

Verification

At the head, rebased onto feat/horizontal-scaling (where #1733 merged as f24c0a6ba): ruff check, ruff format --check and ty check are clean, run as CI runs them (uv run --frozen --all-extras ty check --project packages/server); the full server suite without integration tests passes (2128); and in test containers the integration tests pass against PostgreSQL and Qdrant (520 passed, 6 skipped). They cover the vector store (the registry on PostgreSQL included), the vector-store semantic storage, event memory, long-term memory and the resource manager. The 6 skipped are three Neo4j-specific semantic storage cases, which skip on the other backends, and three long-term memory tests, for want of a NebulaGraph server.

Commit 38 changes the purge document only: its measured costs are 2026-10-06 figures, each naming the commit whose code was measured.

Commit 37 renames the count attempts_without_progress, after what resets it, and states the base backoff parameter's rule alone, with no change in behavior; the registry tests pass on SQLite and PostgreSQL, and ty check is clean.

Commit 36 counts each attempt in the claim, as job queues count attempts when work is taken, in place of a later claim counting a round that never ended. The claim is one plain UPDATE; consecutive_failed_rounds is consecutive_attempts (attempts_without_progress from commit 37); claimed_at is the open claim's time and last_failed_at when a round last raised. A round that never ended is claimed again after its lease and the backoff, as the next attempt; a retry logs "attempt N of 10", and a raised round's error carries a note naming the tombstone and the attempt. Breaking any of 13 writes or conditions (the count, the backoff from the lease's end, the retry warning, the error's note, the cancel's take-back, a reset, a raise's date or end, the dead-letter report or bound) fails a registry test, on SQLite and PostgreSQL. Its database cost, alternated with commit 35's on 1,000,000 tombstones not yet due (AC power): a claimed round is unchanged (SQLite 1.44-1.46 vs 1.48-1.53 ms, PostgreSQL 2.06-2.59 vs 2.09-2.52 ms), and a claim that reads past tombstones backing off costs at most 4% more at 1,000 and 10,000 of them, and 7% (SQLite) and 3% (PostgreSQL) more at 100,000.

Commits 34 and 35 rename and restructure, with no change in behavior: the queue's failed_rounds is consecutive_failed_rounds, the purge round is split into steps named for what each acts on, and the claim returns a typed result. Breaking the cancellation path, the never-ended-round path, or the elapsed-time check each fails registry tests, and the claim's database cost is unchanged within run-to-run variation (no-op rounds, alternated 8 times, AC power).

Commits 29-33, after the fourth review round:

  • Every new registry test fails, on SQLite and on PostgreSQL, with the behavior it pins removed: the lease, its parameter and its generation fence; the counting of a round that never ended; the cancelled round's release; the base backoff parameter; a deletion and a reservation that each commit as one transaction; namespace scoping; cancel's live guard. Against the claim held across the round (commit 28's), the SQLite writability test fails with "database is locked" and the PostgreSQL test finds a backend idle in a transaction.
  • Measured database cost per round (no-op rounds, 500 per case, versions alternated, AC power, 2026-10-06): for a round that found records, the lease at commit 503687c costs 1.52 ms on SQLite and 2.21 ms on PostgreSQL, against 0.78 ms and 1.05 ms for the transaction-held claim at commit f8d091b, so 0.7-1.2 ms more. Commit 36's counting in the claim matches commit 503687c within run-to-run variation (SQLite 1.40 vs 1.41 ms, PostgreSQL 1.80 vs 1.90 ms).
  • A contention probe (8 purgers on separate engines, 300 tombstones with mixed outcomes and failures, 3 runs per dialect) ran every tombstone's rounds exactly as planned, with no two rounds on one tombstone at once.

Commits 27 and 28, the fourth review round: at commit 27, ruff check, ruff format --check and ty check are clean and the full server suite without integration tests passes (2108); its tests, that a failed and a cancelled confirmation each free the name, fail against the code before it, and a confirmation that committed before failing leaves the collection live. Commit 28 changes the purge document only; at it, one run failed an hnswlib slot-reuse search test the stack does not touch, which then passed 10 runs of 10 at the stack's head.

Commits 19-26, the third review round:

  • At commits 20-26, ruff check, ruff format --check and ty check are clean, and the full server suite without integration tests passes: 2099, 2098, 2099, 2100, 2100, 2101 and 2101 tests. Commit 19 changes a design document only.
  • Each new test fails against the code before its commit: a count past the dead-letter bound is reported; a second SQLite purger waits for the round under way, then claims the next tombstone; a cancel that fails after its creation stopped waiting is still logged; every lifecycle call is tracked.
  • After those runs, commit 22's SQLite test gained a 60-second busy timeout on its two engines, and commits 23-26 were restacked on it; the head's tree differs from the verified one by those three test lines only, and the test passes at the head.

Commits 13-18, the second review round:

  • At commits 13, 14 and 16, ruff check, ruff format --check and ty check are clean, and the full server suite without integration tests passes: 2093, 2097 and 2097 tests.
  • Commit 15 changes the purge document only.
  • At commit 17, ruff check is clean and the registry tests pass on SQLite (35 passed, 3 skipped). Commit 18 touches none of its files, so the head's run covers it otherwise, its claim on PostgreSQL included.
  • Commit 18 is the head.

Commits 1-12, rebased onto #1733's head b05becefa: each is patch-identical to the one verified below, apart from context lines, except commit 3, whose import of require_declared_types follows the function into utils.py. At commit 3, ruff check, ruff format --check and ty check are clean, and the full server suite without integration tests passes (2083).

At each commit, before the rebase:

  • ruff check, ruff format --check and ty check are clean, run as CI runs them.
  • The full server suite without integration tests passes at commits 1-6: 2031, 2064, 2075, 2079, 2082 and 2082 tests.
  • Commit 7 adds the design documents only.
  • Commits 8 and 9 change docstrings and one design-document line only. At each, ruff check and ruff format --check are clean.
  • At commit 10, ruff check, ruff format --check and ty check are clean, and the full server suite without integration tests passes (2083). Its new service-locator test fails against the locator before it.
  • At commit 11, ruff check, ruff format --check and ty check are clean, the full server suite without integration tests passes (2084), and the registry tests pass on SQLite and PostgreSQL (69, 3 skipped).
  • At commit 12, a rename (register → reserve, mark_live → confirm, a pending registration's unregister → cancel, PendingRegistration → Reservation, LiveRegistration → Registration): ruff check, ruff format --check and ty check are clean, the full server suite without integration tests passes (2084), and in test containers the integration tests pass against PostgreSQL and Qdrant 1.17.0 (493 passed, 6 skipped).

🤖 Generated with Claude Code

This was referenced Oct 1, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/vector-store-collection-registry-speedkick branch from 4356fb2 to 5ec495c Compare October 1, 2026 22:30
@edwinyyyu edwinyyyu added the horizontal scaling Wrong or unsafe when more than one server process serves the same backends (replicas or workers) label Oct 1, 2026
@edwinyyyu
edwinyyyu marked this pull request as ready for review October 1, 2026 23:46
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 1, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 2, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@edwinyyyu
edwinyyyu force-pushed the feat/vector-store-collection-registry-speedkick branch from e917697 to 34d197e Compare October 2, 2026 03:19
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…t claim

The purge document's backoff and interference figures timed earlier
forms of the claim, some of them the query alone. They are now
measurements from 2026-10-06, each naming the commit whose code was
measured: the backoff from commit 96a8b5b's claim, as whole calls, and
the interference from commit 503687c's lease, whose rounds cost the same
database time as commit 96a8b5b's.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
… and the contract

MemMachine#1734 names the registry's handles for their holders: `reserve` answers
a Reservation, whose `confirm` answers a Registration, and whose `cancel`
gives the name back. The Qdrant handle takes a Registration, and the
collection lifecycle contract's racing winners reserve, then confirm.
Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@edwinyyyu
edwinyyyu merged commit 5a84657 into MemMachine:feat/horizontal-scaling Oct 6, 2026
39 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
… and the contract

MemMachine#1734 names the registry's handles for their holders: `reserve` answers
a Reservation, whose `confirm` answers a Registration, and whose `cancel`
gives the name back. The Qdrant handle takes a Registration, and the
collection lifecycle contract's racing winners reserve, then confirm.
Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit that referenced this pull request Oct 6, 2026
…n registry (#1735)

* Test the Qdrant store against Qdrant 1.19.1

The integration tests ran Qdrant 1.17.0. The Qdrant store's design,
which follows in this stack, was measured against 1.19.1: its one-shot
purge of an incarnation by filter, and the per-tenant index layout. 1.17.0
predates the filter-resolution fence that 1.19.0 added to filter deletes
(qdrant#9678), whose extra cost 1.19.1 no longer shows, so the tests
could not see the behavior the store is tuned for.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound every Qdrant request by a configured timeout

The Qdrant client was built with qdrant-client's default timeout, so how
long a write to Qdrant can be in flight was nothing the configuration
stated. The purge that follows in this stack waits out a retention longer
than any write can be in flight, and the request timeout is the part of
that time the store controls.

`QdrantConf.request_timeout_seconds`, a positive whole number of seconds
defaulting to 30, is passed to the client. The sample configurations and
the configuration docs show it. The tests' Qdrant clients are built by
one fixture with the configuration's default, so they run with the
timeout a configured store has.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Move the Qdrant store onto the collection registry

The Qdrant store kept its catalog in a `__registry` collection per
namespace and serialized creating and deleting a collection with locks in
one process. A collection's name was the tenant discriminator on its
points, so a handle kept writing into a collection deleted and created
again under its name (#1563), and a write in flight during a deletion
outlived it.

`QdrantVectorStore` is now a `RegistryBackedVectorStore`, and any number
of processes sharing its registry may serve its collections:
- Its catalog is the collection registry in the relational database
  `QdrantConf.collection_registry` names. The `__registry` collections,
  `registry_replication_factor` and the process-local locks go. The
  database manager builds and starts the registry before it opens the
  client, so a registry database it cannot resolve leaves no client open.
- Every point carries its collection's incarnation in `sys-incarnation`,
  in place of the name, under the same tenant index, and every search
  filters on it.
- A point's id is `uuid5(incarnation, record UUID)`, and the record UUID
  is kept in the payload as `sys-record_uuid`, which a search returns.
  The logical collections sharing a native collection share its id space:
  with the record UUID as the id, an upsert of a UUID another collection
  held replaced that collection's point.
- Preparing a collection's storage creates its native collection and
  payload indexes, each under its own already-exists guard, so a creation
  that failed part way is completed by the next.
- A purge round looks for one point under the incarnation and, finding
  one, deletes the incarnation's points with one filter-delete.
  `QdrantConf.tombstone_retention_seconds`, a day by default, is the
  retention, and the configuration refuses one below 10 x
  `request_timeout_seconds` + 300 seconds.
- An upsert is halved only when Qdrant or a proxy refuses it as sent,
  with a 400 or a 413. Any other error raises at once: a timed-out upsert
  may still be applied, and sending it again adds load to a server
  already too slow.

The collection lifecycle contract (`collection_lifecycle_contract.py`),
which a store's tests mix in with hooks that read the backend directly,
runs on Qdrant: stale handles, a collection created again starting empty,
creation races and their outcomes, failed preparations, the purge, a
write landing under a dead incarnation, and the registry lookups each
operation makes. The store's own tests check what Qdrant holds by
scrolling the incarnation past the store.

Breaking: existing Qdrant data is orphaned, since its points carry names
and its catalog is in the `__registry` collections; no migration is
included. Every Qdrant store needs a relational database for its
registry, which the samples, the Helm chart, the configuration wizard
and the configuration docs name.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Document how the Qdrant store meets the shared contracts

`design/qdrant_vector_store.md` records the Qdrant store's layout, its
derived point ids and why they are one-way, filtered-search correctness,
the purge by filter, and its consistency on one node and replicated,
with the measurements behind each choice.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Build Qdrant handles from live registrations

The collection registry now answers registrations: a PendingRegistration
from register, which its creator marks live, and a LiveRegistration from
resolve and mark_live, whose require_current is a handle's fence. The
Qdrant handle takes its live registration in place of the namespace,
name, incarnation, configuration and registry lookup, and
_build_collection_handle builds it from one.

The collection lifecycle contract follows:
- the tests that fail a check or count checks patch the registration
  type's require_current, where they replaced the handle's lookup;
- the racing winners register and mark live through a pending
  registration, and the winner is found with resolve;
- churn counts VectorStoreCollectionDeletedError, which a creation
  undone by a concurrent deletion now raises, among the domain's
  outcomes.
The Qdrant tests that build a handle on a mocked client give it a live
registration that stays current.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use the registry's reservations and registrations in the Qdrant store and the contract

#1734 names the registry's handles for their holders: `reserve` answers
a Reservation, whose `confirm` answers a Registration, and whose `cancel`
gives the name back. The Qdrant handle takes a Registration, and the
collection lifecycle contract's racing winners reserve, then confirm.
Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the lifecycle contract's stale check of a query with limit 0

A limit at or below zero is now refused as invalid input, checked
before the handle's liveness, so it no longer stands for a query with
nothing to send to the backend; the query with no vectors still does.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say to drop the Qdrant collections before upgrading, and list the registry

The user docs did not say what happens to existing Qdrant data. The
store keeps its native collections' names, so points left in one stay,
invisible and never purged, and dropping the collection after the
upgrade drops the new data too; databases.mdx now says to drop the
collections before upgrading. The Helm README's configuration summary
also lists collection_registry, which the configmap template sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Run every Qdrant store test against a real Qdrant server

Qdrant's local mode answers differently from a server: it ignores
payload indexes and raises its own exceptions where a server answers
404 or 409 over REST or NOT_FOUND or ALREADY_EXISTS over gRPC. The
store fixture's client is now a testcontainers Qdrant over REST or
gRPC, marked integration, so every store test exercises what a
deployment runs. The tests that build a handle on a mocked client stay
in the default suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Give the wiring tests' collection registry a file database

The Qdrant wiring tests gave the collection registry an in-memory
SQLite database, which aiosqlite serves from one shared connection;
the segment store and the SQLite vector stores refuse such an engine,
and the registry is to do the same. Each test now puts the registry in
a SQLite file under its own tmp_path. What each test asserts is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse engines that share one connection or hold SQLite in memory

The registry's arbitration rests on each operation running in a transaction
of its own: a reservation's insert and queue check, a deletion's row and
tombstone, a purge claim. On an engine whose pool is StaticPool, every
session shares one connection, so concurrent operations would run inside one
another's transactions. On in-memory SQLite each connection gets a separate
database, so the registry's state would not be shared even within one
process. A SQLite configuration whose path is ":memory:" produces the former
(aiosqlite uses StaticPool for it). The segment store and both SQLite vector
stores already refuse both; the registry now does too, with the same
messages.

The dialect and SQLite-version refusal tests now use file engines, so each
checks only its own refusal. The new tests fail with either check removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Pin the Qdrant client's timeout wiring with a non-default value

The wiring test asserted that the client got a 30 s timeout, which is
also QdrantConf's default, so a manager that ignored
request_timeout_seconds and passed 30 passed it. It now configures 7 s
and asserts the client is built with 7; hardcoding 30, or dropping the
timeout from the client's arguments, fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound the lifecycle contract's drain and settle before empty checks

The contract drained purges with an unbounded while loop, so a store
whose purge round kept reporting records hung the suite. The drain now
fails the test once 1,000 rounds have each found something, far more
than the few deleted collections a contract test leaves; a Qdrant
round that always reports records now fails the purge tests instead of
hanging them.

The two tests that check that a collection created again under a
deleted one's name starts empty now call the store's settle hook first.
On a store whose reads lag its writes, a query could miss the old
life's record for that reason alone, and the check would pass even if
the new life could reach it; after settling it fails only when the new
life cannot. Dropping the incarnation filter from Qdrant's search
fails both.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check a stale upsert by what it stores, not by its registry reads

The contract counted require_current calls: two per upsert, one per
query or delete. That pinned how a handle checks, not what the check
guarantees, and a store that checked differently but kept the
guarantee failed it. The test that replaces it pins the state: an
upsert through a handle whose collection is already deleted raises
StaleError, and the dead life's records, read past the store, are
exactly those it held before. Removing the handle's liveness check
before its write lets the record land under the dead incarnation and
fails it.

The rest of what the counts stood for is pinned by state elsewhere in
the contract: a stale query and a stale delete raise, and an upsert
whose collection is deleted between its check and its write raises.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that filtered queries and namespaces keep Qdrant collections apart

Two collections of one namespace and configuration share a native
collection, and the existing isolation tests either queried without a
filter or used a filter only the querying collection's records met, so
a search whose property filter displaced the incarnation filter, or sat
beside it under `should`, passed them all.

- A filtered query returns only its own collection's records: both
  collections hold records meeting every filter (a comparison, In, Or,
  Not(IsNull) on an undeclared property, And), and each query must
  return exactly its own. Replacing the incarnation filter with the
  property filter, or putting both under `should`, fails it.
- Collections of two namespaces keep their records in separate storage,
  as VectorStore states: an upsert into one namespace's collection adds
  to that namespace's storage and leaves the other's count unchanged,
  under one name and configuration. Building the native collection's
  name without the namespace fails it.

The lifecycle contract's count_stored hook now uses the module's
_count_stored helper, which the namespace test also reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test Qdrant purge rounds on missing storage and on a write after a round

- A tombstone whose storage was never made is retired: a creation
  whose storage preparation fails before touching Qdrant cancels its
  reservation, queuing a tombstone for a native collection that does
  not exist. Its round must find nothing and retire it, raising
  nothing: the first purge call runs a round and the next finds none
  due. On Qdrant 1.19.1 the round's scroll of the missing collection
  raises UnexpectedResponse 404 over REST and AioRpcError NOT_FOUND
  over gRPC (a filter-delete raises the same). Removing the not-found
  handling around the scroll fails it on both transports; classifying
  only REST's 400, or gRPC's INVALID_ARGUMENT, as not found fails it
  on that transport.
- A write landing after a purge round is reclaimed by the next: after
  a round deletes a deleted collection's records, a write through the
  stale handle's backend path, past its liveness checks, lands under
  the dead incarnation, as a write in flight across the deletion does.
  Draining must leave nothing under the incarnation. A round that
  reports nothing found after deleting retires the tombstone on its
  first round and leaves the late write in place, failing it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that a Qdrant upsert raises when a point is refused on its own

An upsert refused with a 400 or 413 is halved until its halves fit or a
single point is refused, and the store's docstring says the single
point's refusal raises. The halving test refused only batches over a
size, so every point was eventually accepted. The new test refuses one
point whenever it is sent: the upsert must raise the refusal, with its
status, rather than return as if the point were accepted. Returning,
instead of raising, once halving reaches a single point fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test ranking, scores and threshold direction for every similarity metric

The suite queried cosine collections only, so a metric mapped to the
wrong Qdrant distance, or a threshold kept on the wrong side for a
distance, went unseen. For each of COSINE, DOT, EUCLIDEAN and
MANHATTAN, on a real Qdrant over REST and gRPC, five vectors that the
four metrics rank in four different orders, with no ties, are queried
against one vector:

- matches come best first, highest for a similarity and lowest for a
  distance;
- each score is what QueryMatch defines: cosine similarity, dot
  product, Euclidean distance or Manhattan distance;
- a threshold halfway between the second and third best scores keeps
  exactly the best two: scores above it for a similarity, below it for
  a distance.

Mapping EUCLIDEAN to Qdrant's Manhattan distance or back, COSINE to
dot product or back, dropping the threshold, negating it for distances,
or applying it as a similarity for every metric fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test a seeded Qdrant operation sequence against a model

Two collections of one namespace and configuration, so sharing a native
collection, take 200 seeded steps over a pool of 12 sorted uuid4s:
upserts and re-upserts that change or drop properties (one of them
undeclared), deletes of a collection's own UUIDs, the other
collection's and absent ones, filtered queries (comparisons, In, IsNull,
Not, And, Or) with a limit of the whole pool, deletion and re-creation
under the same name, and purge drains. After every step, each
collection's stored record UUIDs, read past the store, and an
unfiltered query's matches (each record once, scored by its latest
vector, best first) agree with a model of each collection's records; a
query step's filtered matches agree with the model's filter; a drain
leaves exactly the model's records in the native collection.

Merging an upsert into the stored point's payload, so that a dropped or
changed property keeps its old value, fails it; so does deriving point
ids from the record UUID alone, searching without the incarnation
filter, and a purge round that never deletes what it finds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test concurrent churn across two Qdrant stores sharing one registry

Two stores, one on a REST and one on a gRPC client, each with its own
registry object and engine on one registry database (a SQLite file, and
PostgreSQL), serve eight seeded workers. Each worker owns four record
UUIDs, so its records in each collection life follow from its own
operations, and takes 50 steps over three collection names: upserts and
deletes of its records, queries filtered on its own records, and
deletion of a collection, half the time followed by its creation. A
purger per store drains whenever a collection is deleted.

- Every operation returns or raises a documented domain error (stale
  handle, already exists, pending, deleted, attempts exhausted); any
  backend or database error fails the test, and a deadline turns a
  deadlock into a failure.
- A query whose collection was live throughout it, which a delete of
  nothing confirms afterwards, returns exactly its worker's records;
  any query returns none of the worker's other lives' records.
- After quiescing and draining, each live collection's stored records
  and an unfiltered query equal what its workers wrote and kept, and a
  deleted life keeps no record but one whose upsert raised as stale: a
  write in flight across a deletion can land after a round found the
  incarnation empty, the race the tombstone retention closes, and the
  test runs with no retention.

Searching without the incarnation filter, deriving point ids from the
record UUID alone, misclassifying REST's 409 on an existing native
collection, a purge round that never deletes, and a purge round that
never returns each fail it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Make the Qdrant creation race tests race, and drop a dead index check

The two-workers test checked that the native collection ended up
indexed, but only the reservation's winner prepares storage, on a
native collection no one else touches, so that check could not fail
the way its docstring said; it is gone, and the docstring says what the
test does. Its creators now reserve the name together, behind a
barrier, so the loser always loses at the reservation rather than
finding the winner already pending or live. It also writes a record
through one handle and reads it through the other, so agreeing on one
collection is shown by state. An open-or-create that raises on losing
the reservation, and a strict create that treats a taken name as
created, each fail it.

A new test races what the old one meant to: two workers create two
collections of one namespace and configuration at once, both
preparations held at a barrier until both start, so both create the
shared native collection together and one finds it already there. Both
creations must succeed and each collection must hold its own records,
over REST and over gRPC. Checking whether the native collection exists
before creating it, without guarding the create, fails it on both
transports while every sequential test passes; so does misclassifying
the transport's already-exists answer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Keep a Qdrant match that scores exactly the score threshold

Qdrant rounds a query's score_threshold to single precision and keeps only
scores strictly better than it, on every metric and transport, so a match
scoring exactly the threshold was dropped, even a perfect cosine match under
a threshold of 1.0. The SQLite stores and Milvus keep it, and the contract
did not say which.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the adjacent single-precision value on the worse side
(numpy's nextafter in float32), so Qdrant keeps a score equal to the
threshold, and checks the caller's threshold on the returned scores, so a
score Qdrant's rounding let through but that falls short of the threshold as
returned is dropped. A threshold beyond single precision is not sent; the
store's check applies it.

Sending the threshold, rather than only checking it in Python, keeps Qdrant
from returning matches the threshold cuts. Measured against Qdrant 1.19.1
(4 collections of 5,000 768-dimension points sharing one native collection,
300 queries per cell with the two variants alternated, a threshold passing a
quarter of the limit; AC power), p50 per query, threshold sent vs checked in
Python only: REST 2.89 vs 3.04 ms at limit 10, 2.85 vs 2.96 at 40, 3.13 vs
3.42 at 100; gRPC 2.86 vs 2.93 ms at 10, 3.36 vs 3.80 at 40, 3.33 vs 4.23 at
100. Both returned identical matches.

The new test, on every metric and on REST and gRPC, sets the threshold to a
match's reported score (the match is kept) and one step better (it is
dropped). It fails with the threshold sent unchanged, stepped the wrong way,
without the store's check, or with that check strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that the SQLite stores keep a match scoring exactly the threshold

The contract says a match scoring exactly the score threshold is returned.
Both SQLite stores already keep it; each now has the test the Qdrant store
has, on every metric it supports (cosine, dot and Euclidean on the USearch
store; cosine and Euclidean on sqlite-vec): a threshold equal to a match's
reported score keeps the match, and one a step better drops it. Each fails
with its store's threshold check made strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use serial commas in the Qdrant document, and shorten the threshold comment

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…SQL registry, with an incarnation per collection life and a purge (MemMachine#1734)

* State a collection's lifetime in the vector store contract

The contract had one process manage a collection, with the consumer
sharding names across processes, and said nothing of what deleting a
collection means for the handles and writes that outlive it. It now
states what a store guarantees across a deletion, and leaves which
processes may share a collection to each store:
- A collection created again under a deleted one's (namespace, name)
  starts empty, and reclaiming the deleted one's storage leaves it
  untouched.
- A handle whose collection was deleted may raise
  `VectorStoreCollectionHandleStaleError`, or, on a store that cannot
  tell, act on a collection created again under the same name.
- `delete_collection` makes the collection unreachable when it returns. A
  store that reclaims storage later does so in
  `purge_deleted_collections`, which does a bounded amount of work per
  call, returns whether it ran a round, and is safe to call from several
  processes at once.
- Creating a collection that is being created raises the already-exists
  error; opening one whose creation has not completed raises
  `VectorStoreCollectionPendingError`; a store that gives up after
  repeated attempts that made no progress raises
  `VectorStoreAttemptsExhaustedError`.
- A record's UUID names it in its collection only.

Every store keeps reclaiming storage in `delete_collection`, so
`purge_deleted_collections` returns False on each, and none raises the
pending or stale-handle error yet. Each store states which processes may
share a collection: the SQLite store's engine lives in the process that
opened the collection, sqlite-vec's database file in one host, and the
Qdrant and Milvus stores serialize creating and deleting a collection
within a process only. On each, a handle used after its collection is
deleted acts on a collection created again under its name.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Add a SQL collection registry for vector stores whose backend cannot arbitrate one

Qdrant and Milvus hold only points: the stores keep their catalogs of
logical collections as points and entities, written without a
transaction, and serialize creating and deleting a collection with locks
in one process. Two processes creating or deleting one collection race,
and a write in flight when its collection is deleted can land after the
deletion, in a collection created again under the same name.

`VectorStoreCollectionRegistry` is that catalog, arbitrated across every
process that shares it:
- `register` mints a fresh incarnation, the value a collection's records
  carry, under a (namespace, name) no live or pending collection holds.
  The collection is pending until `mark_live` marks its storage prepared.
- `unregister` makes a collection unreachable when it returns and queues
  its incarnation as a tombstone; `unregister_incarnation` does the same
  for a caller holding the incarnation, and leaves a collection registered
  since under the same name.
- `claim_purgeable_incarnation` hands one due tombstone to a purge round.
  A tombstone comes due once a retention, longer than any write can be in
  flight, has passed since the deletion, and is removed when a round finds
  no records. Its incarnation is not minted again before then, so a
  collection created again starts empty and no purge reclaims its records.

`SQLAlchemyVectorStoreCollectionRegistry` keeps it in PostgreSQL or SQLite
(3.35 or later, for RETURNING): a table of collections keyed by vector
store, namespace and name, and a queue of tombstones claimed in the order
they come due on the database clock. The primary key arbitrates
registration, a conditional update marks a collection live,
unregistration is one transaction, and on PostgreSQL a claim is a row
lock. A round that raises backs its tombstone off, doubling up to a
bound; after ten consecutive failed rounds the tombstone is
dead-lettered: kept, its incarnation reserved, skipped by claims, and
reported in an error log.

No store uses it yet.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Add RegistryBackedVectorStore, a base for stores whose collections a registry arbitrates

A store on the collection registry keeps the same lifecycle whatever its
backend, so the lifecycle lives in a base and a subclass supplies the
backend calls.

`RegistryBackedVectorStore` registers a collection pending, prepares its
storage (`_prepare_storage`), and marks it live; only a live collection
opens, and opening a pending one raises
`VectorStoreCollectionPendingError`. A collection deleted while its
storage is prepared is not marked live, so its creation counts as one
followed by a deletion. A preparation that raises or is cancelled
unregisters the incarnation, shielded so that a cancelled creation still
frees the name; a collection the registry cannot unregister stays
pending until it is deleted. `open_or_create_collection` retries a
bounded number of times, a second apart: it opens a pending collection
once another creator marks it live, creates again when it loses the name
or the mark, and raises the pending error or
`VectorStoreAttemptsExhaustedError` when the attempts run out.
`delete_collection` unregisters, and each `purge_deleted_collections`
call runs one round (`_purge_round`) on a due tombstone.

`RegistryBackedVectorStoreCollection` is a handle bound to one
incarnation. Each operation checks its inputs, then that the collection
registered under its name still carries that incarnation: `upsert` before
and after its backend call, `query` before, and `delete` after. A write
that raced the deletion raises `VectorStoreCollectionHandleStaleError`
instead of reporting success, and the purge reclaims whatever it landed.
A subclass implements `_upsert`, `_query` and `_delete`.

`require_identifiers` checks a namespace and a name together.

The tests drive the lifecycle through a store whose storage preparation
each test controls, on a SQLite registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Run each vector store's purge from the resource manager

A store that reclaims a deleted collection's storage later does so only
when `purge_deleted_collections` is called. The resource manager starts
one sweeper per vector store the first time it hands the store out, as it
does for segment stores. A sweeper calls the purge again after a short
pause while rounds keep running, after the idle interval once nothing is
due, and logs a call that raises and retries it on the next tick.
Sweepers in other processes may run the same store's purge at the same
time, which the contract allows. Closing the manager cancels its
sweepers, and a sweeper holds no reference to the manager.

Every store returns False from the purge as of this commit, so a sweeper
idles.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Open the event backend's collection a racing creator won

The event backend's service locator opens a session's collection, or
creates it when there is none and opens it after. Two callers building
one session's memory at once can both find none: the loser's create
raises the already-exists error, and an open while the winner is still
preparing the collection raises the pending error. Either failed the
build.

The locator now retries open-or-create on either error, up to ten
attempts a second apart: it opens the winner's collection once it is
live, and creates the collection itself if the winner's creation was
undone. A collection that stays pending through every attempt fails the
build with a RuntimeError naming the partition.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Let the segment store's callers decide to mint again, not the insert

The collection registry follows the segment store's incarnation design,
with its retry decision in `register`, which counts the attempts. The
segment store had that decision in its callee: `_insert_partition_row`
logged "re-minting" and "retrying with a fresh incarnation", and its
error was documented as "retry with a fresh incarnation". The insert now
raises the error saying what happened (the incarnation awaits purge, or
the insert failed with no row under the key), and `create_partition` and
`_open_or_create_partition`, which count the attempts, log the rejection
and mint another. No behavior changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Document the vector store's horizontal scaling design

Five shared documents, made with every backend in mind:
- horizontal scaling: the problem, the goals, a collection's lifecycle,
  and what each store guarantees;
- the collection registry: its tables, operations, creation races,
  handles and their fencing;
- purge: tombstones, the retention, the claim, backoff, dead-lettering
  and the sweeper, with the claim's measured cost;
- consistency: what a query sees of earlier writes, and why the contract
  has no `get`;
- isolation: between collections, and record UUIDs and their reuse.

They describe the design the Qdrant and Milvus stores complete when they
move onto the registry; each store's own document comes with that move.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Put "pending" before the noun in the registry's docstrings

"Mark the collection registered, pending, under an incarnation live" set
"pending" off in commas, which made the sentence hard to parse. It and
its kind now read "the pending collection with the given incarnation",
"a new pending collection" and "registered as pending", in the registry
ABC, the registry-backed base and the horizontal scaling design.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* State what mark_live does with an incarnation that is not pending

The contract said what a successful call does, but not what happens when
no collection carries the incarnation or the one that does is already
live. In both cases nothing changes and the call returns False, which is
what the SQLAlchemy registry's conditional UPDATE does; the return value
now reads as whether this call marked a collection live.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Raise from mark_live when the collection is no longer pending, not a bool

mark_live answered whether it marked the collection live. Its one caller
that read the answer, open-or-create, used False to mean that a deletion
had removed the collection while its storage was prepared, and
create_collection ignored it, so a creation a deletion undid reported
success. A method named for an action now succeeds or raises, as the
registry's other methods do.

mark_live takes the collection's (namespace, name) beside its
incarnation, matches on all three, and raises
VectorStoreCollectionDeletedError, naming the collection, when no
pending collection carries the incarnation. create_collection lets it
propagate and states it in the VectorStore contract; open-or-create and
the event backend's service locator catch it and create again.

Tests: a creation whose collection is deleted during preparation raises;
a second mark of one incarnation raises; the service locator creates
again after a deletion undid its creation (fails against the locator
before this commit).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Answer registrations from the collection registry: pending from register, live from resolve

The registry had two addressing schemes. register, get and unregister
took a (namespace, name); mark_live and unregister_incarnation took the
incarnation register had minted, which every caller carried back by hand,
and get answered a record whose `live` field went stale as soon as the
creator's own mark_live ran. Any incarnation fit any call.

The registry is now addressed by name alone: register, resolve,
unregister and claim_purgeable_incarnation. What acts on one life of a
collection is a registration, a handle bound to its incarnation, split by
state:
- register answers a PendingRegistration, whose creator marks it live
  (mark_live, which answers the LiveRegistration) or abandons it
  (unregister, this life only);
- resolve answers a LiveRegistration, or None when no collection holds the
  name, and raises VectorStoreCollectionPendingError when the collection is
  pending; a LiveRegistration's require_current raises
  VectorStoreCollectionHandleStaleError once the collection is deleted.
A registration's fields (namespace, name, configuration, incarnation)
never change; the state is conveyed by which outcome a call takes, so an
open is still one read and the fence one. A live collection offers no
mark_live and no unregister: it is deleted by name only.

The pending error carries the collection's configuration, so
open-or-create still refuses a pending collection of another
configuration without waiting for it. A store's collection handle is
built from its live registration, and its fence is
`registration.require_current()`. The SQLAlchemy registry supplies both
registration types; the name-keyed and incarnation-keyed unregistrations
share one transaction. The registry design document records why handles,
why two types, and why no state fields.

Tests: the registry's tests drive registrations (SQLite and PostgreSQL),
with a new test that a live registration is current until its collection
is deleted; the base's tests patch the pending registration's unregister
where they patched the registry's.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the registry's handles for their holders: a reservation and a registration

PendingRegistration and LiveRegistration named the collection's state
when the registry answered them, which a deletion by anyone falsifies:
a "live" registration may have been deleted since. What does not change
is what each handle's holder may do, so the handles are named for that,
in the words of two known patterns: Try-Confirm/Cancel for the creator,
and the stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's
  hold on the name while it prepares the collection's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks
  the collection live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(namespace, name)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the collection's state, in the store
contract and VectorStoreCollectionPendingError. The base store's
creation flow, its task set and log text, the tests and the design
documents follow; the registry design records why the names are roles.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Leave a confirmed collection alone when its reservation is cancelled

Reservation.cancel() deleted the row carrying the reservation's
incarnation whatever its state, while confirm() acted only on a pending
row. A cancel issued after a confirmation that committed but whose
answer was lost would have tombstoned a live collection outside the
deletion by name, the one path that ends a live collection. cancel()
now acts only while the row is pending, as confirm() does; once the
reservation is confirmed it does nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse a query limit that is not positive in registry-backed handles

The handle answered empty results for a limit at or below zero, where
the query contract now refuses it. It raises ValueError with the other
input checks, before the liveness check and whether or not there are
query vectors, so _query is called with a positive limit.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* State how long the purge claim holds its transaction

The claim's SELECT ... FOR UPDATE SKIP LOCKED transaction stays open,
idle on PostgreSQL, through the backend's deletion, so a server-side
idle_in_transaction_session_timeout shorter than a round fails the
round, and one that keeps doing so dead-letters the tombstone. The
purge document states the requirement beside the measured round
durations.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Chain the locator's give-up error to the last error it retried

When a session's collection stayed pending, or kept losing races, the
retry loop gave up with a RuntimeError that carried no cause, so the
log lost the pending error's registered_at, the time an operator needs
to tell an abandoned creation from a slow one. The RuntimeError is now
raised from the last error the loop caught; its type and message are
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Never shift the purge backoff by a negative count

The claim computes 1 << (failed_rounds - 1) for each candidate row,
rows at 0 failures included, where the shift count is -1: defined on
SQLite, undefined behavior inside PostgreSQL's int4shl. The row's first
disjunct made it claimable regardless, so no claim went wrong, but the
expression was undefined there. The count is now clamped at 0 with a
CASE, which both dialects evaluate alike; backoffs after a failure are
unchanged, and SQLite's plan for the claim is the same index range on
(vector_store_name, enqueued_at).

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Chain open-or-create's give-up error to the race it last lost

When every attempt lost the reservation or the confirmation to another
creator or a deleter, open_or_create_collection gave up with a
VectorStoreAttemptsExhaustedError that carried no cause. It is now
raised from the last lost race, as the event backend's locator chains
its own give-up error; its type and message are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the limit of 0 from the registry doc's no-op calls

A query whose limit is not positive is refused with the other invalid
inputs, before the handle checks its liveness, so it is no longer an
example of a call with nothing to send that still checks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Report a dead-lettered tombstone at or past the bound

The error log fired only when a failed round brought the count to
exactly the bound. The claim takes only tombstones under the bound, so a
tombstone is dead-lettered once its count reaches the bound or goes
past it, as when two purgers fail one tombstone at once; the log now
fires in either case.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Run a purge round inside the registry instead of handing out a claim

claim_purgeable_incarnation() was an async context manager, which cannot
see what its body returns, so a round reported whether it found records
through PurgeClaim.any_records_found: a field that started None and a
guard that raised when a round left it unset. run_purge_round() takes
the round as a callable instead. The registry claims the tombstone that
came due first, calls the round with its namespace, configuration and
incarnation, and records what the round returns under the claim. A
round that raises still counts as a failed round. PurgeClaim, its None
state and the guard go; the vector store's _purge_round hook is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Claim a purge round on SQLite with a write, as the segment store does

On SQLite the claim was a SELECT whose FOR UPDATE SKIP LOCKED the
dialect drops, and the driver defers BEGIN until a write, so the claim
ran outside any transaction: two processes sharing a SQLite registry
could claim one tombstone at once. The round recorded afterwards then
decided the failed-round reset from what it read at claim time, so a
round that succeeded could skip the reset after a racing round failed.

On SQLite the claim is now the segment store's: an UPDATE ... RETURNING
on the oldest eligible tombstone, which opens the write transaction, so
purgers serialize at the claim and rounds run one at a time.
PostgreSQL keeps FOR UPDATE SKIP LOCKED. The cost on SQLite is that
its single write lock is held across the round's remote deletion: the
registry's other writers wait for it, and past the driver's busy
timeout they fail with a locked-database error. The purge document
states it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Report a failed reservation cancel from the cancel's own task

A creation whose storage preparation failed or was cancelled starts
the reservation's cancel as a shielded task and logged its failure
around the await. A creation cancelled again stops awaiting it, so a
cancel that then failed went unlogged in the store, surfacing only as
asyncio's "Task exception was never retrieved" without the collection's
name. The task's done-callback now reports the failure, whoever is
still awaiting it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check identifiers with the shared helper in every store

utils.require_identifiers checked the namespace and name of the
registry-backed base's lifecycle calls, while the four stores kept the
check inline: Qdrant and Milvus with the helper's own body, the SQLite
stores with one combined check whose message, "Invalid namespace ... or
name ...", named neither the rule nor which identifier broke it. Every
store's lifecycle calls now use the helper, so an invalid identifier is
reported the same way everywhere, with the rule it breaks.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Track open_collection like the other lifecycle calls

The registry-backed store's create, open-or-create, delete and purge
ran under the store's operation tracker, and open_collection, which
resolves the name in the registry and is the first call the event
backend's locator makes on every open, did not.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say to drop the native collections before upgrading

The deployment note said existing Qdrant and Milvus data is orphaned.
The native collections keep their names, so the upgrade's new data
lands beside the old: a Qdrant collection kept through it holds old
points no search sees and no purge reclaims, and dropping it afterwards
drops the new data too; an existing Milvus collection has the earlier
schema, which the store cannot prepare. The note now says to drop the
collections before upgrading, and why.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Cancel the reservation when confirming a creation fails or is cancelled

A creation cancelled its reservation when preparing the collection's
storage raised or was cancelled, so the name was free for the next
attempt, but not when the confirmation that follows did: a failure or a
cancellation while confirming left the collection pending, its name
taken until someone deleted it. Preparation and confirmation now share
one cleanup, the same shielded, logged cancel. The cancel acts only on
a pending collection, so a confirmation that committed before its
failure or cancellation was observed stands.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say which writers wait on SQLite's lock during a purge round

The purge document said the registry's other writers wait while a
SQLite round holds the write lock. The lock covers the whole database
file, so every store writing to it waits, and under the configuration
wizard's defaults the episode store, session manager, segment store and
configuration database share the registry's file. The round holds the
lock across remote calls bounded by the store's request timeout, and
past SQLite's default 5 s busy timeout, which the server does not
change, those writers fail with a locked-database error.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Wrap the registry doc's creation-cleanup paragraph

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Lease the registry's purge claims instead of holding a transaction across the round

The purge claim held its transaction for the whole round. On SQLite, whose
write lock covers the whole database file, every writer to the database
waited out the round's remote calls to Qdrant or Milvus, each bounded by the
request timeout, and failed past SQLite's 5 s busy timeout with "database is
locked"; with the registry in one SQLite file beside the other stores, as the
configuration wizard sets up, that reached every store. On PostgreSQL the
session sat idle in a transaction for the round, holding back vacuum.

The claim is now a lease. One UPDATE ... RETURNING, committed at once, takes
the oldest due tombstone that no unexpired claim holds (picked FOR UPDATE
SKIP LOCKED on PostgreSQL; SQLite drops the clause and serializes the
write), stamps its claimed_at with the database's now() and increments its
claim_generation. The round runs with no transaction open. Its outcome is
recorded in a short transaction: a round that found nothing removes the
tombstone under any claim; one that found records or raised ends its claim,
clearing or counting its failed rounds, only while claim_generation is still
its claim's. A round that outlasted its lease therefore neither ends the
claim taken after it nor counts a failure against it, and logs a warning.

The lease lets purgers split a backlog; correctness rests, as before, on rounds
being safe to repeat and to run on two purgers at once. A claim holds until
its round ends or purge_lease_seconds (a new registry parameter, default
300) has passed since claimed_at, applied when a claim is decided, on the
database clock, as the retention and the backoff are.

Behavior changes:
- SQLite purgers run rounds on different tombstones at once, where they ran
  one at a time, and the database stays writable during a round.
- A round that found records always writes, to end its claim.
- A cancelled or crashed round's tombstone is claimed again once its lease
  has passed, rather than at once.
- An error recording a round's outcome is not counted as a failed round.

collection_registry_gc gains claimed_at and claim_generation. The test that
pinned SQLite's one-at-a-time rounds is replaced. The new tests: SQLite
stays writable during a round to a connection with a 0.1 s busy timeout (the
old claim made it raise "database is locked"); no PostgreSQL backend is idle
in a transaction during a round; a claimed tombstone is skipped until its
round ends or its lease passes; a round that outlasted its lease leaves the
claim after it; purgers on separate engines claim each tombstone once; and
the lease runs by the database clock. Removing any one lease mechanism fails
at least one of them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the purge backoff's first delay base_purge_retry_backoff_seconds

purge_retry_backoff_seconds read as the backoff itself, though it is the
delay after the first failure, doubled for each further one. The base_
prefix says so and mirrors max_purge_retry_backoff_seconds, the cap. No
configuration sets it: the stores construct the registry with its default.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Count a purge round that never ended as failed, and end a cancelled round's claim

A round whose purger died wrote nothing after its claim, so it was never
counted: a tombstone whose round killed its purger every time was claimed
again each time its lease passed and never dead-lettered, and nothing tied
the crash to the tombstone.

The claim now counts such a round, as job queues count attempts. A claim
that finds the oldest due tombstone's previous claim still open, its lease
passed, runs no round: it counts that round as failed, as of when it was
claimed (failed_rounds + 1, last_failed_at = the old claimed_at), ends the
claim, bumps the generation so the old round's writes are fenced, and logs
a warning naming the incarnation and the claim time, plus the dead-letter
error when the count reaches the bound. One UPDATE handles both cases: its
SET reads the row as it was, on PostgreSQL and SQLite alike, and leaves
claimed_at null for an open claim, which RETURNING reports. No column is
added.

A cancelled round now ends its claim uncounted, in a write shielded from the
cancellation, as a cancelled creation cancels its reservation, so its
tombstone is claimable at once and a shutdown does not look like a crash. A
release that fails is logged by its own task, and the round is then counted
once its lease passes.

Behavior changes:
- A round whose purger died counts as a failed round once its lease passes,
  and its tombstone backs off from the time the round was claimed.
- A round whose outcome could not be written counts the same way.
- A cancelled round's tombstone is claimable at once, rather than after the
  lease.

Tests, each failing when the behavior it pins is removed: a cancelled round
ends its claim uncounted and logs nothing; a round that never ended keeps its
tombstone until the lease passes, then is counted, reported and backed off
from its claim time, and a lost release is reported; rounds that never end
dead-letter the tombstone; purgers on four engines that find one unended claim
count it once; and the overrun tests now expect the stale round counted once.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test the registry's wiring, atomicity, isolation and behavior under concurrency

New tests, run on SQLite and PostgreSQL; each fails, on both, with the
behavior it pins removed:
- A changed lease applies to claims already held, with a one-minute lease
  against a five-minute claim on both sides of the boundary (fails when the
  lease parameter is ignored; the other tests use the default).
- The first purge backoff comes from its parameter, and doubles (fails when
  the parameter is ignored).
- A deletion whose tombstone cannot be queued changes nothing: the
  collection stays registered and nothing is queued (fails when the delete
  and the queue insert commit separately, which would leave a deleted
  incarnation without a tombstone).
- A reservation whose purge-queue check fails leaves the name free (fails
  when the insert commits before the check).
- A round whose outcome cannot be written propagates the error uncounted,
  holds its claim until the lease passes, and is then counted once (fails
  when a write error counts as a failed round at once).
- The same name in two namespaces is two collections that resolve, delete
  and purge apart (fails when the namespace is left out of a lookup).
- A reservation whose minted incarnation collides with a tombstone being
  purged mints again at once (fails with the claim held across the round:
  SQLite's write lock, or PostgreSQL's row lock under the queue check).
- Creators, deleters, readers and purgers on four engines race over shared
  names: only documented refusals, no database error or deadlock, no
  overlapping rounds, no incarnation both registered and queued, and the
  purge drains (fails without the lease).
- Seeded sequences of every operation agree with a model of the contract
  step by step (fails when a cancel ends a live collection, a failed
  tombstone is claimed during its backoff, or a round that found records
  removes its tombstone).

The fault-injection tests fail one statement on the engine by its prefix,
naming the registry's tables.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the purge cutoffs and the consecutive count, and shorten comments

- `_now_less(period, key)` is now `_database_time_ago(seconds)`. It takes a
  number of seconds or a SQL expression, so the retention, the lease, and
  the per-tombstone backoff cutoffs go through one helper, with anonymous
  bound parameters in place of the named ones. The registry keeps its
  durations as seconds.
- The purge queue's `failed_rounds` is now `consecutive_failed_rounds`, the
  count it is: a round that finds records resets it. The dead-letter log
  names the new column.
- The purge code's comments are shorter, and the design documents and
  docstrings this stack adds use serial commas.

The claim's database cost is unchanged within run-to-run variation
(no-op rounds, 500 per case, alternated 8 times, AC power): SQLite 1.52 vs
1.54 ms and 1.66 vs 1.60 ms, PostgreSQL 2.23 vs 2.45 ms and 2.14 vs 2.16 ms,
for rounds that found records and found nothing.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Split the purge round into named steps, and return a typed tombstone claim

`run_purge_round` now reads as its sequence: claim the oldest due tombstone,
report a round that never ended, run the round, then record its outcome,
count its failure, or end its claim after a cancellation. Each step is a
method named for what it acts on:

- `_claim_oldest_due_tombstone` builds the claim and returns a
  `_TombstoneClaim` (what the round needs, and the claim's generation), an
  `_UnendedPurgeRound` (an expired claim, counted as failed), or None, in
  place of one result row whose fields meant different things in each case.
- `_has_elapsed(seconds, since=column)` states each condition the claim
  checks (the retention since the deletion, the backoff since the last
  failure, the lease since the claim), with `_purge_retry_backoff_seconds`
  computing each tombstone's backoff. It replaces `_database_time_ago` and
  `_backoff_cutoff`; the SQL is unchanged.
- `_report_unended_purge_round`, `_count_failed_purge_round`,
  `_report_dead_lettered_tombstone`, `_end_tombstone_claim_after_cancellation`,
  `_end_tombstone_claim`, and `_record_purge_round` take the claim.
- `_insert` is `_insert_pending_collection`; the registry's and the store's
  task sets are `_tombstone_claim_endings` and `_reservation_cancellations`.

The namespace test's purge loop is bounded, so a claim that never stops
fails it instead of hanging it. Breaking the cancellation path, the
unended-round path, or the elapsed check each fails registry tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Count each purge attempt in its claim

A purge round whose purger dies writes nothing after its claim. The claim
that later found such a claim past its lease counted the round, ran no
round of its own, and reported it. That took a claim statement writing
every column one of two ways, a second kind of claim result
(`_UnendedPurgeRound`), and `last_failed_at` holding a claim time in that
case so that `RETURNING` could report it.

Each claim now counts its attempt, as job queues do when work is taken:
`consecutive_failed_rounds` is `consecutive_attempts`, incremented by the
claim's plain `UPDATE`. A round that finds records resets it, a round that
raises ends its claim and sets `last_failed_at`, a cancelled round takes
its attempt back, and a round whose purger died writes nothing. Each column
has one meaning: `claimed_at` is the open claim's time, and
`last_failed_at` is when a round last raised.

- A tombstone whose round never ended is claimed again once its lease, and
  then the backoff, have passed. The backoff runs from `last_failed_at`
  after a raise and from the lease's end after a round that never ended.
- A claim of a second or later attempt logs a warning naming the
  incarnation and the attempt ("attempt 3 of 10"). A raised round's error
  carries a note with both, so the sweeper's log of it names the tombstone.
  Nothing is logged later about an attempt that never ended.
- After 10 attempts, claims skip the tombstone. A last attempt that raises
  logs the dead-letter error; one whose purger died is not reported, since
  nothing runs on the tombstone after it.

New tests: a round that never ended is retried as the next attempt only
after its lease and the backoff, with the retry warning; racing purgers
retry it once; a cancelled retry of it leaves the tombstone claimable;
resetting a dead-lettered tombstone's attempts returns it at once; and a
raised round's error names the tombstone. The test of a count past the
dead-letter bound is gone, since claims stop at the bound. Breaking the
count, the backoff from the lease's end, the retry warning, the error's
note, the cancel's take-back, the reset, a raise's date or end, or the
dead-letter report or bound each fails a registry test.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the purge count attempts_without_progress, and tighten the purge docstrings

`consecutive_attempts` did not say what resets it. The column counts the
purge rounds claimed since a round last made progress, finding records
and deleting them, the open one included and a cancelled round's attempt
taken back: it is `attempts_without_progress`, and the dead-letter bound
is `_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The dead-letter error and the
design documents name the new column.

The base backoff's docstring and description state the rule alone:
seconds before a tombstone is claimed again after a failed round, doubled
for each further failed round in a row. Where the wait is measured from
stays in the purge design document. A raise on the last attempt is
described as reported, the claims skipping the tombstone from then on, and
the column and claim comments are shorter.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Replace the purge claim's measured costs with figures from the current claim

The purge document's backoff and interference figures timed earlier
forms of the claim, some of them the query alone. They are now
measurements from 2026-10-06, each naming the commit whose code was
measured: the backoff from commit 96a8b5b's claim, as whole calls, and
the interference from commit 503687c's lease, whose rounds cost the same
database time as commit 96a8b5b's.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…n registry (MemMachine#1735)

* Test the Qdrant store against Qdrant 1.19.1

The integration tests ran Qdrant 1.17.0. The Qdrant store's design,
which follows in this stack, was measured against 1.19.1: its one-shot
purge of an incarnation by filter, and the per-tenant index layout. 1.17.0
predates the filter-resolution fence that 1.19.0 added to filter deletes
(qdrant#9678), whose extra cost 1.19.1 no longer shows, so the tests
could not see the behavior the store is tuned for.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound every Qdrant request by a configured timeout

The Qdrant client was built with qdrant-client's default timeout, so how
long a write to Qdrant can be in flight was nothing the configuration
stated. The purge that follows in this stack waits out a retention longer
than any write can be in flight, and the request timeout is the part of
that time the store controls.

`QdrantConf.request_timeout_seconds`, a positive whole number of seconds
defaulting to 30, is passed to the client. The sample configurations and
the configuration docs show it. The tests' Qdrant clients are built by
one fixture with the configuration's default, so they run with the
timeout a configured store has.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Move the Qdrant store onto the collection registry

The Qdrant store kept its catalog in a `__registry` collection per
namespace and serialized creating and deleting a collection with locks in
one process. A collection's name was the tenant discriminator on its
points, so a handle kept writing into a collection deleted and created
again under its name (MemMachine#1563), and a write in flight during a deletion
outlived it.

`QdrantVectorStore` is now a `RegistryBackedVectorStore`, and any number
of processes sharing its registry may serve its collections:
- Its catalog is the collection registry in the relational database
  `QdrantConf.collection_registry` names. The `__registry` collections,
  `registry_replication_factor` and the process-local locks go. The
  database manager builds and starts the registry before it opens the
  client, so a registry database it cannot resolve leaves no client open.
- Every point carries its collection's incarnation in `sys-incarnation`,
  in place of the name, under the same tenant index, and every search
  filters on it.
- A point's id is `uuid5(incarnation, record UUID)`, and the record UUID
  is kept in the payload as `sys-record_uuid`, which a search returns.
  The logical collections sharing a native collection share its id space:
  with the record UUID as the id, an upsert of a UUID another collection
  held replaced that collection's point.
- Preparing a collection's storage creates its native collection and
  payload indexes, each under its own already-exists guard, so a creation
  that failed part way is completed by the next.
- A purge round looks for one point under the incarnation and, finding
  one, deletes the incarnation's points with one filter-delete.
  `QdrantConf.tombstone_retention_seconds`, a day by default, is the
  retention, and the configuration refuses one below 10 x
  `request_timeout_seconds` + 300 seconds.
- An upsert is halved only when Qdrant or a proxy refuses it as sent,
  with a 400 or a 413. Any other error raises at once: a timed-out upsert
  may still be applied, and sending it again adds load to a server
  already too slow.

The collection lifecycle contract (`collection_lifecycle_contract.py`),
which a store's tests mix in with hooks that read the backend directly,
runs on Qdrant: stale handles, a collection created again starting empty,
creation races and their outcomes, failed preparations, the purge, a
write landing under a dead incarnation, and the registry lookups each
operation makes. The store's own tests check what Qdrant holds by
scrolling the incarnation past the store.

Breaking: existing Qdrant data is orphaned, since its points carry names
and its catalog is in the `__registry` collections; no migration is
included. Every Qdrant store needs a relational database for its
registry, which the samples, the Helm chart, the configuration wizard
and the configuration docs name.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Document how the Qdrant store meets the shared contracts

`design/qdrant_vector_store.md` records the Qdrant store's layout, its
derived point ids and why they are one-way, filtered-search correctness,
the purge by filter, and its consistency on one node and replicated,
with the measurements behind each choice.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Build Qdrant handles from live registrations

The collection registry now answers registrations: a PendingRegistration
from register, which its creator marks live, and a LiveRegistration from
resolve and mark_live, whose require_current is a handle's fence. The
Qdrant handle takes its live registration in place of the namespace,
name, incarnation, configuration and registry lookup, and
_build_collection_handle builds it from one.

The collection lifecycle contract follows:
- the tests that fail a check or count checks patch the registration
  type's require_current, where they replaced the handle's lookup;
- the racing winners register and mark live through a pending
  registration, and the winner is found with resolve;
- churn counts VectorStoreCollectionDeletedError, which a creation
  undone by a concurrent deletion now raises, among the domain's
  outcomes.
The Qdrant tests that build a handle on a mocked client give it a live
registration that stays current.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use the registry's reservations and registrations in the Qdrant store and the contract

MemMachine#1734 names the registry's handles for their holders: `reserve` answers
a Reservation, whose `confirm` answers a Registration, and whose `cancel`
gives the name back. The Qdrant handle takes a Registration, and the
collection lifecycle contract's racing winners reserve, then confirm.
Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the lifecycle contract's stale check of a query with limit 0

A limit at or below zero is now refused as invalid input, checked
before the handle's liveness, so it no longer stands for a query with
nothing to send to the backend; the query with no vectors still does.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say to drop the Qdrant collections before upgrading, and list the registry

The user docs did not say what happens to existing Qdrant data. The
store keeps its native collections' names, so points left in one stay,
invisible and never purged, and dropping the collection after the
upgrade drops the new data too; databases.mdx now says to drop the
collections before upgrading. The Helm README's configuration summary
also lists collection_registry, which the configmap template sets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Run every Qdrant store test against a real Qdrant server

Qdrant's local mode answers differently from a server: it ignores
payload indexes and raises its own exceptions where a server answers
404 or 409 over REST or NOT_FOUND or ALREADY_EXISTS over gRPC. The
store fixture's client is now a testcontainers Qdrant over REST or
gRPC, marked integration, so every store test exercises what a
deployment runs. The tests that build a handle on a mocked client stay
in the default suite.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Give the wiring tests' collection registry a file database

The Qdrant wiring tests gave the collection registry an in-memory
SQLite database, which aiosqlite serves from one shared connection;
the segment store and the SQLite vector stores refuse such an engine,
and the registry is to do the same. Each test now puts the registry in
a SQLite file under its own tmp_path. What each test asserts is
unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse engines that share one connection or hold SQLite in memory

The registry's arbitration rests on each operation running in a transaction
of its own: a reservation's insert and queue check, a deletion's row and
tombstone, a purge claim. On an engine whose pool is StaticPool, every
session shares one connection, so concurrent operations would run inside one
another's transactions. On in-memory SQLite each connection gets a separate
database, so the registry's state would not be shared even within one
process. A SQLite configuration whose path is ":memory:" produces the former
(aiosqlite uses StaticPool for it). The segment store and both SQLite vector
stores already refuse both; the registry now does too, with the same
messages.

The dialect and SQLite-version refusal tests now use file engines, so each
checks only its own refusal. The new tests fail with either check removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Pin the Qdrant client's timeout wiring with a non-default value

The wiring test asserted that the client got a 30 s timeout, which is
also QdrantConf's default, so a manager that ignored
request_timeout_seconds and passed 30 passed it. It now configures 7 s
and asserts the client is built with 7; hardcoding 30, or dropping the
timeout from the client's arguments, fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound the lifecycle contract's drain and settle before empty checks

The contract drained purges with an unbounded while loop, so a store
whose purge round kept reporting records hung the suite. The drain now
fails the test once 1,000 rounds have each found something, far more
than the few deleted collections a contract test leaves; a Qdrant
round that always reports records now fails the purge tests instead of
hanging them.

The two tests that check that a collection created again under a
deleted one's name starts empty now call the store's settle hook first.
On a store whose reads lag its writes, a query could miss the old
life's record for that reason alone, and the check would pass even if
the new life could reach it; after settling it fails only when the new
life cannot. Dropping the incarnation filter from Qdrant's search
fails both.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check a stale upsert by what it stores, not by its registry reads

The contract counted require_current calls: two per upsert, one per
query or delete. That pinned how a handle checks, not what the check
guarantees, and a store that checked differently but kept the
guarantee failed it. The test that replaces it pins the state: an
upsert through a handle whose collection is already deleted raises
StaleError, and the dead life's records, read past the store, are
exactly those it held before. Removing the handle's liveness check
before its write lets the record land under the dead incarnation and
fails it.

The rest of what the counts stood for is pinned by state elsewhere in
the contract: a stale query and a stale delete raise, and an upsert
whose collection is deleted between its check and its write raises.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that filtered queries and namespaces keep Qdrant collections apart

Two collections of one namespace and configuration share a native
collection, and the existing isolation tests either queried without a
filter or used a filter only the querying collection's records met, so
a search whose property filter displaced the incarnation filter, or sat
beside it under `should`, passed them all.

- A filtered query returns only its own collection's records: both
  collections hold records meeting every filter (a comparison, In, Or,
  Not(IsNull) on an undeclared property, And), and each query must
  return exactly its own. Replacing the incarnation filter with the
  property filter, or putting both under `should`, fails it.
- Collections of two namespaces keep their records in separate storage,
  as VectorStore states: an upsert into one namespace's collection adds
  to that namespace's storage and leaves the other's count unchanged,
  under one name and configuration. Building the native collection's
  name without the namespace fails it.

The lifecycle contract's count_stored hook now uses the module's
_count_stored helper, which the namespace test also reads.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test Qdrant purge rounds on missing storage and on a write after a round

- A tombstone whose storage was never made is retired: a creation
  whose storage preparation fails before touching Qdrant cancels its
  reservation, queuing a tombstone for a native collection that does
  not exist. Its round must find nothing and retire it, raising
  nothing: the first purge call runs a round and the next finds none
  due. On Qdrant 1.19.1 the round's scroll of the missing collection
  raises UnexpectedResponse 404 over REST and AioRpcError NOT_FOUND
  over gRPC (a filter-delete raises the same). Removing the not-found
  handling around the scroll fails it on both transports; classifying
  only REST's 400, or gRPC's INVALID_ARGUMENT, as not found fails it
  on that transport.
- A write landing after a purge round is reclaimed by the next: after
  a round deletes a deleted collection's records, a write through the
  stale handle's backend path, past its liveness checks, lands under
  the dead incarnation, as a write in flight across the deletion does.
  Draining must leave nothing under the incarnation. A round that
  reports nothing found after deleting retires the tombstone on its
  first round and leaves the late write in place, failing it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that a Qdrant upsert raises when a point is refused on its own

An upsert refused with a 400 or 413 is halved until its halves fit or a
single point is refused, and the store's docstring says the single
point's refusal raises. The halving test refused only batches over a
size, so every point was eventually accepted. The new test refuses one
point whenever it is sent: the upsert must raise the refusal, with its
status, rather than return as if the point were accepted. Returning,
instead of raising, once halving reaches a single point fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test ranking, scores and threshold direction for every similarity metric

The suite queried cosine collections only, so a metric mapped to the
wrong Qdrant distance, or a threshold kept on the wrong side for a
distance, went unseen. For each of COSINE, DOT, EUCLIDEAN and
MANHATTAN, on a real Qdrant over REST and gRPC, five vectors that the
four metrics rank in four different orders, with no ties, are queried
against one vector:

- matches come best first, highest for a similarity and lowest for a
  distance;
- each score is what QueryMatch defines: cosine similarity, dot
  product, Euclidean distance or Manhattan distance;
- a threshold halfway between the second and third best scores keeps
  exactly the best two: scores above it for a similarity, below it for
  a distance.

Mapping EUCLIDEAN to Qdrant's Manhattan distance or back, COSINE to
dot product or back, dropping the threshold, negating it for distances,
or applying it as a similarity for every metric fails it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test a seeded Qdrant operation sequence against a model

Two collections of one namespace and configuration, so sharing a native
collection, take 200 seeded steps over a pool of 12 sorted uuid4s:
upserts and re-upserts that change or drop properties (one of them
undeclared), deletes of a collection's own UUIDs, the other
collection's and absent ones, filtered queries (comparisons, In, IsNull,
Not, And, Or) with a limit of the whole pool, deletion and re-creation
under the same name, and purge drains. After every step, each
collection's stored record UUIDs, read past the store, and an
unfiltered query's matches (each record once, scored by its latest
vector, best first) agree with a model of each collection's records; a
query step's filtered matches agree with the model's filter; a drain
leaves exactly the model's records in the native collection.

Merging an upsert into the stored point's payload, so that a dropped or
changed property keeps its old value, fails it; so does deriving point
ids from the record UUID alone, searching without the incarnation
filter, and a purge round that never deletes what it finds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test concurrent churn across two Qdrant stores sharing one registry

Two stores, one on a REST and one on a gRPC client, each with its own
registry object and engine on one registry database (a SQLite file, and
PostgreSQL), serve eight seeded workers. Each worker owns four record
UUIDs, so its records in each collection life follow from its own
operations, and takes 50 steps over three collection names: upserts and
deletes of its records, queries filtered on its own records, and
deletion of a collection, half the time followed by its creation. A
purger per store drains whenever a collection is deleted.

- Every operation returns or raises a documented domain error (stale
  handle, already exists, pending, deleted, attempts exhausted); any
  backend or database error fails the test, and a deadline turns a
  deadlock into a failure.
- A query whose collection was live throughout it, which a delete of
  nothing confirms afterwards, returns exactly its worker's records;
  any query returns none of the worker's other lives' records.
- After quiescing and draining, each live collection's stored records
  and an unfiltered query equal what its workers wrote and kept, and a
  deleted life keeps no record but one whose upsert raised as stale: a
  write in flight across a deletion can land after a round found the
  incarnation empty, the race the tombstone retention closes, and the
  test runs with no retention.

Searching without the incarnation filter, deriving point ids from the
record UUID alone, misclassifying REST's 409 on an existing native
collection, a purge round that never deletes, and a purge round that
never returns each fail it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Make the Qdrant creation race tests race, and drop a dead index check

The two-workers test checked that the native collection ended up
indexed, but only the reservation's winner prepares storage, on a
native collection no one else touches, so that check could not fail
the way its docstring said; it is gone, and the docstring says what the
test does. Its creators now reserve the name together, behind a
barrier, so the loser always loses at the reservation rather than
finding the winner already pending or live. It also writes a record
through one handle and reads it through the other, so agreeing on one
collection is shown by state. An open-or-create that raises on losing
the reservation, and a strict create that treats a taken name as
created, each fail it.

A new test races what the old one meant to: two workers create two
collections of one namespace and configuration at once, both
preparations held at a barrier until both start, so both create the
shared native collection together and one finds it already there. Both
creations must succeed and each collection must hold its own records,
over REST and over gRPC. Checking whether the native collection exists
before creating it, without guarding the create, fails it on both
transports while every sequential test passes; so does misclassifying
the transport's already-exists answer.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Keep a Qdrant match that scores exactly the score threshold

Qdrant rounds a query's score_threshold to single precision and keeps only
scores strictly better than it, on every metric and transport, so a match
scoring exactly the threshold was dropped, even a perfect cosine match under
a threshold of 1.0. The SQLite stores and Milvus keep it, and the contract
did not say which.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the adjacent single-precision value on the worse side
(numpy's nextafter in float32), so Qdrant keeps a score equal to the
threshold, and checks the caller's threshold on the returned scores, so a
score Qdrant's rounding let through but that falls short of the threshold as
returned is dropped. A threshold beyond single precision is not sent; the
store's check applies it.

Sending the threshold, rather than only checking it in Python, keeps Qdrant
from returning matches the threshold cuts. Measured against Qdrant 1.19.1
(4 collections of 5,000 768-dimension points sharing one native collection,
300 queries per cell with the two variants alternated, a threshold passing a
quarter of the limit; AC power), p50 per query, threshold sent vs checked in
Python only: REST 2.89 vs 3.04 ms at limit 10, 2.85 vs 2.96 at 40, 3.13 vs
3.42 at 100; gRPC 2.86 vs 2.93 ms at 10, 3.36 vs 3.80 at 40, 3.33 vs 4.23 at
100. Both returned identical matches.

The new test, on every metric and on REST and gRPC, sets the threshold to a
match's reported score (the match is kept) and one step better (it is
dropped). It fails with the threshold sent unchanged, stepped the wrong way,
without the store's check, or with that check strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that the SQLite stores keep a match scoring exactly the threshold

The contract says a match scoring exactly the score threshold is returned.
Both SQLite stores already keep it; each now has the test the Qdrant store
has, on every metric it supports (cosine, dot and Euclidean on the USearch
store; cosine and Euclidean on sqlite-vec): a threshold equal to a match's
reported score keeps the match, and one a step better drops it. Each fails
with its store's threshold check made strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use serial commas in the Qdrant document, and shorten the threshold comment

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
@edwinyyyu edwinyyyu mentioned this pull request Oct 7, 2026
26 tasks
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
MemMachine#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit that referenced this pull request Oct 9, 2026
…n registry, against a Milvus server (#1736)

* Move the Milvus store onto the collection registry, against a Milvus server

The Milvus store kept its catalog in a `memmachine_<namespace>__registry`
collection per namespace, whose `insert` does not enforce primary-key
uniqueness, so two creators of one name both succeeded, and it serialized
creating and deleting a collection with locks in one process. A
collection's name was the partition key on its entities, so a handle kept
writing into a collection deleted and created again under its name, and a
write in flight during a deletion outlived it. Each call ran the sync
client on a thread of the event loop's default executor, so calls waiting
on Milvus queued every other `to_thread` call of the process behind them.

`MilvusVectorStore` is now a `RegistryBackedVectorStore`, and any number
of processes sharing its registry may serve its collections. The store
was rewritten rather than adapted; what changes:
- The catalog is the collection registry in the relational database
  `MilvusConf.collection_registry` names, built and started by the
  database manager before it opens the client. The registry collections
  and the process-local locks go.
- Every entity carries its collection's incarnation as the partition key,
  with partition-key isolation, and in its primary key,
  `"{incarnation}:{record_uuid}"`.
- Declared properties are typed, indexed fields (`_p_<name>`, a datetime
  as TIMESTAMPTZ with its UTC offset beside it in `_tz_<name>`), and
  undeclared ones live in the JSON field, still filterable. Negation is
  the complement, as on Qdrant: a negated condition holds where the
  property has no value.
- The vector index is HNSW_SQ with 4-bit codes and FP16 refinement, what
  AUTOINDEX builds on CPU from Milvus 2.6.10, named so every server builds
  the same; a search rescores `limit x 8` candidates. Scores are the
  server's.
- Reads run at Milvus's default consistency level, Bounded, which the
  store states: a query reflects every write made at least the server's
  `common.gracefulTime` before it. `MilvusConf.consistency_level` goes.
- The store calls Milvus through pymilvus's `AsyncMilvusClient`, and
  bounds every request by `MilvusConf.request_timeout_seconds`.
- A purge round lists a batch of the incarnation's primary keys, at most
  `MilvusConf.purge_batch_size`, and deletes them.
  `MilvusConf.tombstone_retention_seconds` is the retention, refused below
  10 x `request_timeout_seconds` + 300 seconds as for Qdrant. A declared
  string's VARCHAR length is `MilvusConf.max_varchar_length`. Limits the
  server configures stay the server's.
- A delete raises unless Milvus accepted every key sent.

Milvus Lite is dropped. It is a separate embedded engine that scores,
indexes and enforces collection properties differently, so a store tested
against it is not tested against what production runs; MemMachine's
local, single-node backend is the SQLite vector store. `MilvusConf.uri`
defaults to `http://localhost:19530`, a URI with no scheme, which pymilvus
reads as a Lite file, is refused, and the milvus extra no longer installs
milvus-lite (the lock drops it with the packages only it required).

The store's tests, including the collection lifecycle contract, run
against a Milvus 2.6.24 server container as integration tests, and read
past the store at Strong to check what Milvus holds.

Breaking: existing Milvus data is orphaned, and an existing native Milvus
collection has to be dropped; no migration is included. Every Milvus
store needs a relational database for its registry, which the samples,
the configuration wizard and the configuration docs name.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Document how the Milvus store meets the shared contracts

`design/milvus_vector_store.md` records the Milvus store's layout: the
shared native collection and partition-key tenancy, the composite key,
the index and why it was chosen, declared properties as typed fields, the
purge in batches, consistency levels, and the async client, with the
measurements behind each choice.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Build Milvus handles from live registrations

As for Qdrant: the Milvus handle takes its live registration in place of
the namespace, name, incarnation, configuration and registry lookup, and
_build_collection_handle builds it from one. The test that builds a
handle on a mocked client gives it a live registration that stays
current.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Take a Registration in the Milvus handle

#1734 renames the registry's LiveRegistration to Registration. Names only.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say to drop the Milvus collections before upgrading

The Milvus store keeps the native collection's name from main, so on
an upgraded server it finds main's collection, whose schema it cannot
prepare, and the first request of every new session fails. The user
docs' upgrade note now names Milvus beside Qdrant.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that Milvus queries and deletes stay within their collection

Two collections of one namespace and configuration share a native
collection. test_a_query_returns_only_its_own_collections_records gives
both records that match the same filters and checks that each answers
with its own records alone, unfiltered, on a declared property, on an
undeclared one, and under a negation. It fails when _query lets the
property filter replace the incarnation filter, ORs the two, or drops
the incarnation filter: Milvus 2.6.24 refuses each of those searches
under partition key isolation, and with isolation off as well the query
returns the other collection's records.

test_same_uuid_can_exist_in_different_logical_collections now also
deletes the shared UUID through one collection and checks that the
other still holds and returns its record. It fails when _delete
deletes by a record_uuid filter in place of the composite key.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test the Milvus score threshold and ranking for every metric

test_a_threshold_keeps_the_matches_within_it_best_first upserts three
records at increasing distance from the query, out of order, and
queries with a threshold between the middle and the farthest: the two
nearest match, best first, with the metric's scores. For Euclidean the
distances are 1, 2 and 5 and the threshold 3, between the middle
distance and its square; dot product and cosine take their own
thresholds.

It fails when the threshold is applied to Milvus's squared Euclidean
distance before the square root, when every metric keeps scores at or
above the threshold, when every metric keeps scores at or below it,
when the matches are sorted in reverse, and when every metric's score
is the square root of Milvus's distance.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that the Milvus store reads at Bounded or stronger

test_the_store_reads_at_bounded_or_stronger creates a native collection,
writes, queries, deletes the collection and drains the purge, then
checks that describe_collection reports Bounded and that no search,
query, get or hybrid search the store made names a level other than
Bounded or Strong. The purge's listing and the tombstone retention rely
on reads lagging writes by at most common.gracefulTime.

It fails when the native collection is created at Session or
Eventually, when the store's search names Eventually, and when the
purge round's listing names Session.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check random Milvus filters against the in-memory evaluator

test_random_filters_agree_with_the_model writes 48 records whose
properties cover every type, declared and undeclared, with datetimes at
offsets from -12:00 to +14:00, instants a microsecond and a second
apart, strings with quotes, backslashes and non-ASCII letters, and each
property missing a quarter of the time. 200 seeded filter trees of
Comparison, In, IsNull, And, Or and Not, with values of their
properties' types, must select the UUIDs evaluate_filter selects, after
one settle. The trees compare with != only through Not(=): on a missing
property the store's != holds, as the complement of =, while
evaluate_filter's does not.

It fails when a declared datetime is written without its offset, when
a negated condition no longer holds on a missing property, when
negation is not pushed through And and Or by De Morgan's laws, and when
undeclared datetimes are encoded at their own offset in place of UTC.

Dropping the store's own UTC normalization of datetime filter literals
is an equivalent mutant: Comparison and In normalize datetime values to
UTC when they are built, so the store's call never changes a value.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Check a random sequence of Milvus operations against a model

test_a_random_sequence_of_operations_agrees_with_a_model runs 60 seeded
steps on two collections sharing a native collection: upserts of new
and replacing records, deletes of present and absent records, deletion
and recreation of a collection, which then reuses its UUIDs, and purge
drains. Twelve UUIDs serve both collections and every life of each.
After every step each collection holds exactly the model's records,
once each, read at Strong; its settled queries, unfiltered and under
random filters, select what evaluate_filter selects; a replaced
record's old property values no longer select it; and after a drain no
deleted life keeps a record. It fails when _upsert inserts in place of
upserting, when the primary key leaves out the incarnation, and, by
timing out its bounded drain, when the purge round never deletes.

test_upsert_removes_stale_filter_fields now also checks that a filter
on the replaced value finds nothing, which inserting in place of
upserting fails. test_upsert_calls_native_upsert compared the entity
the store built with itself; the model test checks the state an upsert
leaves instead, so it is removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test Milvus purge rounds on a dropped collection and on two purgers

The store fixture now builds its store with _started_store over a
per-test registry database, and other_store builds one as another
process would beside it: its own Milvus client and registry
connections, over the same Milvus and registry database.

test_a_purge_round_on_a_dropped_native_collection_retires_the_tombstone
deletes a collection, drops its native collection, and drains the
purge: the round raises nothing, and the next call finds nothing due.
It fails when _purge_round no longer checks that the native collection
exists, since Milvus refuses the listing.

test_purgers_on_two_stores_reclaim_the_deleted_collections_alone
deletes four collections of seven records each beside a live one in the
same native collection, and drains both stores' purges at once in
batches of 3. Every deleted collection ends empty, the live one keeps
its seven records, and neither store finds anything due. It fails when
a round reports finding nothing after deleting its batch, and when the
listing leaves out the incarnation filter.

test_a_purge_round_milvus_does_not_accept_in_full_raises gives a store
a mocked client whose delete accepts one of the two keys the round
listed: the purge raises. It fails when _purge_round no longer checks
the count. test_a_delete_milvus_does_not_accept_in_full_raises now has
Milvus accept one of two keys, so a check for zero alone fails it, and
matches the exception's type rather than its wording.

The drain in these tests is bounded, so a purge that never ends fails
the test instead of hanging it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test concurrent preparations of one Milvus native collection

Milvus 2.6.24 answers two or more identical create requests for one
collection, sent at once from separate clients, with success every
time (four clients, five trials), and identical concurrent index
requests likewise; a create whose schema differs gets code 1100,
"create duplicate collection with different parameters". So a real
server cannot make a racing creation take the store's already-exists
branch.

test_a_creation_that_loses_the_native_collection_to_another_completes_it
covers that branch with a mocked client: the native collection is
missing when checked and its create is refused as already existing.
The creation indexes and loads the native collection, and the
collection opens. It fails when the already-exists tolerance is
dropped, and when a creation that loses the create returns without
indexing and loading.

test_a_creation_finding_the_native_collection_mid_creation_completes_it
runs two stores on separate clients and registry connections. The
first creates the native collection and is held before its indexes
until the second's creation of another collection, which finds the
native collection present, has returned and been written to and
queried. Both collections end usable with every index. It fails when a
creation that finds the native collection present skips its index and
load steps.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test Milvus settings with values other than their defaults

The configuration, database manager and store tests set the request
timeout, the declared-string length and the purge batch size to values
other than their defaults, so a default passed in place of the
configured value fails them.

- test_every_request_carries_the_timeout runs at a 17-second timeout.
  It fails when the store or the handles it builds use 30 in place of
  the configured timeout.
- test_milvus_client_kwargs_forwarded expects the configured 7-second
  timeout on the client, and test_milvus_creates_vector_store expects
  7, 2048 and 500 in the store's parameters. Each fails when
  async_get_milvus_client passes the default for that setting.
- test_parse_valid_storage_dict parses max_varchar_length 2048 and
  purge_batch_size 500 from the configuration. It fails when either
  field reads another key, so the configured value is ignored.

test_milvus_conf_rejects_a_length_or_batch_size_that_is_not_positive and
test_a_setting_that_is_not_positive_is_refused check that MilvusConf
and MilvusVectorStoreParams refuse 0 for each of these settings. Each
fails when that field's positivity constraint is dropped.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test concurrent churn across two Milvus stores with purgers running

test_churn_across_two_stores_with_purgers_keeps_every_collection_exact
runs six workers, three on each of two stores with their own clients
and registry connections over one registry database, for 40 seeded
steps each on three collections sharing a native collection: upserts
and deletes of records the worker alone owns, owner-filtered queries,
and deletion and recreation of the shared collections. Each store's
purger drains, in batches of 2, after every collection deletion, and
each purge round begins at a settled read, as the tombstone retention
would give it. A ledger records, by the incarnation each operation
reached, the records upserts and deletes that returned leave, the
records sent, and the upserts that raced their collection's deletion.

Only the documented outcomes of a lost race may be raised, and the
whole churn is bounded, so a deadlock fails it. During the churn, a
query returns only records its owner sent to its incarnation. Once the
churn is quiet, settled and drained, each live collection holds
exactly its ledger's records, once each, with their latest values, and
its settled owner and tag filters select them; the deleted
collections keep only records of upserts that raced their deletion.
In three runs the churn made 16 to 18 collection lives, 7 to 10
upserts raced a deletion, and queries, deletes and recreations met
stale handles, existing names and deleted reservations.

It fails when _upsert inserts in place of upserting, when the primary
key leaves out the incarnation, when the purge listing leaves out the
incarnation filter, when a round reports finding nothing after its
batch, and, by timing out, when the purge round never deletes.

The two-purgers test now builds its batched stores with the same
_with_purge_batch_size helper.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Give the Milvus wiring tests' collection registry a file database

The Milvus wiring tests built their collection registry on the module's
in-memory SQLite configuration, which the Qdrant wiring tests' change below
replaced with a file per test, since the registry refuses an in-memory or
single-connection engine. Each Milvus wiring test now puts the registry in a
SQLite file under its own tmp_path, as the Qdrant ones do. What each test
asserts is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Test that the Milvus store keeps a match scoring exactly the threshold

The contract says a match scoring exactly the score threshold is returned,
and the Milvus store already keeps it. It now has the test the Qdrant and
SQLite stores have, on cosine, dot and Euclidean: a threshold equal to a
match's reported score keeps the match, and one a step better drops it. It
fails with the store's threshold check made strict.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Use serial commas in the Milvus document and the store's comments

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the Milvus filter helpers for what they produce

The module's expression helpers had names that did not say what they make:
`_expr_string`, `_literal`, `_declared_literal`, `_fits`, `_absent`,
`_condition`, and `_milvus_filter` are now `_expression_string_literal`,
`_property_value_literal`, `_declared_property_value_literal`,
`_is_comparable_with_declared_type`, `_property_absent_expression`,
`_condition_expression`, and `_filter_expression`; the handle's `_score` is
`_score_from_distance`. No behavior changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Send Milvus filter strings as UTF-8, not ASCII escapes

json.dumps escapes a character outside the Basic Multilingual Plane as a
UTF-16 surrogate pair, which Milvus's expression parser refuses, so a
filter on a string such as "café 😀" failed to parse on a declared and an
undeclared property alike. The literal now keeps its characters as UTF-8.

The new test filters on such a string on both kinds of property; it fails
with the ASCII escaping.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Compare an undeclared Milvus property only with values of a comparable type

An undeclared property's condition compared only its stored value, so a
string equal to an undeclared datetime's stored text matched it, where the
declared path, the other stores' JSON properties, and the in-memory
evaluator match nothing. The condition now also requires the stored type
tag to be one the value compares with, by the rule the declared path
already uses: a bool with a bool, an int or float with either, and any other
value with its own type.

The new test filters an undeclared datetime by its stored text, and an
undeclared float by an int; it fails without the type check.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Index and load a part-made Milvus collection before a purge round lists it

A creation that failed after creating the native collection, before its
indexes or its load, leaves it unindexed and unloaded, and its cancelled
collection's tombstone is due for purge. The round's listing query then
raised "collection not loaded" on every claim, so the tombstone was
dead-lettered unless another creation of the same namespace and
configuration completed the collection first. A round now creates the
missing indexes and loads the collection, the steps preparation already
takes, before it lists.

The new test fails a creation at its index step, then drains the purge: it
fails with "collection not loaded" without the change.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the Milvus store's already-exists guard, which no supported server reaches

Preparation caught a create refused as "already exist" by matching the
message text. Milvus answers a create of an existing collection with the
same schema with success, and refuses one with another schema as "create
duplicate collection with different parameters" (code 1100, Milvus
2.6.24), which the text did not match, so the branch ran only under the
mocked client of its own test. The guard, its message match, and that test
go; the concurrent preparation tests cover racing creators against a real
server.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse a metric Milvus does not support before reserving the collection

The store checked the metric while preparing storage, after the registry
had reserved the name, so a refused configuration cost a reservation and
its cancellation, and stayed pending until deleted if the cancellation
failed. create_collection and open_or_create_collection now check it first.

The test, for both operations, finds nothing to purge after the refusal; it
fails with the check back in preparation.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Bound the Milvus connectivity check by the request timeout

validate_milvus_client listed collections with no timeout, and the
client's own timeout bounds only connecting, so a Milvus that accepts the
connection and never answers hung validation, the one request
request_timeout_seconds did not bound. The check now passes the configured
timeout, on a store's first build and in build_all's validation.

The new test checks both paths pass it; it fails with the timeout dropped.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Require testcontainers 4.15.0, the first with the community modules

The shared conftest imports MilvusContainer from testcontainers.community,
which 4.14.2, the declared floor, does not have, so an install at the floor
could not import the conftest or collect any server test. The lock already
pins 4.15.0; only the floor moves.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Dispose a mocked-client purge test's registry engine when it fails

The test disposed the registry engine it opened only after its last
assertion, so a failure left the engine and its aiosqlite thread open for
the rest of the session.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop the Milvus store's re-sort of search results

The store sorted each query's matches by score, though Milvus returns hits
best first and the one transform the store applies, the square root of a
Euclidean distance, keeps their order. On Milvus 2.6.24, 3,000 random
vectors in each metric gave 450 of 450 result lists in order, filtered and
not, at limit 100.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Say that a Milvus native collection keeps its VARCHAR length

The store sizes a declared string's VARCHAR field when it creates the
native collection, and a later preparation that finds the collection there
leaves it as it is, so a changed max_varchar_length applies only to native
collections created afterward. The setting's description, the
configuration docs, and the design document now say so.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Match no declared Milvus int property with a float filter value

The declared path counted an int and a float as comparable either way, so
a float filter value on a declared int property reached Milvus as a float
literal against an INT64 field, which Milvus refuses to parse ("cannot
cast value to Int64", code 1100). A float now compares only with a float
property, and matches no int one, as a value of another type matches
nothing; an int still compares with a float property by value.

The new test filters a declared float property with an int and a declared
int property with a float; it fails with the float counted as comparable
with an int property.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Create Milvus native collections at Bounded by name

The store's reads run at the collection's consistency level, and the
purge's listing and the tombstone retention rely on that level being
Bounded or stronger. The store named no level when it created a native
collection, so the level was whatever pymilvus defaults to, which a
pymilvus release could change without the store noticing. The store now
creates each native collection at Bounded.

The test checks the create request names Bounded; it fails with the
keyword removed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Report Milvus store operation latencies to a metrics factory

The database manager built the Milvus store's parameters without a
metrics factory, and MilvusConf had no metrics_factory_id, so a
deployment could not configure one and every timed Milvus operation
recorded nothing, where the Qdrant and Neo4j stores report theirs.
MilvusConf takes metrics_factory_id, and the manager passes the factory
it names, as it does for Qdrant.

The wiring test checks the factory reaches the store's parameters; it
fails with the factory not passed.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name Milvus's cap on a record's undeclared properties

A record's undeclared properties are stored in one JSON field, and Milvus
refuses an upsert whose JSON exceeds its common.JSONMaxLength setting,
65,536 bytes by default (internal/proxy/validate_util.go and
pkg/util/paramtable/component_param.go at v2.6.24). The design document's
list of server-configured limits and the database docs now name it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse a Milvus Lite URI by its .db suffix, as pymilvus selects Lite

The configuration refused a URI with no "://" as a Milvus Lite file. pymilvus
serves a URI ending in .db with Lite and passes a unix: URI to the server as
a socket address, so a file:// URI ending in .db passed the check and reached
Lite, and a unix: socket URI was refused as a Lite file. The check now refuses
a URI ending in .db, and its message names the Lite file.

The tests refuse ./milvus.db, milvus.db, and a file:// URI ending in .db, and
accept http, https, and unix: URIs; the file:// and unix: cases fail with the
"://" check.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Write a declared Milvus datetime as its UTC instant

The entity builder wrote a declared datetime's text with its own offset,
while both filter literals convert to UTC first. Milvus refuses a
TIMESTAMPTZ string whose offset has a seconds component ("invalid timezone
name", code 1100 on 2.6.24), as local mean time and other historical
offsets have, so such a record could not be upserted. The builder now
writes the UTC instant; the offset field already keeps the offset written.

The new test upserts a datetime at +00:19:32, reads back its instant and
offset, and matches it by its instant; it fails with the original offset
written, at Milvus's refusal.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Halve a Milvus upsert refused as too large, as the Qdrant store does

Milvus's proxy refuses a request over proxy.grpc.serverMaxRecvSize (64 MiB
unless configured) with gRPC's RESOURCE_EXHAUSTED status, before writing any
of it, and pymilvus raises that status as it is: on 2.6.24 a 72 MB upsert was
refused in 0.2 s with nothing stored. So a batch the Qdrant store accepts by
halving it failed outright on Milvus. The store now halves an upsert refused
with RESOURCE_EXHAUSTED until its halves fit or a single entity is refused.
Any other error raises at once, since a timed-out request may still be
applied.

The store imports grpc for the status, so the milvus extra declares grpcio:
1.59.0 is the first release with Python 3.12 wheels, and grpc.aio.AioRpcError
and grpc.StatusCode are older.

The integration test upserts 1,200 records with a 60,000-character property,
about 72 MB, and fails with RESOURCE_EXHAUSTED without the halving. The
mocked-client tests pin the halving, which its test fails without, a single
refused entity raising, and a timeout or another refusal sent once.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name Milvus's datetime offset fields for their unit

A declared datetime's UTC offset was stored in an INT32 field named
_tz_<key>, which left its unit to the code. The field is now
_tz_offset_seconds_<key>, under the prefix _TZ_OFFSET_SECONDS_FIELD_PREFIX;
its value is unchanged, the offset in seconds. The schema test and the two
datetime round-trip tests read the new names.

A property key is at most 32 bytes, so a declared property's field names are
at most 51 characters, within Milvus's proxy.maxNameLength (255 unless
configured). A new test declares a datetime under a 32-byte key, writes it,
reads both fields back, and matches it by its instant; the design document
states the bound.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Keep a declared Milvus datetime as its UTC instant alone

A declared datetime was stored as its UTC instant in a TIMESTAMPTZ field and
its UTC offset in an INT32 field beside it. A filter compares only instants,
and a search answers with record UUIDs and scores, so nothing reads the
offset back; the store the vector records are derived from keeps it. The
offset fields go: every declared property, a datetime included, is one
field, `_p_<key>`, of those proxy.maxFieldNum allows. Undeclared properties
keep their type-tagged JSON, offset included.

The schema test asserts the native collection's exact field set, the
datetime round-trip tests check that the stored instant is the written one,
in UTC, and the longest-key test checks the one field.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…er, live from resolve

The registry is addressed by partition key only. `register(key, schema)`
answers a PendingRegistration, whose `mark_live()` answers the same life
as a LiveRegistration or raises VectorStorePartitionDeletedError when the
partition was deleted meanwhile, and whose `unregister()` abandons that
life alone. `resolve(key)` answers the LiveRegistration of the live
partition, None when there is none, and raises
VectorStorePartitionPendingError, which now carries the pending
partition's schema, when it is pending. A LiveRegistration's
`require_current()` raises VectorStorePartitionHandleStaleError once its
partition is deleted. `get`, `mark_live(incarnation)`,
`unregister_incarnation` and RegisteredPartition are gone, so deletion is
by key or by a pending registration, never through a live one. The
SQLAlchemy registry supplies frozen-dataclass registrations holding the
engine and the vector store name, and a module function runs the
unregistration transaction.

The base handle takes its live registration, fences on
`require_current()`, and `_partition_handle` builds a handle from one.
`create_partition` lets VectorStorePartitionDeletedError propagate, and
the VectorStore interface names it; open-or-create catches it and creates
again, refuses a pending partition of another schema at once from the
pending error's schema, and re-raises the last pending error. A
`get_partition` of a pending partition of another schema still reports the
mismatch first. The Qdrant and Milvus handles take the registration in
place of the key, the incarnation and the registry lookup.

Tests follow: the registry's tests answer registrations and gain one for
`require_current`; the base's tests patch the pending registration type's
`unregister` and expect the deleted error from a creation a deletion
undid; the lifecycle contract patches the live registration type's
`require_current`, registers its racing winners through pending
registrations, and counts the deleted error among churn's outcomes; the
Qdrant and Milvus tests that build a handle on a mocked client give it a
registration that stays current.

The same change as MemMachine#1734's 735c649, f88aaa8, e25be60 and
b2bb2d8 and the handle halves of MemMachine#1735's 9096edd and MemMachine#1736's
5582cf4, for this PR's partition registry. The event-backend locator
change has no counterpart: the locator here opens its partition with
open-or-create, which creates again itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…ain open-or-create's give-up

Two review changes to MemMachine#1734's registry and base store, for this PR's
partitions:

- `Reservation.cancel()` acts only while the partition is pending, as
  `confirm()` does. A creator whose confirmation committed but whose
  answer was lost, and which then cancels, leaves the live partition
  alone: only a deletion by key ends it. The registry contract, the
  SQLAlchemy registry and the registry design document say so, and a
  test cancels after a confirmation.
- Open-or-create's `VectorStoreAttemptsExhaustedError` is raised from the
  race it last lost, the last `VectorStorePartitionAlreadyExistsError` or
  `VectorStorePartitionDeletedError` it caught. A partition that stays
  pending still raises the pending error itself. A base-store test checks
  the cause.

The same changes as MemMachine#1734's 9daf7e7 and 93a85da, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…r is cancelled

A creation cancelled its reservation only when preparing the partition's
storage raised. A confirmation that raised, or a creation cancelled while
it confirmed, left the partition pending until someone deleted it.

The base store's creation step now prepares the storage and confirms the
reservation together, and cancels the reservation if either raises or the
creation is cancelled, shielded as before; the cancel's task reports its
own failure. The cancel acts only on a pending partition, so a
confirmation that committed before its failure was observed stands.
create_partition and open-or-create both go through it; open-or-create
still takes a confirmation's VectorStorePartitionDeletedError as a race to
create again. Base-store tests cover a failed confirmation, a cancelled
one, and one that committed before failing. The registry design document
says so, and the purge document says which writers wait on SQLite's lock
during a purge round.

The same changes as MemMachine#1734's f4c585e and 4dde7c8, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
A round that never ended was counted by the next claim, which took a
CASE-shaped update that both recorded the lost round and ended its claim;
the count lived apart from the claims that made it.

Each claim now counts its attempt: the claim is a plain update that
increments `consecutive_attempts` (renamed from `consecutive_failed_rounds`),
stamps `claimed_at`, and bumps the generation, and `_TombstoneClaim.attempt`
carries the count. A tombstone is claimable when it has no attempts, when no
claim is open and the backoff has passed since its last failure, or when an
open claim has outlived its lease plus the backoff. A raised round ends its
claim and stamps `last_failed_at`; a cancelled round ends its claim and takes
its attempt back; a round that found records resets the attempts. After
`_MAX_PURGE_ATTEMPTS` attempts a tombstone is dead-lettered, and a last
attempt that raises is reported. A retry logs which attempt it is, and a
raised round's error names its incarnation and attempt. The registry tests,
the purge design document, and the registry design document follow.

The same change as MemMachine#1734's 1732038, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
`consecutive_attempts` counted the purge rounds claimed since a round last
found records, but its name said only that they were in a row. It is
`attempts_without_progress`, and `_MAX_PURGE_ATTEMPTS` is
`_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The column's comment, the claim's
attempt, the backoff parameter's description, the dead-letter log, the
class docstring, the tests, the purge design document, and the registry
design document's table follow; nothing else changes.

The same change as MemMachine#1734's 96a8b5b, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…urrent claim

The measured cost of the claim came from earlier forms of it, which held a
transaction across the round. The section now gives the figures measured on
2026-10-06 against the claim this registry ports, naming the MemMachine#1734 commits
they were taken at: the backoff scan with 1k, 10k, and 100k tombstones
backing off, and interactive throughput and liveness latency beside two
sweepers on PostgreSQL and on SQLite.

The same change as MemMachine#1734's bab379b, for this PR's purge document.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

horizontal scaling Wrong or unsafe when more than one server process serves the same backends (replicas or workers)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants