Skip to content

[sqlite store fixes 1/7] Publish vector index files atomically (but not durably) - #1460

Open
edwinyyyu wants to merge 29 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:atomic_index_swap
Open

edwinyyyu wants to merge 29 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:atomic_index_swap

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Jun 18, 2026 •

Copy link
Copy Markdown
Contributor

Problem

SQLiteVectorStore persists each collection's index by calling the search engine's save(), which writes directly to the final path. A crash partway through leaves a truncated file on disk.

The torn file is only half of it. _save_collection_index trims the applied _PendingOperationRows as soon as save() returns, and that log is the only other copy of those vectors — the records table has no vector column. So the trim is safe at exactly one instant: when the index at the published path holds them. Getting that instant wrong is silent, because a missing vector is indistinguishable from a vector that simply doesn't match.

Fix

A shared index_persistence helper writes the index to a sibling temp file, flushes it, and swaps it into place with os.replace; both engines adopt it and neither gains anything else. load clears a temp left by an interrupted save, so a crash leaks at most one stale file per index.

After this, a save either publishes a complete index or leaves the previous one exactly where it was. Nothing else changes: no schema change, no migration, no rename of index_path, and no new file at the base path.

What a save does not promise

It does not promise that the publication is durable. That is a decision rather than an oversight, so the reasoning is below.

A rename changes a directory entry, not file data, and the two have very different guarantees:

make it durable available
file contents fsync / F_FULLFSYNC / FlushFileBuffers everywhere, documented
directory entry fsync on a directory file descriptor POSIX only, and not on every filesystem
  • POSIX closes the gap with a parent-directory fsync, but only best-effort — network and FUSE filesystems are where that bites.
  • Windows has no equivalent at all. os.fsync is _commit, which is FlushFileBuffers, which is for file data; you cannot open a directory to fsync it. MOVEFILE_WRITE_THROUGH does not help either: its documented guarantee is scoped to "a move performed as a copy and delete operation", the cross-volume path, and says nothing about flushing NTFS metadata for a same-volume rename. The decisive evidence is SQLite's — its VFS threads a directory-sync flag through every commit-relevant directory operation, unixDelete honors it, and winDelete declares the same parameter /* Not used on win32 */. SQLite cares about this more than almost any software and makes no attempt on Windows.

So a power failure can roll the rename back after save returned, while the trim behind it stays committed.

Why that gap stays open

Closing it means never using a directory operation as the commit point: two index slots plus a generation record written into a file that already exists, in the spirit of SQLite's PERSIST journal mode. An earlier revision of this PR implemented exactly that, and it worked. It was still the wrong trade, for four reasons.

The window is narrow, and only one kind of failure lands in it. A clean shutdown, an unhandled exception, an OOM kill, a SIGKILL — none of these lose a rename, because the kernel still owns the page cache and writes it back. Those failures are already covered: the pending log holds everything the engine has not been checkpointed with, and startup replays it. What is left is the machine itself dying between the rename and writeback — a power cut, a host failure, a panic — plus the filesystems where a directory fsync is best-effort or unavailable anyway.

What it costs is recall, not correctness. The records table is the authority on what exists, and every query hit resolves through it, so a reverted publication cannot produce a wrong or stale result — only fewer results. Concretely: the records survive, get still returns them, query stops finding them, and the exposure is bounded by the operations applied since the previous checkpoint, i.e. at most save_threshold. Re-upserting an affected record repairs it. That is the direction this store already tolerates, and it is repairable by the same ingest path that created the record.

Detecting it is not cheap enough to be worth it. Comparing the index's size against the record count is the obvious check and it does not work: row_ids are not reused (#1469 makes that unconditional; before it only the highest row's id can come back), so deleting a record and upserting the same uuid again moves it to a new id. A rollback spanning that pair leaves the index holding the old id and missing the new one — one extra, one missing, identical count — and a re-embedding workload produces that shape constantly. A check that does work needs an id-set digest maintained by every engine, or a full reconcile scan at every boot, which is most expensive precisely where the index is large. Neither earns its keep against a bounded recall gap.

And the guarantee is not free to hold. It cannot be delegated to a rename, so it has to become a layout: two slots and a generation record per index, which every engine has to implement and every future engine has to be audited against. Engines that persist by writing one file — which is what both engines here do, and what most ANN libraries expose — stop being usable as they ship. Existing indexes stop loading until they are cleared and re-ingested, which is a migration this revision no longer asks anyone to perform. Fewer guarantees, less machinery to keep correct, and a wider set of engines that can plug in unmodified.

The direction that would cost — a published index that will not parse — stays closed: the atomic swap prevents it, the pre-swap fsync narrows the window where a rename outlives the data behind it, and a saved-but-unloadable index remains a loud IndexLoadError rather than a silently empty rebuild.

Nothing here detects a lost publication for the caller. A deployment that needs every record searchable after a power failure must be able to re-ingest.

Where the contract is stated

The guarantee is only useful if callers can find it, so it is written where each of them looks:

  • VectorSearchEngine.save — returning means a later load reads this index or the one it replaced, never a mixture; it does not mean the publication survives a power failure.
  • index_persistence — the mechanism, and why a stronger guarantee is not portable.
  • sqlite_vector_store — what survives a process crash, what a power failure costs, and that re-ingest is the repair.
  • _save_collection_index — why the save-then-trim order is the whole protocol.

One consumer needed more than a docstring. VectorStoreSemanticStorage.update_feature reads a feature's stored embedding back when the caller updates the feature without supplying one — the only place in the server that depends on the index still holding a vector. It reported a lost vector as Vector record not found, naming the feature, which points at the wrong thing and hides the repair; it now separates a record that is genuinely absent from one whose embedding is gone, and the second says to pass a fresh embedding.

Tests

  • test_index_persistence.py: the swap publishes completely or not at all, a failed write leaves the previous index intact and removes the temp, and a stale temp is cleared on load.
  • Per-engine (hnswlib + usearch): a save leaves no temp file behind, and a load clears a stale temp file and still finds every key.
  • Store-level: test_a_reverted_publication_costs_search_not_records reconstructs a lost publication deterministically — restore the previous index bytes after the trim has committed — and pins the direction it fails in: the record still resolves by uuid, and only search loses it. Missing and corrupt indexes each still raise IndexLoadError once index_saved is set.

vector_store suites pass (284 passed, 140 integration deselected) and semantic_memory suites pass (345 passed, 2 skipped, 999 integration deselected); ruff and ty clean.

Changed from the earlier revision

This PR previously shipped the two-slot generation-record protocol and its breaking on-disk change. That guarantee has been dropped in favor of the atomic swap plus an explicit contract, so the breaking-change notice no longer applies: indexes written by the current code keep loading, and there is nothing to migrate.

🤖 Generated with Claude Code

Port

This branch carries this change as merged into speedkick (#1588, 21105d6e6), the one commit cc5311ef9; the branch's earlier commits are replaced. Adaptation: the commit sits on [search results], as the rest of this stage must, though on speedkick #1588 merged before #1598: its engine tests construct the engines without a metric, the reverted-publication test reads the surviving row through the records table since get is gone, the module docstring says no read reaches a stored vector, and its semantic-storage change (reporting a lost embedding apart from a missing record) has nothing to apply to, since #1733 removed the read-back it reworded.

Stack

21 open PRs: three independent PRs, and the vector store tree of short parallel branches. Every PR in the tree has feat/horizontal-scaling as its GitHub base, and the independent PRs have main. The branches are in a fork, and a pull request can target only this repository's branches, so the on column gives the order the PRs build on each other. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main
— #1786 Refuse a filter value of the wrong type for a datetime column, and answer an invalid list filter with 422 (port of #1620) main
— #1792 Refuse property values that some store refuses or alters where an episode enters main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1702 and #1670 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (merged into feat/horizontal-scaling) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs feat/horizontal-scaling
[vector store scale-out 3/6] #1734 (merged into feat/horizontal-scaling) Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge feat/horizontal-scaling
[vector store scale-out 4/6] #1735 (merged into feat/horizontal-scaling) Move the Qdrant store onto the collection registry feat/horizontal-scaling
[vector store scale-out 5/6] #1736 (merged into feat/horizontal-scaling) Move the Milvus store onto the collection registry, against a Milvus server feat/horizontal-scaling
— #1775 (merged into feat/horizontal-scaling) Accept attempts exhausted in the lifecycle churn contract, and say which Qdrant operations filter on the incarnation feat/horizontal-scaling
— #1779 (merged into feat/horizontal-scaling) Classify Qdrant errors by status code alone feat/horizontal-scaling
— #1813 (merged into feat/horizontal-scaling) Create Qdrant collections with strict mode off feat/horizontal-scaling
— #1788 Refuse a repeated record UUID or a non-finite property value at upsert, and state the datetime property contract feat/horizontal-scaling
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1702 Keep undeclared properties out of the vector store feat/horizontal-scaling
[user properties 2/2] #1670 Remove per-project filterable properties (port of #1606) #1702
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1670
[session storage 1/2] #1622 Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1627
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1628
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1663
[sqlite store fixes 1/7] #1460 (this PR) Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[milvus options] #1741 Let a deployment tune a Milvus collection's vector index and its searches #1618

This PR is its one commit, 8b9b271ae, stacked on #1663. #1469 is stacked on it.

Verification

At this PR's head 8b9b271ae, on 2026-10-09, on feat/horizontal-scaling with #1813 merged: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean.

Before the rebase onto #1813's merge: At this PR's head ec03a19be, on 2026-10-09: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean; the SQLite stores' and search engines' tests pass, 233 tests. The full suites last passed at b8f302699, on 2026-10-08, before #1628 moved beneath this PR: the server suite without integration tests, 2108 tests, and the integration tests of the vector stores, the resource manager, episodic memory, and semantic storage, 668 tests, against PostgreSQL 16, Neo4j, Qdrant 1.19.1, and Milvus 2.6.24 in containers..

@edwinyyyu edwinyyyu mentioned this pull request Jul 2, 2026
12 of 26 tasks
@github-actions

github-actions Bot commented Aug 3, 2026

Copy link
Copy Markdown
Contributor

This pull request has been automatically marked as stale because it has not had recent activity. It will be closed in 14 days if no further activity occurs. If you are still working on this, please push a commit or leave a comment. Reviewers: please respond, or add the keep-open label if this PR should be held open for a longer review cycle.

@github-actions github-actions Bot added the Stale label Aug 3, 2026
@edwinyyyu edwinyyyu added the keep-open Prevents the auto-close task from closing this issue. label Aug 3, 2026
@edwinyyyu edwinyyyu changed the title Atomically swap vector search engine index files on save Atomically and durably swap vector search engine index files on save Aug 5, 2026
@edwinyyyu edwinyyyu changed the title Atomically and durably swap vector search engine index files on save Let the engine own index durability, not the vector store Aug 6, 2026
@edwinyyyu edwinyyyu changed the title Let the engine own index durability, not the vector store Publish vector index files atomically, not durably Aug 12, 2026
@edwinyyyu edwinyyyu mentioned this pull request Aug 13, 2026
10 of 15 tasks
@edwinyyyu edwinyyyu changed the title Publish vector index files atomically, not durably Publish vector index files atomically, but not durably Sep 8, 2026
@edwinyyyu edwinyyyu changed the title Publish vector index files atomically, but not durably Publish vector index files atomically (but not durably) Sep 8, 2026
@edwinyyyu edwinyyyu changed the title Publish vector index files atomically (but not durably) [speedkick port 7/17] Publish vector index files atomically (but not durably) (port of #1588) Sep 16, 2026
edwinyyyu and others added 28 commits October 9, 2026 17:04
…ck) (MemMachine#1606)

* Regenerate the OpenAPI document under the locked FastAPI

`docs/openapi.json` predates the FastAPI release in `uv.lock`
(0.141.1), whose `ValidationError` component carries `input` and `ctx`;
regenerating the document with `docs/tools/generate_openapi.py` adds the
two fields and changes nothing else. Separate from the API changes above
it so their diffs of this file show only what they change.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

* Remove per-project filterable properties

A project could declare `properties_schema`, a set of caller property keys
with types, on its long-term memory configuration; the event backend merged
it into the vector store collection's indexed schema and rejected filters on
any other `m.<key>`. That let a tenant create database resources (indexes,
columns) by naming them in a request, which is what forced per-collection
native resources named by a hash of their schema on the backends that limit
them.

The option is removed from the server configuration, the project API and
the memory-configuration API, the Python SDK, the sample configurations,
the configuration docs and the OpenAPI document. A filter may name any
`m.<key>`; the stores evaluate it on the properties they hold. What a store
indexes is decided by the deployment, not per project.

A breaking API change on `speedkick`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Rebased onto MemMachine#1631: the per-project schema also leaves MemMachine#1631's service
locator, which creates the session's collection in a retry loop, and the
commented option goes from the event sample configuration MemMachine#1698 added.

---------

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
(cherry picked from commit a8322a7)
…res' lookup to get_partition

A vector store's logical collection becomes a partition, the segment
store's word for the same thing, and both stores' lookup is get_partition,
answering None like a Python get. Identifiers only, produced by the script
below; the (namespace, name) identity, the per-partition config and every
docstring are as they were, and the next change gives them their meaning.
The native clients' create_collection and delete_collection keep their
names.

The vocabulary also reaches what sits under the rename: the registry
package, its modules, classes and tables, the `collection_registry` option
of a Qdrant or Milvus backend and the store parameter and attribute that
carry it (`partition_registry`), the stale-handle and pending errors, the
registry-backed base and its handle lookup, the purge method
(`purge_deleted_partitions`, the segment store's name for it) and
open-or-create.

```sh
set -e
cd "$(git rev-parse --show-toplevel)"
git mv packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_collection.py \
       packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_partition.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_lifecycle_contract.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_lifecycle_contract.py
git mv packages/server/src/memmachine_server/common/vector_store/collection_registry \
       packages/server/src/memmachine_server/common/vector_store/partition_registry
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/partition_registry.py
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_partition_registry.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_registry \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry
git mv packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_collection_registry.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_partition_registry.py
git ls-files -z 'packages/server/*.py' 'docs/*.mdx' 'sample_configs/*.sample' 'sample_configs/*.yml' 'deployments/helm/templates/*.yaml' | xargs -0 perl -0pi -e '
  s/VectorStoreCollection(?!Config)/VectorStorePartition/g;
  s/in_memory_vector_store_collection/in_memory_vector_store_partition/g;
  s/collection_lifecycle_contract/partition_lifecycle_contract/g;
  s/CollectionLifecycleContract/PartitionLifecycleContract/g;
  s/collection_registry/partition_registry/g;
  s/SQLAlchemyCollectionRegistry/SQLAlchemyPartitionRegistry/g;
  s/CollectionRegistry/PartitionRegistry/g;
  s/RegisteredCollection/RegisteredPartition/g;
  s/get_registered_collection/get_registered_partition/g;
  s/_build_collection_handle/_build_partition_handle/g;
  s/\bCollectionT\b/PartitionT/g;
  s/register_collection\b/register_partition/g;
  s/\b_Collection\b/_Partition/g;
  s/CollectionRow/PartitionRow/g;
  s/vector_store_collection(?!_schema|_namespace)/vector_store_partition/g;
  s/open_or_create_collection/open_or_create_partition/g;
  s/open_collection/get_partition/g;
  s/_purge_deleted_collections_forever/_purge_deleted_vector_store_partitions_forever/g;
  s/purge_deleted_collections/purge_deleted_partitions/g;
  s/def create_collection\(/def create_partition(/g;
  s/def delete_collection\(/def delete_partition(/g;
  s/\.create_collection\((\s*namespace=)/.create_partition($1/g;
  s/\.delete_collection\((\s*namespace=)/.delete_partition($1/g;
  s/\.create_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.create_partition/g;
  s/\.delete_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.delete_partition/g;
  s/"create_collection"/"create_partition"/g;
  s/"delete_collection"/"delete_partition"/g;
  s/"open_or_create_collection"/"open_or_create_partition"/g;
  s/only delete_collection is invoked/only delete_partition is invoked/g;
  s/test_delete_collection_/test_delete_partition_/g;
  s/open_partition/get_partition/g;
  s/(vector_store_partition \(VectorStorePartition\):\n\s+)Vector store collection\./$1Vector store partition./g;
'
uv run ruff check --fix --quiet packages/server
uv run ruff format --quiet packages/server
```

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…uilt by the composition root

A vector store was a factory of logical collections, each identified by
a (namespace, name) pair and created with its own dimensions, metric and
schema; Qdrant and Milvus shared one native collection among logical
collections of equal configuration, under a name derived from a hash of
that configuration, and the registry mapped (namespace, name) to it.

A store is now one collection: `VectorStore(vector_store_name,
vector_dimensions, similarity_metric, indexed_properties)` names its one
native collection (or its tables and index files) at construction, every
partition of it shares the store's dimensions, metric and schema, and
`provision()` creates the store's durable resources idempotently, before
`startup`. `create_partition(key)`, `open_or_create_partition(key)`,
`get_partition(key)` and `delete_partition(key)` take a string key; a
partition is the records carrying its incarnation in the store's native
collection (Qdrant, Milvus) or a pair of tables (the SQLite stores), and
the registry records what each partition was created under, so a store
built with other dimensions, another metric or another schema raises
`VectorStorePartitionSchemaMismatchError` instead of reading columns and
vectors that are not there. "Collection" names only the native Qdrant or
Milvus collection. Vector store names match `[a-z0-9_]+` and are at most
32 bytes, the rule partition and property keys follow, so every name
works on every backend: the Qdrant store names its collection by the
vector store name, and the Milvus store by `sys_` followed by it, since
Milvus requires a leading letter or underscore and names its own
internals with a leading underscore. The hash-derived native names go,
and with them the per-partition config.

The partition registry is keyed by partition key within one store: its
tables, `partition_registry_pt` and `partition_registry_gc`, are shared
by every store on a database and keyed by vector store name, so the
registries of several stores share one database and nothing else.
`mark_live` and `unregister_incarnation` take the incarnation alone. On
the registry-backed base, the storage a store's partitions share is
prepared by `provision()` (`_prepare_storage()`), and the storage a
partition keeps of its own by `_prepare_partition_storage(key,
incarnation)`, between its registration and its mark; the Qdrant and
Milvus stores keep nothing per partition. A purge round works on one
incarnation in the store's own collection.

`DatabaseManager.get_vector_store(backend, vector_store_name=,
vector_dimensions=, similarity_metric=, indexed_properties=)` builds and
caches one store per (backend, vector store name), provisioning its
registry and then the store; asking for a name again with other
dimensions, another metric or other keys is a
`VectorStoreConfigurationError`. The event backend's store is named by a
UUIDv5 of its embedder id in the UUIDv5 of its backend's key in a namespace
fixed for the event backend, written as 32 hexadecimal digits, so neither
the key nor the id is constrained; stores on different backends never share
a name, so their registries stay apart when the backends share a registry
database, since a store name identifies one store among all the stores
whose registries share it.
Semantic memory's one store is named `semantic_memory` whatever its
embedder and holds every org's features in one partition,
`semantic_memory`, as main's one collection does.
The event backend opens a session's partition with
`open_or_create_partition`, which waits on one another worker is still
creating.

The SQLite stores change shape only: their registry tables become
`vector_store_sqlite_pt` and `vector_store_sqlite_vec_pt`, keyed by
vector store name and partition key, the pending-operation log is keyed
the same way, and per-partition table names embed the vector store name.
No migration is provided; existing SQLite vector data is orphaned.

Rebuilt as one change on MemMachine#1631: this PR's earlier history carried copies
of MemMachine#1631's commits as of `38684b385`, its own commits on them, and
mirrors of MemMachine#1631's later commits in its terms. Its code is that history's
final tree, which was verified, merged with main and without MemMachine#1624's
commits. The design documents are still MemMachine#1631's and describe collections.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…up, not a separate provision

A vector store and its partition registry each had a provision() that
created their durable resources, run by the composition root before
startup(). The split between provisioning a store and starting it is
MemMachine#1570's to make for every store at once, so here startup() does both
again, as on main: a store's startup prepares the storage its
partitions share (the native collection and its indexes, or the SQLite
tables), and the registry's startup creates its tables, as MemMachine#1631's does.
The database manager starts the registry, then the store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…er, live from resolve

The registry is addressed by partition key only. `register(key, schema)`
answers a PendingRegistration, whose `mark_live()` answers the same life
as a LiveRegistration or raises VectorStorePartitionDeletedError when the
partition was deleted meanwhile, and whose `unregister()` abandons that
life alone. `resolve(key)` answers the LiveRegistration of the live
partition, None when there is none, and raises
VectorStorePartitionPendingError, which now carries the pending
partition's schema, when it is pending. A LiveRegistration's
`require_current()` raises VectorStorePartitionHandleStaleError once its
partition is deleted. `get`, `mark_live(incarnation)`,
`unregister_incarnation` and RegisteredPartition are gone, so deletion is
by key or by a pending registration, never through a live one. The
SQLAlchemy registry supplies frozen-dataclass registrations holding the
engine and the vector store name, and a module function runs the
unregistration transaction.

The base handle takes its live registration, fences on
`require_current()`, and `_partition_handle` builds a handle from one.
`create_partition` lets VectorStorePartitionDeletedError propagate, and
the VectorStore interface names it; open-or-create catches it and creates
again, refuses a pending partition of another schema at once from the
pending error's schema, and re-raises the last pending error. A
`get_partition` of a pending partition of another schema still reports the
mismatch first. The Qdrant and Milvus handles take the registration in
place of the key, the incarnation and the registry lookup.

Tests follow: the registry's tests answer registrations and gain one for
`require_current`; the base's tests patch the pending registration type's
`unregister` and expect the deleted error from a creation a deletion
undid; the lifecycle contract patches the live registration type's
`require_current`, registers its racing winners through pending
registrations, and counts the deleted error among churn's outcomes; the
Qdrant and Milvus tests that build a handle on a mocked client give it a
registration that stays current.

The same change as MemMachine#1734's 735c649, f88aaa8, e25be60 and
b2bb2d8 and the handle halves of MemMachine#1735's 9096edd and MemMachine#1736's
5582cf4, for this PR's partition registry. The event-backend locator
change has no counterpart: the locator here opens its partition with
open-or-create, which creates again itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nd their partitions

The design documents came from MemMachine#1733-MemMachine#1736, where a store holds logical
collections addressed by namespace and name, each with its own
configuration. Here a store is one collection, with its dimensions, metric
and declared schema fixed at construction, and a partition is one tenant's
records in it, addressed by key. The documents say so:

- the collection registry document becomes the partition registry
  document: a registry belongs to one store and is addressed by partition
  key; its tables are `partition_registry_pt` and `partition_registry_gc`,
  and a tombstone needs no location, since the incarnation alone finds a
  dead partition's records in the store's native collection; a partition
  created under another schema is refused; `startup` creates the tables, as
  every store's startup creates its durable resources; a decision records
  that a store is one collection;
- the Qdrant and Milvus documents lay out one native collection per store,
  named by the vector store name (`sys_` and the name on Milvus), created at
  startup, where a partition's creation makes nothing in the backend;
- the overview, isolation, consistency and purge documents speak of
  partitions, `purge_deleted_partitions` and `settle(partition)`, and the
  overview describes the store and its partitions.

The measurements and their conditions are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ain open-or-create's give-up

Two review changes to MemMachine#1734's registry and base store, for this PR's
partitions:

- `Reservation.cancel()` acts only while the partition is pending, as
  `confirm()` does. A creator whose confirmation committed but whose
  answer was lost, and which then cancels, leaves the live partition
  alone: only a deletion by key ends it. The registry contract, the
  SQLAlchemy registry and the registry design document say so, and a
  test cancels after a confirmation.
- Open-or-create's `VectorStoreAttemptsExhaustedError` is raised from the
  race it last lost, the last `VectorStorePartitionAlreadyExistsError` or
  `VectorStorePartitionDeletedError` it caught. A partition that stays
  pending still raises the pending error itself. A base-store test checks
  the cause.

The same changes as MemMachine#1734's 9daf7e7 and 93a85da, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…r is cancelled

A creation cancelled its reservation only when preparing the partition's
storage raised. A confirmation that raised, or a creation cancelled while
it confirmed, left the partition pending until someone deleted it.

The base store's creation step now prepares the storage and confirms the
reservation together, and cancels the reservation if either raises or the
creation is cancelled, shielded as before; the cancel's task reports its
own failure. The cancel acts only on a pending partition, so a
confirmation that committed before its failure was observed stands.
create_partition and open-or-create both go through it; open-or-create
still takes a confirmation's VectorStorePartitionDeletedError as a race to
create again. Base-store tests cover a failed confirmation, a cancelled
one, and one that committed before failing. The registry design document
says so, and the purge document says which writers wait on SQLite's lock
during a purge round.

The same changes as MemMachine#1734's f4c585e and 4dde7c8, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The query contract said nothing of a match scoring exactly the threshold,
and Qdrant dropped it: the server compares its threshold in single precision
and keeps only scores strictly better than it.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the server the adjacent single-precision value on
the worse side, and applies the caller's threshold exactly itself; a
threshold beyond single precision is not sent. Both SQLite stores already
keep the match, and each now tests it on every metric it supports.

The same changes as MemMachine#1735's ccba117 and 6bbc131, for this PR's stores.
MemMachine#1736's cc29d27, Milvus's test of the same, is ported with the store
tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…cycle contract's drain

MemMachine#1735's and MemMachine#1736's latest tests, for this PR's partitions:

- Every Qdrant store test runs against a Qdrant server, over REST and gRPC:
  local mode ignores payload indexes and raises its own exceptions where a
  server answers not-found or already-exists, so it is no longer a fixture.
  The tests on a mocked client stay in the default suite.
- Both stores gain tests that partitions and stores of different names keep
  their records apart, that purge rounds reclaim a write landing after a
  round and drain two stores' tombstones alone, that ranking, scores and the
  threshold follow every similarity metric, that a point or entity the
  server refuses on its own raises, that concurrent creations and startups
  agree, that two stores churning one registry keep every partition exact,
  and that seeded operation sequences, and on Milvus random filters, agree
  with a model. The Milvus tests also pin its read consistency and its
  settings with values other than their defaults.
- The lifecycle contract's drain fails after a bounded number of rounds, it
  settles before checking that a new life is empty, and it checks a stale
  upsert by what the partition holds rather than by its registry reads.

The same changes as MemMachine#1735's 30c30f4, e9e71b5, 6c2362b, 4d3241c,
836db21, 252bcd2, 5e8c049, ef8d3ab, 276a034, 4fb5bd4 and
2aabd3d, and MemMachine#1736's 509c235, 7f1eb3b, 6e312e0, 6d1ed1e,
25e7948, 6f36877, d43e9e6, 6dfbf59 and cc29d27, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The base store's purge contract says a round that finds the incarnation's
storage missing returns False, but neither store's round checked: once its
native collection was dropped from outside, every round on it raised,
counted against the tombstone, and dead-lettered it after ten.

A Qdrant round that the server answers not found, over REST or gRPC, and a
Milvus round that finds no native collection, now return False: the
collection is gone with everything in it, so the tombstone is retired. Each
store has a test that drops its native collection and drains the purge.

The same handling as MemMachine#1735's and MemMachine#1736's stores, whose rounds already
checked for a missing native collection; MemMachine#1736's 2575af0 tests it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A round that never ended was counted by the next claim, which took a
CASE-shaped update that both recorded the lost round and ended its claim;
the count lived apart from the claims that made it.

Each claim now counts its attempt: the claim is a plain update that
increments `consecutive_attempts` (renamed from `consecutive_failed_rounds`),
stamps `claimed_at`, and bumps the generation, and `_TombstoneClaim.attempt`
carries the count. A tombstone is claimable when it has no attempts, when no
claim is open and the backoff has passed since its last failure, or when an
open claim has outlived its lease plus the backoff. A raised round ends its
claim and stamps `last_failed_at`; a cancelled round ends its claim and takes
its attempt back; a round that found records resets the attempts. After
`_MAX_PURGE_ATTEMPTS` attempts a tombstone is dead-lettered, and a last
attempt that raises is reported. A retry logs which attempt it is, and a
raised round's error names its incarnation and attempt. The registry tests,
the purge design document, and the registry design document follow.

The same change as MemMachine#1734's 1732038, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
`consecutive_attempts` counted the purge rounds claimed since a round last
found records, but its name said only that they were in a row. It is
`attempts_without_progress`, and `_MAX_PURGE_ATTEMPTS` is
`_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The column's comment, the claim's
attempt, the backoff parameter's description, the dead-letter log, the
class docstring, the tests, the purge design document, and the registry
design document's table follow; nothing else changes.

The same change as MemMachine#1734's 96a8b5b, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…urrent claim

The measured cost of the claim came from earlier forms of it, which held a
transaction across the round. The section now gives the figures measured on
2026-10-06 against the claim this registry ports, naming the MemMachine#1734 commits
they were taken at: the backoff scan with 1k, 10k, and 100k tombstones
backing off, and interactive throughput and liveness latency beside two
sweepers on PostgreSQL and on SQLite.

The same change as MemMachine#1734's bab379b, for this PR's purge document.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's review round, for this PR's Milvus store:

- Filter strings reach Milvus as UTF-8, so a character outside the Basic
  Multilingual Plane parses, and an undeclared property's condition requires
  the stored type tag to be one the value compares with.
- Startup's already-exists guard goes, with its mocked-client test: Milvus
  answers a create of an existing collection with the same schema with
  success.
- The store no longer re-sorts search results, which Milvus returns best
  first.
- The mocked-client purge test disposes its registry engine when it fails.

The same changes as MemMachine#1736's 1a7fffa, d6fc219, 9fd7f76, 6fcd8c5, and
c264b3d, for this PR's store. MemMachine#1736's 28e8c50 and b3ef4f1 have no
counterpart here: startup prepares the store's one native collection before
any purge round runs, and the store refuses an unsupported metric at
construction, before anything is reserved.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The declared path counted an int and a float as comparable either way, so
a float filter value on a declared int property reached Milvus as a float
literal against an INT64 field, which Milvus refuses to parse ("cannot
cast value to Int64", code 1100). A float now compares only with a float
property, and matches no int one, as a value of another type matches
nothing; an int still compares with a float property by value.

The new test filters a declared float property with an int and a declared
int property with a float; it fails with the float counted as comparable
with an int property.

The same change as MemMachine#1736's 31fc7c4, for this PR's store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…y name

MemMachine#1736's baabf48 creates Milvus native collections at Bounded by name,
and this PR's base carries that into startup's create, so the read level
the purge and the tombstone retention rely on no longer comes from
pymilvus's default. The consistency test now spies create_collection and
checks the level it names, as MemMachine#1736's test does; it fails with the level
dropped from the create. The design document and the store's docstring
say the store creates its one native collection at Bounded.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's f56686f halves a Milvus upsert refused with RESOURCE_EXHAUSTED,
and this PR's base carries the halving into the partition handle's
`_upsert`. Its tests come here in partition terms: the integration test
upserts 1,200 records with a 60,000-character property, about 72 MB, and
fails with RESOURCE_EXHAUSTED without the halving; the mocked-client tests
pin the halving, a single refused entity raising, and a timeout or another
refusal sent once, on a handle `_partition_on` builds, which the mocked
delete test now uses too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nt alone

MemMachine#1736's 49b9ec2 drops the offset field beside a declared Milvus
datetime, keeping its UTC instant alone, and this PR's base carries that
into the store. The tests follow in partition terms: the native
collection's fields are exactly the fixed ones and one per declared
property; a declared datetime is stored as its instant in UTC; and a store
whose declared datetime has the longest property key, 32 bytes, writes its
one field, reads it back, and matches the record by its instant, the test
MemMachine#1736's b015229 adds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The test of MemMachine#1813, merged into this PR's base, in partition terms: the
server defaults new collections to strict mode, the store's collection
reads back with it off, and a filter on a partition's unindexed property
is served.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
A partition stored every property of a record and filtered on any key,
which made a caller's arbitrary keys part of the store's schema: the
SQLite stores kept them in a JSON column and filtered with json_extract,
Milvus the keys its schema did not declare in a JSON field, and a filter
on a key the store never indexed scanned. Since EventMemory routes a
filter on an undeclared key to the segment store, the vector store need
not hold undeclared keys at all.

A partition now stores the properties its store declares and no others.
`upsert` raises UndeclaredPropertyKeyError before anything is sent for a
record naming an undeclared key, and PropertyTypeMismatchError for a
value of another type than its key declares; `query` raises
UndeclaredPropertyKeyError for a filter naming an undeclared key and
UnsupportedFilterError for a node outside the partition's
`supported_filter_nodes`. Both SQLite stores keep one typed, indexed,
nullable column per declared key on the records table (sql_columns.py);
sqlite-vec 0.1.9 rejects NULL in a vec0 metadata column and a declared
key is optional per record, so that store keeps the columns on the
records table and hands the KNN a `rowid IN (SELECT ...)` allowlist,
evaluating the filter during the search instead of after it. Qdrant
drops the JSON copy and keeps a payload field per declared key; Milvus
drops its JSON field and keeps its typed field per declared key.
Datetimes are stored as microseconds since the epoch where a backend has
no datetime type.

Since every key a filter may name is now indexed, the Qdrant store
creates its collection in strict mode (`unindexed_filtering_retrieve`
and `_update` false, Qdrant Cloud's default): a filter on an unindexed key is refused by the server instead
of scanned for. A leaf whose value is of another type than its key
declares matches nothing, as on the SQL stores; the Qdrant compiler
answers it with a filter no point satisfies, since the server would
refuse the condition for the field's index, and Milvus no longer
compares an int with a float key. Local mode does not record the
setting, so a unit test checks the request and integration tests the
server's answer.

`declared_schema_contract.py` states the contract every backend's test
module runs: which records a filtered search admits, over fixtures small
enough that every backend searches them exactly, checked after each
upsert so an approximate index fails on recall, by name, and not on the
filter.

On the registry-backed stores, the checks are the base handle's: `upsert`
runs require_declared_properties in place of the type check it ran, and
`query` runs require_supported_filter against the subclass's
`supported_filter_nodes`, which each subclass now implements. A datetime
column on the SQLite stores has a `tz_<key>` column beside it holding the
UTC offset in seconds, written with the value and read by no filter, so a
stored datetime is the value written, as on Milvus, Qdrant and the segment
store; the microseconds column alone would keep only the instant.

The Qdrant and Milvus design documents describe the declared-only
properties and Qdrant's strict mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A declared datetime on the SQLite stores had a `tz_<key>` column beside its
microseconds column, holding the UTC offset in seconds that no filter
reads. The vector stores keep a datetime property's instant only, as
MemMachine#1736 now does for Milvus and MemMachine#1788 for Qdrant, and the segment store
keeps the offset, so the column, `offset_column_name`, and the value
written to it go: each declared property is one column, and a datetime
stays microseconds since the epoch.

The tests read a stored datetime back as its instant in UTC, and the
roundtrip test checks that the stored instant equals the written one.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Every embedder MemMachine ships produces vectors meant to be compared by
cosine -- OpenAI hard-coded it, Bedrock defaulted to it, SentenceTransformer
only reported what the model declared -- while every layer that touched a
score paid for the other three metrics in direction flags, threshold
directions, and per-backend tables mapping the enum onto native metric
names. `SimilarityMetric` is gone; scores are cosine similarities in
[-1, 1] and the names say so: `QueryMatch.score` and `SearchMatch.score`
become `cosine_similarity`, and `query(score_threshold=)` becomes
`query(min_cosine_similarity=)`, which no longer needs a direction to be
meaningful and is refused when not finite, as the threshold was. A store
and its partitions no longer have a `similarity_metric`, the schema a
partition is registered under no longer records one, and a search engine
factory takes the dimensions alone. The Bedrock embedder's
`similarity_metric` config key and semantic memory's
`vector_similarity_metric` go with it, and the install and configuration
docs drop them.

The vector graph stores carried a metric per stored embedding, as a
companion property beside every vector; that is gone and `Node.embeddings`
holds plain vectors. NebulaGraph's `cosine()` cannot take `APPROXIMATE` and
its vector indexes offer only L2 and IP. Cosine similarity between unit
vectors is their inner product, so embeddings are normalized on the way in
and compared with `inner_product()` against an IP index.

The cosine half of MemMachine#1598 (`abf92a3a4` on speedkick), re-derived on the
one-collection store: the rest of MemMachine#1598 (queries answer record UUIDs and
scores, no `get`, semantic memory's `vector_uuid`) and all of MemMachine#1603 are in
registry-backed base and its partition handle lose the metric, the stores'
own threshold checks become `require_valid_min_cosine_similarity`, and the
design documents describe cosine scoring and a schema without a metric.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…emMachine#1588)

* Atomically swap vector search engine index files on save

SQLiteVectorStore persists each collection's index by calling the search
engine's save(), which wrote directly to the final path. A crash mid-write
left a truncated/corrupt file. Because index_saved=True makes the on-disk
index a durable contract (missing/corrupt is a hard IndexLoadError, not a
silent empty rebuild), an interrupted save could render a collection
unrecoverable.

Write the index to a sibling temp file and swap it into place with
os.replace (atomic on POSIX and Windows on the same filesystem), so a reader
sees either the old or new index, never a partial write; a failed save leaves
the previous index intact. Leftover temp files are cleared on load so a crash
does not leak them across restarts.

Implemented in the engines (shared index_persistence helper) rather than in
SQLiteVectorStore/SQLiteVectorStoreCollection, since the index save location
and number of files written differ across engine implementations.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* Make the index swap durable, not only atomic

The swap protects a reader from a torn index, but the vector store also
trims its pending-operation log once `save` returns -- and that log is the
only other copy of those vectors, since the records table stores no vector
column. So the swap reaching disk is load-bearing rather than a bonus:

- fsync the parent directory after the replace, since POSIX `rename(2)`
  leaves the new directory entry in the page cache. Best-effort and ignored
  on failure, matching SQLite's `unixSync`; a no-op on Windows, which has no
  equivalent operation.
- stop swallowing a failed fsync of the temp file. SQLite draws the same
  line -- a file fsync failure raises SQLITE_IOERR_FSYNC while a directory
  fsync failure is ignored -- and `EIO` means the writeback already failed
  and the dirty pages were dropped, which is exactly when the save must not
  be reported as committed. The existing cleanup then leaves the previous
  index in place with the log untrimmed, so the next save retries.
- use F_FULLFSYNC on macOS, where plain `fsync` leaves the data in the
  drive's volatile write cache, falling back when a filesystem refuses it.

State the resulting obligation on `VectorSearchEngine.save` itself, since
that is what the store now relies on: replace atomically, then make the
replacement as durable as the platform allows. An engine whose backend
already implements the whole protocol can delegate to it and skip these
helpers; the rest use `atomic_index_write`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let the engine own index durability, not the vector store

The pending log holds the only durable copy of a vector between
checkpoints -- the records table has no vector column -- so trimming it is
safe only at an instant when the index provably holds those vectors. The
temp-write + rename protocol this PR shipped could not provide that
instant. A rename changes a directory entry, and Windows exposes no way
to flush one: os.fsync is _commit, which is FlushFileBuffers, which is
for file data, and you cannot open a directory to fsync it. The decisive
evidence is SQLite's own -- it threads a directory-sync flag through
every commit-relevant directory operation, honors it in unixDelete, and
declares it /* Not used on win32 */ in winDelete. So os.replace could
return, _save_collection_index could commit its trim durably behind it,
and a power cut could still roll the rename back: records forward, index
back, no copy of the difference left. MOVEFILE_WRITE_THROUGH is not a
fix; its documented guarantee covers copy-and-delete (cross-volume)
moves, not same-volume renames.

Take SQLite's answer, which was not to harden the directory operation
but to stop using one as a commit point (PERSIST commits by zeroing a
header, TRUNCATE by truncating, WAL by appending frames).

A base path now expands into two index slots plus a generation record
each, created once and thereafter only overwritten. A checkpoint writes
the index over the inactive slot and flushes it, then writes that slot's
generation record and flushes that. The record is the commit, and it is
a write into a file that already exists. It holds the generation and its
bitwise complement, so a torn write reads as absent rather than as some
other generation -- all or nothing without needing single-sector
atomicity from the hardware. load takes the highest believable
generation, and deliberately does not fall back to the older slot when
the published index will not parse: the log was trimmed against the
newer one, so the older is stale by exactly the ops that can no longer
be replayed.

Both backends already write straight to the path they are given, which
is what this protocol wants -- verified that repeated saves preserve the
inode and leave no stray files -- so no engine gains a temp file, a
buffer, or a rename.

Durability is entirely the engine's, including which artifact is live.
The store keeps no slot pointer, manifest, or generation, so no schema
change and no migration: what remains is one rule, never trim past what
save says is durable, and _save_collection_index already had that order.
index_path becomes index_base_path since it no longer names a file, and
discarding a collection asks the engine layer which files that covers.

BREAKING CHANGE: an index written by the previous protocol is not
published under the new one, so a collection with index_saved=True
raises IndexLoadError until its index directory is cleared and the
records re-ingested.

Anomaly tests walk every crash point in the publish sequence by
constructing the on-disk state each would leave, plus one that pins the
ordering itself (a failed index write must publish nothing) since
state-based tests cannot observe it. Verified against three deliberate
breaks -- dropping the complement check, writing the record first, and
reusing one slot instead of alternating -- each caught.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Publish the index atomically, and stop promising durability

The two-slot generation-record protocol bought a guarantee we have
decided not to make: that a save survives a power failure. Every engine
would have to implement and maintain that protocol, and the failure it
buys out is bounded -- search recall for the records applied since the
last checkpoint, repaired by re-ingesting them. The direction that
actually costs, a published index that will not parse, is closed by the
atomic swap on its own.

So this returns to the temp-file-plus-rename publication and spends the
difference on stating the contract instead of strengthening it: `save`
publishes atomically, never durably; the store trims the pending log
behind a publication a power failure can revert; a record whose vector
is lost that way still resolves by uuid, is absent from search until it
is upserted again, and nothing here detects the gap for the caller.

Reverts the durability and engine-owned-publication commits, keeps the
atomic swap, and adds a store-level test that reconstructs a reverted
publication deterministically -- restore the previous index bytes after
the trim -- to pin the direction it fails in.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Report a lost embedding as lost, not as a missing feature

`update_feature` reads the stored embedding back when a caller updates a
feature without supplying one, and that is the only place in the server
that depends on the index still holding a vector. With publication now
atomic rather than durable, a power failure can leave a feature whose
row is intact and whose vector is not -- a state this path reported as
"Vector record not found", which points the caller at the wrong thing
and hides the repair.

Split the two cases. A record that is genuinely absent keeps the old
message; a record whose embedding the index no longer holds says so and
names the fix, which is to pass a fresh embedding.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let a failed fsync fail the save, and cut the essay around it

`_flush_to_disk` wrapped its fsync in `contextlib.suppress(OSError)` and
called itself best-effort. A failed fsync is exactly the evidence that the
bytes are not safe to publish -- on Linux an EIO from fsync means writeback
failed, reported once and then cleared -- so swallowing it and renaming
anyway published a file we had positive evidence was bad. The safeguard
cost something and, in the one case it existed for, guaranteed nothing.
Nothing tested it either.

Let it propagate. `atomic_index_write` already unlinks the temp and
re-raises, so a failed flush now leaves the previously published index
standing, which is the correct outcome. A test pins that.

The fsync is not best-effort, and the docstring should not have said so: it
rules out a class rather than narrowing a window. Because the flush
completes before the rename is issued, and a durable write does not
un-happen, the new name can never appear over incomplete bytes. What the
missing directory fsync costs is the other direction -- the rename may not
survive, so the publish reverts -- and that is the benign one this store
already accepts.

The module docstring was 76 lines against 49 of everything else, most of it
argument rather than documentation: a walk through SQLite's `unixDelete` /
`winDelete` sync-flag handling, and a rejected two-slot commit protocol.
That is the PR's case for the design, not something to re-read every time
someone opens a 20-line module, and the PR body carries it. What a reader
here needs is the guarantee, the non-guarantee, and the cost.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Say why the swap is a rename, and what the reopened fd cannot see

Two things the module was silent on.

Why a rename at all. The stronger answer is to put the commit point inside
the file, where an fsync reaches it portably -- SQLite never renames, and
commits by truncating or zeroing its rollback journal, or in WAL mode by
appending frames whose checksums make a torn tail self-identifying. Both
need the writer to own the file format. A search engine owns its own and
exposes `save(path)`, so above that call a rename is the only atomicity
primitive left, and an engine whose format already commits that way needs
none of this. Worth saying, because "why not do the better thing" is the
first question the module invites.

What the reopened descriptor cannot see. Flushing is fine on a fresh fd --
dirty pages belong to the file, not to the descriptor that dirtied them --
but error reporting is not: Linux hands a writeback error to descriptors
open when it was recorded, so one recorded between the engine's close and
this open is never reported and the save proceeds on bytes already known
bad. Same shape as the 2018 PostgreSQL fsync report. It cannot be closed
from here: the engine writes through its own descriptor and closes it
before returning, and closing the window needs an engine that writes
through a handle the caller supplies.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fsync the index on a descriptor that predates the write

The fsync was on a descriptor opened after the engine had written and
closed its own, which flushes correctly -- dirty pages belong to the file,
not to the descriptor that dirtied them -- but reports nothing useful.
Linux samples the writeback error sequence when a file is opened, so a
descriptor opened after an error was recorded never learns of it: the fsync
returns success and the save publishes bytes already known bad. Same shape
as the 2018 PostgreSQL fsync report.

Open the temp before yielding it and hold it across the caller's write, so
the descriptor predates the bytes and any error from writing them is
reported here, where it fails the save.

That assumes the caller writes in place. Both engines do -- verified: the
inode is unchanged across `save_index` and `save`, and the held descriptor
sees the written size -- but it is their behaviour, not their contract. An
engine that built a file of its own and renamed it over the temp would
leave this descriptor on an orphaned inode, and the fsync would report on a
file nobody is about to publish. So it is checked before the fsync, and a
mismatch fails the save rather than passing it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* State the in-place rule where the caller reads it

The descriptor held across the body only flushes what the body wrote if the
body writes the yielded path in place, and that requirement was recorded in
`_flush_to_disk` -- a private function nobody writing an engine opens. It
belongs on `atomic_index_write`, which is the API they use, alongside what
happens when it is broken: an `OSError` and no publication, so the mistake
surfaces at the first save rather than at a power cut.

`_flush_to_disk` keeps the mechanism -- why the descriptor has to predate
the write -- and now points at the rule instead of restating it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Ask the drive to flush on macOS, where fsync does not

`os.fsync` is not the same guarantee on all three platforms this ships to.
On Linux it flushes to the device, and on Windows `FlushFileBuffers` does
the same. Darwin's `fsync` explicitly does not: it returns once the data
reaches the drive, which may hold it in a volatile write cache. So on macOS
the ordering this module is built on -- data durable before the rename is
issued -- did not hold at the device, which is exactly the case it claims
to rule out.

`F_FULLFSYNC` asks the drive to flush that cache. Filesystems that cannot
refuse it, and there `fsync` is the most that can be asked, so a refusal
falls back; any other error is a write failure and propagates, as before.

The flush-failure test patched `os.fsync`, which Darwin no longer reaches.
It patches the module's own `_fsync` instead, which every platform does.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Stop guessing which errnos mean "this drive cannot do that"

The fallback from `F_FULLFSYNC` to `fsync` was gated on an errno allowlist
-- ENOTSUP, EOPNOTSUPP, EINVAL -- so that a genuine write failure would
propagate rather than be quietly downgraded. Probing the actual returns on
Darwin shows the list is both incomplete and partly invented:

  ENOTSUP=45  EOPNOTSUPP=102   distinct here, so both are needed
  /dev/null   F_FULLFSYNC -> ENODEV(19), while fsync succeeds
  pipe/socket F_FULLFSYNC -> EBADF(9)
  EINVAL      never came from F_FULLFSYNC at all; it came from fsync

So ENODEV -- a real refusal, on a path anyone can reproduce -- would have
raised instead of falling back, and EINVAL was in the list by analogy
rather than evidence. What a network mount answers is not knowable from
here, which makes the whole list a guess that fails closed on whatever it
missed. This codebase does not classify driver errors by guessing, and
this was that.

Fall through on any failure instead. It is not a suppression: `fsync` runs
on the same descriptor and raises in its turn, so a flush that cannot
happen still fails the save. What the fallback gives up is the drive-cache
flush -- the guarantee this had before `F_FULLFSYNC` was asked for at all.
That is also what SQLite does with this same call, for the same reason.

The test drives it through `/dev/null`, which refuses with ENODEV and
accepts `fsync`; it fails against the allowlist and passes without it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fail the save when the drive cannot be told to flush

`F_FULLFSYNC` failing fell through to `fsync`, on the reasoning that a
refusal is a statement about the filesystem rather than a write failure.
That reasoning does not survive asking what `fsync` alone actually buys on
Darwin.

Against a process or kernel crash it is enough: the data has left the OS
for the drive before the rename is issued, so the rename dies in the page
cache and the old index stands. Against power loss it is not. The data
sits in the drive's volatile cache, the rename's metadata joins it moments
later, and nothing orders them -- and the rename is a few bytes against an
index of megabytes, so a drive flushing as it pleases can easily put the
new name on media while the bytes behind it are still queued. That is the
torn publication this module exists to prevent, in precisely the scenario
its docstring is about.

So the fallback answered a request for ordering with a flush that does not
provide it, and said nothing. A filesystem that cannot order data ahead of
a rename is not one to publish an index onto; raise, and let the operator
point `index_directory` at storage that can.

This is also the simpler code. Refusal and failure now take the same path,
so no errno is inspected -- there is no line to draw and no list to get
wrong, which is what the previous two revisions kept getting wrong in
opposite directions. The `/dev/null` test went with the fallback it pinned.

The earlier defence of falling back rested on network and FUSE mounts
being a realistic home for an index directory. They are not.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
(cherry picked from commit 21105d6)
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…chine#1672 to MemMachine#1676) into the final state

Their upsert keeps the delete and fresh insert, writing the declared
columns, and refuses a repeated UUID as MemMachine#1788 does; open-or-create stays
out, as MemMachine#1625 removes it. Their tests create their partition as a session
does and build filters with the closed union of MemMachine#1616.

Co-Authored-By: Claude Opus 5.5 <[email protected]>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

keep-open Prevents the auto-close task from closing this issue. Stale

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant