Skip to content

[final state] The vector store tree combined, for reference (not for merge) - #1814

Draft
edwinyyyu wants to merge 55 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:integration/vector-store-final-state
Draft

edwinyyyu wants to merge 55 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:integration/vector-store-final-state

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

One branch holding every open PR of the vector store tree, merged and consistent, so the end state can be read and tested as a whole. Not for merge. The tree's PRs merge on their own. This branch is the reference they are split from and checked against: a split is consistent when its last PR's tree equals this one.

Included, at their heads on 2026-10-09: #1788, #1702, #1670, #1627, #1622, #1625, #1663, #1460, #1469, #1672, #1673, #1674, #1675, #1676, #1618, #1741, #1628, and #1616.

Not included: the event memory stack (#1684 and above), the PRs on main (#1624, #1786, and #1792), and operator-declared indexed properties, which have no PR.

How the conflicts were resolved

The branch starts at #1616's head, which carries #1702, #1670, #1627, #1663, and #1628, and merges each other leaf with a merge commit. Every PR in it is at its head restacked onto feat/horizontal-scaling at 15f2c0686, where #1813 is merged. Where two PRs meet, the final state takes the later design:

Verification

Rebuilt on 2026-10-09 from the PR heads restacked onto 15f2c0686, #1813's merge: #1616 13d104245, #1625 cd75d4cb9, #1788 8257b8bda, #1676 a070372c4, and #1741 2c78deb61. The tree at the head 3e822399d is identical to the previous head 587dde06c, so the results below hold for it. At the restacked leaves, ruff, ruff format, and ty are clean, and their store suites pass, unit and integration, against Qdrant 1.19.1 with Cloud's strict-mode defaults and Milvus 2.6.24.

At a40e31a86, on 2026-10-09: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean; the server and client suites without integration tests, 2535 passed; the integration tests of the vector stores, the resource manager, episodic memory, and semantic storage, 826 passed, 1 failed, and 13 skipped, against PostgreSQL 16 (pgvector), Neo4j 5.23, Qdrant 1.19.1, and Milvus 2.6.24 in containers. The failure was a Milvus test that the upsert validation merge had left without its asyncio mark. The head 3fab491a3 restores the mark and changes nothing else; there, ruff check is clean and the Milvus store's test file, unit and integration, 143 passed. At the head 587dde06c, after the re-merge of #1813, the Qdrant store's and the database manager's tests, unit and integration, 365 passed against the Qdrant container defaulting new collections to strict mode.

Stack

Directly on feat/horizontal-scaling, as a reference beside the vector store tree.

🤖 Generated with Claude Code

edwinyyyu and others added 27 commits October 9, 2026 17:04
Two versions of one record in one batch are ambiguous, and the stores
disagreed on them: SQLiteVectorStore and Qdrant kept the last,
SQLiteVecVectorStore failed partway with a raw sqlite3 UNIQUE
constraint error, and Milvus refused the batch. Every store now raises
ValueError naming the UUID before anything is sent to the backend, from
one helper beside the other upsert input checks; the check costs one
pass over the batch. The VectorStoreCollection.upsert docstring says
so, and the registry-backed _upsert hook may rely on distinct UUIDs.

Each store's test upserts a batch repeating a UUID after another
record, expects the error, and finds nothing stored. It fails with the
check removed: SQLiteVectorStore and Qdrant raise nothing,
SQLiteVecVectorStore raises sqlite3's OperationalError, and Milvus
2.6.24 raises its MilvusException.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The stores disagreed on nan, inf, and -inf as property values: Qdrant
over REST stores them as null, so the property reads back absent, and
over gRPC refuses them with a raw error; Milvus 2.6.24 refuses one in a
declared float field, takes one in an undeclared property, and cannot
parse a filter on it; and the SQLite stores store them. Every store now
raises ValueError naming the property key before anything is sent to
the backend, for declared and undeclared properties alike, as the
query path already refuses a non-finite query vector coordinate or
score threshold. The VectorStoreCollection.upsert docstring says so,
and the registry-backed _upsert hook may rely on finite values.

Each store's test upserts a finite record and one with a non-finite
value of each kind, on a declared and an undeclared property, expects
the error, and finds nothing stored. It fails with the check removed:
the SQLite stores, Qdrant over REST, and Milvus on an undeclared
property raise nothing; Qdrant over gRPC raises grpc's AioRpcError; and
Milvus on a declared property raises its MilvusException.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A filter compares datetime values by their instant, for equality and
ordering alike, and a store keeps a datetime property's UTC offset as
written, taking a naive datetime as UTC. Every store behaves so: the
SQLite stores keep the offset beside the UTC instant in their
type-tagged JSON, Qdrant keeps the RFC 3339 text with its offset in the
payload, and Milvus keeps a declared datetime's offset in a field beside
its TIMESTAMPTZ and an undeclared one's in its properties JSON. The
VectorStoreCollection docstring says so.

Each store's tests write a datetime at +05:30, declared and undeclared,
read its offset back from the backend, and filter on it at other
offsets: the same instant at -08:00 matches =, <=, and >=, a later
instant at an earlier wall-clock time matches <, and an earlier instant
at a later wall-clock time matches >. The read-back tests fail with the
offset dropped at write (the type-tagged JSON's offset written as 0,
Qdrant's payload converted to UTC, or Milvus's offset field written as
0). The filter tests fail with the filter node reading a value's
wall-clock time as UTC, and on the SQLite stores also with the
wall-clock text stored in place of the UTC instant.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The event backend wrote every property of an event into its vector record
and mapped the caller's whole filter onto the vector store, so a user key
that a deployment never declared was both stored and filtered there: on
Qdrant and Milvus as unindexed payload a filtered query scans for. The
segment store already holds every property and already receives the whole
filter for the context windows, so the vector side only duplicated work the
segment store does anyway.

The vector record now carries the keys the collection declares:
EventMemory's reserved timestamp, the `_`-prefixed system properties an
adapter stamps on the event, and the keys a project's `properties_schema`
declares. The vector store is queried with the conjuncts of the filter
that name only such fields; a conjunct is dropped whole when any field
under it is undeclared, so dropping only ever widens the vector search,
and the segment store narrows it back on the windows. An undeclared key
never reaches the vector store, so a tenant's undeclared properties cannot
shape what it stores or scans.

`filter_fields` joins the filter parser: every field name a tree addresses.

Rebased onto MemMachine#1631, where the vector record no longer carries the segment
uuid (the segment store maps a derivative to its segment).

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A vector store's properties exist for filters, which compare a datetime by
its instant, and a search answers with record UUIDs and scores, so no
caller reads a stored offset back. The VectorStoreCollection docstring now
says a store keeps a datetime property's instant, taking a naive datetime
as UTC, and may drop its UTC offset; a filter compares datetime values by
their instant, as before.

Qdrant's read-back test checks the stored instant, not its offset. The
tests of the type-tagged JSON, which keeps the offset (the SQLite stores and
Milvus's undeclared properties), are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…ck) (MemMachine#1606)

* Regenerate the OpenAPI document under the locked FastAPI

`docs/openapi.json` predates the FastAPI release in `uv.lock`
(0.141.1), whose `ValidationError` component carries `input` and `ctx`;
regenerating the document with `docs/tools/generate_openapi.py` adds the
two fields and changes nothing else. Separate from the API changes above
it so their diffs of this file show only what they change.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

* Remove per-project filterable properties

A project could declare `properties_schema`, a set of caller property keys
with types, on its long-term memory configuration; the event backend merged
it into the vector store collection's indexed schema and rejected filters on
any other `m.<key>`. That let a tenant create database resources (indexes,
columns) by naming them in a request, which is what forced per-collection
native resources named by a hash of their schema on the backends that limit
them.

The option is removed from the server configuration, the project API and
the memory-configuration API, the Python SDK, the sample configurations,
the configuration docs and the OpenAPI document. A filter may name any
`m.<key>`; the stores evaluate it on the properties they hold. What a store
indexes is decided by the deployment, not per project.

A breaking API change on `speedkick`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Rebased onto MemMachine#1631: the per-project schema also leaves MemMachine#1631's service
locator, which creates the session's collection in a retry loop, and the
commented option goes from the event sample configuration MemMachine#1698 added.

---------

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
(cherry picked from commit a8322a7)
A datetime written to Qdrant keeps its UTC offset only to the minute, over
REST and gRPC alike: on 1.19.1, one written at -00:44:30 reads back at
-00:44. The store wrote a datetime property at its own offset, so one at an
offset with a seconds component, as local mean time had until 1972, was
stored up to 59 seconds off its instant, and a filter on that instant
missed it: the store did not keep the instant the collection's contract
promises. The store now converts the datetime to UTC before writing it,
which applies the whole offset to the instant, so nothing is left for
Qdrant to cut.

The read-back test checks that the payload holds the written instant in
UTC. The new filter test writes a datetime at -00:44:30, under a declared
key and an undeclared one, over REST and gRPC, and filters on its instant
with =, <, and >. Without the change, all 8 of their cases fail.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…res' lookup to get_partition

A vector store's logical collection becomes a partition, the segment
store's word for the same thing, and both stores' lookup is get_partition,
answering None like a Python get. Identifiers only, produced by the script
below; the (namespace, name) identity, the per-partition config and every
docstring are as they were, and the next change gives them their meaning.
The native clients' create_collection and delete_collection keep their
names.

The vocabulary also reaches what sits under the rename: the registry
package, its modules, classes and tables, the `collection_registry` option
of a Qdrant or Milvus backend and the store parameter and attribute that
carry it (`partition_registry`), the stale-handle and pending errors, the
registry-backed base and its handle lookup, the purge method
(`purge_deleted_partitions`, the segment store's name for it) and
open-or-create.

```sh
set -e
cd "$(git rev-parse --show-toplevel)"
git mv packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_collection.py \
       packages/server/server_tests/memmachine_server/common/vector_store/in_memory_vector_store_partition.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_lifecycle_contract.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_lifecycle_contract.py
git mv packages/server/src/memmachine_server/common/vector_store/collection_registry \
       packages/server/src/memmachine_server/common/vector_store/partition_registry
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/partition_registry.py
git mv packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_collection_registry.py \
       packages/server/src/memmachine_server/common/vector_store/partition_registry/sqlalchemy_partition_registry.py
git mv packages/server/server_tests/memmachine_server/common/vector_store/collection_registry \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry
git mv packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_collection_registry.py \
       packages/server/server_tests/memmachine_server/common/vector_store/partition_registry/test_sqlalchemy_partition_registry.py
git ls-files -z 'packages/server/*.py' 'docs/*.mdx' 'sample_configs/*.sample' 'sample_configs/*.yml' 'deployments/helm/templates/*.yaml' | xargs -0 perl -0pi -e '
  s/VectorStoreCollection(?!Config)/VectorStorePartition/g;
  s/in_memory_vector_store_collection/in_memory_vector_store_partition/g;
  s/collection_lifecycle_contract/partition_lifecycle_contract/g;
  s/CollectionLifecycleContract/PartitionLifecycleContract/g;
  s/collection_registry/partition_registry/g;
  s/SQLAlchemyCollectionRegistry/SQLAlchemyPartitionRegistry/g;
  s/CollectionRegistry/PartitionRegistry/g;
  s/RegisteredCollection/RegisteredPartition/g;
  s/get_registered_collection/get_registered_partition/g;
  s/_build_collection_handle/_build_partition_handle/g;
  s/\bCollectionT\b/PartitionT/g;
  s/register_collection\b/register_partition/g;
  s/\b_Collection\b/_Partition/g;
  s/CollectionRow/PartitionRow/g;
  s/vector_store_collection(?!_schema|_namespace)/vector_store_partition/g;
  s/open_or_create_collection/open_or_create_partition/g;
  s/open_collection/get_partition/g;
  s/_purge_deleted_collections_forever/_purge_deleted_vector_store_partitions_forever/g;
  s/purge_deleted_collections/purge_deleted_partitions/g;
  s/def create_collection\(/def create_partition(/g;
  s/def delete_collection\(/def delete_partition(/g;
  s/\.create_collection\((\s*namespace=)/.create_partition($1/g;
  s/\.delete_collection\((\s*namespace=)/.delete_partition($1/g;
  s/\.create_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.create_partition/g;
  s/\.delete_collection(?=\s*=\s*AsyncMock|\.assert_|\.side_effect|\.await_count)/.delete_partition/g;
  s/"create_collection"/"create_partition"/g;
  s/"delete_collection"/"delete_partition"/g;
  s/"open_or_create_collection"/"open_or_create_partition"/g;
  s/only delete_collection is invoked/only delete_partition is invoked/g;
  s/test_delete_collection_/test_delete_partition_/g;
  s/open_partition/get_partition/g;
  s/(vector_store_partition \(VectorStorePartition\):\n\s+)Vector store collection\./$1Vector store partition./g;
'
uv run ruff check --fix --quiet packages/server
uv run ruff format --quiet packages/server
```

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…uilt by the composition root

A vector store was a factory of logical collections, each identified by
a (namespace, name) pair and created with its own dimensions, metric and
schema; Qdrant and Milvus shared one native collection among logical
collections of equal configuration, under a name derived from a hash of
that configuration, and the registry mapped (namespace, name) to it.

A store is now one collection: `VectorStore(vector_store_name,
vector_dimensions, similarity_metric, indexed_properties)` names its one
native collection (or its tables and index files) at construction, every
partition of it shares the store's dimensions, metric and schema, and
`provision()` creates the store's durable resources idempotently, before
`startup`. `create_partition(key)`, `open_or_create_partition(key)`,
`get_partition(key)` and `delete_partition(key)` take a string key; a
partition is the records carrying its incarnation in the store's native
collection (Qdrant, Milvus) or a pair of tables (the SQLite stores), and
the registry records what each partition was created under, so a store
built with other dimensions, another metric or another schema raises
`VectorStorePartitionSchemaMismatchError` instead of reading columns and
vectors that are not there. "Collection" names only the native Qdrant or
Milvus collection. Vector store names match `[a-z0-9_]+` and are at most
32 bytes, the rule partition and property keys follow, so every name
works on every backend: the Qdrant store names its collection by the
vector store name, and the Milvus store by `sys_` followed by it, since
Milvus requires a leading letter or underscore and names its own
internals with a leading underscore. The hash-derived native names go,
and with them the per-partition config.

The partition registry is keyed by partition key within one store: its
tables, `partition_registry_pt` and `partition_registry_gc`, are shared
by every store on a database and keyed by vector store name, so the
registries of several stores share one database and nothing else.
`mark_live` and `unregister_incarnation` take the incarnation alone. On
the registry-backed base, the storage a store's partitions share is
prepared by `provision()` (`_prepare_storage()`), and the storage a
partition keeps of its own by `_prepare_partition_storage(key,
incarnation)`, between its registration and its mark; the Qdrant and
Milvus stores keep nothing per partition. A purge round works on one
incarnation in the store's own collection.

`DatabaseManager.get_vector_store(backend, vector_store_name=,
vector_dimensions=, similarity_metric=, indexed_properties=)` builds and
caches one store per (backend, vector store name), provisioning its
registry and then the store; asking for a name again with other
dimensions, another metric or other keys is a
`VectorStoreConfigurationError`. The event backend's store is named by a
UUIDv5 of its embedder id in the UUIDv5 of its backend's key in a namespace
fixed for the event backend, written as 32 hexadecimal digits, so neither
the key nor the id is constrained; stores on different backends never share
a name, so their registries stay apart when the backends share a registry
database, since a store name identifies one store among all the stores
whose registries share it.
Semantic memory's one store is named `semantic_memory` whatever its
embedder and holds every org's features in one partition,
`semantic_memory`, as main's one collection does.
The event backend opens a session's partition with
`open_or_create_partition`, which waits on one another worker is still
creating.

The SQLite stores change shape only: their registry tables become
`vector_store_sqlite_pt` and `vector_store_sqlite_vec_pt`, keyed by
vector store name and partition key, the pending-operation log is keyed
the same way, and per-partition table names embed the vector store name.
No migration is provided; existing SQLite vector data is orphaned.

Rebuilt as one change on MemMachine#1631: this PR's earlier history carried copies
of MemMachine#1631's commits as of `38684b385`, its own commits on them, and
mirrors of MemMachine#1631's later commits in its terms. Its code is that history's
final tree, which was verified, merged with main and without MemMachine#1624's
commits. The design documents are still MemMachine#1631's and describe collections.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…up, not a separate provision

A vector store and its partition registry each had a provision() that
created their durable resources, run by the composition root before
startup(). The split between provisioning a store and starting it is
MemMachine#1570's to make for every store at once, so here startup() does both
again, as on main: a store's startup prepares the storage its
partitions share (the native collection and its indexes, or the SQLite
tables), and the registry's startup creates its tables, as MemMachine#1631's does.
The database manager starts the registry, then the store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…er, live from resolve

The registry is addressed by partition key only. `register(key, schema)`
answers a PendingRegistration, whose `mark_live()` answers the same life
as a LiveRegistration or raises VectorStorePartitionDeletedError when the
partition was deleted meanwhile, and whose `unregister()` abandons that
life alone. `resolve(key)` answers the LiveRegistration of the live
partition, None when there is none, and raises
VectorStorePartitionPendingError, which now carries the pending
partition's schema, when it is pending. A LiveRegistration's
`require_current()` raises VectorStorePartitionHandleStaleError once its
partition is deleted. `get`, `mark_live(incarnation)`,
`unregister_incarnation` and RegisteredPartition are gone, so deletion is
by key or by a pending registration, never through a live one. The
SQLAlchemy registry supplies frozen-dataclass registrations holding the
engine and the vector store name, and a module function runs the
unregistration transaction.

The base handle takes its live registration, fences on
`require_current()`, and `_partition_handle` builds a handle from one.
`create_partition` lets VectorStorePartitionDeletedError propagate, and
the VectorStore interface names it; open-or-create catches it and creates
again, refuses a pending partition of another schema at once from the
pending error's schema, and re-raises the last pending error. A
`get_partition` of a pending partition of another schema still reports the
mismatch first. The Qdrant and Milvus handles take the registration in
place of the key, the incarnation and the registry lookup.

Tests follow: the registry's tests answer registrations and gain one for
`require_current`; the base's tests patch the pending registration type's
`unregister` and expect the deleted error from a creation a deletion
undid; the lifecycle contract patches the live registration type's
`require_current`, registers its racing winners through pending
registrations, and counts the deleted error among churn's outcomes; the
Qdrant and Milvus tests that build a handle on a mocked client give it a
registration that stays current.

The same change as MemMachine#1734's 735c649, f88aaa8, e25be60 and
b2bb2d8 and the handle halves of MemMachine#1735's 9096edd and MemMachine#1736's
5582cf4, for this PR's partition registry. The event-backend locator
change has no counterpart: the locator here opens its partition with
open-or-create, which creates again itself.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nd their partitions

The design documents came from MemMachine#1733-MemMachine#1736, where a store holds logical
collections addressed by namespace and name, each with its own
configuration. Here a store is one collection, with its dimensions, metric
and declared schema fixed at construction, and a partition is one tenant's
records in it, addressed by key. The documents say so:

- the collection registry document becomes the partition registry
  document: a registry belongs to one store and is addressed by partition
  key; its tables are `partition_registry_pt` and `partition_registry_gc`,
  and a tombstone needs no location, since the incarnation alone finds a
  dead partition's records in the store's native collection; a partition
  created under another schema is refused; `startup` creates the tables, as
  every store's startup creates its durable resources; a decision records
  that a store is one collection;
- the Qdrant and Milvus documents lay out one native collection per store,
  named by the vector store name (`sys_` and the name on Milvus), created at
  startup, where a partition's creation makes nothing in the backend;
- the overview, isolation, consistency and purge documents speak of
  partitions, `purge_deleted_partitions` and `settle(partition)`, and the
  overview describes the store and its partitions.

The measurements and their conditions are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ain open-or-create's give-up

Two review changes to MemMachine#1734's registry and base store, for this PR's
partitions:

- `Reservation.cancel()` acts only while the partition is pending, as
  `confirm()` does. A creator whose confirmation committed but whose
  answer was lost, and which then cancels, leaves the live partition
  alone: only a deletion by key ends it. The registry contract, the
  SQLAlchemy registry and the registry design document say so, and a
  test cancels after a confirmation.
- Open-or-create's `VectorStoreAttemptsExhaustedError` is raised from the
  race it last lost, the last `VectorStorePartitionAlreadyExistsError` or
  `VectorStorePartitionDeletedError` it caught. A partition that stays
  pending still raises the pending error itself. A base-store test checks
  the cause.

The same changes as MemMachine#1734's 9daf7e7 and 93a85da, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…r is cancelled

A creation cancelled its reservation only when preparing the partition's
storage raised. A confirmation that raised, or a creation cancelled while
it confirmed, left the partition pending until someone deleted it.

The base store's creation step now prepares the storage and confirms the
reservation together, and cancels the reservation if either raises or the
creation is cancelled, shielded as before; the cancel's task reports its
own failure. The cancel acts only on a pending partition, so a
confirmation that committed before its failure was observed stands.
create_partition and open-or-create both go through it; open-or-create
still takes a confirmation's VectorStorePartitionDeletedError as a race to
create again. Base-store tests cover a failed confirmation, a cancelled
one, and one that committed before failing. The registry design document
says so, and the purge document says which writers wait on SQLite's lock
during a purge round.

The same changes as MemMachine#1734's f4c585e and 4dde7c8, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The query contract said nothing of a match scoring exactly the threshold,
and Qdrant dropped it: the server compares its threshold in single precision
and keeps only scores strictly better than it.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the server the adjacent single-precision value on
the worse side, and applies the caller's threshold exactly itself; a
threshold beyond single precision is not sent. Both SQLite stores already
keep the match, and each now tests it on every metric it supports.

The same changes as MemMachine#1735's ccba117 and 6bbc131, for this PR's stores.
MemMachine#1736's cc29d27, Milvus's test of the same, is ported with the store
tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…cycle contract's drain

MemMachine#1735's and MemMachine#1736's latest tests, for this PR's partitions:

- Every Qdrant store test runs against a Qdrant server, over REST and gRPC:
  local mode ignores payload indexes and raises its own exceptions where a
  server answers not-found or already-exists, so it is no longer a fixture.
  The tests on a mocked client stay in the default suite.
- Both stores gain tests that partitions and stores of different names keep
  their records apart, that purge rounds reclaim a write landing after a
  round and drain two stores' tombstones alone, that ranking, scores and the
  threshold follow every similarity metric, that a point or entity the
  server refuses on its own raises, that concurrent creations and startups
  agree, that two stores churning one registry keep every partition exact,
  and that seeded operation sequences, and on Milvus random filters, agree
  with a model. The Milvus tests also pin its read consistency and its
  settings with values other than their defaults.
- The lifecycle contract's drain fails after a bounded number of rounds, it
  settles before checking that a new life is empty, and it checks a stale
  upsert by what the partition holds rather than by its registry reads.

The same changes as MemMachine#1735's 30c30f4, e9e71b5, 6c2362b, 4d3241c,
836db21, 252bcd2, 5e8c049, ef8d3ab, 276a034, 4fb5bd4 and
2aabd3d, and MemMachine#1736's 509c235, 7f1eb3b, 6e312e0, 6d1ed1e,
25e7948, 6f36877, d43e9e6, 6dfbf59 and cc29d27, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The base store's purge contract says a round that finds the incarnation's
storage missing returns False, but neither store's round checked: once its
native collection was dropped from outside, every round on it raised,
counted against the tombstone, and dead-lettered it after ten.

A Qdrant round that the server answers not found, over REST or gRPC, and a
Milvus round that finds no native collection, now return False: the
collection is gone with everything in it, so the tombstone is retired. Each
store has a test that drops its native collection and drains the purge.

The same handling as MemMachine#1735's and MemMachine#1736's stores, whose rounds already
checked for a missing native collection; MemMachine#1736's 2575af0 tests it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A round that never ended was counted by the next claim, which took a
CASE-shaped update that both recorded the lost round and ended its claim;
the count lived apart from the claims that made it.

Each claim now counts its attempt: the claim is a plain update that
increments `consecutive_attempts` (renamed from `consecutive_failed_rounds`),
stamps `claimed_at`, and bumps the generation, and `_TombstoneClaim.attempt`
carries the count. A tombstone is claimable when it has no attempts, when no
claim is open and the backoff has passed since its last failure, or when an
open claim has outlived its lease plus the backoff. A raised round ends its
claim and stamps `last_failed_at`; a cancelled round ends its claim and takes
its attempt back; a round that found records resets the attempts. After
`_MAX_PURGE_ATTEMPTS` attempts a tombstone is dead-lettered, and a last
attempt that raises is reported. A retry logs which attempt it is, and a
raised round's error names its incarnation and attempt. The registry tests,
the purge design document, and the registry design document follow.

The same change as MemMachine#1734's 1732038, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
`consecutive_attempts` counted the purge rounds claimed since a round last
found records, but its name said only that they were in a row. It is
`attempts_without_progress`, and `_MAX_PURGE_ATTEMPTS` is
`_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The column's comment, the claim's
attempt, the backoff parameter's description, the dead-letter log, the
class docstring, the tests, the purge design document, and the registry
design document's table follow; nothing else changes.

The same change as MemMachine#1734's 96a8b5b, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…urrent claim

The measured cost of the claim came from earlier forms of it, which held a
transaction across the round. The section now gives the figures measured on
2026-10-06 against the claim this registry ports, naming the MemMachine#1734 commits
they were taken at: the backoff scan with 1k, 10k, and 100k tombstones
backing off, and interactive throughput and liveness latency beside two
sweepers on PostgreSQL and on SQLite.

The same change as MemMachine#1734's bab379b, for this PR's purge document.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's review round, for this PR's Milvus store:

- Filter strings reach Milvus as UTF-8, so a character outside the Basic
  Multilingual Plane parses, and an undeclared property's condition requires
  the stored type tag to be one the value compares with.
- Startup's already-exists guard goes, with its mocked-client test: Milvus
  answers a create of an existing collection with the same schema with
  success.
- The store no longer re-sorts search results, which Milvus returns best
  first.
- The mocked-client purge test disposes its registry engine when it fails.

The same changes as MemMachine#1736's 1a7fffa, d6fc219, 9fd7f76, 6fcd8c5, and
c264b3d, for this PR's store. MemMachine#1736's 28e8c50 and b3ef4f1 have no
counterpart here: startup prepares the store's one native collection before
any purge round runs, and the store refuses an unsupported metric at
construction, before anything is reserved.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The declared path counted an int and a float as comparable either way, so
a float filter value on a declared int property reached Milvus as a float
literal against an INT64 field, which Milvus refuses to parse ("cannot
cast value to Int64", code 1100). A float now compares only with a float
property, and matches no int one, as a value of another type matches
nothing; an int still compares with a float property by value.

The new test filters a declared float property with an int and a declared
int property with a float; it fails with the float counted as comparable
with an int property.

The same change as MemMachine#1736's 31fc7c4, for this PR's store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…y name

MemMachine#1736's baabf48 creates Milvus native collections at Bounded by name,
and this PR's base carries that into startup's create, so the read level
the purge and the tombstone retention rely on no longer comes from
pymilvus's default. The consistency test now spies create_collection and
checks the level it names, as MemMachine#1736's test does; it fails with the level
dropped from the create. The design document and the store's docstring
say the store creates its one native collection at Bounded.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu and others added 28 commits October 9, 2026 17:04
MemMachine#1736's f56686f halves a Milvus upsert refused with RESOURCE_EXHAUSTED,
and this PR's base carries the halving into the partition handle's
`_upsert`. Its tests come here in partition terms: the integration test
upserts 1,200 records with a 60,000-character property, about 72 MB, and
fails with RESOURCE_EXHAUSTED without the halving; the mocked-client tests
pin the halving, a single refused entity raising, and a timeout or another
refusal sent once, on a handle `_partition_on` builds, which the mocked
delete test now uses too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nt alone

MemMachine#1736's 49b9ec2 drops the offset field beside a declared Milvus
datetime, keeping its UTC instant alone, and this PR's base carries that
into the store. The tests follow in partition terms: the native
collection's fields are exactly the fixed ones and one per declared
property; a declared datetime is stored as its instant in UTC; and a store
whose declared datetime has the longest property key, 32 bytes, writes its
one field, reads it back, and matches the record by its instant, the test
MemMachine#1736's b015229 adds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The test of MemMachine#1813, merged into this PR's base, in partition terms: the
server defaults new collections to strict mode, the store's collection
reads back with it off, and a filter on a partition's unindexed property
is served.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…quest

The event backend created a session's vector store collection and segment
store partition on the first request that opened the session, so a search
or a write for an unknown session created storage as a side effect, and
the service locator was the only place that knew both stores' create paths.

The owner is the session. Every path that creates a session row runs
through EpisodicMemoryManager._create_session, which inserts the row and,
when the row is new, creates the session's partitions in its segment store
and its vector store (create_episodic_memory_storage); an equivalent
re-create accepts the row and leaves the storage as it is. The request
path binds handles with the stores' lookups and raises
SessionPartitionMissingError when a partition is absent: a session without
its storage is broken, not new. Deleting a session with no open instance
deletes its partitions by key, so a session whose storage was never fully
created can still be deleted. MemMachine.create_session goes through the
manager for the same reason.

The semantic manager owns its one collection and creates it, once, at the
storage's first use. With that, nothing calls the stores' open-or-create.

The API is unchanged: the manager's open-or-create still creates a session
a memory request names, now through the same path.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Nothing calls them since a session's storage is created with the session:
`open_or_create_partition` leaves the vector store interface, its
registry-backed base and the two SQLite stores; `open_or_create_partition`
and `close_partition` leave the segment store interface and its
implementation, and SegmentStorePartitionConfigMismatchError, which only
the segment store's open-or-create raised, goes with them. A store creates
on `create_partition`, strictly, and looks up on `get_partition`, answering
None; create-if-absent is the owner's, where the key's provenance is known.
With open-or-create go its bounded wait on a partition another caller is
still creating, where `get_partition` raises VectorStorePartitionPendingError
for one, as before, and its retry after a creation a deletion undid, which
`create_partition` reports with VectorStorePartitionDeletedError.

Source changes are deletions only. The tests that exercised open-or-create
as a fixture use a test-side create-if-absent instead, and the tests of its
own semantics go; the refusal of a pending partition of another schema is
tested through `get_partition`, which refuses it the same way. The
partition registry design document drops open-or-create too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A partition stored every property of a record and filtered on any key,
which made a caller's arbitrary keys part of the store's schema: the
SQLite stores kept them in a JSON column and filtered with json_extract,
Milvus the keys its schema did not declare in a JSON field, and a filter
on a key the store never indexed scanned. Since EventMemory routes a
filter on an undeclared key to the segment store, the vector store need
not hold undeclared keys at all.

A partition now stores the properties its store declares and no others.
`upsert` raises UndeclaredPropertyKeyError before anything is sent for a
record naming an undeclared key, and PropertyTypeMismatchError for a
value of another type than its key declares; `query` raises
UndeclaredPropertyKeyError for a filter naming an undeclared key and
UnsupportedFilterError for a node outside the partition's
`supported_filter_nodes`. Both SQLite stores keep one typed, indexed,
nullable column per declared key on the records table (sql_columns.py);
sqlite-vec 0.1.9 rejects NULL in a vec0 metadata column and a declared
key is optional per record, so that store keeps the columns on the
records table and hands the KNN a `rowid IN (SELECT ...)` allowlist,
evaluating the filter during the search instead of after it. Qdrant
drops the JSON copy and keeps a payload field per declared key; Milvus
drops its JSON field and keeps its typed field per declared key.
Datetimes are stored as microseconds since the epoch where a backend has
no datetime type.

Since every key a filter may name is now indexed, the Qdrant store
creates its collection in strict mode (`unindexed_filtering_retrieve`
and `_update` false, Qdrant Cloud's default): a filter on an unindexed key is refused by the server instead
of scanned for. A leaf whose value is of another type than its key
declares matches nothing, as on the SQL stores; the Qdrant compiler
answers it with a filter no point satisfies, since the server would
refuse the condition for the field's index, and Milvus no longer
compares an int with a float key. Local mode does not record the
setting, so a unit test checks the request and integration tests the
server's answer.

`declared_schema_contract.py` states the contract every backend's test
module runs: which records a filtered search admits, over fixtures small
enough that every backend searches them exactly, checked after each
upsert so an approximate index fails on recall, by name, and not on the
filter.

On the registry-backed stores, the checks are the base handle's: `upsert`
runs require_declared_properties in place of the type check it ran, and
`query` runs require_supported_filter against the subclass's
`supported_filter_nodes`, which each subclass now implements. A datetime
column on the SQLite stores has a `tz_<key>` column beside it holding the
UTC offset in seconds, written with the value and read by no filter, so a
stored datetime is the value written, as on Milvus, Qdrant and the segment
store; the microseconds column alone would keep only the instant.

The Qdrant and Milvus design documents describe the declared-only
properties and Qdrant's strict mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A declared datetime on the SQLite stores had a `tz_<key>` column beside its
microseconds column, holding the UTC offset in seconds that no filter
reads. The vector stores keep a datetime property's instant only, as
MemMachine#1736 now does for Milvus and MemMachine#1788 for Qdrant, and the segment store
keeps the offset, so the column, `offset_column_name`, and the value
written to it go: each declared property is one column, and a datetime
stays microseconds since the epoch.

The tests read a stored datetime back as its instant in UTC, and the
roundtrip test checks that the stored instant equals the written one.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Every embedder MemMachine ships produces vectors meant to be compared by
cosine -- OpenAI hard-coded it, Bedrock defaulted to it, SentenceTransformer
only reported what the model declared -- while every layer that touched a
score paid for the other three metrics in direction flags, threshold
directions, and per-backend tables mapping the enum onto native metric
names. `SimilarityMetric` is gone; scores are cosine similarities in
[-1, 1] and the names say so: `QueryMatch.score` and `SearchMatch.score`
become `cosine_similarity`, and `query(score_threshold=)` becomes
`query(min_cosine_similarity=)`, which no longer needs a direction to be
meaningful and is refused when not finite, as the threshold was. A store
and its partitions no longer have a `similarity_metric`, the schema a
partition is registered under no longer records one, and a search engine
factory takes the dimensions alone. The Bedrock embedder's
`similarity_metric` config key and semantic memory's
`vector_similarity_metric` go with it, and the install and configuration
docs drop them.

The vector graph stores carried a metric per stored embedding, as a
companion property beside every vector; that is gone and `Node.embeddings`
holds plain vectors. NebulaGraph's `cosine()` cannot take `APPROXIMATE` and
its vector indexes offer only L2 and IP. Cosine similarity between unit
vectors is their inner product, so embeddings are normalized on the way in
and compared with `inner_product()` against an IP index.

The cosine half of MemMachine#1598 (`abf92a3a4` on speedkick), re-derived on the
one-collection store: the rest of MemMachine#1598 (queries answer record UUIDs and
scores, no `get`, semantic memory's `vector_uuid`) and all of MemMachine#1603 are in
registry-backed base and its partition handle lose the metric, the stores'
own threshold checks become `require_valid_min_cosine_similarity`, and the
design documents describe cosine scoring and a schema without a metric.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…emMachine#1588)

* Atomically swap vector search engine index files on save

SQLiteVectorStore persists each collection's index by calling the search
engine's save(), which wrote directly to the final path. A crash mid-write
left a truncated/corrupt file. Because index_saved=True makes the on-disk
index a durable contract (missing/corrupt is a hard IndexLoadError, not a
silent empty rebuild), an interrupted save could render a collection
unrecoverable.

Write the index to a sibling temp file and swap it into place with
os.replace (atomic on POSIX and Windows on the same filesystem), so a reader
sees either the old or new index, never a partial write; a failed save leaves
the previous index intact. Leftover temp files are cleared on load so a crash
does not leak them across restarts.

Implemented in the engines (shared index_persistence helper) rather than in
SQLiteVectorStore/SQLiteVectorStoreCollection, since the index save location
and number of files written differ across engine implementations.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* Make the index swap durable, not only atomic

The swap protects a reader from a torn index, but the vector store also
trims its pending-operation log once `save` returns -- and that log is the
only other copy of those vectors, since the records table stores no vector
column. So the swap reaching disk is load-bearing rather than a bonus:

- fsync the parent directory after the replace, since POSIX `rename(2)`
  leaves the new directory entry in the page cache. Best-effort and ignored
  on failure, matching SQLite's `unixSync`; a no-op on Windows, which has no
  equivalent operation.
- stop swallowing a failed fsync of the temp file. SQLite draws the same
  line -- a file fsync failure raises SQLITE_IOERR_FSYNC while a directory
  fsync failure is ignored -- and `EIO` means the writeback already failed
  and the dirty pages were dropped, which is exactly when the save must not
  be reported as committed. The existing cleanup then leaves the previous
  index in place with the log untrimmed, so the next save retries.
- use F_FULLFSYNC on macOS, where plain `fsync` leaves the data in the
  drive's volatile write cache, falling back when a filesystem refuses it.

State the resulting obligation on `VectorSearchEngine.save` itself, since
that is what the store now relies on: replace atomically, then make the
replacement as durable as the platform allows. An engine whose backend
already implements the whole protocol can delegate to it and skip these
helpers; the rest use `atomic_index_write`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let the engine own index durability, not the vector store

The pending log holds the only durable copy of a vector between
checkpoints -- the records table has no vector column -- so trimming it is
safe only at an instant when the index provably holds those vectors. The
temp-write + rename protocol this PR shipped could not provide that
instant. A rename changes a directory entry, and Windows exposes no way
to flush one: os.fsync is _commit, which is FlushFileBuffers, which is
for file data, and you cannot open a directory to fsync it. The decisive
evidence is SQLite's own -- it threads a directory-sync flag through
every commit-relevant directory operation, honors it in unixDelete, and
declares it /* Not used on win32 */ in winDelete. So os.replace could
return, _save_collection_index could commit its trim durably behind it,
and a power cut could still roll the rename back: records forward, index
back, no copy of the difference left. MOVEFILE_WRITE_THROUGH is not a
fix; its documented guarantee covers copy-and-delete (cross-volume)
moves, not same-volume renames.

Take SQLite's answer, which was not to harden the directory operation
but to stop using one as a commit point (PERSIST commits by zeroing a
header, TRUNCATE by truncating, WAL by appending frames).

A base path now expands into two index slots plus a generation record
each, created once and thereafter only overwritten. A checkpoint writes
the index over the inactive slot and flushes it, then writes that slot's
generation record and flushes that. The record is the commit, and it is
a write into a file that already exists. It holds the generation and its
bitwise complement, so a torn write reads as absent rather than as some
other generation -- all or nothing without needing single-sector
atomicity from the hardware. load takes the highest believable
generation, and deliberately does not fall back to the older slot when
the published index will not parse: the log was trimmed against the
newer one, so the older is stale by exactly the ops that can no longer
be replayed.

Both backends already write straight to the path they are given, which
is what this protocol wants -- verified that repeated saves preserve the
inode and leave no stray files -- so no engine gains a temp file, a
buffer, or a rename.

Durability is entirely the engine's, including which artifact is live.
The store keeps no slot pointer, manifest, or generation, so no schema
change and no migration: what remains is one rule, never trim past what
save says is durable, and _save_collection_index already had that order.
index_path becomes index_base_path since it no longer names a file, and
discarding a collection asks the engine layer which files that covers.

BREAKING CHANGE: an index written by the previous protocol is not
published under the new one, so a collection with index_saved=True
raises IndexLoadError until its index directory is cleared and the
records re-ingested.

Anomaly tests walk every crash point in the publish sequence by
constructing the on-disk state each would leave, plus one that pins the
ordering itself (a failed index write must publish nothing) since
state-based tests cannot observe it. Verified against three deliberate
breaks -- dropping the complement check, writing the record first, and
reusing one slot instead of alternating -- each caught.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Publish the index atomically, and stop promising durability

The two-slot generation-record protocol bought a guarantee we have
decided not to make: that a save survives a power failure. Every engine
would have to implement and maintain that protocol, and the failure it
buys out is bounded -- search recall for the records applied since the
last checkpoint, repaired by re-ingesting them. The direction that
actually costs, a published index that will not parse, is closed by the
atomic swap on its own.

So this returns to the temp-file-plus-rename publication and spends the
difference on stating the contract instead of strengthening it: `save`
publishes atomically, never durably; the store trims the pending log
behind a publication a power failure can revert; a record whose vector
is lost that way still resolves by uuid, is absent from search until it
is upserted again, and nothing here detects the gap for the caller.

Reverts the durability and engine-owned-publication commits, keeps the
atomic swap, and adds a store-level test that reconstructs a reverted
publication deterministically -- restore the previous index bytes after
the trim -- to pin the direction it fails in.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Report a lost embedding as lost, not as a missing feature

`update_feature` reads the stored embedding back when a caller updates a
feature without supplying one, and that is the only place in the server
that depends on the index still holding a vector. With publication now
atomic rather than durable, a power failure can leave a feature whose
row is intact and whose vector is not -- a state this path reported as
"Vector record not found", which points the caller at the wrong thing
and hides the repair.

Split the two cases. A record that is genuinely absent keeps the old
message; a record whose embedding the index no longer holds says so and
names the fix, which is to pass a fresh embedding.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let a failed fsync fail the save, and cut the essay around it

`_flush_to_disk` wrapped its fsync in `contextlib.suppress(OSError)` and
called itself best-effort. A failed fsync is exactly the evidence that the
bytes are not safe to publish -- on Linux an EIO from fsync means writeback
failed, reported once and then cleared -- so swallowing it and renaming
anyway published a file we had positive evidence was bad. The safeguard
cost something and, in the one case it existed for, guaranteed nothing.
Nothing tested it either.

Let it propagate. `atomic_index_write` already unlinks the temp and
re-raises, so a failed flush now leaves the previously published index
standing, which is the correct outcome. A test pins that.

The fsync is not best-effort, and the docstring should not have said so: it
rules out a class rather than narrowing a window. Because the flush
completes before the rename is issued, and a durable write does not
un-happen, the new name can never appear over incomplete bytes. What the
missing directory fsync costs is the other direction -- the rename may not
survive, so the publish reverts -- and that is the benign one this store
already accepts.

The module docstring was 76 lines against 49 of everything else, most of it
argument rather than documentation: a walk through SQLite's `unixDelete` /
`winDelete` sync-flag handling, and a rejected two-slot commit protocol.
That is the PR's case for the design, not something to re-read every time
someone opens a 20-line module, and the PR body carries it. What a reader
here needs is the guarantee, the non-guarantee, and the cost.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Say why the swap is a rename, and what the reopened fd cannot see

Two things the module was silent on.

Why a rename at all. The stronger answer is to put the commit point inside
the file, where an fsync reaches it portably -- SQLite never renames, and
commits by truncating or zeroing its rollback journal, or in WAL mode by
appending frames whose checksums make a torn tail self-identifying. Both
need the writer to own the file format. A search engine owns its own and
exposes `save(path)`, so above that call a rename is the only atomicity
primitive left, and an engine whose format already commits that way needs
none of this. Worth saying, because "why not do the better thing" is the
first question the module invites.

What the reopened descriptor cannot see. Flushing is fine on a fresh fd --
dirty pages belong to the file, not to the descriptor that dirtied them --
but error reporting is not: Linux hands a writeback error to descriptors
open when it was recorded, so one recorded between the engine's close and
this open is never reported and the save proceeds on bytes already known
bad. Same shape as the 2018 PostgreSQL fsync report. It cannot be closed
from here: the engine writes through its own descriptor and closes it
before returning, and closing the window needs an engine that writes
through a handle the caller supplies.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fsync the index on a descriptor that predates the write

The fsync was on a descriptor opened after the engine had written and
closed its own, which flushes correctly -- dirty pages belong to the file,
not to the descriptor that dirtied them -- but reports nothing useful.
Linux samples the writeback error sequence when a file is opened, so a
descriptor opened after an error was recorded never learns of it: the fsync
returns success and the save publishes bytes already known bad. Same shape
as the 2018 PostgreSQL fsync report.

Open the temp before yielding it and hold it across the caller's write, so
the descriptor predates the bytes and any error from writing them is
reported here, where it fails the save.

That assumes the caller writes in place. Both engines do -- verified: the
inode is unchanged across `save_index` and `save`, and the held descriptor
sees the written size -- but it is their behaviour, not their contract. An
engine that built a file of its own and renamed it over the temp would
leave this descriptor on an orphaned inode, and the fsync would report on a
file nobody is about to publish. So it is checked before the fsync, and a
mismatch fails the save rather than passing it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* State the in-place rule where the caller reads it

The descriptor held across the body only flushes what the body wrote if the
body writes the yielded path in place, and that requirement was recorded in
`_flush_to_disk` -- a private function nobody writing an engine opens. It
belongs on `atomic_index_write`, which is the API they use, alongside what
happens when it is broken: an `OSError` and no publication, so the mistake
surfaces at the first save rather than at a power cut.

`_flush_to_disk` keeps the mechanism -- why the descriptor has to predate
the write -- and now points at the rule instead of restating it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Ask the drive to flush on macOS, where fsync does not

`os.fsync` is not the same guarantee on all three platforms this ships to.
On Linux it flushes to the device, and on Windows `FlushFileBuffers` does
the same. Darwin's `fsync` explicitly does not: it returns once the data
reaches the drive, which may hold it in a volatile write cache. So on macOS
the ordering this module is built on -- data durable before the rename is
issued -- did not hold at the device, which is exactly the case it claims
to rule out.

`F_FULLFSYNC` asks the drive to flush that cache. Filesystems that cannot
refuse it, and there `fsync` is the most that can be asked, so a refusal
falls back; any other error is a write failure and propagates, as before.

The flush-failure test patched `os.fsync`, which Darwin no longer reaches.
It patches the module's own `_fsync` instead, which every platform does.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Stop guessing which errnos mean "this drive cannot do that"

The fallback from `F_FULLFSYNC` to `fsync` was gated on an errno allowlist
-- ENOTSUP, EOPNOTSUPP, EINVAL -- so that a genuine write failure would
propagate rather than be quietly downgraded. Probing the actual returns on
Darwin shows the list is both incomplete and partly invented:

  ENOTSUP=45  EOPNOTSUPP=102   distinct here, so both are needed
  /dev/null   F_FULLFSYNC -> ENODEV(19), while fsync succeeds
  pipe/socket F_FULLFSYNC -> EBADF(9)
  EINVAL      never came from F_FULLFSYNC at all; it came from fsync

So ENODEV -- a real refusal, on a path anyone can reproduce -- would have
raised instead of falling back, and EINVAL was in the list by analogy
rather than evidence. What a network mount answers is not knowable from
here, which makes the whole list a guess that fails closed on whatever it
missed. This codebase does not classify driver errors by guessing, and
this was that.

Fall through on any failure instead. It is not a suppression: `fsync` runs
on the same descriptor and raises in its turn, so a flush that cannot
happen still fails the save. What the fallback gives up is the drive-cache
flush -- the guarantee this had before `F_FULLFSYNC` was asked for at all.
That is also what SQLite does with this same call, for the same reason.

The test drives it through `/dev/null`, which refuses with ENODEV and
accepts `fsync`; it fails against the allowlist and passes without it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fail the save when the drive cannot be told to flush

`F_FULLFSYNC` failing fell through to `fsync`, on the reasoning that a
refusal is a statement about the filesystem rather than a write failure.
That reasoning does not survive asking what `fsync` alone actually buys on
Darwin.

Against a process or kernel crash it is enough: the data has left the OS
for the drive before the rename is issued, so the rename dies in the page
cache and the old index stands. Against power loss it is not. The data
sits in the drive's volatile cache, the rename's metadata joins it moments
later, and nothing orders them -- and the rename is a few bytes against an
index of megabytes, so a drive flushing as it pleases can easily put the
new name on media while the bytes behind it are still queued. That is the
torn publication this module exists to prevent, in precisely the scenario
its docstring is about.

So the fallback answered a request for ordering with a flush that does not
provide it, and said nothing. A filesystem that cannot order data ahead of
a rename is not one to publish an index onto; raise, and let the operator
point `index_directory` at storage that can.

This is also the simpler code. Refusal and failure now take the same path,
so no errno is inspected -- there is no line to draw and no list to get
wrong, which is what the previous two revisions kept getting wrong in
opposite directions. The `/dev/null` test went with the fallback it pinned.

The earlier defence of falling back rested on network and FUSE mounts
being a realistic home for an index directory. They are not.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
(cherry picked from commit 21105d6)
…kend

The filter language is a closed union: Equals, Ordering, In, IsNull,
n-ary And and Or, and Not. Every compiler is an exhaustive match. The
parser maps `!=` and `<>` to Not(Equals), `NOT IN` to Not(In) and
`IS NOT NULL` to Not(IsNull), so a negation means one thing.

A leaf on a field holding no value is false, IsNull on it is true, and
Not is the ordinary complement over the record set. The SQL compilers
render Not as `NOT COALESCE(x, FALSE)`; they compiled `~inner`, which is
three-valued and dropped rows holding no value. Milvus already pushes
Not to the leaves by De Morgan, since its own `not` excludes entities
lacking the field, and Qdrant's must_not already admits them. A
predicate matches only a value of the compared type. The graph stores
and the Neo4j semantic storage get the union only; their NOT over a
comparison on a missing property stays three-valued, since no backend
here can test a change.

Where the compiled output changes: common/filter/sql_filter_util.py and
common/vector_store/sql_columns.py (COALESCE under Not), and the
parser's mapping of `!=`. Where behavior changes: the in-memory test
collection's evaluator matches values by type, the declared-schema
contract gains the complement law and the typed-match test, and two
semantic-storage tests that ordered string values go, since ordering
strings is not expressible in the union. An `In` over no values admits
nothing: every compiler renders it as false, so EventMemory's empty id
lists (session, source, block kind) keep meaning "nothing", and its
complement admits everything. Every other file in this commit changes
node names, import paths and dispatch shape only, with the same output
as before.

The tests the rebuilt base adds use the closed union too, and the segment
store's random context-read test models the new semantics: a condition on
a property with no value is false and a negation is the complement, where
the model had followed SQL's three-valued logic.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Never reuse a row id in SQLiteVectorStore

The records table used a plain INTEGER PRIMARY KEY, which is SQLite's
rowid, assigned as max(rowid) + 1: deleting the highest row freed its id
for the very next insert. query() scores keys in the search engine and
resolves them to rows in a second step, holding nothing in between, so a
reused id let a record that was never scored come back wearing the score
of the record that was. Nothing about that result looks wrong: the record
exists and the score is in range.

Declare the table with sqlite_autoincrement=True so ids are never reused.
A stale engine key then matches no row and is dropped.

Both tests fail without the flag: one pins the id policy directly, the
other parks a query between scoring and row lookup, retires the scored
record, inserts another, and asserts the query returns nothing.

First half of MemMachine#1468.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
A leaf whose value was of another type than its key declares matched
nothing on every store, each compiler rendering it as a filter no record
satisfies, so a caller who compared a key with the wrong type got an
empty result and no sign of the mistake; and `score > 3` on a float key
matched nothing, since the parser reads `3` as an int. A NaN or an
infinity reached the backends too: it matched every held value on
Qdrant's REST transport, a NaN failed on its gRPC transport, and either
failed on Milvus.

`bind_filter` replaces `require_supported_filter` at every store's query
and returns the filter the store compiles in place of the caller's. It
still refuses an undeclared key and a node outside the store's
`supported_filter_nodes`, and in the filter it returns every value is of
its key's declared type: an int for a float key becomes the float it
equals, and a membership test of ints on a float key the disjunction of
those floats' equalities, since a membership test holds ints or strs.
Any other value of another type than its key declares, a bool for a
number, and a float that is not finite raise PropertyTypeMismatchError,
which `upsert` already raises for a record; the VectorStorePartition
query documents it.

The compilers' branches that rendered a mistyped leaf as no match are
unreachable and go: the Qdrant store's no-match filter, the Milvus
store's filtering of values by type, and the SQL columns compiler's
false branches, keeping only an `In` over no values as false. None of
the three needs the declared schema any more. The Qdrant and Milvus
design documents describe the binding.

The declared-schema contract replaces its no-match test with tests every
store runs: a value of another type and a float that is not finite each
raise from `query`, and an int compares with a float key as the float it
equals, each filter and its complement. The first two fail with the
check in `_require_value` removed, the non-finite cases also with its
finiteness clause removed, and the third fails on Qdrant, whose float
index refuses an int match, with the int-to-float conversion removed
from `_bound_value` or from the membership branch of `_bound_filter`.
The Milvus and Qdrant modules' own no-match tests go, the Qdrant one
with the strict-mode reason it stated.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…(speedkick) (MemMachine#1612)

Own the search engine's concurrency in the store, not in each engine

Each engine wrapped its five methods in its own read-write lock and the
abstract class promised concurrent use. The store is the only caller,
and it already knows which calls must exclude which: a search shares the
engine with other searches, and everything else runs alone. Holding that
lock in the store puts every exclusion in the one file that reads the
engine, and lets a rule the engines could not express hold: a rewrite's
remove and add are one step to a reader, where before a search could run
between them and see neither version.

Engines drop their lock and keep only the index calls; the abstract
class now states that the owner serializes. The store keeps one
read-write lock per collection beside the engine, kept for the store's
lifetime like the engine's other per-collection state, and takes it at
every engine call: searches on the read side; mutations, loads, and the
index save on the write side. A rewrite's remove and add sit under one
hold. The save's trim runs after the lock is released, so readers wait
for the file write and never for SQL.

The row-id regression test from MemMachine#1589 parked inside a wrapper engine's
search, outside the real engine's lock; under the store's lock that
parks the read side, and the writes it then awaits cannot proceed. It
now parks where it meant to, between the engine search and the row
lookup. One new test pins the one-step rewrite; it fails against the
engines' own locks.

The turbovec engine in flight carries the same lock and needs the same
subtraction.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
A record could not hold an int for a float key: upsert refused it as a
type mismatch, though a filter may compare a float key with an int. A
float property that is not finite reached the backends instead, and
each kept it its own way: the SQLite stores held a NaN as no value and
an infinity as written, Qdrant's REST transport held either as no
value, and Qdrant's gRPC transport and Milvus refused the upsert.

`bind_record` replaces `require_declared_properties` at every store's
upsert and returns the record the store writes in place of the
caller's: its keys are declared and every value is of its key's
declared type, an int for a float key becoming the float it equals. Any
other value of another type, a bool for a number, and a float that is
not finite raise PropertyTypeMismatchError, by the same rule the bound
filter follows. The in-memory test partition binds its records too, and
the VectorStorePartition upsert and the error document the rule.

The declared-schema contract checks the rule on every store: an int
written to a float key is stored as the float it equals, read past the
store through a new `stored_value` hook each store's test module
supplies, and matched by a float filter, an int one, and a membership
test of ints; and each mistyped or non-finite value raises from
`upsert`. The per-store tests of refused values lose their
int-for-float case. The first test fails with the int-for-float clause
removed from `_require_value`, and on Qdrant, which keeps an int
payload as an int, with `bind_record` returning the caller's record;
the second fails with the finiteness clause removed, and its bool cases
with a bool counted as an int.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…edkick) (MemMachine#1607)

Serialize a collection's writes so the engine sees them in order

A write commits to SQLite and only then applies to the search engine, so
two writers to one uuid could reach the engine in the opposite order to
the one they committed in: an upsert overtaking a delete re-adds a vector
for a record that is gone, and two upserts inverting leave the engine
serving the older vector. Never reusing a row id does not cover this,
because an upsert of an existing uuid keeps its row id.

A save in that window costs a write outright: it publishes the index and
trims every applied log row, and a write that applied after the index was
written is then in neither.

A per-collection asyncio.Lock now spans a write from SQL commit through
engine apply, mark-applied, and any save it triggers; shutdown's save
takes it too. The lock belongs to the store, not to a collection handle:
a handle is constructed per open_collection call, so several can address
one collection, and only a shared lock serializes them. Readers are
untouched.

Three tests fail without the lock, each interleaving made deterministic
by gating the engine: an upsert overtaking a delete of its uuid, a save
trimming a write it did not publish, and that overtake across two
handles, which a per-handle lock passes. The rest pin behavior the lock
must preserve: an upsert surviving a delete of another record, disjoint
concurrent upserts and deletes, writes racing a checkpoint, and batches
that name one uuid twice.

Fixes MemMachine#1468.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
Every backend holds an int in 64 signed bits, and each answered a wider
one its own way. A record holding one failed to write on the SQLite
stores, on Qdrant's gRPC transport, and on Milvus, and Qdrant's REST
transport stored it as the largest int64, which an equality filter on
that value then matched. A filter holding one failed on the SQLite
stores and on Milvus, failed an equality on Qdrant, and answered an
ordering on Qdrant with its bound clamped, so `>= 2**63` matched the
largest int64.

The binding both `bind_record` and `bind_filter` apply now refuses an
int outside -2**63 to 2**63 - 1 for any key with
PropertyTypeMismatchError, before anything is sent; the bounds are
MIN_INT_VALUE and MAX_INT_VALUE beside the error. An int for a float key
is held to the same range, so its float always exists.

The declared-schema contract adds out-of-range ints to its mistyped
records and filters, and a test that every store compares the int64
extremes with filters at those extremes. The mistyped cases fail with
the range clause removed from `_require_value`, and the extremes test
with the range made exclusive.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
The declared-schema contract gains a test that every store matches
strings holding accented letters, an emoji, quotes, a backslash, and
control characters, by equality and by membership. It fails on Milvus
with `ensure_ascii=False` removed from `_expression_string_literal`, the
UTF-8 literal [vector store scale-out 6/6] ports from MemMachine#1736's 1a7fffa.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…MemMachine#1609)

Take SQLite's write lock at BEGIN, not at the first write

`delete` resolves row ids and then writes, in one transaction. Under
SQLite's default deferred BEGIN the write lock is taken at the first
write, so another writer can take it between that read and the write,
and the transaction that read first is the one that loses: its own write
cannot upgrade, and SQLite reports that immediately rather than waiting
out busy_timeout, because waiting could only deadlock. Reproduced in both
journal modes:

    journal_mode  BEGIN      other writer  our write
    delete        deferred   shut out      fails
    delete        IMMEDIATE  shut out      ok
    wal           deferred   commits       fails
    wal           IMMEDIATE  shut out      ok

The store now emits BEGIN explicitly and lets a transaction ask for
BEGIN IMMEDIATE, which every write path does. The mode is chosen per
transaction, not per engine: a hook that asked for IMMEDIATE
unconditionally would make every read take the write lock, and two
readers would then serialize against each other.

Two tests fail without it. One holds a competing lock across a
read-then-write transaction and asserts that transaction completes;
asserting instead that the other writer is excluded passes either way,
because a deferred BEGIN shuts it out too, later and by a different lock.
The other races four creates of one name: under a deferred BEGIN the
losers read no stored config, go on to CREATE TABLE, and fail there with
"table already exists" instead of VectorStoreCollectionAlreadyExistsError.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…tization

`QdrantConf` gains `hnsw_config`, `optimizers_config` and
`quantization_config`, plain mappings mirroring qdrant-client's
`HnswConfigDiff`, `OptimizersConfigDiff` and `QuantizationConfig`, so
qdrant-client stays optional for configuration parsing; the store's
params validate them against qdrant's own models. They apply to the
store's one collection, created at startup.

`m` must be 0 or unset: the collection is multi-tenant and disables the
global graph in favor of per-partition payload indexing, so a deployment
tunes `payload_m`, which defaults to 16 as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The sample configurations gain a commented block with the three keys,
checked against qdrant-client's models.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ches

`MilvusConf` gains `vector_index`, a plain mapping of `index_type`,
`params` (build parameters) and `search_params` (parameters every search
passes), so pymilvus stays optional for configuration parsing; the
store's params coerce it into `MilvusVectorIndex`. pymilvus has no models
for index or search parameters, so the unit is the whole spec: build and
search parameters come from one author for one index type. The index
applies when startup creates the collection's vector index, and the
search parameters to every query. Unset selects HNSW_SQ searched with
refine_k 8, as before.

`params` may not set `index_type` or `metric_type`, which pymilvus would
let override the spec's own index type and the store's cosine metric.
Milvus checks the parameters the index type takes when it creates the
index and when a search reaches an indexed segment.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
The sample configurations gain a commented `vector_index` block under the
Milvus alternative, checked against `MilvusConf` and `MilvusVectorIndex`.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…he core final state

Open-or-create leaves the vector store interface; the service locator's
tests build the cosine-only search engine of MemMachine#1663, and the SQLite stores'
tests keep the cosine-only metric tests of MemMachine#1663.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
Every store refuses a batch naming a record UUID twice. A non-finite float
is already refused by the declared schema's type check, so its separate
check and tests go, and the datetime tests run on the declared key alone,
since an undeclared key is refused. The store keeps the strict mode of
MemMachine#1628 over the one MemMachine#1813 merged into the base, and MemMachine#1813's test of it off
stays out.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…chine#1672 to MemMachine#1676) into the final state

Their upsert keeps the delete and fresh insert, writing the declared
columns, and refuses a repeated UUID as MemMachine#1788 does; open-or-create stays
out, as MemMachine#1625 removes it. Their tests create their partition as a session
does and build filters with the closed union of MemMachine#1616.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…achine#1741) into the final state

A Qdrant collection is created with the deployment's HNSW, optimizer and
quantization options and with the strict mode of MemMachine#1628.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@edwinyyyu
edwinyyyu force-pushed the integration/vector-store-final-state branch from 587dde0 to 3e82239 Compare October 10, 2026 00:31

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant