Skip to content

[sqlite store fixes 7/7] Give every write a fresh row id, so a key names one version (port of #1610) - #1676

Draft
edwinyyyu wants to merge 35 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:port/fresh-row-id-per-write-main
Draft

edwinyyyu wants to merge 35 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:port/fresh-row-id-per-write-main

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Port

Copy of #1610, merged into speedkick as acb4f9aa2, on main: re-derived onto the one-collection store of #1627, whose per-partition tables are the origin's shape, so the change is the origin's; with [search results] below, a record always carries a vector, as in the origin. With [declared schema 1/2] (#1628) below too, the fresh insert writes the record's declared columns, and the tests read a committed name from its declared column.

Purpose of the change

An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them. A rewrite of the very record being returned could land between those reads, pairing one version's score with another version's filter verdict. Concretely, with a record that fails the filter at version 1 and passes at version 2: the engine scores version 1, the rewrite commits, the filter check reads version 2's properties and passes, and version 1's score comes back for a record that only qualifies as version 2. The write lock (#1607) does not help, because reads take no write lock.

Every write now takes a fresh row id: the previous row is deleted and a new one inserted in the same transaction, a delete for the old key is staged beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version, and the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version. A key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT (#1589) remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order.

A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often.

Cost

Medians of three runs on one machine, records per second, 64-dimensional vectors, file-backed store with an index directory:

batch save threshold insert before → after rewrite before → after
500 1000 20477 → 19290 16450 → 12902
500 none 21228 → 22486 19013 → 15899
1 none 341 → 345 316 → 274

Inserts move within run-to-run noise, in both directions. Rewrites cost 13 to 22 percent more; the upper end is the configuration where the doubled log rows double the number of index saves. MemMachine's callers write vector records once and rarely rewrite them, so the insert path is the one that matters, and the read path is untouched.

Tests

test_a_query_cannot_pair_a_score_with_a_later_version drives the case above deterministically: a wrapper around the query's key filter parks the engine's worker thread between scoring and the filter check, a rewrite commits in the gap, and the check must not admit the old score. test_a_rewrite_moves_the_record_to_a_new_row_id pins the mechanism. The mirror case, an old verdict admitting a new score, cannot be reproduced with the USearch engine because its read lock spans the whole search including overfetch rounds, so the mechanism test is what covers it.

Stack

21 open PRs: three independent PRs, and the vector store tree of short parallel branches. Every PR in the tree has feat/horizontal-scaling as its GitHub base, and the independent PRs have main. The branches are in a fork, and a pull request can target only this repository's branches, so the on column gives the order the PRs build on each other. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main
— #1786 Refuse a filter value of the wrong type for a datetime column, and answer an invalid list filter with 422 (port of #1620) main
— #1792 Refuse property values that some store refuses or alters where an episode enters main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1702 and #1670 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (merged into feat/horizontal-scaling) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs feat/horizontal-scaling
[vector store scale-out 3/6] #1734 (merged into feat/horizontal-scaling) Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge feat/horizontal-scaling
[vector store scale-out 4/6] #1735 (merged into feat/horizontal-scaling) Move the Qdrant store onto the collection registry feat/horizontal-scaling
[vector store scale-out 5/6] #1736 (merged into feat/horizontal-scaling) Move the Milvus store onto the collection registry, against a Milvus server feat/horizontal-scaling
— #1775 (merged into feat/horizontal-scaling) Accept attempts exhausted in the lifecycle churn contract, and say which Qdrant operations filter on the incarnation feat/horizontal-scaling
— #1779 (merged into feat/horizontal-scaling) Classify Qdrant errors by status code alone feat/horizontal-scaling
— #1813 (merged into feat/horizontal-scaling) Create Qdrant collections with strict mode off feat/horizontal-scaling
— #1788 Refuse a repeated record UUID or a non-finite property value at upsert, and state the datetime property contract feat/horizontal-scaling
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1702 Keep undeclared properties out of the vector store feat/horizontal-scaling
[user properties 2/2] #1670 Remove per-project filterable properties (port of #1606) #1702
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1670
[session storage 1/2] #1622 Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1627
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1628
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1663
[sqlite store fixes 1/7] #1460 Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 (this PR) Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[milvus options] #1741 Let a deployment tune a Milvus collection's vector index and its searches #1618

This PR is its one commit, a070372c4, stacked on #1675.

Verification

At this PR's head a070372c4, on 2026-10-09, on feat/horizontal-scaling with #1813 merged: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean; the SQLite store's and the search engines' tests pass, 181 tests.

Before the rebase onto #1813's merge: At this PR's head 102203e25, on 2026-10-09: ruff check and ruff format --check clean; ty check clean as CI runs it (uv run --frozen --all-extras ty check --project packages/server); uv lock --check clean; the SQLite stores' and search engines' tests pass, 257 tests. The full suites last passed at f687f68f9, on 2026-10-08, before #1628 moved beneath this PR: the server suite without integration tests, 2132 tests, and the integration tests of the vector stores, the resource manager, episodic memory, and semantic storage, 668 tests, against PostgreSQL 16, Neo4j, Qdrant 1.19.1, and Milvus 2.6.24 in containers.

🤖 Generated with Claude Code

https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

This was referenced Sep 17, 2026
@edwinyyyu
edwinyyyu force-pushed the port/fresh-row-id-per-write-main branch 2 times, most recently from 595707e to 3723e9d Compare September 17, 2026 18:18
@edwinyyyu edwinyyyu changed the title [sqlite store fixes 8/8] Give every write a fresh row id, so a key names one version (port of #1610) [sqlite store fixes 7/7] Give every write a fresh row id, so a key names one version (port of #1610) Sep 17, 2026
@edwinyyyu
edwinyyyu force-pushed the port/fresh-row-id-per-write-main branch from 3723e9d to 0aaceb9 Compare September 17, 2026 18:27
@edwinyyyu
edwinyyyu force-pushed the port/fresh-row-id-per-write-main branch from 0aaceb9 to b7d2bb4 Compare September 17, 2026 19:32
@edwinyyyu edwinyyyu changed the title [sqlite store fixes 7/7] Give every write a fresh row id, so a key names one version (port of #1610) [sqlite store fixes 6/6] Give every write a fresh row id, so a key names one version (port of #1610) Sep 17, 2026
@edwinyyyu
edwinyyyu force-pushed the port/fresh-row-id-per-write-main branch 2 times, most recently from b21343e to 411472d Compare September 18, 2026 19:40
edwinyyyu and others added 28 commits October 9, 2026 17:04
…gistration

PendingRegistration and LiveRegistration named the partition's state when
the registry answered them, which a deletion by anyone falsifies: a "live"
registration may have been deleted since. What does not change is what
each handle's holder may do, so the handles are named for that, in the
words of two known patterns: Try-Confirm/Cancel for the creator, and the
stale handle for everyone else.

- `register` is `reserve`, and answers a `Reservation`: the creator's hold
  on the key while it prepares the partition's storage.
- `PendingRegistration.mark_live` is `Reservation.confirm`, which marks the
  partition live and answers its `Registration`.
- `PendingRegistration.unregister` is `Reservation.cancel`.
- `LiveRegistration` is `Registration`; `resolve`, `require_current` and
  `unregister(partition_key)` keep their names.
- The fields both share sit in a private base, `_RegistryEntry`.

"Pending" stays the word for the partition's state, in the store contract
and VectorStorePartitionPendingError. The base store's creation flow, its
task set and log text, the Qdrant and Milvus handles, the tests and the
design documents follow; the registry design records why the names are
roles.

The same change as MemMachine#1734's ffa954c, MemMachine#1735's 87edbad and MemMachine#1736's
8e65391, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…ain open-or-create's give-up

Two review changes to MemMachine#1734's registry and base store, for this PR's
partitions:

- `Reservation.cancel()` acts only while the partition is pending, as
  `confirm()` does. A creator whose confirmation committed but whose
  answer was lost, and which then cancels, leaves the live partition
  alone: only a deletion by key ends it. The registry contract, the
  SQLAlchemy registry and the registry design document say so, and a
  test cancels after a confirmation.
- Open-or-create's `VectorStoreAttemptsExhaustedError` is raised from the
  race it last lost, the last `VectorStorePartitionAlreadyExistsError` or
  `VectorStorePartitionDeletedError` it caught. A partition that stays
  pending still raises the pending error itself. A base-store test checks
  the cause.

The same changes as MemMachine#1734's 9daf7e7 and 93a85da, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…1734's other review changes

MemMachine#1734's latest review changes, for this PR's partition registry and stores:

- `run_purge_round(purge_round)` replaces `claim_purgeable_incarnation()`.
  The registry claims the oldest due tombstone, calls the round with its
  incarnation inside the claim's transaction, and records what the round
  returns. `PurgeClaim` and its `any_records_found` guard go; `PurgeRound`
  takes only the incarnation, since a tombstone here carries nothing else.
  On SQLite the claim is an `UPDATE ... RETURNING` of the oldest eligible
  row, so purgers serialize at the claim; PostgreSQL keeps `FOR UPDATE SKIP
  LOCKED`. A failure count at or past the dead-letter bound is reported.
- A reservation's cancel reports its own failure from its task, so a
  creation cancelled again still has the failure logged.
- `get_partition` runs under the tracker like the other lifecycle calls,
  and the SQLite stores check partition keys with the shared
  `require_partition_key`.
- The registry and purge documents follow. The upgrade notes state what
  holds for these stores: they name their native collections by vector
  store name, which no earlier release did, so an existing Qdrant or Milvus
  collection is never read or purged, and can be dropped before or after
  upgrading. The Milvus design document's consequence, which said an
  existing collection has to be dropped, says the same, and the Helm
  README lists `partition_registry` among the Qdrant store's keys.

The same changes as MemMachine#1734's d041785, cd13045, 35b57b7, 0a68f37,
d05556b, 118d23b, 5225933 and 783967e, MemMachine#1735's a40bc50 and
MemMachine#1736's caefa7b, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…r is cancelled

A creation cancelled its reservation only when preparing the partition's
storage raised. A confirmation that raised, or a creation cancelled while
it confirmed, left the partition pending until someone deleted it.

The base store's creation step now prepares the storage and confirms the
reservation together, and cancels the reservation if either raises or the
creation is cancelled, shielded as before; the cancel's task reports its
own failure. The cancel acts only on a pending partition, so a
confirmation that committed before its failure was observed stands.
create_partition and open-or-create both go through it; open-or-create
still takes a confirmation's VectorStorePartitionDeletedError as a race to
create again. Base-store tests cover a failed confirmation, a cancelled
one, and one that committed before failing. The registry design document
says so, and the purge document says which writers wait on SQLite's lock
during a purge round.

The same changes as MemMachine#1734's f4c585e and 4dde7c8, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…saction across the round

A purge round ran inside the claim's transaction: on PostgreSQL a row lock
held idle while the backend deleted, and on SQLite the database file's write
lock, which every store sharing the file waited on for the whole round.

- The claim is a lease. One committed `UPDATE ... RETURNING` takes the
  oldest due tombstone that no unexpired claim holds, stamping `claimed_at`
  and incrementing `claim_generation`; the round runs with no transaction
  open; the writes that end the claim are conditioned on its generation, so
  a round that outlasted its lease cannot end the claim taken after it. A
  round that found nothing removes the tombstone under any claim.
  `purge_lease_seconds` (default 300) sets the lease.
- A claim that finds the previous claim unended past its lease runs no
  round: it counts that round as failed, as of when it was claimed, and
  logs it. A cancelled round ends its claim uncounted, in a shielded write
  whose task reports its own failure.
- `purge_retry_backoff_seconds` is `base_purge_retry_backoff_seconds`, the
  first delay the backoff doubles.
- The registry refuses an engine on StaticPool or in-memory SQLite, whose
  connections do not arbitrate as separate transactions; the wiring tests
  give their registry a file database.

The registry tests cover the lease, the outcome writes and their fence,
deletion and reservation atomicity under injected faults, two registries
sharing a database, collision without waiting, churn across engines, and
random operation sequences against a model. The purge and registry design
documents describe the lease and the alternatives considered.

The same changes as MemMachine#1734's f1e9de8, 054b079, a4e2b24 and 900a926,
MemMachine#1735's 3d7761e and 81b1bf5, and MemMachine#1736's 1b94f7f, for this PR's
partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The query contract said nothing of a match scoring exactly the threshold,
and Qdrant dropped it: the server compares its threshold in single precision
and keeps only scores strictly better than it.

The contract now says a match scoring exactly the threshold is returned.
The Qdrant store sends the server the adjacent single-precision value on
the worse side, and applies the caller's threshold exactly itself; a
threshold beyond single precision is not sent. Both SQLite stores already
keep the match, and each now tests it on every metric it supports.

The same changes as MemMachine#1735's ccba117 and 6bbc131, for this PR's stores.
MemMachine#1736's cc29d27, Milvus's test of the same, is ported with the store
tests.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…cycle contract's drain

MemMachine#1735's and MemMachine#1736's latest tests, for this PR's partitions:

- Every Qdrant store test runs against a Qdrant server, over REST and gRPC:
  local mode ignores payload indexes and raises its own exceptions where a
  server answers not-found or already-exists, so it is no longer a fixture.
  The tests on a mocked client stay in the default suite.
- Both stores gain tests that partitions and stores of different names keep
  their records apart, that purge rounds reclaim a write landing after a
  round and drain two stores' tombstones alone, that ranking, scores and the
  threshold follow every similarity metric, that a point or entity the
  server refuses on its own raises, that concurrent creations and startups
  agree, that two stores churning one registry keep every partition exact,
  and that seeded operation sequences, and on Milvus random filters, agree
  with a model. The Milvus tests also pin its read consistency and its
  settings with values other than their defaults.
- The lifecycle contract's drain fails after a bounded number of rounds, it
  settles before checking that a new life is empty, and it checks a stale
  upsert by what the partition holds rather than by its registry reads.

The same changes as MemMachine#1735's 30c30f4, e9e71b5, 6c2362b, 4d3241c,
836db21, 252bcd2, 5e8c049, ef8d3ab, 276a034, 4fb5bd4 and
2aabd3d, and MemMachine#1736's 509c235, 7f1eb3b, 6e312e0, 6d1ed1e,
25e7948, 6f36877, d43e9e6, 6dfbf59 and cc29d27, for this PR's
partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The base store's purge contract says a round that finds the incarnation's
storage missing returns False, but neither store's round checked: once its
native collection was dropped from outside, every round on it raised,
counted against the tombstone, and dead-lettered it after ten.

A Qdrant round that the server answers not found, over REST or gRPC, and a
Milvus round that finds no native collection, now return False: the
collection is gone with everything in it, so the tombstone is retired. Each
store has a test that drops its native collection and drains the purge.

The same handling as MemMachine#1735's and MemMachine#1736's stores, whose rounds already
checked for a missing native collection; MemMachine#1736's 2575af0 tests it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…n, and use serial commas

MemMachine#1734's, MemMachine#1735's, and MemMachine#1736's latest round, for this PR's partition registry
and stores; no behavior changes:

- The queue's `failed_rounds` column is `consecutive_failed_rounds`, in the
  code, the tests, the documents, and the dead-letter log's hint.
- The registry keeps its durations as seconds. `_has_elapsed(seconds,
  since=)` answers whether a duration has passed on the database clock, for
  the retention, the backoff, and the lease, and
  `_purge_retry_backoff_seconds()` computes each tombstone's capped backoff.
- `run_purge_round` runs named steps: `_claim_oldest_due_tombstone` answers a
  `_TombstoneClaim`, an `_UnendedPurgeRound`, or None, and the round's
  outcome goes to `_count_failed_purge_round`,
  `_end_tombstone_claim_after_cancellation`, or `_record_purge_round`.
  `_insert` is `_insert_pending_partition`, `_claim_releases`
  `_tombstone_claim_endings`, and the base store's `_cancellations`
  `_reservation_cancellations`.
- The Milvus filter helpers are named for what they produce, comments are
  shorter, and lists in the documents, comments, and docstrings take a
  serial comma.

The same changes as MemMachine#1734's 2a3d87c and 503687c, MemMachine#1735's 106e106, and
MemMachine#1736's 3681d09 and bc6d4f9, for this PR's partitions.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A round that never ended was counted by the next claim, which took a
CASE-shaped update that both recorded the lost round and ended its claim;
the count lived apart from the claims that made it.

Each claim now counts its attempt: the claim is a plain update that
increments `consecutive_attempts` (renamed from `consecutive_failed_rounds`),
stamps `claimed_at`, and bumps the generation, and `_TombstoneClaim.attempt`
carries the count. A tombstone is claimable when it has no attempts, when no
claim is open and the backoff has passed since its last failure, or when an
open claim has outlived its lease plus the backoff. A raised round ends its
claim and stamps `last_failed_at`; a cancelled round ends its claim and takes
its attempt back; a round that found records resets the attempts. After
`_MAX_PURGE_ATTEMPTS` attempts a tombstone is dead-lettered, and a last
attempt that raises is reported. A retry logs which attempt it is, and a
raised round's error names its incarnation and attempt. The registry tests,
the purge design document, and the registry design document follow.

The same change as MemMachine#1734's 1732038, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
`consecutive_attempts` counted the purge rounds claimed since a round last
found records, but its name said only that they were in a row. It is
`attempts_without_progress`, and `_MAX_PURGE_ATTEMPTS` is
`_MAX_PURGE_ATTEMPTS_WITHOUT_PROGRESS`. The column's comment, the claim's
attempt, the backoff parameter's description, the dead-letter log, the
class docstring, the tests, the purge design document, and the registry
design document's table follow; nothing else changes.

The same change as MemMachine#1734's 96a8b5b, for this PR's partition registry.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…urrent claim

The measured cost of the claim came from earlier forms of it, which held a
transaction across the round. The section now gives the figures measured on
2026-10-06 against the claim this registry ports, naming the MemMachine#1734 commits
they were taken at: the backoff scan with 1k, 10k, and 100k tombstones
backing off, and interactive throughput and liveness latency beside two
sweepers on PostgreSQL and on SQLite.

The same change as MemMachine#1734's bab379b, for this PR's purge document.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's review round, for this PR's Milvus store:

- Filter strings reach Milvus as UTF-8, so a character outside the Basic
  Multilingual Plane parses, and an undeclared property's condition requires
  the stored type tag to be one the value compares with.
- Startup's already-exists guard goes, with its mocked-client test: Milvus
  answers a create of an existing collection with the same schema with
  success.
- The store no longer re-sorts search results, which Milvus returns best
  first.
- The mocked-client purge test disposes its registry engine when it fails.

The same changes as MemMachine#1736's 1a7fffa, d6fc219, 9fd7f76, 6fcd8c5, and
c264b3d, for this PR's store. MemMachine#1736's 28e8c50 and b3ef4f1 have no
counterpart here: startup prepares the store's one native collection before
any purge round runs, and the store refuses an unsupported metric at
construction, before anything is reserved.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The declared path counted an int and a float as comparable either way, so
a float filter value on a declared int property reached Milvus as a float
literal against an INT64 field, which Milvus refuses to parse ("cannot
cast value to Int64", code 1100). A float now compares only with a float
property, and matches no int one, as a value of another type matches
nothing; an int still compares with a float property by value.

The new test filters a declared float property with an int and a declared
int property with a float; it fails with the float counted as comparable
with an int property.

The same change as MemMachine#1736's 31fc7c4, for this PR's store.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…y name

MemMachine#1736's baabf48 creates Milvus native collections at Bounded by name,
and this PR's base carries that into startup's create, so the read level
the purge and the tombstone retention rely on no longer comes from
pymilvus's default. The consistency test now spies create_collection and
checks the level it names, as MemMachine#1736's test does; it fails with the level
dropped from the create. The design document and the store's docstring
say the store creates its one native collection at Bounded.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
MemMachine#1736's f56686f halves a Milvus upsert refused with RESOURCE_EXHAUSTED,
and this PR's base carries the halving into the partition handle's
`_upsert`. Its tests come here in partition terms: the integration test
upserts 1,200 records with a 60,000-character property, about 72 MB, and
fails with RESOURCE_EXHAUSTED without the halving; the mocked-client tests
pin the halving, a single refused entity raising, and a timeout or another
refusal sent once, on a handle `_partition_on` builds, which the mocked
delete test now uses too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…nt alone

MemMachine#1736's 49b9ec2 drops the offset field beside a declared Milvus
datetime, keeping its UTC instant alone, and this PR's base carries that
into the store. The tests follow in partition terms: the native
collection's fields are exactly the fixed ones and one per declared
property; a declared datetime is stored as its instant in UTC; and a store
whose declared datetime has the longest property key, 32 bytes, writes its
one field, reads it back, and matches the record by its instant, the test
MemMachine#1736's b015229 adds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
The test of MemMachine#1813, merged into this PR's base, in partition terms: the
server defaults new collections to strict mode, the store's collection
reads back with it off, and a filter on a partition's unindexed property
is served.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
A partition stored every property of a record and filtered on any key,
which made a caller's arbitrary keys part of the store's schema: the
SQLite stores kept them in a JSON column and filtered with json_extract,
Milvus the keys its schema did not declare in a JSON field, and a filter
on a key the store never indexed scanned. Since EventMemory routes a
filter on an undeclared key to the segment store, the vector store need
not hold undeclared keys at all.

A partition now stores the properties its store declares and no others.
`upsert` raises UndeclaredPropertyKeyError before anything is sent for a
record naming an undeclared key, and PropertyTypeMismatchError for a
value of another type than its key declares; `query` raises
UndeclaredPropertyKeyError for a filter naming an undeclared key and
UnsupportedFilterError for a node outside the partition's
`supported_filter_nodes`. Both SQLite stores keep one typed, indexed,
nullable column per declared key on the records table (sql_columns.py);
sqlite-vec 0.1.9 rejects NULL in a vec0 metadata column and a declared
key is optional per record, so that store keeps the columns on the
records table and hands the KNN a `rowid IN (SELECT ...)` allowlist,
evaluating the filter during the search instead of after it. Qdrant
drops the JSON copy and keeps a payload field per declared key; Milvus
drops its JSON field and keeps its typed field per declared key.
Datetimes are stored as microseconds since the epoch where a backend has
no datetime type.

Since every key a filter may name is now indexed, the Qdrant store
creates its collection in strict mode (`unindexed_filtering_retrieve`
and `_update` false, Qdrant Cloud's default): a filter on an unindexed key is refused by the server instead
of scanned for. A leaf whose value is of another type than its key
declares matches nothing, as on the SQL stores; the Qdrant compiler
answers it with a filter no point satisfies, since the server would
refuse the condition for the field's index, and Milvus no longer
compares an int with a float key. Local mode does not record the
setting, so a unit test checks the request and integration tests the
server's answer.

`declared_schema_contract.py` states the contract every backend's test
module runs: which records a filtered search admits, over fixtures small
enough that every backend searches them exactly, checked after each
upsert so an approximate index fails on recall, by name, and not on the
filter.

On the registry-backed stores, the checks are the base handle's: `upsert`
runs require_declared_properties in place of the type check it ran, and
`query` runs require_supported_filter against the subclass's
`supported_filter_nodes`, which each subclass now implements. A datetime
column on the SQLite stores has a `tz_<key>` column beside it holding the
UTC offset in seconds, written with the value and read by no filter, so a
stored datetime is the value written, as on Milvus, Qdrant and the segment
store; the microseconds column alone would keep only the instant.

The Qdrant and Milvus design documents describe the declared-only
properties and Qdrant's strict mode.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A declared datetime on the SQLite stores had a `tz_<key>` column beside its
microseconds column, holding the UTC offset in seconds that no filter
reads. The vector stores keep a datetime property's instant only, as
MemMachine#1736 now does for Milvus and MemMachine#1788 for Qdrant, and the segment store
keeps the offset, so the column, `offset_column_name`, and the value
written to it go: each declared property is one column, and a datetime
stays microseconds since the epoch.

The tests read a stored datetime back as its instant in UTC, and the
roundtrip test checks that the stored instant equals the written one.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Every embedder MemMachine ships produces vectors meant to be compared by
cosine -- OpenAI hard-coded it, Bedrock defaulted to it, SentenceTransformer
only reported what the model declared -- while every layer that touched a
score paid for the other three metrics in direction flags, threshold
directions, and per-backend tables mapping the enum onto native metric
names. `SimilarityMetric` is gone; scores are cosine similarities in
[-1, 1] and the names say so: `QueryMatch.score` and `SearchMatch.score`
become `cosine_similarity`, and `query(score_threshold=)` becomes
`query(min_cosine_similarity=)`, which no longer needs a direction to be
meaningful and is refused when not finite, as the threshold was. A store
and its partitions no longer have a `similarity_metric`, the schema a
partition is registered under no longer records one, and a search engine
factory takes the dimensions alone. The Bedrock embedder's
`similarity_metric` config key and semantic memory's
`vector_similarity_metric` go with it, and the install and configuration
docs drop them.

The vector graph stores carried a metric per stored embedding, as a
companion property beside every vector; that is gone and `Node.embeddings`
holds plain vectors. NebulaGraph's `cosine()` cannot take `APPROXIMATE` and
its vector indexes offer only L2 and IP. Cosine similarity between unit
vectors is their inner product, so embeddings are normalized on the way in
and compared with `inner_product()` against an IP index.

The cosine half of MemMachine#1598 (`abf92a3a4` on speedkick), re-derived on the
one-collection store: the rest of MemMachine#1598 (queries answer record UUIDs and
scores, no `get`, semantic memory's `vector_uuid`) and all of MemMachine#1603 are in
registry-backed base and its partition handle lose the metric, the stores'
own threshold checks become `require_valid_min_cosine_similarity`, and the
design documents describe cosine scoring and a schema without a metric.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…emMachine#1588)

* Atomically swap vector search engine index files on save

SQLiteVectorStore persists each collection's index by calling the search
engine's save(), which wrote directly to the final path. A crash mid-write
left a truncated/corrupt file. Because index_saved=True makes the on-disk
index a durable contract (missing/corrupt is a hard IndexLoadError, not a
silent empty rebuild), an interrupted save could render a collection
unrecoverable.

Write the index to a sibling temp file and swap it into place with
os.replace (atomic on POSIX and Windows on the same filesystem), so a reader
sees either the old or new index, never a partial write; a failed save leaves
the previous index intact. Leftover temp files are cleared on load so a crash
does not leak them across restarts.

Implemented in the engines (shared index_persistence helper) rather than in
SQLiteVectorStore/SQLiteVectorStoreCollection, since the index save location
and number of files written differ across engine implementations.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* Make the index swap durable, not only atomic

The swap protects a reader from a torn index, but the vector store also
trims its pending-operation log once `save` returns -- and that log is the
only other copy of those vectors, since the records table stores no vector
column. So the swap reaching disk is load-bearing rather than a bonus:

- fsync the parent directory after the replace, since POSIX `rename(2)`
  leaves the new directory entry in the page cache. Best-effort and ignored
  on failure, matching SQLite's `unixSync`; a no-op on Windows, which has no
  equivalent operation.
- stop swallowing a failed fsync of the temp file. SQLite draws the same
  line -- a file fsync failure raises SQLITE_IOERR_FSYNC while a directory
  fsync failure is ignored -- and `EIO` means the writeback already failed
  and the dirty pages were dropped, which is exactly when the save must not
  be reported as committed. The existing cleanup then leaves the previous
  index in place with the log untrimmed, so the next save retries.
- use F_FULLFSYNC on macOS, where plain `fsync` leaves the data in the
  drive's volatile write cache, falling back when a filesystem refuses it.

State the resulting obligation on `VectorSearchEngine.save` itself, since
that is what the store now relies on: replace atomically, then make the
replacement as durable as the platform allows. An engine whose backend
already implements the whole protocol can delegate to it and skip these
helpers; the rest use `atomic_index_write`.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let the engine own index durability, not the vector store

The pending log holds the only durable copy of a vector between
checkpoints -- the records table has no vector column -- so trimming it is
safe only at an instant when the index provably holds those vectors. The
temp-write + rename protocol this PR shipped could not provide that
instant. A rename changes a directory entry, and Windows exposes no way
to flush one: os.fsync is _commit, which is FlushFileBuffers, which is
for file data, and you cannot open a directory to fsync it. The decisive
evidence is SQLite's own -- it threads a directory-sync flag through
every commit-relevant directory operation, honors it in unixDelete, and
declares it /* Not used on win32 */ in winDelete. So os.replace could
return, _save_collection_index could commit its trim durably behind it,
and a power cut could still roll the rename back: records forward, index
back, no copy of the difference left. MOVEFILE_WRITE_THROUGH is not a
fix; its documented guarantee covers copy-and-delete (cross-volume)
moves, not same-volume renames.

Take SQLite's answer, which was not to harden the directory operation
but to stop using one as a commit point (PERSIST commits by zeroing a
header, TRUNCATE by truncating, WAL by appending frames).

A base path now expands into two index slots plus a generation record
each, created once and thereafter only overwritten. A checkpoint writes
the index over the inactive slot and flushes it, then writes that slot's
generation record and flushes that. The record is the commit, and it is
a write into a file that already exists. It holds the generation and its
bitwise complement, so a torn write reads as absent rather than as some
other generation -- all or nothing without needing single-sector
atomicity from the hardware. load takes the highest believable
generation, and deliberately does not fall back to the older slot when
the published index will not parse: the log was trimmed against the
newer one, so the older is stale by exactly the ops that can no longer
be replayed.

Both backends already write straight to the path they are given, which
is what this protocol wants -- verified that repeated saves preserve the
inode and leave no stray files -- so no engine gains a temp file, a
buffer, or a rename.

Durability is entirely the engine's, including which artifact is live.
The store keeps no slot pointer, manifest, or generation, so no schema
change and no migration: what remains is one rule, never trim past what
save says is durable, and _save_collection_index already had that order.
index_path becomes index_base_path since it no longer names a file, and
discarding a collection asks the engine layer which files that covers.

BREAKING CHANGE: an index written by the previous protocol is not
published under the new one, so a collection with index_saved=True
raises IndexLoadError until its index directory is cleared and the
records re-ingested.

Anomaly tests walk every crash point in the publish sequence by
constructing the on-disk state each would leave, plus one that pins the
ordering itself (a failed index write must publish nothing) since
state-based tests cannot observe it. Verified against three deliberate
breaks -- dropping the complement check, writing the record first, and
reusing one slot instead of alternating -- each caught.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Publish the index atomically, and stop promising durability

The two-slot generation-record protocol bought a guarantee we have
decided not to make: that a save survives a power failure. Every engine
would have to implement and maintain that protocol, and the failure it
buys out is bounded -- search recall for the records applied since the
last checkpoint, repaired by re-ingesting them. The direction that
actually costs, a published index that will not parse, is closed by the
atomic swap on its own.

So this returns to the temp-file-plus-rename publication and spends the
difference on stating the contract instead of strengthening it: `save`
publishes atomically, never durably; the store trims the pending log
behind a publication a power failure can revert; a record whose vector
is lost that way still resolves by uuid, is absent from search until it
is upserted again, and nothing here detects the gap for the caller.

Reverts the durability and engine-owned-publication commits, keeps the
atomic swap, and adds a store-level test that reconstructs a reverted
publication deterministically -- restore the previous index bytes after
the trim -- to pin the direction it fails in.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Report a lost embedding as lost, not as a missing feature

`update_feature` reads the stored embedding back when a caller updates a
feature without supplying one, and that is the only place in the server
that depends on the index still holding a vector. With publication now
atomic rather than durable, a power failure can leave a feature whose
row is intact and whose vector is not -- a state this path reported as
"Vector record not found", which points the caller at the wrong thing
and hides the repair.

Split the two cases. A record that is genuinely absent keeps the old
message; a record whose embedding the index no longer holds says so and
names the fix, which is to pass a fresh embedding.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>

* Let a failed fsync fail the save, and cut the essay around it

`_flush_to_disk` wrapped its fsync in `contextlib.suppress(OSError)` and
called itself best-effort. A failed fsync is exactly the evidence that the
bytes are not safe to publish -- on Linux an EIO from fsync means writeback
failed, reported once and then cleared -- so swallowing it and renaming
anyway published a file we had positive evidence was bad. The safeguard
cost something and, in the one case it existed for, guaranteed nothing.
Nothing tested it either.

Let it propagate. `atomic_index_write` already unlinks the temp and
re-raises, so a failed flush now leaves the previously published index
standing, which is the correct outcome. A test pins that.

The fsync is not best-effort, and the docstring should not have said so: it
rules out a class rather than narrowing a window. Because the flush
completes before the rename is issued, and a durable write does not
un-happen, the new name can never appear over incomplete bytes. What the
missing directory fsync costs is the other direction -- the rename may not
survive, so the publish reverts -- and that is the benign one this store
already accepts.

The module docstring was 76 lines against 49 of everything else, most of it
argument rather than documentation: a walk through SQLite's `unixDelete` /
`winDelete` sync-flag handling, and a rejected two-slot commit protocol.
That is the PR's case for the design, not something to re-read every time
someone opens a 20-line module, and the PR body carries it. What a reader
here needs is the guarantee, the non-guarantee, and the cost.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Say why the swap is a rename, and what the reopened fd cannot see

Two things the module was silent on.

Why a rename at all. The stronger answer is to put the commit point inside
the file, where an fsync reaches it portably -- SQLite never renames, and
commits by truncating or zeroing its rollback journal, or in WAL mode by
appending frames whose checksums make a torn tail self-identifying. Both
need the writer to own the file format. A search engine owns its own and
exposes `save(path)`, so above that call a rename is the only atomicity
primitive left, and an engine whose format already commits that way needs
none of this. Worth saying, because "why not do the better thing" is the
first question the module invites.

What the reopened descriptor cannot see. Flushing is fine on a fresh fd --
dirty pages belong to the file, not to the descriptor that dirtied them --
but error reporting is not: Linux hands a writeback error to descriptors
open when it was recorded, so one recorded between the engine's close and
this open is never reported and the save proceeds on bytes already known
bad. Same shape as the 2018 PostgreSQL fsync report. It cannot be closed
from here: the engine writes through its own descriptor and closes it
before returning, and closing the window needs an engine that writes
through a handle the caller supplies.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fsync the index on a descriptor that predates the write

The fsync was on a descriptor opened after the engine had written and
closed its own, which flushes correctly -- dirty pages belong to the file,
not to the descriptor that dirtied them -- but reports nothing useful.
Linux samples the writeback error sequence when a file is opened, so a
descriptor opened after an error was recorded never learns of it: the fsync
returns success and the save publishes bytes already known bad. Same shape
as the 2018 PostgreSQL fsync report.

Open the temp before yielding it and hold it across the caller's write, so
the descriptor predates the bytes and any error from writing them is
reported here, where it fails the save.

That assumes the caller writes in place. Both engines do -- verified: the
inode is unchanged across `save_index` and `save`, and the held descriptor
sees the written size -- but it is their behaviour, not their contract. An
engine that built a file of its own and renamed it over the temp would
leave this descriptor on an orphaned inode, and the fsync would report on a
file nobody is about to publish. So it is checked before the fsync, and a
mismatch fails the save rather than passing it quietly.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* State the in-place rule where the caller reads it

The descriptor held across the body only flushes what the body wrote if the
body writes the yielded path in place, and that requirement was recorded in
`_flush_to_disk` -- a private function nobody writing an engine opens. It
belongs on `atomic_index_write`, which is the API they use, alongside what
happens when it is broken: an `OSError` and no publication, so the mistake
surfaces at the first save rather than at a power cut.

`_flush_to_disk` keeps the mechanism -- why the descriptor has to predate
the write -- and now points at the rule instead of restating it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Ask the drive to flush on macOS, where fsync does not

`os.fsync` is not the same guarantee on all three platforms this ships to.
On Linux it flushes to the device, and on Windows `FlushFileBuffers` does
the same. Darwin's `fsync` explicitly does not: it returns once the data
reaches the drive, which may hold it in a volatile write cache. So on macOS
the ordering this module is built on -- data durable before the rename is
issued -- did not hold at the device, which is exactly the case it claims
to rule out.

`F_FULLFSYNC` asks the drive to flush that cache. Filesystems that cannot
refuse it, and there `fsync` is the most that can be asked, so a refusal
falls back; any other error is a write failure and propagates, as before.

The flush-failure test patched `os.fsync`, which Darwin no longer reaches.
It patches the module's own `_fsync` instead, which every platform does.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Stop guessing which errnos mean "this drive cannot do that"

The fallback from `F_FULLFSYNC` to `fsync` was gated on an errno allowlist
-- ENOTSUP, EOPNOTSUPP, EINVAL -- so that a genuine write failure would
propagate rather than be quietly downgraded. Probing the actual returns on
Darwin shows the list is both incomplete and partly invented:

  ENOTSUP=45  EOPNOTSUPP=102   distinct here, so both are needed
  /dev/null   F_FULLFSYNC -> ENODEV(19), while fsync succeeds
  pipe/socket F_FULLFSYNC -> EBADF(9)
  EINVAL      never came from F_FULLFSYNC at all; it came from fsync

So ENODEV -- a real refusal, on a path anyone can reproduce -- would have
raised instead of falling back, and EINVAL was in the list by analogy
rather than evidence. What a network mount answers is not knowable from
here, which makes the whole list a guess that fails closed on whatever it
missed. This codebase does not classify driver errors by guessing, and
this was that.

Fall through on any failure instead. It is not a suppression: `fsync` runs
on the same descriptor and raises in its turn, so a flush that cannot
happen still fails the save. What the fallback gives up is the drive-cache
flush -- the guarantee this had before `F_FULLFSYNC` was asked for at all.
That is also what SQLite does with this same call, for the same reason.

The test drives it through `/dev/null`, which refuses with ENODEV and
accepts `fsync`; it fails against the allowlist and passes without it.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

* Fail the save when the drive cannot be told to flush

`F_FULLFSYNC` failing fell through to `fsync`, on the reasoning that a
refusal is a statement about the filesystem rather than a write failure.
That reasoning does not survive asking what `fsync` alone actually buys on
Darwin.

Against a process or kernel crash it is enough: the data has left the OS
for the drive before the rename is issued, so the rename dies in the page
cache and the old index stands. Against power loss it is not. The data
sits in the drive's volatile cache, the rename's metadata joins it moments
later, and nothing orders them -- and the rename is a few bytes against an
index of megabytes, so a drive flushing as it pleases can easily put the
new name on media while the bytes behind it are still queued. That is the
torn publication this module exists to prevent, in precisely the scenario
its docstring is about.

So the fallback answered a request for ordering with a flush that does not
provide it, and said nothing. A filesystem that cannot order data ahead of
a rename is not one to publish an index onto; raise, and let the operator
point `index_directory` at storage that can.

This is also the simpler code. Refusal and failure now take the same path,
so no errno is inspected -- there is no line to draw and no list to get
wrong, which is what the previous two revisions kept getting wrong in
opposite directions. The `/dev/null` test went with the fallback it pinned.

The earlier defence of falling back rested on network and FUSE mounts
being a realistic home for an index directory. They are not.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
(cherry picked from commit 21105d6)
Never reuse a row id in SQLiteVectorStore

The records table used a plain INTEGER PRIMARY KEY, which is SQLite's
rowid, assigned as max(rowid) + 1: deleting the highest row freed its id
for the very next insert. query() scores keys in the search engine and
resolves them to rows in a second step, holding nothing in between, so a
reused id let a record that was never scored come back wearing the score
of the record that was. Nothing about that result looks wrong: the record
exists and the score is in range.

Declare the table with sqlite_autoincrement=True so ids are never reused.
A stale engine key then matches no row and is dropped.

Both tests fail without the flag: one pins the id policy directly, the
other parks a query between scoring and row lookup, retires the scored
record, inserts another, and asserts the query returns nothing.

First half of MemMachine#1468.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…(speedkick) (MemMachine#1612)

Own the search engine's concurrency in the store, not in each engine

Each engine wrapped its five methods in its own read-write lock and the
abstract class promised concurrent use. The store is the only caller,
and it already knows which calls must exclude which: a search shares the
engine with other searches, and everything else runs alone. Holding that
lock in the store puts every exclusion in the one file that reads the
engine, and lets a rule the engines could not express hold: a rewrite's
remove and add are one step to a reader, where before a search could run
between them and see neither version.

Engines drop their lock and keep only the index calls; the abstract
class now states that the owner serializes. The store keeps one
read-write lock per collection beside the engine, kept for the store's
lifetime like the engine's other per-collection state, and takes it at
every engine call: searches on the read side; mutations, loads, and the
index save on the write side. A rewrite's remove and add sit under one
hold. The save's trim runs after the lock is released, so readers wait
for the file write and never for SQL.

The row-id regression test from MemMachine#1589 parked inside a wrapper engine's
search, outside the real engine's lock; under the store's lock that
parks the read side, and the writes it then awaits cannot proceed. It
now parks where it meant to, between the engine search and the row
lookup. One new test pins the one-step rewrite; it fails against the
engines' own locks.

The turbovec engine in flight carries the same lock and needs the same
subtraction.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…edkick) (MemMachine#1607)

Serialize a collection's writes so the engine sees them in order

A write commits to SQLite and only then applies to the search engine, so
two writers to one uuid could reach the engine in the opposite order to
the one they committed in: an upsert overtaking a delete re-adds a vector
for a record that is gone, and two upserts inverting leave the engine
serving the older vector. Never reusing a row id does not cover this,
because an upsert of an existing uuid keeps its row id.

A save in that window costs a write outright: it publishes the index and
trims every applied log row, and a write that applied after the index was
written is then in neither.

A per-collection asyncio.Lock now spans a write from SQL commit through
engine apply, mark-applied, and any save it triggers; shutdown's save
takes it too. The lock belongs to the store, not to a collection handle:
a handle is constructed per open_collection call, so several can address
one collection, and only a shared lock serializes them. Readers are
untouched.

Three tests fail without the lock, each interleaving made deterministic
by gating the engine: an upsert overtaking a delete of its uuid, a save
trimming a write it did not publish, and that overtake across two
handles, which a per-handle lock passes. The rest pin behavior the lock
must preserve: an upsert surviving a delete of another record, disjoint
concurrent upserts and deletes, writes racing a checkpoint, and batches
that name one uuid twice.

Fixes MemMachine#1468.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…MemMachine#1609)

Take SQLite's write lock at BEGIN, not at the first write

`delete` resolves row ids and then writes, in one transaction. Under
SQLite's default deferred BEGIN the write lock is taken at the first
write, so another writer can take it between that read and the write,
and the transaction that read first is the one that loses: its own write
cannot upgrade, and SQLite reports that immediately rather than waiting
out busy_timeout, because waiting could only deadlock. Reproduced in both
journal modes:

    journal_mode  BEGIN      other writer  our write
    delete        deferred   shut out      fails
    delete        IMMEDIATE  shut out      ok
    wal           deferred   commits       fails
    wal           IMMEDIATE  shut out      ok

The store now emits BEGIN explicitly and lets a transaction ask for
BEGIN IMMEDIATE, which every write path does. The mode is chosen per
transaction, not per engine: a hook that asked for IMMEDIATE
unconditionally would make every read take the write lock, and two
readers would then serialize against each other.

Two tests fail without it. One holds a competing lock across a
read-then-write transaction and asserts that transaction completes;
asserting instead that the other writer is excluded passes either way,
because a deferred BEGIN shuts it out too, later and by a different lock.
The other races four creates of one name: under a deferred BEGIN the
losers read no stored config, go on to CREATE TABLE, and fail there with
"table already exists" instead of VectorStoreCollectionAlreadyExistsError.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…chine#1672 to MemMachine#1676) into the final state

Their upsert keeps the delete and fresh insert, writing the declared
columns, and refuses a repeated UUID as MemMachine#1788 does; open-or-create stays
out, as MemMachine#1625 removes it. Their tests create their partition as a session
does and build filters with the closed union of MemMachine#1616.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@edwinyyyu
edwinyyyu force-pushed the port/fresh-row-id-per-write-main branch from 102203e to a070372 Compare October 10, 2026 00:32

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant