Repository navigation
Give every write a fresh row id, so a key names one version (speedkick) - #1610
Merged
edwinyyyu merged 1 commit intoSep 11, 2026
Conversation
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-versioned-keys-speedkick
branch
3 times, most recently
from
September 10, 2026 22:25
bae0d8b to
b9d2aff
Compare
This was referenced Sep 10, 2026
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-versioned-keys-speedkick
branch
9 times, most recently
from
September 11, 2026 01:01
04c3a05 to
135052a
Compare
An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.
Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.
A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.
Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:
batch save threshold insert before / after rewrite before / after
500 1000 20477 / 19290 16450 / 12902
500 none 21228 / 22486 19013 / 15899
1 none 341 / 345 316 / 274
Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-versioned-keys-speedkick
branch
from
September 11, 2026 03:17
135052a to
eada6db
Compare
This was referenced Sep 14, 2026
Merged
Closed
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 25, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This was referenced Oct 1, 2026
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 1, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 2, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 2, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 3, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 3, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 6, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 6, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 6, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 6, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 7, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This was referenced Oct 7, 2026
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 7, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 7, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 8, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 9, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 9, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 9, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 9, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 9, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 10, 2026
…k) (MemMachine#1610) Give every write a fresh row id, so a key names one version An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them, and a rewrite of the very record being returned could land between those reads: the score of one version paired with the filter verdict of another, or, when the overfetch loop rescored a key whose verdict was already cached, the other way round. No lock covers this, because reads deliberately take none. Every write now deletes the previous row and inserts a new one in the same transaction, stages a delete for the old key beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version: the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version, and a key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order. A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often; that test's expectation changes accordingly. Measured on this machine, medians of three runs, records per second, 64-dimensional vectors, file-backed store with an index directory: batch save threshold insert before / after rewrite before / after 500 1000 20477 / 19290 16450 / 12902 500 none 21228 / 22486 19013 / 15899 1 none 341 / 345 316 / 274 Inserts move within run-to-run noise, in both directions. Rewrites cost 13-22% more, the upper end where the doubled log rows double the saves. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose of the change
An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them. A rewrite of the very record being returned could land between those reads, pairing one version's score with another version's filter verdict. Concretely, with a record that fails the filter at version 1 and passes at version 2: the engine scores version 1, the rewrite commits, the filter check reads version 2's properties and passes, and version 1's score comes back for a record that only qualifies as version 2. The write lock (#1607) does not help, because reads deliberately take none.
Every write now takes a fresh row id: the previous row is deleted and a new one inserted in the same transaction, a delete for the old key is staged beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version, and the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version. A key whose version has been rewritten resolves to no row and is dropped.
AUTOINCREMENT(#1589) remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order.A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often.
Cost
Medians of three runs on one machine, records per second, 64-dimensional vectors, file-backed store with an index directory:
Inserts move within run-to-run noise, in both directions. Rewrites cost 13 to 22 percent more; the upper end is the configuration where the doubled log rows double the number of index saves. MemMachine's callers write vector records once and rarely rewrite them, so the insert path is the one that matters, and the read path is untouched.
Tests
test_a_query_cannot_pair_a_score_with_a_later_versiondrives the case above deterministically: a wrapper around the query's key filter parks the engine's worker thread between scoring and the filter check, a rewrite commits in the gap, and the check must not admit the old score.test_a_rewrite_moves_the_record_to_a_new_row_idpins the mechanism. The mirror case, an old verdict admitting a new score, cannot be reproduced with the USearch engine because its read lock spans the whole search including overfetch rounds, so the mechanism test is what covers it.Verification
test_sqlite_vector_store.py: 93 passed. Against Take SQLite's write lock at BEGIN, not at the first write (speedkick) #1609's store, the two new tests and the re-derived save-threshold test fail and the rest pass.pytest packages/server/server_tests1923 passed, 3 skipped;ty check --project packages/serverall checks passed;ruff check/ruff format --checkclean.Stacked on #1609, so the diff here includes #1612, #1607, #1608 and #1609 until they merge.
🤖 Generated with Claude Code
https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn