Skip to content

Give every write a fresh row id, so a key names one version (speedkick) - #1610

Merged
edwinyyyu merged 1 commit into
MemMachine:speedkickfrom
edwinyyyu:fix/sqlite-vector-store-versioned-keys-speedkick
Sep 11, 2026
Merged

edwinyyyu merged 1 commit into
MemMachine:speedkickfrom
edwinyyyu:fix/sqlite-vector-store-versioned-keys-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

An upsert of an existing uuid kept its row id, so one engine key spanned every version of a record. A query reads the engine, then each candidate's properties, then its uuid, at three instants with nothing held between them. A rewrite of the very record being returned could land between those reads, pairing one version's score with another version's filter verdict. Concretely, with a record that fails the filter at version 1 and passes at version 2: the engine scores version 1, the rewrite commits, the filter check reads version 2's properties and passes, and version 1's score comes back for a record that only qualifies as version 2. The write lock (#1607) does not help, because reads deliberately take none.

Every write now takes a fresh row id: the previous row is deleted and a new one inserted in the same transaction, a delete for the old key is staged beside the upsert for the new, and the engine removes the old key and adds the new. Rows are immutable, so a key names one version, and the score computed under it, the filter verdict for it, and the uuid it resolves to belong to that version. A key whose version has been rewritten resolves to no row and is dropped. AUTOINCREMENT (#1589) remains what keeps a retired key from being reissued, and the write lock what keeps two rewrites of one record in order.

A batch that names a uuid twice is collapsed to its last record before the insert, which the on-conflict update used to do implicitly. The save threshold counts log rows and a rewrite now adds two, so a rewrite-heavy workload checkpoints about twice as often.

Cost

Medians of three runs on one machine, records per second, 64-dimensional vectors, file-backed store with an index directory:

batch save threshold insert before → after rewrite before → after
500 1000 20477 → 19290 16450 → 12902
500 none 21228 → 22486 19013 → 15899
1 none 341 → 345 316 → 274

Inserts move within run-to-run noise, in both directions. Rewrites cost 13 to 22 percent more; the upper end is the configuration where the doubled log rows double the number of index saves. MemMachine's callers write vector records once and rarely rewrite them, so the insert path is the one that matters, and the read path is untouched.

Tests

test_a_query_cannot_pair_a_score_with_a_later_version drives the case above deterministically: a wrapper around the query's key filter parks the engine's worker thread between scoring and the filter check, a rewrite commits in the gap, and the check must not admit the old score. test_a_rewrite_moves_the_record_to_a_new_row_id pins the mechanism. The mirror case, an old verdict admitting a new score, cannot be reproduced with the USearch engine because its read lock spans the whole search including overfetch rounds, so the mechanism test is what covers it.

Verification

  • test_sqlite_vector_store.py: 93 passed. Against Take SQLite's write lock at BEGIN, not at the first write (speedkick) #1609's store, the two new tests and the re-derived save-threshold test fail and the rest pass.
  • On this tree, the top of the stack: pytest packages/server/server_tests 1923 passed, 3 skipped; ty check --project packages/server all checks passed; ruff check / ruff format --check clean.

Stacked on #1609, so the diff here includes #1612, #1607, #1608 and #1609 until they merge.


🤖 Generated with Claude Code

https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

@edwinyyyu
edwinyyyu force-pushed the fix/sqlite-vector-store-versioned-keys-speedkick branch 3 times, most recently from bae0d8b to b9d2aff Compare September 10, 2026 22:25
@edwinyyyu
edwinyyyu force-pushed the fix/sqlite-vector-store-versioned-keys-speedkick branch 9 times, most recently from 04c3a05 to 135052a Compare September 11, 2026 01:01
An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
@edwinyyyu
edwinyyyu force-pushed the fix/sqlite-vector-store-versioned-keys-speedkick branch from 135052a to eada6db Compare September 11, 2026 03:17
@edwinyyyu
edwinyyyu merged commit acb4f9a into MemMachine:speedkick Sep 11, 2026
39 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 1, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 2, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 2, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 3, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 3, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 6, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 8, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 9, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 9, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 9, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 9, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 9, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…k) (MemMachine#1610)

Give every write a fresh row id, so a key names one version

An upsert of an existing uuid kept its row id, so one engine key spanned
every version of a record. A query reads the engine, then each
candidate's properties, then its uuid, at three instants with nothing
held between them, and a rewrite of the very record being returned could
land between those reads: the score of one version paired with the
filter verdict of another, or, when the overfetch loop rescored a key
whose verdict was already cached, the other way round. No lock covers
this, because reads deliberately take none.

Every write now deletes the previous row and inserts a new one in the
same transaction, stages a delete for the old key beside the upsert for
the new, and the engine removes the old key and adds the new. Rows are
immutable, so a key names one version: the score computed under it, the
filter verdict for it, and the uuid it resolves to belong to that
version, and a key whose version has been rewritten resolves to no row
and is dropped. AUTOINCREMENT remains what keeps a retired key from being
reissued, and the write lock what keeps two rewrites of one record in
order.

A batch that names a uuid twice is collapsed to its last record before
the insert, which the on-conflict update used to do implicitly. The save
threshold counts log rows and a rewrite now adds two, so a rewrite-heavy
workload checkpoints about twice as often; that test's expectation
changes accordingly.

Measured on this machine, medians of three runs, records per second,
64-dimensional vectors, file-backed store with an index directory:

    batch  save threshold  insert before / after  rewrite before / after
    500    1000            20477 / 19290          16450 / 12902
    500    none            21228 / 22486          19013 / 15899
    1      none              341 /   345            316 /   274

Inserts move within run-to-run noise, in both directions. Rewrites cost
13-22% more, the upper end where the doubled log rows double the saves.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant