Repository navigation
Serialize a collection's writes so the engine sees them in order (speedkick) - #1607
Merged
edwinyyyu merged 1 commit intoSep 11, 2026
Conversation
This was referenced Sep 10, 2026
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-write-lock-speedkick
branch
from
September 10, 2026 21:45
dc7effd to
f7c4b24
Compare
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-write-lock-speedkick
branch
2 times, most recently
from
September 10, 2026 22:25
222315e to
eef5925
Compare
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-write-lock-speedkick
branch
3 times, most recently
from
September 11, 2026 00:03
72afec2 to
cd50fe7
Compare
A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu
force-pushed
the
fix/sqlite-vector-store-write-lock-speedkick
branch
from
September 11, 2026 00:06
cd50fe7 to
22d7873
Compare
This was referenced Sep 14, 2026
Merged
Closed
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 17, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This was referenced Sep 17, 2026
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 17, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 17, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 17, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 18, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 18, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 21, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 21, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Sep 25, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This was referenced Sep 29, 2026
Open
This was referenced Oct 7, 2026
edwinyyyu
added a commit
to edwinyyyu/MemMachine
that referenced
this pull request
Oct 10, 2026
…edkick) (MemMachine#1607) Serialize a collection's writes so the engine sees them in order A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone, and two upserts inverting leave the engine serving the older vector. Never reusing a row id does not cover this, because an upsert of an existing uuid keeps its row id. A save in that window costs a write outright: it publishes the index and trims every applied log row, and a write that applied after the index was written is then in neither. A per-collection asyncio.Lock now spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed per open_collection call, so several can address one collection, and only a shared lock serializes them. Readers are untouched. Three tests fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of its uuid, a save trimming a write it did not publish, and that overtake across two handles, which a per-handle lock passes. The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice. Fixes MemMachine#1468. Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Co-authored-by: Claude Fable 5.1 <[email protected]>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose of the change
A write commits to SQLite and only then applies to the search engine, so two writers to one uuid could reach the engine in the opposite order to the one they committed in: an upsert overtaking a delete re-adds a vector for a record that is gone (it wins result slots and is dropped from them, and the next save publishes it for good), and two upserts inverting leave the engine serving the older vector. Never reusing a row id (#1589) does not cover this: an upsert of an existing uuid keeps its row id.
A save falling in that window costs a write outright. A save publishes the index and then trims every applied log row, and the log is the only other copy of those vectors, so a write that applied after the index was written is then in neither.
A per-collection
asyncio.Locknow spans a write from SQL commit through engine apply, mark-applied, and any save it triggers; shutdown's save takes it too. The lock belongs to the store, not to a collection handle: a handle is constructed peropen_collectioncall, so several can address one collection, and only a shared lock serializes them. Readers are untouched.Fixes #1468.
Tests
Three fail without the lock, each interleaving made deterministic by gating the engine: an upsert overtaking a delete of the same uuid, a save trimming a write it did not publish, and that overtake across two handles on one collection, which a per-handle lock passes. The interleaving tests poll for what another task has committed; the deadline sits between polls rather than in a timeout around them, because cancelling a query mid-flight leaves its read transaction open on the pooled connection and blocks the next commit.
The rest pin behavior the lock must preserve: an upsert surviving a delete of another record, disjoint concurrent upserts and deletes, writes racing a checkpoint, and batches that name one uuid twice.
Verification
test_sqlite_vector_store.py: 81 passed. Against Never reuse a row id in SQLiteVectorStore (speedkick) #1589's store, the 3 tests above fail and the other 78 pass.ruff check/ruff format --check: clean.Stacked on #1612, which moved the engine's lock into the store, so the diff here includes it until it merges. The stack is #1612 → this → #1608 → #1609 → #1610.
🤖 Generated with Claude Code
https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn