Skip to content

Refuse a pending row replay cannot honor, instead of dropping it (speedkick) - #1608

Merged
edwinyyyu merged 1 commit into
MemMachine:speedkickfrom
edwinyyyu:fix/sqlite-vector-store-replay-refusal-speedkick
Sep 11, 2026
Merged

edwinyyyu merged 1 commit into
MemMachine:speedkickfrom
edwinyyyu:fix/sqlite-vector-store-replay-refusal-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Replay matched operation_type == "upsert" and vector is not None and let everything else fall through. An upsert row with no vector, a vector that is not a whole number of float32s (numpy's ValueError about buffer sizes escaped as is), or a row with an unknown operation_type was skipped, and because the skipped row reached neither the engine remove set nor the mark-applied update, it survived the restart to be skipped again on the next one. A decodable vector of the wrong width was the mild case: it reached the engine, which refused it with its own dimension error at startup.

Between a write returning and the next index save, the log holds the only copy of the vector. So the outcome was a record that exists in SQLite, cannot be found by any search, and reports nothing, permanently.

That is damage to a durable record, not a state to heal. Replay now raises PendingOperationCorruptError naming the collection, the row and the fault, leaving the log intact for whoever repairs it. Both error types' docstrings now say what a caller should do, since the two remedies are opposites: read the cause of an IndexLoadError before choosing between rebuilding the index and fixing the disk; never clear the log to get past a PendingOperationCorruptError.

Tests

Four corrupt a log row each way and assert the restart refuses. The rest of TestPendingLogStates pins what replay guarantees for intact rows: a rewritten uuid replays its last write, an upsert then delete stays deleted, a failed save leaves the write replayable, and the save threshold counts log rows rather than writes.

Verification

Stacked on #1607, so the diff here includes #1612 and #1607 until they merge.


🤖 Generated with Claude Code

https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
@edwinyyyu
edwinyyyu force-pushed the fix/sqlite-vector-store-replay-refusal-speedkick branch from af6e8cc to 8ceae3d Compare September 11, 2026 00:14
@edwinyyyu
edwinyyyu merged commit 72e4dde into MemMachine:speedkick Sep 11, 2026
39 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 17, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 18, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…edkick) (MemMachine#1608)

Refuse a pending row replay cannot honor, instead of dropping it

Replay matched `operation_type == "upsert" and vector is not None` and
let everything else fall through. An upsert row with no vector, a vector
that is not a whole number of float32s, or an unknown operation_type was
skipped, and a decodable vector of the wrong width reached the engine,
which refused it with its own error at startup. The skipped rows were
the worse case: and because the skipped row reached neither the engine remove
set nor the mark-applied update it survived the restart to be skipped
again on the next one. Between a write returning and the next index save
the log holds the only copy of the vector, so the outcome was a record
that exists in SQLite and can never be found by search.

That is damage to a durable record, not a state to heal. Replay now
raises PendingOperationCorruptError for all four, naming the collection,
the row and the fault, and leaves the log intact for whoever repairs it. Both error
types' docstrings now say what a caller should do: read the cause of an
IndexLoadError before choosing a remedy, and never clear the log to get
past a PendingOperationCorruptError.

Four tests corrupt a log row each way and assert the restart refuses.
The rest of TestPendingLogStates pins what replay guarantees for intact
rows: a rewritten uuid replays its last write, an upsert then delete
stays deleted, a failed save leaves the write replayable, and the save
threshold counts log rows rather than writes.

Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn

Co-authored-by: Claude Fable 5.1 <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant