You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Episode add is not atomic across the store, episodic memory, and semantic history; partial failures are unrecoverable without a full scan #1738
MemMachine.add_episodes (packages/server/src/memmachine_server/main/memmachine.py) performs three independent writes per request: the episode store insert, the episodic-memory write (short-term buffer plus the long-term index), and the semantic history registration. On main the store insert runs first and the other two run in parallel after it. On #1707 all three run in parallel. Neither arrangement is atomic, and neither leaves a record that recovery could read without scanning.
Failure shapes
Any subset of the three writes can fail while the others commit. The request returns an error whenever one does, and the minted episode ids are returned only on success, so the caller never learns which ids were written.
Store fails, memories succeed (only possible with concurrent writes). The episodic index holds an entry with no store row. Event-backend search hydrates from the store and drops it; declarative search returns it from the index. The semantic history row is removed by ingestion's grace-period delist.
A memory write fails, store succeeds (possible under both orderings). The row exists but is unsearchable, or never reaches semantic ingestion. The client retries, the retry mints new ids, and the content is now stored twice with the first copy half-registered.
Why it is unrecoverable today
A failed request's ids reach no caller and no durable record. The failure log line now includes them (feat: persist UUID episode identifiers #1707), but a process killed mid-request logs nothing, and logs are not queryable state.
Nothing enumerates index entries without a store row, or store rows without an index entry or history row. Finding either set is an anti-join across the episode store and the vector or segment store, a scan of both.
The only self-healing in place is the semantic grace period, which is driven by the per-set ingestion pass and therefore bounded, but covers only the semantic history row.
What scan-free recovery needs
A durable intent record written before the fan-out, so recovery reads a small indexed set instead of scanning:
Pending marker on the store row. Insert the row with a pending state in the same request, run the memory writes, then mark it committed. A partial index on the pending state lets a sweeper read only rows older than a threshold, re-drive the memory writes idempotently by uid or delete by uid, and mark them. Cost: one update per add. This keeps read-your-writes on the episodic index for the success path.
Outbox consumed by the memories. The store row is the intent; the memory writes become asynchronous consumers with idempotent upserts by uid. Recovery is replay. Cost: a search issued right after an add may not see the episode yet.
Either design makes the write ordering a pure latency question. Under both, "episode missing from the store" means "deleted", so ingestion's grace period and its storage support can be removed.
Related
#1707 (the concurrent-write discussion on the review thread), #1716.
Where
MemMachine.add_episodes(packages/server/src/memmachine_server/main/memmachine.py) performs three independent writes per request: the episode store insert, the episodic-memory write (short-term buffer plus the long-term index), and the semantic history registration. Onmainthe store insert runs first and the other two run in parallel after it. On #1707 all three run in parallel. Neither arrangement is atomic, and neither leaves a record that recovery could read without scanning.Failure shapes
Any subset of the three writes can fail while the others commit. The request returns an error whenever one does, and the minted episode ids are returned only on success, so the caller never learns which ids were written.
Why it is unrecoverable today
What scan-free recovery needs
A durable intent record written before the fan-out, so recovery reads a small indexed set instead of scanning:
Either design makes the write ordering a pure latency question. Under both, "episode missing from the store" means "deleted", so ingestion's grace period and its storage support can be removed.
Related
#1707 (the concurrent-write discussion on the review thread), #1716.
🤖 Written by Claude Fable 5.1 via Claude Code; posted from @edwinyyyu's account.