You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Stale segment-store handles need cache eviction and an API status mapping #1571
Status 2026-10-02. The segment store fencing this issue builds on is on main (#1661, port of #1548); the vector store gains the same fencing in #1734 -> #1736, so its stale-handle error needs the same mapping and eviction. Tracked under #1755 (step 5): map the stale errors at the API and evict the cached instance, or build instances per request once short-term memory is process-scoped (#1757).
#1545 gives segment-store partition handles fencing: a handle bound to a deleted partition (or a deleted-and-recreated one) raises SegmentStorePartitionHandleStaleError on every data operation instead of acting on the successor tenant. Two consequences are left for the layers above the store, which have no handling for that error today:
No domain-error-to-status mapping. A session drop concurrent with an in-flight add_episodes/search makes add_segments (or a read) raise the stale-handle error, which propagates through EventMemory and LongTermMemory to the API layer as an unhandled 500. SegmentStoreAttemptsExhaustedError (raised when partition creation makes no progress under churn) reaches the API the same way. There is no central layer mapping domain errors to statuses to extend; mapping stale-handle to 404 vs 409 is a product decision.
Cross-replica cached handles never recover. Replica A deletes session s1; replica B's MemoryInstanceCache still holds an EpisodicMemory whose LongTermMemory wraps a handle bound to the dead incarnation. Every later call on replica B raises the stale-handle error until LRU eviction, because episodic_memory_manager.delete_episodic_memory's cache invalidation is process-local. Without a reopen-on-stale path, even a correct status code would be permanent for that cached instance. This is not a regression: on the previous partitioned layout the cached handle referenced dropped tables and failed equally permanently, just with a driver error instead of a typed one -- Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) #1545 makes the condition recognizable, which is what a recovery path needs.
Suggested shape: treat SegmentStorePartitionHandleStaleError as the signal to evict the cached EpisodicMemory for that session (and re-resolve on the next request), and map it -- along with SegmentStoreAttemptsExhaustedError -- to an explicit status at the API boundary.
Investigated and written by Claude (Claude Code), filed from the account of the user who commissioned the investigation.
Status 2026-10-02. The segment store fencing this issue builds on is on main (#1661, port of #1548); the vector store gains the same fencing in #1734 -> #1736, so its stale-handle error needs the same mapping and eviction. Tracked under #1755 (step 5): map the stale errors at the API and evict the cached instance, or build instances per request once short-term memory is process-scoped (#1757).
#1545 gives segment-store partition handles fencing: a handle bound to a deleted partition (or a deleted-and-recreated one) raises
SegmentStorePartitionHandleStaleErroron every data operation instead of acting on the successor tenant. Two consequences are left for the layers above the store, which have no handling for that error today:No domain-error-to-status mapping. A session drop concurrent with an in-flight
add_episodes/searchmakesadd_segments(or a read) raise the stale-handle error, which propagates throughEventMemoryandLongTermMemoryto the API layer as an unhandled 500.SegmentStoreAttemptsExhaustedError(raised when partition creation makes no progress under churn) reaches the API the same way. There is no central layer mapping domain errors to statuses to extend; mapping stale-handle to 404 vs 409 is a product decision.Cross-replica cached handles never recover. Replica A deletes session
s1; replica B'sMemoryInstanceCachestill holds anEpisodicMemorywhoseLongTermMemorywraps a handle bound to the dead incarnation. Every later call on replica B raises the stale-handle error until LRU eviction, becauseepisodic_memory_manager.delete_episodic_memory's cache invalidation is process-local. Without a reopen-on-stale path, even a correct status code would be permanent for that cached instance. This is not a regression: on the previous partitioned layout the cached handle referenced dropped tables and failed equally permanently, just with a driver error instead of a typed one -- Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) #1545 makes the condition recognizable, which is what a recovery path needs.Suggested shape: treat
SegmentStorePartitionHandleStaleErroras the signal to evict the cachedEpisodicMemoryfor that session (and re-resolve on the next request), and map it -- along withSegmentStoreAttemptsExhaustedError-- to an explicit status at the API boundary.Investigated and written by Claude (Claude Code), filed from the account of the user who commissioned the investigation.