Skip to content

Stale segment-store handles need cache eviction and an API status mapping #1571

Description

@edwinyyyu

Status 2026-10-02. The segment store fencing this issue builds on is on main (#1661, port of #1548); the vector store gains the same fencing in #1734 -> #1736, so its stale-handle error needs the same mapping and eviction. Tracked under #1755 (step 5): map the stale errors at the API and evict the cached instance, or build instances per request once short-term memory is process-scoped (#1757).

#1545 gives segment-store partition handles fencing: a handle bound to a deleted partition (or a deleted-and-recreated one) raises SegmentStorePartitionHandleStaleError on every data operation instead of acting on the successor tenant. Two consequences are left for the layers above the store, which have no handling for that error today:

  1. No domain-error-to-status mapping. A session drop concurrent with an in-flight add_episodes/search makes add_segments (or a read) raise the stale-handle error, which propagates through EventMemory and LongTermMemory to the API layer as an unhandled 500. SegmentStoreAttemptsExhaustedError (raised when partition creation makes no progress under churn) reaches the API the same way. There is no central layer mapping domain errors to statuses to extend; mapping stale-handle to 404 vs 409 is a product decision.

  2. Cross-replica cached handles never recover. Replica A deletes session s1; replica B's MemoryInstanceCache still holds an EpisodicMemory whose LongTermMemory wraps a handle bound to the dead incarnation. Every later call on replica B raises the stale-handle error until LRU eviction, because episodic_memory_manager.delete_episodic_memory's cache invalidation is process-local. Without a reopen-on-stale path, even a correct status code would be permanent for that cached instance. This is not a regression: on the previous partitioned layout the cached handle referenced dropped tables and failed equally permanently, just with a driver error instead of a typed one -- Overhaul segment store: shared tables with incarnation-scoped tenant keys (fixes #1544, #1546, #1549) #1545 makes the condition recognizable, which is what a recovery path needs.

Suggested shape: treat SegmentStorePartitionHandleStaleError as the signal to evict the cached EpisodicMemory for that session (and re-resolve on the next request), and map it -- along with SegmentStoreAttemptsExhaustedError -- to an explicit status at the API boundary.

Investigated and written by Claude (Claude Code), filed from the account of the user who commissioned the investigation.

Metadata

Metadata

Assignees

No one assigned

    Labels

    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)recoveryA failure or crash leaves durable state nothing repairs: partial writes, lost jobs, no retry

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions