Skip to content

add_episodes writes the episode store before, not alongside, the episodic and semantic writes #1634

Description

@wanghy73

What happened

MemMachine.add_episodes (packages/server/src/memmachine_server/main/memmachine.py:735, as of 98f3cde speedkick) runs in two serial stages:

  1. await episode_storage.add_episodes(...) (lines 754-759): one INSERT into the episode store, awaited alone.
  2. Only after that returns: the episodic write (episodic_session.add_memory_episodes) and the semantic write (semantic_session_manager.add_message) run concurrently via asyncio.gather (lines 761-788).

Each add request therefore pays the episode-store round trip before any memory write starts. Episodic and semantic are already concurrent with each other; the episode store is not concurrent with them.

Why it cannot just be added to the gather

Both later stages consume the stored Episode, specifically its uid, and that uid is generated by the database. SqlAlchemyEpisodeStore.add_episodes inserts with insert(Episode).returning(Episode) and builds the returned models from the persisted rows (common/episode_store/episode_sqlalchemy_store.py:228-236). Neither the episodic write nor SemanticSessionManager._add_single_episode (which records episode.uid via add_message_to_sets, semantic_memory/semantic_session_manager.py:129) can start until the id exists.

Options

  • Assign episode ids in the application (UUIDv7 or similar) when building the Episode from the EpisodeEntry. The three writes can then run in one gather/TaskGroup. This changes Episode.id from an autoincrement integer (get_episode, get_episodes and delete_episodes all parse ids with int()) and needs a migration plus a decision about existing integer ids.
  • Keep the order but decide what happens on partial failure. Today a failure in the second stage leaves the episode row committed: reproduced with types: ["semantic"] while semantic memory is disabled, where the request returns 500 but the episode is stored. If the writes become concurrent, the same question applies in more combinations, so the failure contract should be settled together with the change.

No latency measured yet. The cost is one episode-store INSERT round trip per add request, on the critical path.

Related

Investigated and written by Claude (Claude Code), filed from the account of the user who commissioned the investigation.

Activity

  1. edwinyyyu commented on Oct 3, 2026

    @edwinyyyu
    Contributor

    The ordering question here is decided by the fix for #1738 (episode add is not atomic across the three writes): with a durable record of a partial write (an outbox or a per-session watermark advanced by a replay job), the store write no longer has to complete before the other two start, and the serial stage this issue measures goes with it. Both are tracked under #1756; this stays open as the performance side of that change.


    🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performanceIssues relating to MemMachine performance

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions