Skip to content

Semantic background ingestion runs in every replica with no claim, so N replicas ingest the same messages N times #1745

Description

@edwinyyyu

What happened

Every server process starts the semantic ingestion loop: main/memmachine.py calls semantic_service.start() from MemMachine.start whenever semantic memory is enabled, and semantic_memory/semantic_memory.py:148-156 creates the task. Each tick the loop lists the dirty sets (:795-802) and hands them to process_set_ids. _process_single_set (semantic_memory/semantic_ingestion.py:122-132) reads the five oldest un-ingested messages of the set with a plain SELECT, runs the LLM per message and category (:215-220), inserts features with plain add_feature inserts (:294-301), and only then marks the messages ingested (:270-273). Consolidation in the same module runs unclaimed as well.

Nothing claims a set or a message: no FOR UPDATE SKIP LOCKED, no lease, no leader role. With N replicas every replica selects the same sets and the same five messages, so the LLM and embedding work is done N times and the ADD commands insert N copies of each feature row. Concurrent consolidations each insert their own merged features.

What is already safe: the work list is re-read from the database each tick, mark_messages_ingested is an idempotent UPDATE, the history primary key is (set_id, history_id), and purge_ingested_rows deletes only fully ingested sets. Nothing is lost; it is multiplied.

Expected

A set is ingested by one replica at a time. Either the loop claims a set (a claimable row per set, SKIP LOCKED, released or expired on completion) or ingestion runs under a designated role so one process runs the loop. #1251 (closed) reduced the cost of the per-worker poll query and #1699 asks for the interval to be configurable; neither addresses the duplication.

Notes

Code read at d6068cdbf (main), paths under packages/server/src/memmachine_server/. Not reproduced against a live LLM; the duplication follows from the absence of any claim in the read path. Semantic memory has lower priority than the session lifecycle (#1655); this issue records the gap so the fix lands with the shared claim primitive rather than as a one-off.


🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Metadata

Metadata

Assignees

No one assigned

    Labels

    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions