Skip to content

Semantic config cache is per process and on by default: stale categories across replicas, and a stale replica overwrites a set's embedder and language model #1747

Description

@edwinyyyu

What happened

CachingSemanticConfigStorage (semantic_memory/config_store/caching_semantic_config_storage.py) caches set configs, registered set ids, categories, tags, set-type categories and set-type lists per process (:40-48) with no TTL. Its docstring (:21-24) says entries are invalidated on write operations, which means writes through the same instance. semantic_memory.with_config_cache defaults to true (common/configuration/__init__.py:154-157). #1730 covers the two instances inside one process; this issue is the cross-replica case, which no in-process invalidation fixes.

Three consequences:

  • A category, tag or set type created, disabled or deleted on replica A is invisible to replica B's searches and to B's background ingestion, which keeps extracting with B's cached categories and prompts (semantic_memory/semantic_memory.py:505 reads the cached config for ingestion).
  • Lost update. _set_id_resource (semantic_memory/semantic_memory.py:505-515) calls set_setid_config(set_id=..., embedder_name=<default>) whenever its cached config has no embedder. The upsert (semantic_memory/config_store/config_store_sqlalchemy.py:276-285) writes both embedder_name and language_model_name, and this call passes no language model. After configure_set(embedder_name=X, llm_name=Y) on replica A, the first ingestion or search on replica B that still holds the pre-configure config resets the set to the default embedder and a null language model, and B embeds the set with the wrong model. The same statement erases llm_name on a single process whenever embedder_name is unset and llm_name is set.
  • _registered_set_id_set_types (:41) is cleared only by the local delete_set_type_id, so after a set type is deleted on A, B skips re-registration for the set ids it remembers.

Expected

The cache is off by default for multi-replica deployments, or it is coherent: a version column read per request, or a TTL stated in the configuration docs. Independently, set_setid_config must not overwrite fields it was not given.

Notes

Code read at d6068cdbf (main), paths under packages/server/src/memmachine_server/. Not reproduced with two live replicas; the lost update follows from the upsert writing both columns unconditionally. Lower priority than the session lifecycle (#1655).


🤖 Written by Claude Code (Claude Fable 5.1) on behalf of @edwinyyyu.

Activity

  1. added
    horizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)
    on Oct 2, 2026
  2. added theissue type on Oct 2, 2026
  3. added
    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processes
    on Oct 2, 2026
  4. mikemikimike commented on Oct 9, 2026

    @mikemikimike
    Contributor

    Implemented in #1808.

    The semantic configuration cache is disabled by default so separate server processes read the shared configuration. Partial set configuration updates preserve omitted fields, while an explicit null still clears the supplied field. Regression coverage includes the HTTP endpoint, server entry point, semantic service, cache wrapper, storage updates, and concurrent updates to separate fields.

    All 42 GitHub checks passed for the current PR commit, including server unit and integration tests, client tests, lint, type checks, and installation tests.

  5. mikemikimike commented on Oct 9, 2026

    @mikemikimike
    Contributor

    Update: #1808 has been closed because #1759 already includes the semantic-cache default changes. #1759 does not include the omitted-field update fix; this issue remains open.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    concurrencyRaces, lost updates, unarbitrated read-then-write under concurrent requests or processeshorizontal scalingWrong or unsafe when more than one server process serves the same backends (replicas or workers)

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions