Skip to content

Reject process-local graph backends in multi-process deployments - #1810

Open
mikemikimike wants to merge 1 commit into
MemMachine:mainfrom
mikemikimike:fix/backend-concurrency-scope-1753
Open

mikemikimike wants to merge 1 commit into
MemMachine:mainfrom
mikemikimike:fix/backend-concurrency-scope-1753

Conversation

@mikemikimike

Copy link
Copy Markdown
Contributor

Purpose of the change

Prevent deployments with multiple server processes from silently using graph
backends whose schema, index or semantic-set caches are local to one process.

Issue: #1753

Description

  • Declare PROCESS concurrency support on Neo4j semantic storage and the Neo4j
    and NebulaGraph declarative graph stores.
  • Validate enabled components before starting background services or restoring
    sessions. Enforce the same limit in resource factories and resolved session
    overrides, including AUTO semantic storage's Neo4j fallback.
  • Add deployment scope configuration and an environment override. Multiple
    HTTP processes selected by MEMMACHINE_WORKERS require at least HOST scope;
    deployments on multiple hosts declare CLUSTER.
  • Preserve single-process use, disabled components and other semantic backends.
    Document the configuration and the affected backends' limits.

Fixes/Closes

Fixes #1753

Type of change

  • Bug fix

How Has This Been Tested?

  • Ran the complete server/client non-integration suite in sequential groups:
    2406 passed, 4 skipped. Three existing invalid-credential tests time out
    contacting the OpenAI API; all three fail identically on unchanged main.
  • Ran Ruff check and format check successfully.
  • Re-ran the focused configuration, scope, startup and factory tests on Linux:
    all 92 passed. The POSIX-specific time-zone regression and the 11 local
    resource-manager tests also passed; the same three external API tests
    failed with connection errors.
  • Ran the workflow's ty checks for server Python 3.12/3.14 and client/common
    Python 3.10/3.14 successfully.
  • Reproduced startup through the real configuration model and MemMachine
    start/stop with external resources replaced by test doubles: unchanged main
    accepts each affected backend in CLUSTER scope; this change refuses all
    three before creating background tasks and preserves PROCESS startup.

Project test commands:

pytest packages/server/server_tests packages/client/client_tests
ruff check
ruff format --check
ty check --project packages/server --python-version 3.12
ty check --project packages/server --python-version 3.14
ty check --project packages/client --python-version 3.10
ty check --project packages/client --python-version 3.14
ty check --project packages/common --python-version 3.10
ty check --project packages/common --python-version 3.14

Checklist

  • Added regression tests for backend declarations, startup, overrides,
    disabled components, repeated calls and failure recovery.
  • Updated the configuration documentation.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Declarative memory and the Neo4j semantic backend keep per-process caches that are not refreshed when another replica writes

1 participant