Skip to content

[Feat]: Server tech debt resolution wishlist #1297

Description

@edwinyyyu

Is your feature request related to a problem?

  • horizontal scalability
    Claude:
❯ Rate horizontal scalability of each component.

⏺ Based on what I've read in the codebase:

  ┌──────────────────────────────────┬────────┬───────────────────────────────────────────────────────────────────────────────────┐
  │            Component             │ Rating │                                    Bottleneck                                     │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ MemMachine                       │ B      │ Mostly stateless orchestrator, but holds a singleton _retrieval_agent with a      │
  │                                  │        │ process-local lock. Multiple instances work, minor duplication.                   │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ EpisodeStorage                   │ B+     │ Postgres via SQLAlchemy. Reads scale with replicas. Writes are single-primary.    │
  │                                  │        │ Session-scoped data is naturally shardable.                                       │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ SessionDataManager               │ B+     │ Same as EpisodeStorage — Postgres-backed, shardable by session_key.               │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │                                  │        │ Process-local LRU cache with ref counting. No cross-instance cache coherence.     │
  │ EpisodicMemoryManager            │ D      │ Each server instance builds its own cache, wastes memory, risks stale state. This │
  │                                  │        │  is the main scaling bottleneck.                                                  │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ ShortTermMemory                  │ C      │ Summary state persisted in DB (good), but the in-memory message buffer is         │
  │                                  │        │ process-local. Requires the same instance to serve a session's lifetime.          │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ LongTermMemory/DeclarativeMemory │ A-     │ Stateless facade. All state in external VectorGraphStore. Scales as well as the   │
  │                                  │        │ backing store does.                                                               │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ EventMemory                      │ A-     │ Stateless. All state in external SegmentStore + vector Collection. Same story.    │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ SemanticService                  │ A-     │ Stateless compute (LLM calls). Backed by external SemanticStorage                 │
  │                                  │        │ (Neo4j/PgVector). Scales well.                                                    │
  ├──────────────────────────────────┼────────┼───────────────────────────────────────────────────────────────────────────────────┤
  │ RetrievalAgent                   │ A      │ Stateless LLM-based query refinement. Fully parallelizable.                       │
  └──────────────────────────────────┴────────┴───────────────────────────────────────────────────────────────────────────────────┘
  • configuration
    • really hard to navigate for human, effectively unreviewable
    • includes weird service locator pattern to work around mix of declarative and dynamic configuration, when they should be separate concepts managed by different resource managers (e.g. a resource manager manages only static resources, including dynamic resource managers -- see something like MetricsFactory
    • databases are treated as resources when it does not make sense from a wiring perspective, so the configuration file format is not declarative and bijective with the Python wiring
    • the non-bijective configuration makes it impossible to understand what is a valid config without reading a lot of fragile documentation or just going through the code, which defeats the purpose
❯ Rate the configuration system.

⏺ Honestly? It's overcomplicated.

  What it does well:
  - Session-scoped config stored as JSON in Postgres — durable, queryable
  - Resource IDs (string → object) via resource manager — clean indirection

  What's unnecessarily complex:

  1. Partial/Full config duplication — every config has two classes (LongTermMemoryConf + LongTermMemoryConfPartial,
  ShortTermMemoryConf + ShortTermMemoryConfPartial, EpisodicMemoryConf + EpisodicMemoryConfPartial). The partial exists only for the
  merge system.
  2. The merge system — EpisodicMemoryConfPartial.merge() merges user-provided config with server defaults. It has nested merge
  logic, null coalescing, special cases for session_key/session_id propagation. This is a lot of machinery to do what Pydantic's
  model_copy(update=...) or a simple dict merge would handle.
  3. Config flows through too many layers — Configuration.episodic_memory (partial) → MemMachine._with_default_episodic_memory_conf()
   (merges) → EpisodicMemoryConf (full) → stored in param_data → loaded back → service_locator → Params → constructor. That's 6+
  hops.
  4. _enabled flags alongside nullable fields — long_term_memory: LongTermMemoryConf | None AND long_term_memory_enabled: bool. The
  nullability already signals "not configured"; the bool adds a separate "configured but disabled" state that complicates every
  check.
  • homemade message queue/stream system for semantic ingestion
  • complicated config makes wiring difficult (correct wiring is a multi-day effort even with AI)
  • short-term memory and long-term memory lifecycle are tightly coupled when this reduces scalability and does not need to be the case
    • long-term memory is entirely stateless
    • there can naturally be multiple agents (short-term memories) acting in one channel (long-term memory) e.g. Slack, Discord
  • it is much easier to develop with the memory components directly
    • takes like 20 minutes to tell Claude Code to make a simple chatbot that uses one of the memory classes directly vs. complicated setup required for full server
  • naming nitpicks: technically, episodic memory is a subset of long term memory--not the other way around. short-term memory is not episodic memory.
  • need to understand client timezone
  • mixed ID types
  • auto-increment int IDs are not trivially horizontally scalable
  • no guarantees on crash
  • MCP server is not LLM-friendly and is insecure
  • session data manager's responsibilities are not well-defined
  • source/producer/speaker/etc need to be distinguished because a name is very different from a UID or similar
    • the source used in LLM-friendly formatting should be a name
    • the producer_id is a different concept
  • created_at is semantically confusing
  • retrieval agent is tightly coupled with episodic memory
  • semantic memory hierarchy levels are unintuitive for humans and may be bad for prompts
    • according to @o-love these names were inherited from an older version of the system

Describe the solution you'd like

  • separate memmachine-server into memmachine-server and memmachine-core
    • memmachine-server includes configuration and orchestration, acts as a reference implementation but is not mandatory
    • memmachine-core includes ABCs and implementations, memories
  • separate memmachine-common into memmachine-common and memmachine-api
    • data models for core separated from client/server-only models
  • configs do not use merging (single source of truth)
  • short-term memory is managed by the agent (typical agents have this already, or it's easy to implement) -- not by the server
    • if we need short-term memory, it should have its own lifecycle
  • use a message queue/stream for ingestion
  • disregard backward compatibility
  • API supports multimodality cleanly (see New episodic memory (currently named EventMemory) based on VectorStore and (new) SegmentStore #1199 for example of data models that may work)
  • API should optionally get client timezone
  • also rewrite MCP server (problems identified in [Bug]: Context creeping, even after simple search, #1278)
  • all IDs are UUID (64-bit int is an alternative, but that requires coordination to avoid collisions)
  • operations are transactional or self-healing on restart
  • session data manager is split into configuration manager and short term memory summary store
  • client gets functions for formatting memories
  • memory speaker/source/context need to be separated from metadata
    • clarify semantics of producer/source/etc.
  • episodes have a timestamp or a time range indicating when occurred -- created_at/ingestion time is a different system-defined concept that has no relevance to memory
  • retrieval agent should be able to send queries and/or receive strings or something more generic as context
  • better names for semantic memory hierarchy levels that actually reflect their semantics
    • partition_id → topic → category → attribute → value
  • too many undefined behaviors/no trust in server
  • metadata model is all over the place

Describe alternatives you've considered

No response

Additional context

No response

Activity

  1. added theissue type on Apr 3, 2026
  2. changed the title [-][Feat]: Rewrite server[/-] [+][Feat]: Rewrite server, separation of concerns[/+] on Apr 3, 2026
  3. added
    refactorCode refactoring that doesn't add new features.
    on Apr 4, 2026
  4. edwinyyyu commented on Apr 7, 2026

    @edwinyyyu
    ContributorAuthor
  5. sscargal commented on Apr 7, 2026

    @sscargal
    Contributor

    I agree with the direction. It's going to require a lot of design and ideation before we begin implementation. Likely a candidate for v0.5.0 - v1.0.0? (Future).

  6. changed the title [-][Feat]: Rewrite server, separation of concerns[/-] [+][Feat]: Issues with server[/+] on Apr 7, 2026
  7. changed the title [-][Feat]: Issues with server[/-] [+][Feat]: Server tech debt resolution[/+] on Apr 7, 2026
  8. edwinyyyu commented on Apr 8, 2026

    @edwinyyyu
    ContributorAuthor

    I agree with the direction. It's going to require a lot of design and ideation before we begin implementation. Likely a candidate for v1.0.0?)

    Ideally the current implementation can tide us over so we avoid making too many breaking changes, while we spend more time on a more careful design.

    But it's also having some negative impacts now, e.g. #1304 is ~2000 lines larger than #1199 just for wiring into the configuration system/client-server model. And it's a time sink to review it.

    Also, I'm not sure if the single server model is the right call. It requires stricter coordination and a god API/config which seems to not be working too well right now, especially with default configurations that enable everything.

  9. changed the title [-][Feat]: Server tech debt resolution[/-] [+][Feat]: Server tech debt resolution wishlist[/+] on Apr 8, 2026
  10. maratsultanov2 commented on Apr 8, 2026

    @maratsultanov2

    Привет 🙂

    Спасибо за поднятие важного вопроса о горизонтальной масштабируемости! Учитывая вашу оценку компонентов, возможно, стоит рассмотреть альтернативные паттерны, такие как использование распределенных кешей или асинхронных очередей для разгрузки узких мест. Это может помочь уменьшить время задержки и повысить производительность системы. Возможно, стоит также изучить методы, которые применяются в других проектах. Например, в проекте TreeAngleTap можно найти интересные решения по оптимизации. Удачи с вашими улучшениями!

  11. maratsultanov2 commented on Apr 9, 2026

    @maratsultanov2

    This is a great initiative to address server tech debt and improve horizontal scalability! Your detailed assessment of the components and their ratings is a valuable starting point. To further facilitate this wishlist, it might be helpful to prioritize the components based on their impact on scalability and performance. This way, the team can focus on the most critical areas first. Maybe consider discussing potential solutions or enhancements for each bottleneck identified as well. Looking forward to seeing the improvements unfold! For more context on this issue, feel free to check out TreeAngleTap.

  12. added
    keep-openPrevents the auto-close task from closing this issue.
    on Apr 24, 2026
  13. edwinyyyu commented on Sep 3, 2026

    @edwinyyyu
    ContributorAuthor

    Stale. There's maybe an opportunity to address many of these issues in https://github.com/MemMachine/MemMachine/tree/speedkick.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    keep-openPrevents the auto-close task from closing this issue.performanceIssues relating to MemMachine performancerefactorCode refactoring that doesn't add new features.

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions