Describe the bug
short_term_memory.message_capacity is a length in characters. The product default is 64,000, and docs/open_source/configuration.mdx documents it that way. But the shipped sample configurations, install guides and Helm chart set it to 500, with no unit.
At 500, a single question and answer overflows the short-term window, so every store starts an LLM summary, and every search waits for a summary in progress before it reads. On an otherwise idle host, searches took 7–32 s and stores about 7 s at the median.
Where 500 ships on main:
sample_configs/configuration.event.yml:43
sample_configs/episodic_memory_config.cpu.sample, .gpu.sample, .nebula.sample
deployments/helm/templates/memmachine-configmaps.yaml:45 and deployments/helm/README.md:135
docs/install_guide/install_guide.mdx:115 and :232
docs/install_guide/cloud_deploy/aws_cloudformation.mdx:485
evaluation/retrieval_agent/README.md (5 places)
Two descriptions also call the value a message count, which makes 500 look reasonable:
packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py:59: "The maximum number of messages to summarize."
packages/client/src/memmachine_client/config.py:393: "Maximum number of messages to keep in short-term memory"
Steps to reproduce
- Deploy with any of the files above (short-term memory enabled).
- Store a few chat-sized turns (a question and an answer, about 500 characters together) in one session, searching between them.
- Watch the server's
language_model_* metrics: every store triggers a summary call, and searches in that session take seconds.
- Set
message_capacity: 64000 and repeat: no summary calls while the conversation fits, and searches return in well under a second.
Expected behavior
Shipped configurations use a value that holds a conversation (the documented 64,000, or omit the key), and every description of the setting says it is measured in characters.
Environment
- OS: Linux (Ubuntu 24.04), Docker
- MemMachine Version:
main at c99bc0e (0.3.9+50.gc99bc0e), image built from source; also seen at c08cf26
- Development language version: Python 3.12 (in the image)
- Backend: event memory, PostgreSQL 18, Qdrant 1.19.1
Additional context
Suggested fix: change the value (or remove the key) in every file listed, correct the two descriptions, and consider renaming the setting or rejecting values smaller than one typical message.
Describe the bug
short_term_memory.message_capacityis a length in characters. The product default is 64,000, anddocs/open_source/configuration.mdxdocuments it that way. But the shipped sample configurations, install guides and Helm chart set it to500, with no unit.At 500, a single question and answer overflows the short-term window, so every store starts an LLM summary, and every search waits for a summary in progress before it reads. On an otherwise idle host, searches took 7–32 s and stores about 7 s at the median.
Where
500ships onmain:sample_configs/configuration.event.yml:43sample_configs/episodic_memory_config.cpu.sample,.gpu.sample,.nebula.sampledeployments/helm/templates/memmachine-configmaps.yaml:45anddeployments/helm/README.md:135docs/install_guide/install_guide.mdx:115and:232docs/install_guide/cloud_deploy/aws_cloudformation.mdx:485evaluation/retrieval_agent/README.md(5 places)Two descriptions also call the value a message count, which makes 500 look reasonable:
packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py:59: "The maximum number of messages to summarize."packages/client/src/memmachine_client/config.py:393: "Maximum number of messages to keep in short-term memory"Steps to reproduce
language_model_*metrics: every store triggers a summary call, and searches in that session take seconds.message_capacity: 64000and repeat: no summary calls while the conversation fits, and searches return in well under a second.Expected behavior
Shipped configurations use a value that holds a conversation (the documented 64,000, or omit the key), and every description of the setting says it is measured in characters.
Environment
mainatc99bc0e(0.3.9+50.gc99bc0e), image built from source; also seen atc08cf26Additional context
Suggested fix: change the value (or remove the key) in every file listed, correct the two descriptions, and consider renaming the setting or rejecting values smaller than one typical message.