Describe the bug
With the default configuration, each search and add re-opens its session (PostgreSQL), its segment-store partition (PostgreSQL) and its Qdrant collection, then closes them when the request ends, and the next request on the same session does it all again.
instance_cache_size defaults to 0 (packages/server/src/memmachine_server/episodic_memory/episodic_memory_manager.py:43), so release_ref closes an instance as soon as its last request ends (episodic_memory/instance_lru_cache.py), and every request misses the cache. Each miss runs get_session_info, open_or_create_partition and open_collection. open_collection (common/vector_store/qdrant_vector_store.py:855) records no metric, so its cost does not appear in any histogram; it shows only as time inside http_request_duration_seconds that no child step accounts for.
Steps to reproduce
- Default configuration (
instance_cache_size unset), one server worker.
- Run a sustained search load on a few sessions (20 concurrent clients).
- Compare
session_store_sqlalchemy_latency_seconds{operation="get_session_info"} and segment_store_sqlalchemy_latency_seconds{operation="open_or_create_partition"} counts with the /memories/search count: about one of each per request.
- Measured here: about 90 ms per search in those two calls, plus 33–63 ms per search that no recorded step accounts for. With four workers, about 5 ms plus 3 ms.
Expected behavior
A session in active use keeps its episodic memory between requests, or the documentation explains why caching is off by default. The Qdrant collection open is timed like the other store operations.
Environment
- OS: Linux (Ubuntu 24.04), Docker
- MemMachine Version:
main at c99bc0e (0.3.9+50.gc99bc0e), image built from source; also seen at c08cf26
- Development language version: Python 3.12 (in the image)
- Backend: event memory, PostgreSQL 18, Qdrant 1.19.1
Additional context
Suggested fix: a non-zero default for instance_cache_size (or documentation of why 0 is the default and when to raise it), and an operation tracker around open_collection.
Describe the bug
With the default configuration, each search and add re-opens its session (PostgreSQL), its segment-store partition (PostgreSQL) and its Qdrant collection, then closes them when the request ends, and the next request on the same session does it all again.
instance_cache_sizedefaults to0(packages/server/src/memmachine_server/episodic_memory/episodic_memory_manager.py:43), sorelease_refcloses an instance as soon as its last request ends (episodic_memory/instance_lru_cache.py), and every request misses the cache. Each miss runsget_session_info,open_or_create_partitionandopen_collection.open_collection(common/vector_store/qdrant_vector_store.py:855) records no metric, so its cost does not appear in any histogram; it shows only as time insidehttp_request_duration_secondsthat no child step accounts for.Steps to reproduce
instance_cache_sizeunset), one server worker.session_store_sqlalchemy_latency_seconds{operation="get_session_info"}andsegment_store_sqlalchemy_latency_seconds{operation="open_or_create_partition"}counts with the/memories/searchcount: about one of each per request.Expected behavior
A session in active use keeps its episodic memory between requests, or the documentation explains why caching is off by default. The Qdrant collection open is timed like the other store operations.
Environment
mainatc99bc0e(0.3.9+50.gc99bc0e), image built from source; also seen atc08cf26Additional context
Suggested fix: a non-zero default for
instance_cache_size(or documentation of why 0 is the default and when to raise it), and an operation tracker aroundopen_collection.