Describe the bug
Under load, about one search in 1,500 takes 12–43 s while the p99 is about 1.2 s. MemMachine's own Prometheus histograms put the whole wait in the embedding step (event_memory_query_phase_seconds{phase="embedding"} and embedder_openai_latency_seconds{operation="search_embed"}). Qdrant and PostgreSQL steps never exceeded 2.5 s. The same happens on ingestion (a 16.7 s add, all in ingest_embed).
In a full server log covering a 12.3 s stall, all 4,950 embedding calls in that window returned 200 OK, with no retries and no rate limiting. The provider answered slowly, and nothing cut the call short.
The embedder's client is created with the OpenAI SDK defaults (packages/server/src/memmachine_server/common/resource_manager/embedder_manager.py:212):
client=openai.AsyncOpenAI(
api_key=api_key,
base_url=conf.base_url,
),
That is a 600 s read timeout and 2 internal SDK retries per call. The embedder configuration has no timeout setting.
Steps to reproduce
- Run MemMachine against the OpenAI embeddings API (
text-embedding-3-small).
- Drive a sustained search load (20 concurrent clients for 60 s was enough).
- Watch
embedder_openai_latency_seconds{operation="search_embed"}: occasional requests land above 10 s, and the corresponding /memories/search requests take as long.
Expected behavior
A search embedding has a bounded wait (a few seconds) after which the search fails fast or degrades, rather than holding the request open for as long as the provider takes. Callers that wait for memory before answering (for example a chat that searches before every reply) usually give up after a few seconds anyway.
Environment
- OS: Linux (Ubuntu 24.04), Docker
- MemMachine Version:
main at c99bc0e (0.3.9+50.gc99bc0e), image built from source; also seen at c08cf26
- Development language version: Python 3.12 (in the image)
- Backend: event memory, PostgreSQL 18, Qdrant 1.19.1
- Embedder: OpenAI
text-embedding-3-small
Additional context
Suggested fix: pass timeout= to the client, short for search embeddings and longer for ingestion, and make both configurable in the embedder configuration. Consider max_retries=0 for search embeddings so an SDK retry does not double a stall.
Describe the bug
Under load, about one search in 1,500 takes 12–43 s while the p99 is about 1.2 s. MemMachine's own Prometheus histograms put the whole wait in the embedding step (
event_memory_query_phase_seconds{phase="embedding"}andembedder_openai_latency_seconds{operation="search_embed"}). Qdrant and PostgreSQL steps never exceeded 2.5 s. The same happens on ingestion (a 16.7 s add, all iningest_embed).In a full server log covering a 12.3 s stall, all 4,950 embedding calls in that window returned
200 OK, with no retries and no rate limiting. The provider answered slowly, and nothing cut the call short.The embedder's client is created with the OpenAI SDK defaults (
packages/server/src/memmachine_server/common/resource_manager/embedder_manager.py:212):That is a 600 s read timeout and 2 internal SDK retries per call. The embedder configuration has no timeout setting.
Steps to reproduce
text-embedding-3-small).embedder_openai_latency_seconds{operation="search_embed"}: occasional requests land above 10 s, and the corresponding/memories/searchrequests take as long.Expected behavior
A search embedding has a bounded wait (a few seconds) after which the search fails fast or degrades, rather than holding the request open for as long as the provider takes. Callers that wait for memory before answering (for example a chat that searches before every reply) usually give up after a few seconds anyway.
Environment
mainatc99bc0e(0.3.9+50.gc99bc0e), image built from source; also seen atc08cf26text-embedding-3-smallAdditional context
Suggested fix: pass
timeout=to the client, short for search embeddings and longer for ingestion, and make both configurable in the embedder configuration. Considermax_retries=0for search embeddings so an SDK retry does not double a stall.