Repository navigation
Background ingestion polling causes unbounded database load growth #1251
Description
Activity
- addedpriority: highIssue is urgent or highly impactful. Needs to be addressed as soon as possible.Issue is urgent or highly impactful. Needs to be addressed as soon as possible.
on Mar 23, 2026 Regarding multiple MemMachine processes working on the same database from
2. Every worker runs its own ingestion loop.
I've raised some concerns in different discussions which I want to reiterate.MemMachine as a whole is designed around the idea that there is only one server instance working on the database.
All MemMachine components maintain unbacked (incoherent) caches, global locks, along with other elements that coordinate at the process level.
This means that when reading data from episodic and other memories, data written by other processes will be missing during reads.
Alongside this global locks will be violated, possibly causing issues working with the databases.While the idea has been made of at some point supporting multiple memmachine processes through "session sharding". It isn't something that is currently supported.
Alongside this, some MemMachine components are more stateful, in which data is stored in memory instead of being written to the database. This is noticeable in writes to episodic memory, where a successful post request means that the data has been successfully written to memory, which is then written to the database in the background.
Because of this, any multi-process scaling, whether via uvicorn workers, separate Docker containers, or otherwise; is not currently safe. If we want to support it, we should design and implement that support.
Summary
The semantic memory background ingestion task uses an O(n²) correlated subquery that runs every 2 seconds per worker. Combined with the
set_ingested_historytable being append-only (rows are never deleted after ingestion), this causes database CPU and cost to grow unboundedly over time — even with zero active users.Impact
Production environment (Aurora Serverless v2):
set_ingested_historyDev environment (Aurora Serverless v2):
Root Cause
Three compounding issues:
1. O(n²) correlated subquery in
get_history_set_ids()sqlalchemy_pgvector_semantic.py:681-722generates:The correlated subquery evaluates against every row in the table, even when all rows are already ingested and the result is empty. With 14,055 rows in prod this produces ~197M row comparisons per execution.
2. Every worker runs its own ingestion loop
semantic_memory.py:149spawns a_background_ingestion_taskonstart(), which is called by every uvicorn worker. WithMEMMACHINE_WORKERS=5, the expensive query runs ~2.5 times per second (5 workers × every 2s), all doing identical work.3.
set_ingested_historyis append-onlyRows are inserted on message arrival and marked
ingested=trueafter processing, but never deleted. The table grows with every user interaction, making the polling query progressively slower — even though ingested rows are never relevant again.Suggested Fixes
Rewrite the query — replace the correlated subquery with
GROUP BY/HAVING:Single ingestion loop — ensure only one worker runs the background task (e.g., leader election or env-var guard), not all N workers.
Purge ingested rows — delete rows from
set_ingested_historyafter successful ingestion, or add a periodic cleanup for oldingested=truerows.Add composite index —
(set_id, ingested)onset_ingested_historyto speed up the uningested-row lookups.Add backoff sleep — the loop only sleeps when
dirty_setsis empty. When sets are found but processing fails (e.g., deadlock), it retries immediately in a tight loop.Reproduction
Deploy with
MEMMACHINE_WORKERS > 1, send enough messages to populateset_ingested_history, then observe database CPU climb even after all messages are ingested and no further API traffic occurs.Related