Repository navigation
[Bug]: Search memory got 500 error during load testing: neo4j.exceptions.ConnectionAcquisitionTimeoutError: failed to obtain a connection from the pool within 60.0s (timeout) #940
Description
Activity
Issue:
During load testing with ~100 concurrent episodic memory searches (POST /api/v2/memories/search), about 30% of requests fail with HTTP 500 errors. The traceback shows the error is due toneo4j.exceptions.ConnectionAcquisitionTimeoutError: failed to obtain a connection from the pool within 60.0s (timeout). This is triggered deep inside the vector graph store during episodic memory searches:- The Neo4j driver cannot allocate new connections fast enough under concurrent demand, leading to service-level failures.
Root Cause:
- The default Neo4j connection pool is not sized or tuned for bursty high concurrency.
- Code expects connections to be available instantly; when demand spikes, requests wait up to the pool’s timeout, then error out.
- There is no application-level handling for this particular exception, so the API returns raw Internal Server Error (500), with no actionable message to the user or downstream.
Trace Summary:
- Request enters
/api/v2/memories/search. - Travels through several async layers.
- Eventually calls
neo4j_vector_graph_store.py:search_related_nodes, which defers todriver.execute_query. - Connection pool cannot acquire a connection within 60s → throws
ConnectionAcquisitionTimeoutError. - The error propagates out; no custom handling, so FastAPI/starlette returns a generic 500.
In #917, we introduced the ability for users to adjust the connection pool size for PostgreSQL databases to handle high-volume user connections. Seems we need to do the same for Neo4j.
While increasing the pool size could help, we must also be aware that the system under test (SUT) may not have sufficient system resources (CPU or Memory) to support larger pool sizes. If system resources are exhausted, it could also cause these timeouts.
Neo4j Connection Pool Tuning
Key Tunables
The main tunable for the Neo4j Python driver is the connection pool size. This is controlled via parameters when you initialize the driver, for example, in
neo4j.AsyncGraphDatabase.driver().Common parameters you may want to set:
-
max_connection_pool_size
Maximum number of simultaneous connections. Default is 100. Increase this for high concurrency. -
connection_acquisition_timeout
How long a request waits for a connection to be available, in seconds. Default is 60. You may want to increase or decrease this based on the expected load.
Other less common tunables:
max_connection_lifetimemax_transaction_retry_timeconnection_timeoutfetch_size
See official docs:
https://neo4j.com/docs/api/python-driver/current/api.html#driver-configurationExample Usage
from neo4j import AsyncGraphDatabase driver = AsyncGraphDatabase.driver( bolt_url, auth=(user, password), max_connection_pool_size=250, # Increased from default 100 connection_acquisition_timeout=120 # Wait up to 120s for a pool connection )
-
- addedperformanceIssues relating to MemMachine performanceIssues relating to MemMachine performance
on Jan 12, 2026 @szou-mv Can you capture the CPU and Memory utilization during the tests, please? I'm curious whether you're seeing prolonged 100% CPU utilization or high memory consumption that could indicate a system resource problem.
Have you tested on any other OS (Linux or Windows) using the same load script? If so, do you see the same problem there? If not, this could indicate the problem is specific to macOS, which helps narrow the focus.
@sscargal
I collected the cpu and memory data during load testing (100 concurrent episodic search) on my mac with a 30-second buffer on both ends of the test.data collected using python psutil
system monitor
But I'm a little confused by mac's memory usage because the memory usage number is high before the test.
And I am working on collecting data on linux.
Reacted by Jing@szou-mv Thanks for the info. The memory looks "normal" for Mac. The OS will handle paging. Your CPU charts look like we are not saturating them. The database likely can handle more connections.
Reacted by Jing@sscargal I can reproduce this issue on linux (rocky9 8c16G) with 50 concurrent episodic searches. Half of the requests failed with this issue.)
fail rate table:
round_concurrency ok fail fail_rate elapsed_s avg_response_time_s rps 10 100 0 0.000000 103.3959 8.264743 0.967156 20 100 0 0.000000 130.8846 19.951839 0.764032 30 120 0 0.000000 195.6437 35.796853 0.613360 40 120 0 0.000000 234.4330 57.978499 0.511873 50 52 48 0.480000 145.1151 58.918865 0.689108 60 55 65 0.541667 193.3623 79.929487 0.620597 70 42 98 0.700000 162.7956 70.378121 0.859974 80 35 125 0.781250 162.9615 71.416690 0.981827 90 57 123 0.683333 199.3921 79.017172 0.902744 100 17 83 0.830000 85.8767 75.886645 1.164461 Thanks @szou-mv . Can you test the code changes in PR #958? It exposes
max_connection_pool_sizeandconnection_acquisition_timeoutin the sample_config files. You can increase themax_connection_pool_sizeto 200, for example, and theconnection_acquisition_timeoutto 120 (seconds) to see if this resolves the issue.There may be other factors at play, as the default connection pool size is 100, and your tests fail at 50 users. That's why I was interested in the system resources, in case you were hitting a CPU or Memory limit in the test system.
@sscargal I have tested with latest build, however the results didn't change.
And I noticed something in the the source code.
It seems that for one episodic search, we firstly search for similar nodes, and the limit is 100. And then, we search related nodes for those nodes, which means there could be 100 database queries.
And there are 1000+ memories in my testing environment. Is it possible that the number of matched_derivative_nodes could reach 100 easily, and then we will do search_related_nodes for 100 times, and then the max_connection_pool_size run out for one search? So basically we can only handle the requests one by one?
@malatewang @edwinyyyu Is the observation by szou-mv expected? By default, Neo4j has a connection pool of 100. I added tunables to increase this, but szou-mv is seeing the issue at 50 users. I can see how 50 users, each performing SEARCH memory operations on over 1000+ memories, could easily storm the database connection pool, causing errors like
Too many connections,Connection pool exhausted, ortimeouts.Unbounded parallelism (with no upper concurrency limit) means that if someone raises the limit, the system could launch hundreds or thousands of simultaneous queries. Every user and every search could multiply this effect under load.
One possible option is to use a bounded semaphore or a task pool to cap parallel queries, e.g.:
sem = asyncio.Semaphore(max_parallel) async def throttled_call(...): async with sem: return await db_query(...)
Each matched derivative starts a new query. The QPS is not very good, which I have encountered in my own use. This can be improved as part of a database wrapper redesign.
#1199 and future changes should address performance concerns.
- addedkeep-openPrevents the auto-close task from closing this issue.Prevents the auto-close task from closing this issue.
on May 19, 2026
Describe the bug
During load testing, when there are 100 concurrent episodic memory search. About 30 requests failed with 500 error.
Request example:
Steps to reproduce
Expected behavior
Should avoid this issue or at least return proper error message.
Environment
OS: macOS
memmachine version: c5c5c59
Additional context
No response