Skip to content

[Bug]: Search memory got 500 error during load testing: neo4j.exceptions.ConnectionAcquisitionTimeoutError: failed to obtain a connection from the pool within 60.0s (timeout) #940

Description

@szou-mv

Describe the bug

During load testing, when there are 100 concurrent episodic memory search. About 30 requests failed with 500 error.

Request example:

POST /api/v2/memories/search
payload: {"org_id": "o_260112_1", "project_id": "2", "query": "has the user mentioned any specific database optimization tasks?", "top_k": 9, "types": ["episodic"]}
return code: 500
body: Internal Server Error

Steps to reproduce

Traceback (most recent call last):
  File "/app/.venv/lib/python3.12/site-packages/uvicorn/protocols/http/h11_impl.py", line 403, in run_asgi
    result = await app(  # type: ignore[func-returns-value]
             ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/uvicorn/middleware/proxy_headers.py", line 60, in __call__
    return await self.app(scope, receive, send)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/fastapi/applications.py", line 1139, in __call__
    await super().__call__(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/applications.py", line 107, in __call__
    await self.middleware_stack(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/errors.py", line 186, in __call__
    raise exc
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/errors.py", line 164, in __call__
    await self.app(scope, receive, _send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/base.py", line 191, in __call__
    with recv_stream, send_stream, collapse_excgroups():
                                   ^^^^^^^^^^^^^^^^^^^^
  File "/usr/local/lib/python3.12/contextlib.py", line 158, in __exit__
    self.gen.throw(value)
  File "/app/.venv/lib/python3.12/site-packages/starlette/_utils.py", line 85, in collapse_excgroups
    raise exc
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/base.py", line 193, in __call__
    response = await self.dispatch_func(request, call_next)
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/server/app.py", line 90, in access_log_middleware
    response = await call_next(request)
               ^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/base.py", line 168, in call_next
    raise app_exc from app_exc.__cause__ or app_exc.__context__
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/base.py", line 144, in coro
    await self.app(scope, receive_or_disconnect, send_no_error)
  File "/app/.venv/lib/python3.12/site-packages/starlette/middleware/exceptions.py", line 63, in __call__
    await wrap_app_handling_exceptions(self.app, conn)(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
    raise exc
  File "/app/.venv/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
    await app(scope, receive, sender)
  File "/app/.venv/lib/python3.12/site-packages/fastapi/middleware/asyncexitstack.py", line 18, in __call__
    await self.app(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/routing.py", line 716, in __call__
    await self.middleware_stack(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/routing.py", line 736, in app
    await route.handle(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/routing.py", line 290, in handle
    await self.app(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/fastapi/routing.py", line 119, in app
    await wrap_app_handling_exceptions(app, request)(scope, receive, send)
  File "/app/.venv/lib/python3.12/site-packages/starlette/_exception_handler.py", line 53, in wrapped_app
    raise exc
  File "/app/.venv/lib/python3.12/site-packages/starlette/_exception_handler.py", line 42, in wrapped_app
    await app(scope, receive, sender)
  File "/app/.venv/lib/python3.12/site-packages/fastapi/routing.py", line 105, in app
    response = await f(request)
               ^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/fastapi/routing.py", line 385, in app
    raw_response = await run_endpoint_function(
                   ^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/fastapi/routing.py", line 284, in run_endpoint_function
    return await dependant.call(**values)
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/server/api_v2/router.py", line 295, in search_memories
    return await _search_target_memories(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/server/api_v2/service.py", line 82, in _search_target_memories
    results = await memmachine.query_search(
              ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/main/memmachine.py", line 373, in query_search
    episodic_memory=await episodic_task if episodic_task else None,
                    ^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/main/memmachine.py", line 325, in _search_episodic_memory
    response = await episodic_session.query_memory(
               ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/episodic_memory/episodic_memory.py", line 362, in query_memory
    session_result, scored_long_episodes = await asyncio.gather(
                                           ^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/episodic_memory/long_term_memory/long_term_memory.py", line 162, in search_scored
    await self._declarative_memory.search_scored(
  File "/app/.venv/lib/python3.12/site-packages/memmachine/episodic_memory/declarative_memory/declarative_memory.py", line 385, in search_scored
    for episode_nodes in await asyncio.gather(
                         ^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/memmachine/common/vector_graph_store/neo4j_vector_graph_store.py", line 787, in search_related_nodes
    records, _, _ = await self._driver.execute_query(
                    ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/driver.py", line 947, in execute_query
    return await session._run_transaction(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/work/session.py", line 532, in _run_transaction
    await self._open_transaction(
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/work/session.py", line 398, in _open_transaction
    await self._connect(access_mode=access_mode)
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/work/session.py", line 126, in _connect
    await super()._connect(
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/work/workspace.py", line 184, in _connect
    self._connection = await self._pool.acquire(**acquire_kwargs_)
                       ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/io/_pool.py", line 681, in acquire
    return await self._acquire(
           ^^^^^^^^^^^^^^^^^^^^
  File "/app/.venv/lib/python3.12/site-packages/neo4j/_async/io/_pool.py", line 416, in _acquire
    raise ConnectionAcquisitionTimeoutError(
neo4j.exceptions.ConnectionAcquisitionTimeoutError: failed to obtain a connection from the pool within 60.0s (timeout)

Expected behavior

Should avoid this issue or at least return proper error message.

Environment

OS: macOS
memmachine version: c5c5c59

Additional context

No response

Activity

  1. added theissue type on Jan 12, 2026
  2. sscargal commented on Jan 12, 2026

    @sscargal
    Contributor

    Issue:
    During load testing with ~100 concurrent episodic memory searches (POST /api/v2/memories/search), about 30% of requests fail with HTTP 500 errors. The traceback shows the error is due to neo4j.exceptions.ConnectionAcquisitionTimeoutError: failed to obtain a connection from the pool within 60.0s (timeout). This is triggered deep inside the vector graph store during episodic memory searches:

    • The Neo4j driver cannot allocate new connections fast enough under concurrent demand, leading to service-level failures.

    Root Cause:

    • The default Neo4j connection pool is not sized or tuned for bursty high concurrency.
    • Code expects connections to be available instantly; when demand spikes, requests wait up to the pool’s timeout, then error out.
    • There is no application-level handling for this particular exception, so the API returns raw Internal Server Error (500), with no actionable message to the user or downstream.

    Trace Summary:

    • Request enters /api/v2/memories/search.
    • Travels through several async layers.
    • Eventually calls neo4j_vector_graph_store.py:search_related_nodes, which defers to driver.execute_query.
    • Connection pool cannot acquire a connection within 60s → throws ConnectionAcquisitionTimeoutError.
    • The error propagates out; no custom handling, so FastAPI/starlette returns a generic 500.
  3. sscargal commented on Jan 12, 2026

    @sscargal
    Contributor

    In #917, we introduced the ability for users to adjust the connection pool size for PostgreSQL databases to handle high-volume user connections. Seems we need to do the same for Neo4j.

    While increasing the pool size could help, we must also be aware that the system under test (SUT) may not have sufficient system resources (CPU or Memory) to support larger pool sizes. If system resources are exhausted, it could also cause these timeouts.

    Neo4j Connection Pool Tuning

    Key Tunables

    The main tunable for the Neo4j Python driver is the connection pool size. This is controlled via parameters when you initialize the driver, for example, in neo4j.AsyncGraphDatabase.driver().

    Common parameters you may want to set:

    • max_connection_pool_size
      Maximum number of simultaneous connections. Default is 100. Increase this for high concurrency.

    • connection_acquisition_timeout
      How long a request waits for a connection to be available, in seconds. Default is 60. You may want to increase or decrease this based on the expected load.

    Other less common tunables:

    • max_connection_lifetime
    • max_transaction_retry_time
    • connection_timeout
    • fetch_size

    See official docs:
    https://neo4j.com/docs/api/python-driver/current/api.html#driver-configuration

    Example Usage

    from neo4j import AsyncGraphDatabase
    
    driver = AsyncGraphDatabase.driver(
        bolt_url,
        auth=(user, password),
        max_connection_pool_size=250,         # Increased from default 100
        connection_acquisition_timeout=120    # Wait up to 120s for a pool connection
    )
  4. self-assigned this
    on Jan 12, 2026
  5. sscargal commented on Jan 12, 2026

    @sscargal
    Contributor

    @szou-mv Can you capture the CPU and Memory utilization during the tests, please? I'm curious whether you're seeing prolonged 100% CPU utilization or high memory consumption that could indicate a system resource problem.

    Have you tested on any other OS (Linux or Windows) using the same load script? If so, do you see the same problem there? If not, this could indicate the problem is specific to macOS, which helps narrow the focus.

  6. szou-mv commented on Jan 14, 2026

    @szou-mv
    Author

    @sscargal
    I collected the cpu and memory data during load testing (100 concurrent episodic search) on my mac with a 30-second buffer on both ends of the test.

    data collected using python psutil

    Image Image

    system monitor

    Image Image

    But I'm a little confused by mac's memory usage because the memory usage number is high before the test.

    And I am working on collecting data on linux.

  7. sscargal commented on Jan 14, 2026

    @sscargal
    Contributor

    @szou-mv Thanks for the info. The memory looks "normal" for Mac. The OS will handle paging. Your CPU charts look like we are not saturating them. The database likely can handle more connections.

  8. szou-mv commented on Jan 16, 2026

    @szou-mv
    Author

    @sscargal I can reproduce this issue on linux (rocky9 8c16G) with 50 concurrent episodic searches. Half of the requests failed with this issue.)

    Image

    fail rate table:

    round_concurrency ok fail fail_rate elapsed_s avg_response_time_s rps
    10 100 0 0.000000 103.3959 8.264743 0.967156
    20 100 0 0.000000 130.8846 19.951839 0.764032
    30 120 0 0.000000 195.6437 35.796853 0.613360
    40 120 0 0.000000 234.4330 57.978499 0.511873
    50 52 48 0.480000 145.1151 58.918865 0.689108
    60 55 65 0.541667 193.3623 79.929487 0.620597
    70 42 98 0.700000 162.7956 70.378121 0.859974
    80 35 125 0.781250 162.9615 71.416690 0.981827
    90 57 123 0.683333 199.3921 79.017172 0.902744
    100 17 83 0.830000 85.8767 75.886645 1.164461
  9. sscargal commented on Jan 16, 2026

    @sscargal
    Contributor

    Thanks @szou-mv . Can you test the code changes in PR #958? It exposes max_connection_pool_size and connection_acquisition_timeout in the sample_config files. You can increase the max_connection_pool_size to 200, for example, and the connection_acquisition_timeout to 120 (seconds) to see if this resolves the issue.

    There may be other factors at play, as the default connection pool size is 100, and your tests fail at 50 users. That's why I was interested in the system resources, in case you were hitting a CPU or Memory limit in the test system.

  10. szou-mv commented on Jan 19, 2026

    @szou-mv
    Author

    @sscargal I have tested with latest build, however the results didn't change.
    And I noticed something in the the source code.

    Image

    It seems that for one episodic search, we firstly search for similar nodes, and the limit is 100. And then, we search related nodes for those nodes, which means there could be 100 database queries.

    And there are 1000+ memories in my testing environment. Is it possible that the number of matched_derivative_nodes could reach 100 easily, and then we will do search_related_nodes for 100 times, and then the max_connection_pool_size run out for one search? So basically we can only handle the requests one by one?

  11. sscargal commented on Jan 20, 2026

    @sscargal
    Contributor

    @malatewang @edwinyyyu Is the observation by szou-mv expected? By default, Neo4j has a connection pool of 100. I added tunables to increase this, but szou-mv is seeing the issue at 50 users. I can see how 50 users, each performing SEARCH memory operations on over 1000+ memories, could easily storm the database connection pool, causing errors like Too many connections, Connection pool exhausted, or timeouts.

    Unbounded parallelism (with no upper concurrency limit) means that if someone raises the limit, the system could launch hundreds or thousands of simultaneous queries. Every user and every search could multiply this effect under load.

    One possible option is to use a bounded semaphore or a task pool to cap parallel queries, e.g.:

    sem = asyncio.Semaphore(max_parallel)
    async def throttled_call(...):
        async with sem:
            return await db_query(...)
  12. edwinyyyu commented on Jan 23, 2026

    @edwinyyyu
    Contributor

    Each matched derivative starts a new query. The QPS is not very good, which I have encountered in my own use. This can be improved as part of a database wrapper redesign.

  13. edwinyyyu commented on Apr 10, 2026

    @edwinyyyu
    Contributor

    #1199 and future changes should address performance concerns.

  14. added
    keep-openPrevents the auto-close task from closing this issue.
    on May 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

keep-openPrevents the auto-close task from closing this issue.performanceIssues relating to MemMachine performance

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions