Skip to content

[Bug]: Missing default max_input_length causes unhandled ValueError when ingesting long single messages #1298

Description

@junttang

Describe the bug

When ingesting chat data containing unusually long single messages, the semantic ingestion pipeline fails with a ValueError in the clustering stage if max_input_length is not explicitly configured by the user.

Although max_input_length exists in the sample configuration files, its default value in the code (OpenAIEmbedderConf) is set to None. If a user runs the server without explicitly setting this variable and injects a message that exceeds the limit (or is simply very long):

  1. The text splitting logic (chunk_text_balanced) inside the _embed() method is skipped.
  2. Massive text blocks (aggregated from multiple episodes) are passed directly to cluster_texts().
  3. cluster_texts() raises an unhandled ValueError because the input size exceeds the hard-coded cluster limit (max_total_length_per_cluster = 75,000).

Error log:

INFO:     127.0.0.1:43326 - "POST /api/v2/memories HTTP/1.1" 200 OK
INFO:     127.0.0.1:43326 - "POST /api/v2/memories HTTP/1.1" 500 Internal Server Error
ERROR:    Exception in ASGI application
  + Exception Group Traceback (most recent call last):
  |   File "/path/to/venv/lib/python3.12/site-packages/starlette/_utils.py", line 79, in collapse_excgroups
  |     yield
  |   ... (middleware stack traces omitted for brevity) ...
  |   File "/path/to/project/packages/server/src/memmachine_server/server/api_v2/router.py", line 245, in add_memories
  |     results = await _add_messages_to(...)
  |   File "/path/to/project/packages/server/src/memmachine_server/server/api_v2/service.py", line 66, in _add_messages_to
  |     episode_ids = await memmachine.add_episodes(...)
  |   File "/path/to/project/packages/server/src/memmachine_server/main/memmachine.py", line 659, in add_episodes
  |     await asyncio.gather(*tasks)
  |   File "/path/to/project/packages/server/src/memmachine_server/episodic_memory/episodic_memory.py", line 236, in add_memory_episodes
  |     await asyncio.gather(...)
  |   File "/path/to/project/packages/server/src/memmachine_server/episodic_memory/long_term_memory/long_term_memory.py", line 133, in add_episodes
  |     await self._declarative_memory.add_episodes(...)
  |   File "/path/to/project/packages/server/src/memmachine_server/episodic_memory/declarative_memory/declarative_memory.py", line 148, in add_episodes
  |     derivative_embeddings = await self._embedder.ingest_embed(...)
  |   File "/path/to/project/packages/server/src/memmachine_server/common/embedder/openai_embedder.py", line 114, in ingest_embed
  |     return await self._embed(inputs, max_attempts)
  |   File "/path/to/project/packages/server/src/memmachine_server/common/embedder/openai_embedder.py", line 146, in _embed

  |     chunk_clusters = cluster_texts(...)
  |   File "/path/to/project/packages/server/src/src/memmachine_server/common/utils.py", line 189, in cluster_texts
  |     raise ValueError(
  | ValueError: Text length 348859 exceeds max_total_length_per_cluster 75000
  +------------------------------------

The above exception was the direct cause of the following exception:

Traceback (most recent call last):
  ... (ASGI/Uvicorn stack traces omitted) ...
  File "/path/to/project/packages/server/src/memmachine_server/common/utils.py", line 189, in cluster_texts
    raise ValueError(
ValueError: Text length 348859 exceeds max_total_length_per_cluster 75000

While extremely long messages might be edge cases in standard interactive usage, this becomes critical for handling user inputs without strict client-side validation. Relying solely on manual configuration to prevent crashes poses a stability risk.

Steps to reproduce

  1. Start MemMachine server without explicitly setting max_input_length in cfg.yml.
  2. Prepare a dataset containing at least one very long message (e.g., a conversation turn exceeding 75,000 characters, or large enough to trigger the cluster limit when aggregated).
  3. Inject this data via the add(...) API.
  4. Observe the server crashing with a ValueError.

Note: Ensure message_sentence_chunking is False (default) so full content is passed as single derivatives.

Expected behavior

If None is intended, the documentation better explicitly warn that large batch imports require manual configuration of max_input_length. Otherwise, the system should default to a safe value (e.g., 2048) to ensure text splitting occurs automatically.

Environment

  • OS: Ubuntu 22.04.5 LTS
  • MemMachine Version: v0.3.3
  • Embedding Backend: OpenAI-compatible API (On-premise vLLM)

Additional context

No response

Activity

  1. added theissue type on Apr 6, 2026
  2. changed the title [-][Bug]: Missing default max_input_length causes unhandled ValueError in clustering during large batch ingestion[/-] [+][Bug]: Missing default max_input_length causes unhandled ValueError when ingesting long single messages[/+] on Apr 6, 2026
  3. junttang commented on Apr 6, 2026

    @junttang
    ContributorAuthor

    Update: after further analysis, I confirmed the root cause is the length of individual messages, not the injection speed, and have updated the issue details accordingly.

  4. sscargal commented on Apr 14, 2026

    @sscargal
    Contributor

    Root Cause Analysis

    The bug has a single, well-defined root cause with two contributing design decisions that compound it.

    Primary Root Cause — Missing guard in _embed() before cluster_texts()

    In openai_embedder.py, the _embed() method conditionally chunks input text only when self._max_input_length is not None:

    inputs_chunks = [
        chunk_text_balanced(input_text, self._max_input_length)
        if self._max_input_length is not None
        else [input_text]           # ← entire text passed as-is when None
        for input_text in inputs
    ]
    
    chunks = [chunk for input_chunks in inputs_chunks for chunk in input_chunks]
    chunk_clusters = cluster_texts(         # ← called unconditionally
        chunks,
        self.max_num_inputs_per_request,
        self.max_total_input_length_per_request,   # hard limit: 75,000
    )

    When max_input_length is None, no pre-splitting happens. Any single text whose len() exceeds max_total_input_length_per_request (75,000 characters) is passed directly to cluster_texts(), which then raises unconditionally:

    if text_length > max_total_length_per_cluster:
        raise ValueError(
            f"Text length {text_length} exceeds max_total_length_per_cluster {max_total_length_per_cluster}"
        )

    There is no fallback between "skip chunking" and "call cluster_texts."


    Contributing Factor 1 — max_input_length defaults to None everywhere

    OpenAIEmbedderConf, OpenAIEmbedderParams, AmazonBedrockEmbedderConf, SentenceTransformerEmbedderConf, and the API-level AddOpenAIEmbedderConfig all default max_input_length to None. There is no safe fallback value baked in, so users who don't explicitly set this field get no pre-splitting at all.


    Contributing Factor 2 — cluster_texts() hard-fails instead of splitting oversized texts

    cluster_texts() was designed to pack texts into batches, not to split individual oversized texts. When a single chunk exceeds the cluster limit it immediately raises a ValueError rather than falling back to splitting the offending text further. This makes it a hard failure point for any text that slips through un-chunked.


    The Fix (two complementary approaches)

    Option A (minimal/preferred): In _embed(), always apply a size guard. When self._max_input_length is None, fall back to self.max_total_input_length_per_request as the chunking limit so oversized texts are still split before reaching cluster_texts():

    effective_max = self._max_input_length or self.max_total_input_length_per_request
    inputs_chunks = [
        chunk_text_balanced(input_text, effective_max)
        for input_text in inputs
    ]

    Option B (defensive): Make cluster_texts() resilient by splitting any individual text that exceeds the limit rather than raising immediately (though this is a deeper change and masks mis-configuration).

    Option C (documentation): At minimum, update OpenAIEmbedderConf to use a safe non-None default (e.g. 2048) matching the embedding model's typical token ceiling, so the guard in _embed() always activates.

    The evaluation script at evaluation/retrieval_agent/longmemeval_test.py already works around this by monkey-patching max_total_input_length_per_request to a lower value — confirming the team is aware that large inputs hit this limit but hasn't addressed the upstream root cause.

  5. junttang commented on Apr 15, 2026

    @junttang
    ContributorAuthor

    @sscargal
    Thank you for a clear root cause analysis! It makes perfect sense.

    I agree with your Option A (using max_total_input_length_per_request as a fallback). It seems like the safest approach to prevent crashes without forcing users to manually configure every limit.

    Regarding the fix, would you prefer to handle the implementation yourself (or have the team do it), or should I give it a try?
    I can open a PR if that helps move things forward.

  6. sscargal commented on Apr 15, 2026

    @sscargal
    Contributor

    @junttang Thanks for the feedback. I pushed a proposed fix in PR #1328. If you have cycles to test the fix and provide feedback, it would be very helpful. Thank you.

  7. self-assigned this
    on Apr 15, 2026
  8. junttang commented on Apr 16, 2026

    @junttang
    ContributorAuthor

    @sscargal Thank you for the quick fix.

    I've tested the updated code from your branch, and I can confirm that the issue is now resolved. The server no longer crashes when ingesting long single messages without an explicit max_input_length configuration.

    Confirmed resolved on my end, and I will close the issue. Thanks again!

  9. added a commit that references this issue on Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions