Repository navigation
[Bug]: Missing default max_input_length causes unhandled ValueError when ingesting long single messages #1298
Description
Activity
- changed the title
[-][Bug]: Missing default max_input_length causes unhandled ValueError in clustering during large batch ingestion[/-][+][Bug]: Missing default max_input_length causes unhandled ValueError when ingesting long single messages[/+]on Apr 6, 2026 Update: after further analysis, I confirmed the root cause is the length of individual messages, not the injection speed, and have updated the issue details accordingly.
Root Cause Analysis
The bug has a single, well-defined root cause with two contributing design decisions that compound it.
Primary Root Cause — Missing guard in
_embed()beforecluster_texts()In
openai_embedder.py, the_embed()method conditionally chunks input text only whenself._max_input_length is not None:inputs_chunks = [ chunk_text_balanced(input_text, self._max_input_length) if self._max_input_length is not None else [input_text] # ← entire text passed as-is when None for input_text in inputs ] chunks = [chunk for input_chunks in inputs_chunks for chunk in input_chunks] chunk_clusters = cluster_texts( # ← called unconditionally chunks, self.max_num_inputs_per_request, self.max_total_input_length_per_request, # hard limit: 75,000 )
When
max_input_lengthisNone, no pre-splitting happens. Any single text whoselen()exceedsmax_total_input_length_per_request(75,000 characters) is passed directly tocluster_texts(), which then raises unconditionally:if text_length > max_total_length_per_cluster: raise ValueError( f"Text length {text_length} exceeds max_total_length_per_cluster {max_total_length_per_cluster}" )
There is no fallback between "skip chunking" and "call
cluster_texts."
Contributing Factor 1 —
max_input_lengthdefaults toNoneeverywhereOpenAIEmbedderConf,OpenAIEmbedderParams,AmazonBedrockEmbedderConf,SentenceTransformerEmbedderConf, and the API-levelAddOpenAIEmbedderConfigall defaultmax_input_lengthtoNone. There is no safe fallback value baked in, so users who don't explicitly set this field get no pre-splitting at all.
Contributing Factor 2 —
cluster_texts()hard-fails instead of splitting oversized textscluster_texts()was designed to pack texts into batches, not to split individual oversized texts. When a single chunk exceeds the cluster limit it immediately raises aValueErrorrather than falling back to splitting the offending text further. This makes it a hard failure point for any text that slips through un-chunked.
The Fix (two complementary approaches)
Option A (minimal/preferred): In
_embed(), always apply a size guard. Whenself._max_input_length is None, fall back toself.max_total_input_length_per_requestas the chunking limit so oversized texts are still split before reachingcluster_texts():effective_max = self._max_input_length or self.max_total_input_length_per_request inputs_chunks = [ chunk_text_balanced(input_text, effective_max) for input_text in inputs ]
Option B (defensive): Make
cluster_texts()resilient by splitting any individual text that exceeds the limit rather than raising immediately (though this is a deeper change and masks mis-configuration).Option C (documentation): At minimum, update
OpenAIEmbedderConfto use a safe non-Nonedefault (e.g.2048) matching the embedding model's typical token ceiling, so the guard in_embed()always activates.The evaluation script at
evaluation/retrieval_agent/longmemeval_test.pyalready works around this by monkey-patchingmax_total_input_length_per_requestto a lower value — confirming the team is aware that large inputs hit this limit but hasn't addressed the upstream root cause.@sscargal
Thank you for a clear root cause analysis! It makes perfect sense.I agree with your Option A (using
max_total_input_length_per_requestas a fallback). It seems like the safest approach to prevent crashes without forcing users to manually configure every limit.Regarding the fix, would you prefer to handle the implementation yourself (or have the team do it), or should I give it a try?
I can open a PR if that helps move things forward.@sscargal Thank you for the quick fix.
I've tested the updated code from your branch, and I can confirm that the issue is now resolved. The server no longer crashes when ingesting long single messages without an explicit max_input_length configuration.
Confirmed resolved on my end, and I will close the issue. Thanks again!
- added a commit that references this issue
on Apr 16, 2026 - added a commit that references this issue
on Apr 20, 2026 - added a commit that references this issue
on Sep 18, 2026 - added a commit that references this issue
on Sep 18, 2026
Describe the bug
When ingesting chat data containing unusually long single messages, the semantic ingestion pipeline fails with a ValueError in the clustering stage if
max_input_lengthis not explicitly configured by the user.Although
max_input_lengthexists in the sample configuration files, its default value in the code (OpenAIEmbedderConf) is set to None. If a user runs the server without explicitly setting this variable and injects a message that exceeds the limit (or is simply very long):chunk_text_balanced) inside the_embed()method is skipped.cluster_texts().cluster_texts()raises an unhandled ValueError because the input size exceeds the hard-coded cluster limit (max_total_length_per_cluster= 75,000).Error log:
While extremely long messages might be edge cases in standard interactive usage, this becomes critical for handling user inputs without strict client-side validation. Relying solely on manual configuration to prevent crashes poses a stability risk.
Steps to reproduce
max_input_lengthincfg.yml.Expected behavior
If None is intended, the documentation better explicitly warn that large batch imports require manual configuration of
max_input_length. Otherwise, the system should default to a safe value (e.g., 2048) to ensure text splitting occurs automatically.Environment
Additional context
No response