Skip to content

[Bug]: Context length errors from LLM are not detected, preventing split-and-retry mechanism for episodic summary #1309

Description

@junttang

Describe the bug

I am reporting that the issue raised three weeks ago in #1247 still persists in version v0.3.3.
Please refer to the original issue (#1247) for the detailed bug description, as the context remains the same.

Although a fix was previously merged via PR #1254, my testing indicates that the merged code does not resolve the problem.
The root cause is as follows:

When generating a summary, if the LLM throws a context length error, this error message should ideally be detected by the _is_exceed_context_window_error() function.
However, for some reason, the detection fails. Consequently, the split-and-retry mechanism for summaries is never triggered, causing the process to fail outright.

As mentioned in #1247, this occurs in a specific environment where:

  • Chat history data is being injected rapidly for benchmarking purposes.
  • An on-prem LLM is running (vLLM).

While this specific combination might trigger the issue, this bug poses a risk for possible features involving "bulk chat history loading".
Furthermore, it will affect any users deploying MemMachine with on-prem LLMs.
Therefore, I believe this requires re-evaluation.


Proposed Solution:
At the end of the thread in #1247, I shared a local workaround that completely resolved the issue for me. The solution involves explicitly returning the error message from the LLM response handler, ensuring it is correctly caught by the detection logic.

Could you please review this persistent issue and consider the proposed fix?

Steps to reproduce

Please refer to #1247.

Expected behavior

Please refer to #1247.

Environment

Please refer to #1247.

Additional context

No response

Activity

  1. added theissue type on Apr 8, 2026
  2. changed the title [-][Bug]: Issue #1247 persists: Context length errors from LLM are not detected, preventing split-and-retry mechanism for episodic summary[/-] [+][Bug]: Context length errors from LLM are not detected, preventing split-and-retry mechanism for episodic summary[/+] on Apr 8, 2026
  3. sscargal commented on Apr 17, 2026

    @sscargal
    Contributor

    Thanks for the bug report. We'll investigate and propose a fix.

  4. sscargal commented on Apr 21, 2026

    @sscargal
    Contributor

    Summary of the Bug, the Incomplete Fix, and the Correct Solution

    The Core Bug (Issue #1247)

    The episodic memory summarizer in short_term_memory.py has a split-and-retry mechanism: when a summary LLM call fails due to a context window overflow, _is_exceed_context_window_error() is supposed to detect it and recursively split the input in half, retrying with smaller batches.

    The function works by inspecting the exception's string representation for keywords like "context length", "token limit", "context window", etc.:

    @staticmethod
    def _is_exceed_context_window_error(e: Exception) -> bool:
        """Check if the exception is due to exceeding context window."""
        error_msg = str(e).lower()
        keywords = [
            "context length",
            "context window",
            "maximum input size",
            "input length",
            "token limit",
            "too long",
            "exceeds the maximum",
        ]
        return any(keyword in error_msg for keyword in keywords)

    The problem: The exception that actually reaches _is_exceed_context_window_error() is never the raw openai.BadRequestError — it is an ExternalServiceAPIError that wraps it. And the message passed to ExternalServiceAPIError was constructed without including the original LLM error message:

    error_message = (
        f"[call uuid: {generate_response_call_uuid}] "
        "Giving up generating response "
        f"after failed attempt {attempt} "
        f"due to non-retryable {type(e).__name__}"   # ← only the class name, NO message body
    )
    raise ExternalServiceAPIError(error_message) from e

    So str(ExternalServiceAPIError(...)) contains only:

    "[call uuid: ...] Giving up generating response after failed attempt 1 due to non-retryable BadRequestError"

    None of the context-window keywords are present. _is_exceed_context_window_error() returns False, the split-and-retry branch is never taken, and the process fails with "Summarization failed due to unexpected error" — potentially losing summary data.

    The actual vLLM error message that would have matched — "You passed 202753 input tokens... the model's context length is only 202752 tokens" — is attached as __cause__ on the exception but is never surfaced in the ExternalServiceAPIError message string.


    Why PR #1254 Did Not Fix This

    PR #1254 addressed a different, related gap — it added handling for the edge case where a single oversized episode can't be halved at the episode list level (since there's only one episode). It introduced _split_episode() to split the content of a single episode in half and retry:

    if len(episodes) == 1:
        split_episode = self._split_episode(episodes[0])
        ...
        batches = [[split_episode[0]], [split_episode[1]]]
    else:
        mid = len(episodes) // 2
        batches = [episodes[:mid], episodes[mid:]]

    This is a valuable improvement — but it is downstream of the detection check. The if self._is_exceed_context_window_error(e): guard still has to return True before any of this new splitting logic is reached. Since the root cause (the missing original error message in ExternalServiceAPIError) was never fixed, _is_exceed_context_window_error() still returns False, the if branch is never entered, and all the new splitting logic in PR #1254 is completely unreachable in practice.

    PR #1254 fixed the splitting logic. It never fixed the detection logic. The new unit test in the PR uses a plain ValueError("User prompt exceeds context window") — which does contain the keyword "context window" — so the test passes. But in production, the raised exception is always an ExternalServiceAPIError with a message that contains no such keywords.


    Recommended Fix

    The cleanest and most robust solution is to make _is_exceed_context_window_error walk the full exception chain via __cause__ and __context__, exactly as _is_context_length_exceeded_error already does in semantic_ingestion.py. This finds the keyword in the original BadRequestError where it actually lives, without needing to modify the ExternalServiceAPIError message string at all.

    Change 1 — short_term_memory.py

    Replace the current _is_exceed_context_window_error implementation:

    # BEFORE
    @staticmethod
    def _is_exceed_context_window_error(e: Exception) -> bool:
        """Check if the exception is due to exceeding context window."""
        error_msg = str(e).lower()
        keywords = [
            "context length",
            "context window",
            "maximum input size",
            "input length",
            "token limit",
            "too long",
            "exceeds the maximum",
        ]
        return any(keyword in error_msg for keyword in keywords)
    # AFTER
    @staticmethod
    def _is_exceed_context_window_error(e: Exception) -> bool:
        """Check if the exception is due to exceeding context window.
    
        Walks the full exception chain (__cause__ and __context__) so that
        the original LLM error message is inspected even when the immediate
        exception is a wrapper such as ExternalServiceAPIError.
        """
        keywords = [
            "context length",
            "context window",
            "maximum input size",
            "input length",
            "token limit",
            "too long",
            "exceeds the maximum",
        ]
        seen: set[int] = set()
        current: BaseException | None = e
        while current is not None and id(current) not in seen:
            seen.add(id(current))
            if any(keyword in str(current).lower() for keyword in keywords):
                return True
            current = current.__cause__ or current.__context__
        return False

    This is a self-contained, non-breaking change. No other files need to be modified for the fix itself.


    Change 2 (Optional but Recommended) — Update the unit test

    The existing test in test_short_term_memory.py uses a bare ValueError and therefore does not cover the real production failure mode. A second test case should be added that wraps the context-window error inside an ExternalServiceAPIError, matching what the production code path actually raises:

    @pytest.mark.asyncio
    async def test_summary_exceed_context_window_wrapped_in_external_service_error(
        self, mock_data_manager
    ):
        """
        Regression test for the detection gap that caused PR #1254 to be ineffective.
    
        In production, the LLM adapter raises ExternalServiceAPIError wrapping the
        original openai.BadRequestError. The detection function must walk the exception
        chain to find the context-window keyword in the original cause, not just the
        wrapper message.
        """
        from memmachine_server.common.data_types import ExternalServiceAPIError
    
        class WrappedContextWindowModel(LanguageModel):
            def __init__(self) -> None:
                self.call_count = 0
    
            async def generate_response(
                self,
                system_prompt: str | None = None,
                user_prompt: str | None = None,
                tools: list[dict[str, Any]] | None = None,
                tool_choice: str | dict[str, str] | None = None,
                max_attempts: int = 1,
            ) -> tuple[str, Any]:
                prompt = user_prompt or ""
                if len(prompt) > 10000:
                    # Simulate the real production path: the LLM adapter catches the
                    # openai.BadRequestError and wraps it in ExternalServiceAPIError,
                    # with a message that contains only the exception class name —
                    # NOT the original error text.
                    original = ValueError(
                        "You passed 202753 input tokens. "
                        "The model's context length is only 202752 tokens."
                    )
                    raise ExternalServiceAPIError(
                        "[call uuid: test-uuid] Giving up generating response "
                        "after failed attempt 1 due to non-retryable BadRequestError"
                    ) from original
                self.call_count += 1
                match = re.search(r"summary:(\d+)(?: \d+)?$", prompt)
                prev_summary = int(match.group(1)) if match else 0
                messages = re.findall(r'"([^"]*)"', prompt)
                return (
                    f"summary:{prev_summary + sum(len(m) for m in messages)}",
                    "",
                )
    
            async def generate_response_with_token_usage(
                self,
                system_prompt: str | None = None,
                user_prompt: str | None = None,
                tools: list[dict[str, Any]] | None = None,
                tool_choice: str | dict[str, str] | None = None,
                max_attempts: int = 1,
            ) -> tuple[str, Any, int, int]:
                response, tool_output = await self.generate_response(
                    system_prompt=system_prompt,
                    user_prompt=user_prompt,
                    tools=tools,
                    tool_choice=tool_choice,
                    max_attempts=max_attempts,
                )
                return response, tool_output, 0, 0
    
            async def generate_parsed_response(
                self,
                output_format: type[T],
                system_prompt: str | None = None,
                user_prompt: str | None = None,
                max_attempts: int = 1,
            ) -> T:
                return cast(T, "summary")
    
        model = WrappedContextWindowModel()
        params = ShortTermMemoryConsolidator.Params(
            summary_user_prompt="User prompt: {episodes} {summary} {max_length}",
            summary_system_prompt="System Prompt",
            max_summary_length_words=100,
            session_key="test_session",
            model=model,
            data_manager=mock_data_manager,
        )
        consolidator = ShortTermMemoryConsolidator(params)
    
        oversized_content = "a" * 12000
        episode = create_test_episode(content=oversized_content)
        summary = await consolidator._create_summary("", [episode])
    
        # If detection is broken, call_count == 0 and summary == "" (data loss).
        # If detection works, the episode is split and retried successfully.
        assert model.call_count > 1
        assert summary != ""

    Security Consideration

    No changes are made to the ExternalServiceAPIError message string, so there is no risk of LLM error bodies (or any user content) being inadvertently forwarded into API responses or logs beyond what already exists. The fix is purely internal to the exception chain inspection logic.


    Files to Change

    File Change
    packages/server/src/memmachine_server/episodic_memory/short_term_memory/short_term_memory.py Replace _is_exceed_context_window_error with chain-walking implementation
    packages/server/server_tests/memmachine_server/episodic_memory/short_term_memory/test_short_term_memory.py Add regression test with ExternalServiceAPIError-wrapped cause

    No changes are required to openai_chat_completions_language_model.py or any other LLM adapter.

  5. added this to the v0.3.6 milestone on Apr 21, 2026
  6. self-assigned this
    on Apr 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions