Skip to content

[Bug]: Dead loop to trigger memmachine_server.semantic_memory.semantic_ingestion when LLM reponse is invalid #1270

Description

@pengfeiz901

Describe the bug

This is a related issue with #1265

When the issue in #1265 happened, we could see semantic_ingestion keep called repeatly and report errors.
Does this waste user's token? If yes, it is a critical issue.

memmachine-output.log

Steps to reproduce

  • Configure to use qwen3.5-plus or qwen3.5-flash in configuration.yml
  • add some memory to generate semantic memory and you will see such errors in docker logs.

Expected behavior

MemMachine add/search memory should compatible with qwen3.5 LLM response.

Environment

memmachine 0.3.2
qwen3.5-plus, qwen3.5-flash

Additional context

No response

Activity

  1. added theissue type on Mar 26, 2026
  2. sscargal commented on Apr 14, 2026

    @sscargal
    Contributor

    When Qwen3.5 models respond with <think>...</think> blocks (thinking mode is on by default), the pydantic JSON parser rejects the response. This causes a persistent retry loop where the same messages are re-processed indefinitely, wasting tokens on every cycle.


    Root Cause Analysis

    The bug is the intersection of two separate problems — one in the LLM output (addressed in #1265) and one structural flaw in the ingestion pipeline that makes the impact catastrophic.

    Problem 1 — LLM parsing failure (from #1265)

    Qwen3.5 prepends a <think>...</think> block to its output even when responding to structured JSON requests. The OpenAI-compatible client passes the raw text directly to pydantic's model_validate_json(), which fails immediately since the input starts with <think> rather than {:

    pydantic_core._pydantic_core.ValidationError: 1 validation error for _SemanticFeatureUpdateRes
      Invalid JSON: expected value at line 1 column 1
      [input_value='<think>\nThe user is sha...']
    

    This throws an exception inside llm_feature_update().

    Problem 2 — Failed messages are never marked as ingested (the root cause of the loop)

    In _process_single_set() (semantic_ingestion.py, lines 201–237), when llm_feature_update() raises a non-context_length_exceeded exception, the handler does this:

    except Exception as err:
        if _is_context_length_exceeded_error(err):
            # ... marks message as ingested and continues
            mark_messages.append(message.uid)
            continue
    
        logger.exception(
            "Failed to process message %s for semantic type %s",
            message.uid,
            semantic_category.name,
        )
        if self._debug_fail_loudly:
            raise
    
        continue   # <-- message.uid is NOT added to mark_messages

    The pydantic ValidationError from the Qwen3.5 <think> prefix is not a context_length_exceeded error, so it falls through to the generic handler — which logs, then continues without adding message.uid to mark_messages.

    Because the message is never in mark_messages, mark_messages_ingested() is never called for it. On every subsequent background loop iteration, get_history_set_ids() finds the same set still has uningested messages and schedules it for processing again:

    while not self._is_shutting_down:
        dirty_sets = [
            s
            async for s in self._semantic_storage.get_history_set_ids(
                min_uningested_messages=self._feature_update_message_limit,
                older_than=datetime.now(tz=UTC) - self._feature_time_limit,
            )
        ]
        # dirty_sets will always contain the failing set_id
        ...
        await ingestion_service.process_set_ids(dirty_sets)

    The result is an infinite retry loop — every background cycle re-submits the same message to the LLM, which gives the same bad response, which fails parsing, which logs an error, which leaves the message uningested, and the cycle repeats indefinitely. Each cycle burns real LLM tokens.

    The context_length_exceeded path correctly handles permanent failures (it does add to mark_messages), but the generic error path does not, making any other persistent LLM failure turn into an infinite loop.


    Contributing Factor — No max-retry / skip-after-N-failures logic

    There is no per-message failure counter or skip-after-N-attempts guard in the ingestion pipeline. This means any permanently broken message (not just Qwen3.5 think-mode, but e.g. a malformed message body that always causes an LLM error) will loop forever. The context_length_exceeded branch is the only "skip on permanent failure" path, and it's too narrow.


    Severity

    Critical — confirmed token waste, potential unbounded API cost, and service degradation for any user with a Qwen3.5 model configured. Issue #1265 (the parse fix) was closed, but without fixing the retry loop the root cause of token waste remains.


    Fix Recommendations

    Two complementary fixes are needed:

    Fix A — Strip <think> blocks before parsing (closes the immediate trigger)

    In openai_responses_language_model.py (or in llm_feature_update), strip any <think>...</think> prefix from the LLM response text before passing it to model_validate_json. This is the fix that issue #1265 was supposedly targeting.

    Fix B — Mark messages as "skip" after a persistent parse/LLM failure (closes the loop)

    In process_semantic_type in semantic_ingestion.py, treat pydantic.ValidationError (and any other non-transient LLM response errors) the same way context_length_exceeded is treated — add the message to mark_messages so it is marked ingested and won't be retried forever:

    except Exception as err:
        if _is_context_length_exceeded_error(err):
            ...
            mark_messages.append(message.uid)
            continue
    
        # NEW: also skip on persistent parse failures to avoid infinite retry loops
        if isinstance(err, (ValidationError, ValueError, TypeError)):
            logger.warning(
                "Skipping message %s for semantic type %s due to non-retryable parse error",
                message.uid,
                semantic_category.name,
            )
            if message.uid not in mark_messages:
                mark_messages.append(message.uid)
            continue
    
        logger.exception(...)
        continue

    Alternatively, a general approach would be to add a per-message retry counter across ingestion cycles and skip after N failures, which would catch any future class of permanent errors.

  3. sscargal commented on Apr 15, 2026

    @sscargal
    Contributor

    @pengfeiz901 I have a proposed fix in PR #1327. If you are able to verify it, I welcome feedback. I don't have an easy way to verify myself. Thanks.

  4. github-actions commented on Jun 15, 2026

    @github-actions
    Contributor

    This issue has been automatically marked as stale because it has not had recent activity. It will be closed in 14 days if no further activity occurs. If this issue is still relevant, please comment to keep it open, or add the keep-open label if it should remain open indefinitely (e.g., a roadmap item awaiting a champion). Thank you for your contributions.

  5. github-actions commented on Jun 30, 2026

    @github-actions
    Contributor

    This issue was automatically closed due to inactivity. If it is still relevant, please reopen it or file a new issue referencing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions