Repository navigation
[Bug]: Dead loop to trigger memmachine_server.semantic_memory.semantic_ingestion when LLM reponse is invalid #1270
Description
Activity
When Qwen3.5 models respond with
<think>...</think>blocks (thinking mode is on by default), thepydanticJSON parser rejects the response. This causes a persistent retry loop where the same messages are re-processed indefinitely, wasting tokens on every cycle.
Root Cause Analysis
The bug is the intersection of two separate problems — one in the LLM output (addressed in #1265) and one structural flaw in the ingestion pipeline that makes the impact catastrophic.
Problem 1 — LLM parsing failure (from #1265)
Qwen3.5 prepends a
<think>...</think>block to its output even when responding to structured JSON requests. The OpenAI-compatible client passes the raw text directly topydantic'smodel_validate_json(), which fails immediately since the input starts with<think>rather than{:pydantic_core._pydantic_core.ValidationError: 1 validation error for _SemanticFeatureUpdateRes Invalid JSON: expected value at line 1 column 1 [input_value='<think>\nThe user is sha...']This throws an exception inside
llm_feature_update().Problem 2 — Failed messages are never marked as ingested (the root cause of the loop)
In
_process_single_set()(semantic_ingestion.py, lines 201–237), whenllm_feature_update()raises a non-context_length_exceededexception, the handler does this:except Exception as err: if _is_context_length_exceeded_error(err): # ... marks message as ingested and continues mark_messages.append(message.uid) continue logger.exception( "Failed to process message %s for semantic type %s", message.uid, semantic_category.name, ) if self._debug_fail_loudly: raise continue # <-- message.uid is NOT added to mark_messages
The
pydanticValidationErrorfrom the Qwen3.5<think>prefix is not acontext_length_exceedederror, so it falls through to the generic handler — which logs, thencontinues without addingmessage.uidtomark_messages.Because the message is never in
mark_messages,mark_messages_ingested()is never called for it. On every subsequent background loop iteration,get_history_set_ids()finds the same set still has uningested messages and schedules it for processing again:while not self._is_shutting_down: dirty_sets = [ s async for s in self._semantic_storage.get_history_set_ids( min_uningested_messages=self._feature_update_message_limit, older_than=datetime.now(tz=UTC) - self._feature_time_limit, ) ] # dirty_sets will always contain the failing set_id ... await ingestion_service.process_set_ids(dirty_sets)
The result is an infinite retry loop — every background cycle re-submits the same message to the LLM, which gives the same bad response, which fails parsing, which logs an error, which leaves the message uningested, and the cycle repeats indefinitely. Each cycle burns real LLM tokens.
The
context_length_exceededpath correctly handles permanent failures (it does add tomark_messages), but the generic error path does not, making any other persistent LLM failure turn into an infinite loop.
Contributing Factor — No max-retry / skip-after-N-failures logic
There is no per-message failure counter or skip-after-N-attempts guard in the ingestion pipeline. This means any permanently broken message (not just Qwen3.5 think-mode, but e.g. a malformed message body that always causes an LLM error) will loop forever. The
context_length_exceededbranch is the only "skip on permanent failure" path, and it's too narrow.
Severity
Critical — confirmed token waste, potential unbounded API cost, and service degradation for any user with a Qwen3.5 model configured. Issue #1265 (the parse fix) was closed, but without fixing the retry loop the root cause of token waste remains.
Fix Recommendations
Two complementary fixes are needed:
Fix A — Strip
<think>blocks before parsing (closes the immediate trigger)In
openai_responses_language_model.py(or inllm_feature_update), strip any<think>...</think>prefix from the LLM response text before passing it tomodel_validate_json. This is the fix that issue #1265 was supposedly targeting.Fix B — Mark messages as "skip" after a persistent parse/LLM failure (closes the loop)
In
process_semantic_typeinsemantic_ingestion.py, treatpydantic.ValidationError(and any other non-transient LLM response errors) the same waycontext_length_exceededis treated — add the message tomark_messagesso it is marked ingested and won't be retried forever:except Exception as err: if _is_context_length_exceeded_error(err): ... mark_messages.append(message.uid) continue # NEW: also skip on persistent parse failures to avoid infinite retry loops if isinstance(err, (ValidationError, ValueError, TypeError)): logger.warning( "Skipping message %s for semantic type %s due to non-retryable parse error", message.uid, semantic_category.name, ) if message.uid not in mark_messages: mark_messages.append(message.uid) continue logger.exception(...) continue
Alternatively, a general approach would be to add a per-message retry counter across ingestion cycles and skip after N failures, which would catch any future class of permanent errors.
@pengfeiz901 I have a proposed fix in PR #1327. If you are able to verify it, I welcome feedback. I don't have an easy way to verify myself. Thanks.
github-actions commented
on Jun 15, 2026 on Jun 15, 2026 – with GitHub ActionsContributorMore actionsThis issue has been automatically marked as stale because it has not had recent activity. It will be closed in 14 days if no further activity occurs. If this issue is still relevant, please comment to keep it open, or add the
keep-openlabel if it should remain open indefinitely (e.g., a roadmap item awaiting a champion). Thank you for your contributions.github-actions commented
on Jun 30, 2026 on Jun 30, 2026 – with GitHub ActionsContributorMore actionsThis issue was automatically closed due to inactivity. If it is still relevant, please reopen it or file a new issue referencing this one.
Describe the bug
This is a related issue with #1265
When the issue in #1265 happened, we could see semantic_ingestion keep called repeatly and report errors.
Does this waste user's token? If yes, it is a critical issue.
memmachine-output.log
Steps to reproduce
Expected behavior
MemMachine add/search memory should compatible with qwen3.5 LLM response.
Environment
memmachine 0.3.2
qwen3.5-plus, qwen3.5-flash
Additional context
No response