Repository navigation
[Feat]: Add retry logic for LLM judge and semantic memory consolidation on structured output failures #1332
Description
Activity
Thanks for the feature submission. We'll review and add it to a future milestone.
Reacted by Jace ParkThanks for the detailed and well-researched issue report. After reviewing the current code at both call sites, I can confirm both gaps are real and actionable. Here's a full technical review.
Finding 1 — LLM Judge (
evaluate_llm_judge)The current code at
evaluation/retrieval_agent/llm_judge.py#L156-L158:raw = call_fn(prompt) label = json_repair.loads(raw)["label"] return 1 if label == "CORRECT" else 0
There are actually two distinct failure modes here, not one:
call_fnreturnsNoneor""—json_repair.loads("")returns"", which is not subscriptable, raisingTypeError. The_call_responsesvariant already guards withor ""but_call_chatdoes not (resp.choices[0].message.contentcan beNone).- LLM returns valid JSON but omits the
"label"key —json_repair.loads(raw)["label"]raisesKeyError. This is the primary failure mode you observed (~0.5% with GLM-4.7-Flash).
Additionally, the
main()loop atllm_judge.py#L200has no exception handling aroundevaluate_llm_judge(...), meaning a single failure aborts the entire benchmark run rather than just skipping that judgment.Recommended fix for
evaluate_llm_judge:MAX_JUDGE_RETRIES = 2 def evaluate_llm_judge( question: str, gold_answer: str, generated_answer: str, call_fn: Callable[[str], str], ) -> int | None: """Evaluate a generated answer. Returns 1 (CORRECT), 0 (WRONG), or None on failure.""" prompt = ACCURACY_PROMPT.format( question=question, gold_answer=gold_answer, generated_answer=generated_answer, ) for attempt in range(1, MAX_JUDGE_RETRIES + 2): try: raw = call_fn(prompt) parsed = json_repair.loads(raw or "") if not isinstance(parsed, dict) or "label" not in parsed: raise ValueError(f"Missing 'label' key in response: {raw!r}") label = parsed["label"] return 1 if label == "CORRECT" else 0 except Exception as exc: if attempt <= MAX_JUDGE_RETRIES: logger.warning( "LLM judge attempt %d/%d failed (%s); retrying", attempt, MAX_JUDGE_RETRIES + 1, exc, ) else: logger.error( "LLM judge failed after %d attempts, skipping judgment: %s", MAX_JUDGE_RETRIES + 1, exc, ) return None
And in
main(), update the call site to handleNone:label = evaluate_llm_judge(question, gold_answer, generated_answer, call_fn) if label is None: continue # skip rather than crash the benchmark run
Finding 2 — Semantic Memory Consolidation (
_deduplicate_features)The current code at
semantic_ingestion.py#L381-L391:try: consolidate_resp = await llm_consolidate_features(...) except (ValueError, TypeError): logger.exception("Failed to update features while calling LLM") if self._debug_fail_loudly: raise return
Two confirmed gaps versus the
llm_feature_updatepath inprocess_semantic_type:Behaviour process_semantic_type(llm_feature_update)_deduplicate_features(llm_consolidate_features)Retry on transient parse failure ❌ No retry either, but... ❌ No retry _is_context_length_exceeded_errorcheck✅ Present ❌ Missing pydantic.ValidationErrorcatch✅ Added by PR #1327 ❌ Not present Broad except Exception✅ Present ❌ Only (ValueError, TypeError)Note on PR #1327: That PR's Fix B narrowed the catch in
process_semantic_typetopydantic.ValidationErrorspecifically (on architect guidance thatValueError/TypeErrorcan leak from the embedder and storage layers). The same reasoning applies here — the(ValueError, TypeError)catch in_deduplicate_featuresmay already be too broad in one dimension while being too narrow in another (missingValidationErrorand context-length handling).Recommended fix for
_deduplicate_features:_MAX_CONSOLIDATION_RETRIES = 1 async def _deduplicate_features(self, *, set_id, memories, semantic_category, resources) -> None: consolidate_resp = None last_err: Exception | None = None for attempt in range(1, _MAX_CONSOLIDATION_RETRIES + 2): try: consolidate_resp = await llm_consolidate_features( features=list(memories), model=resources.language_model, consolidate_prompt=semantic_category.prompt.consolidation_prompt, ) break # success except Exception as err: if _is_context_length_exceeded_error(err): logger.warning( "Skipping consolidation for set_id %s / category %s: " "non-retryable context length error", set_id, semantic_category.name, ) if self._debug_fail_loudly: raise return last_err = err if attempt <= _MAX_CONSOLIDATION_RETRIES: logger.warning( "Consolidation LLM attempt %d/%d failed for set_id %s, retrying: %s", attempt, _MAX_CONSOLIDATION_RETRIES + 1, set_id, err, ) else: logger.exception( "Failed to consolidate features for set_id %s after %d attempts", set_id, _MAX_CONSOLIDATION_RETRIES + 1, ) if self._debug_fail_loudly: raise return # ... rest of existing logic unchanged
Design Considerations
A few points worth discussing before implementation:
-
Retry delay: A simple immediate retry is fine for JSON parse failures (these are not rate-limit errors). However, if
call_fn/generate_parsed_responsecan itself raise rate-limit errors (HTTP 429), a short backoff (e.g.,asyncio.sleep(1.0 * attempt)) would be safer. This may be worth checking ingenerate_parsed_response's implementation. -
_call_chatNoneguard: The_call_responsespath already guards withor "", but_call_chatreturnsresp.choices[0].message.contentdirectly, which can beNoneper the OpenAI SDK types. This should be patched toreturn resp.choices[0].message.content or ""for consistency. -
Catch width in
_deduplicate_features: Per the PR fix(semantic-memory): break ingestion retry loop on <think>-prefixed LLM responses #1327 architect note, consider whether the retry loop should catchpydantic.ValidationErrorspecifically (for transient structured-output failures) and keepexcept Exceptionas a separate non-retried fallback — this prevents inadvertently retrying errors from the embedder or storage layers that happen to raiseValueError. -
Metrics/observability: Given the known ~0.5% failure rate on GLM-4.7-Flash, it may be worth adding a counter or log line that is easy to grep for in production to track retry effectiveness over time.
-
Testing: Unit tests should cover: (a) success on retry-after-failure, (b)
Nonereturn after exhausting retries for the judge, (c) context-length-exceeded skips without retry in consolidation, and (d)debug_fail_loudly=Truestill raises even after retries.
Summary
The feature request is well-scoped and confirmed valid by direct code inspection. Both call sites lack the protections that exist elsewhere in the codebase. The proposed single-retry approach is the right default — simple, effective for transient parse failures, and consistent with what was experimentally validated. The main additions beyond what the issue describes are: guarding the
_call_chatNonecase, adding_is_context_length_exceeded_errorto_deduplicate_features, and aligning the exception catch width with the guidance from PR #1327.- added a commit that references this issue
on Apr 21, 2026
Is your feature request related to a problem?
Some LLM call sites using structured output (JSON format) lack retry logic when response parsing fails, causing transient errors to be treated as permanent failures.
Details:
Evaluation w/ LLM judge:
At
evaluation/retrieval_agent/llm_judge.py::evaluate_llm_judge(),In LoCoMo Benchmark with GLM-4.7-Flash model, ~6-8 failures out of 1540 judgments (~0.5%).
Semantic memory feature consolidation:
At
packages/server/src/memmachine_server/semantic_memory/semantic_ingestion.py::_deduplicate_features(),llm_consolidate_features(), the function returns immediately. Retry logic seems to be absent.Describe the solution you'd like
LLM judge retry logic:
Add retry handling in
evaluate_llm_judge()or within thecreate_judge_fn()call functions to catch JSON parsing failures (json_repair.loads(raw)["label"]) and retry the LLM call 1-2 times before marking the judgment as failed.Semantic memory consolidation retry logic
Mirror the approach from PR #1327 for Profile Extraction. For example, add retry logic in
_deduplicate_features()whenllm_consolidate_features()raisesValueError/TypeError, with appropriate handling for non-retryable errors (e.g., context length exceeded).Describe alternatives you've considered
We tested single retry locally for LLM judge in
evaluation/retrieval_agentwith GLM-4.7-Flash and confirmed it resolves the ~0.5% failure rate experimentally.Additional context
Invalid JSON patterns in semantic memory consolidation are categorized as below.