Skip to content

[Feat]: Add retry logic for LLM judge and semantic memory consolidation on structured output failures #1332

Description

@junttang

Is your feature request related to a problem?

Some LLM call sites using structured output (JSON format) lack retry logic when response parsing fails, causing transient errors to be treated as permanent failures.


Details:

Evaluation w/ LLM judge:

At evaluation/retrieval_agent/llm_judge.py::evaluate_llm_judge(),

  raw = call_fn(prompt)
  label = json_repair.loads(raw)["label"]  
  return 1 if label == "CORRECT" else 0
  • LLM returns None or JSON missing the "label" key.
  • Raw SDK calls have no retry handling.

In LoCoMo Benchmark with GLM-4.7-Flash model, ~6-8 failures out of 1540 judgments (~0.5%).


Semantic memory feature consolidation:

At packages/server/src/memmachine_server/semantic_memory/semantic_ingestion.py::_deduplicate_features(),

  • On ValueError/TypeError from llm_consolidate_features(), the function returns immediately. Retry logic seems to be absent.
  • Unlike Profile Extraction, may also lack _is_context_length_exceeded_error handling.

Describe the solution you'd like

LLM judge retry logic:

Add retry handling in evaluate_llm_judge() or within the create_judge_fn() call functions to catch JSON parsing failures (json_repair.loads(raw)["label"]) and retry the LLM call 1-2 times before marking the judgment as failed.


Semantic memory consolidation retry logic

Mirror the approach from PR #1327 for Profile Extraction. For example, add retry logic in _deduplicate_features() when llm_consolidate_features() raises ValueError/TypeError, with appropriate handling for non-retryable errors (e.g., context length exceeded).

Describe alternatives you've considered

Alternative Pros Cons
No retry, fail silently (current) Simple implementation Loses eval data, accumulates redundant memories
Fail loudly (raise exception) Makes failures visible Breaks benchmark runs, requires manual intervention
Retry with exponential backoff Better handling of rate limits More complex, may delay execution
Retry once (proposed) Simple, handles most transient errors May not resolve persistent issues

We tested single retry locally for LLM judge in evaluation/retrieval_agent with GLM-4.7-Flash and confirmed it resolves the ~0.5% failure rate experimentally.

Additional context

Invalid JSON patterns in semantic memory consolidation are categorized as below.

  1. Truncated responses – JSON cuts off before completion (EOF mid-string)
  2. Instruction ignoring – LLM adds preamble text or wraps output in markdown code blocks (```json)
  3. Control characters – Unicode control chars (\u0000–\u001F) embedded in strings
  4. Structural hallucination – Unclosed braces, random text inserted mid-JSON

Activity

  1. added theissue type on Apr 17, 2026
  2. sscargal commented on Apr 17, 2026

    @sscargal
    Contributor

    Thanks for the feature submission. We'll review and add it to a future milestone.

  3. sscargal commented on Apr 21, 2026

    @sscargal
    Contributor

    Thanks for the detailed and well-researched issue report. After reviewing the current code at both call sites, I can confirm both gaps are real and actionable. Here's a full technical review.


    Finding 1 — LLM Judge (evaluate_llm_judge)

    The current code at evaluation/retrieval_agent/llm_judge.py#L156-L158:

    raw = call_fn(prompt)
    label = json_repair.loads(raw)["label"]
    return 1 if label == "CORRECT" else 0

    There are actually two distinct failure modes here, not one:

    1. call_fn returns None or "" — json_repair.loads("") returns "", which is not subscriptable, raising TypeError. The _call_responses variant already guards with or "" but _call_chat does not (resp.choices[0].message.content can be None).
    2. LLM returns valid JSON but omits the "label" key — json_repair.loads(raw)["label"] raises KeyError. This is the primary failure mode you observed (~0.5% with GLM-4.7-Flash).

    Additionally, the main() loop at llm_judge.py#L200 has no exception handling around evaluate_llm_judge(...), meaning a single failure aborts the entire benchmark run rather than just skipping that judgment.

    Recommended fix for evaluate_llm_judge:

    MAX_JUDGE_RETRIES = 2
    
    def evaluate_llm_judge(
        question: str,
        gold_answer: str,
        generated_answer: str,
        call_fn: Callable[[str], str],
    ) -> int | None:
        """Evaluate a generated answer. Returns 1 (CORRECT), 0 (WRONG), or None on failure."""
        prompt = ACCURACY_PROMPT.format(
            question=question,
            gold_answer=gold_answer,
            generated_answer=generated_answer,
        )
        for attempt in range(1, MAX_JUDGE_RETRIES + 2):
            try:
                raw = call_fn(prompt)
                parsed = json_repair.loads(raw or "")
                if not isinstance(parsed, dict) or "label" not in parsed:
                    raise ValueError(f"Missing 'label' key in response: {raw!r}")
                label = parsed["label"]
                return 1 if label == "CORRECT" else 0
            except Exception as exc:
                if attempt <= MAX_JUDGE_RETRIES:
                    logger.warning(
                        "LLM judge attempt %d/%d failed (%s); retrying",
                        attempt,
                        MAX_JUDGE_RETRIES + 1,
                        exc,
                    )
                else:
                    logger.error(
                        "LLM judge failed after %d attempts, skipping judgment: %s",
                        MAX_JUDGE_RETRIES + 1,
                        exc,
                    )
                    return None

    And in main(), update the call site to handle None:

    label = evaluate_llm_judge(question, gold_answer, generated_answer, call_fn)
    if label is None:
        continue  # skip rather than crash the benchmark run

    Finding 2 — Semantic Memory Consolidation (_deduplicate_features)

    The current code at semantic_ingestion.py#L381-L391:

    try:
        consolidate_resp = await llm_consolidate_features(...)
    except (ValueError, TypeError):
        logger.exception("Failed to update features while calling LLM")
        if self._debug_fail_loudly:
            raise
        return

    Two confirmed gaps versus the llm_feature_update path in process_semantic_type:

    Behaviour process_semantic_type (llm_feature_update) _deduplicate_features (llm_consolidate_features)
    Retry on transient parse failure ❌ No retry either, but... ❌ No retry
    _is_context_length_exceeded_error check ✅ Present ❌ Missing
    pydantic.ValidationError catch ✅ Added by PR #1327 ❌ Not present
    Broad except Exception ✅ Present ❌ Only (ValueError, TypeError)

    Note on PR #1327: That PR's Fix B narrowed the catch in process_semantic_type to pydantic.ValidationError specifically (on architect guidance that ValueError/TypeError can leak from the embedder and storage layers). The same reasoning applies here — the (ValueError, TypeError) catch in _deduplicate_features may already be too broad in one dimension while being too narrow in another (missing ValidationError and context-length handling).

    Recommended fix for _deduplicate_features:

    _MAX_CONSOLIDATION_RETRIES = 1
    
    async def _deduplicate_features(self, *, set_id, memories, semantic_category, resources) -> None:
        consolidate_resp = None
        last_err: Exception | None = None
    
        for attempt in range(1, _MAX_CONSOLIDATION_RETRIES + 2):
            try:
                consolidate_resp = await llm_consolidate_features(
                    features=list(memories),
                    model=resources.language_model,
                    consolidate_prompt=semantic_category.prompt.consolidation_prompt,
                )
                break  # success
            except Exception as err:
                if _is_context_length_exceeded_error(err):
                    logger.warning(
                        "Skipping consolidation for set_id %s / category %s: "
                        "non-retryable context length error",
                        set_id,
                        semantic_category.name,
                    )
                    if self._debug_fail_loudly:
                        raise
                    return
                last_err = err
                if attempt <= _MAX_CONSOLIDATION_RETRIES:
                    logger.warning(
                        "Consolidation LLM attempt %d/%d failed for set_id %s, retrying: %s",
                        attempt,
                        _MAX_CONSOLIDATION_RETRIES + 1,
                        set_id,
                        err,
                    )
                else:
                    logger.exception(
                        "Failed to consolidate features for set_id %s after %d attempts",
                        set_id,
                        _MAX_CONSOLIDATION_RETRIES + 1,
                    )
                    if self._debug_fail_loudly:
                        raise
                    return
    
        # ... rest of existing logic unchanged

    Design Considerations

    A few points worth discussing before implementation:

    1. Retry delay: A simple immediate retry is fine for JSON parse failures (these are not rate-limit errors). However, if call_fn / generate_parsed_response can itself raise rate-limit errors (HTTP 429), a short backoff (e.g., asyncio.sleep(1.0 * attempt)) would be safer. This may be worth checking in generate_parsed_response's implementation.

    2. _call_chat None guard: The _call_responses path already guards with or "", but _call_chat returns resp.choices[0].message.content directly, which can be None per the OpenAI SDK types. This should be patched to return resp.choices[0].message.content or "" for consistency.

    3. Catch width in _deduplicate_features: Per the PR fix(semantic-memory): break ingestion retry loop on <think>-prefixed LLM responses #1327 architect note, consider whether the retry loop should catch pydantic.ValidationError specifically (for transient structured-output failures) and keep except Exception as a separate non-retried fallback — this prevents inadvertently retrying errors from the embedder or storage layers that happen to raise ValueError.

    4. Metrics/observability: Given the known ~0.5% failure rate on GLM-4.7-Flash, it may be worth adding a counter or log line that is easy to grep for in production to track retry effectiveness over time.

    5. Testing: Unit tests should cover: (a) success on retry-after-failure, (b) None return after exhausting retries for the judge, (c) context-length-exceeded skips without retry in consolidation, and (d) debug_fail_loudly=True still raises even after retries.


    Summary

    The feature request is well-scoped and confirmed valid by direct code inspection. Both call sites lack the protections that exist elsewhere in the codebase. The proposed single-retry approach is the right default — simple, effective for transient parse failures, and consistent with what was experimentally validated. The main additions beyond what the issue describes are: guarding the _call_chat None case, adding _is_context_length_exceeded_error to _deduplicate_features, and aligning the exception catch width with the guidance from PR #1327.

  4. self-assigned this
    on Apr 21, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions