Skip to content

[Bug]: Sometimes the LLM is stuck in an infinite loop when ingesting semantic memories #974

Description

@mwqgithub

Describe the bug

Sometimes the LLM is stuck in an infinite loop when ingesting semantic memories.

Steps to reproduce

{
    "messages": [
        {
            "role": "system",
            "content": "\n        Your job is to handle memory extraction for a memory system, one which takes the form of a profile recording details relevant to the tags below.\n        You will receive a profile and a user's query to the chat system, your job is to update that profile by extracting or inferring information about the user from the query.\n        A profile is a two-level key-value store. We call the outer key the *tag*, and the inner key the *feature*. Together, a *tag* and a *feature* are associated with one or several *value*s.\n\n        \n    IMPORTANT: Extract ALL personal information, even basic facts like names, ages, locations, etc. Do not consider any personal information as \"irrelevant\" - names, basic demographics, and simple facts are valuable profile data.\n\n\n        How to construct profile entries:\n        - Entries should be atomic. They should communicate a single discrete fact.\n        - Entries should be as short as possible without corrupting meaning. Be careful when leaving out prepositions, qualifiers, negations, etc. Some modifiers will be longer range, find the best way to compactify such phrases.\n        - You may see entries which violate the above rules, those are \"consolidated memories\". Don't rewrite those.\n        - Think of yourself as performing the role of a wide, early layer in a neural network, doing \"edge detection\" in many places in parallel to present as many distinct intermediate features as you possibly can given raw, unprocessed input.\n\n        The tags you are looking for include:\n        \t- Assistant Response Preferences: How the user prefers the assistant to communicate (style, tone, structure, data format).\n\t- Notable Past Conversation Topic Highlights: Recurring or significant discussion themes.\n\t- Helpful User Insights: Key insights that help personalize assistant behavior.\n\t- User Interaction Metadata: Behavioral/technical metadata about platform use.\n\t- Political Views, Likes and Dislikes: Explicit opinions or stated preferences.\n\t- Psychological Profile: Personality characteristics or traits.\n\t- Communication Style: Describes the user's communication tone and pattern.\n\t- Learning Preferences: Preferred modes of receiving information.\n\t- Cognitive Style: How the user processes information or makes decisions.\n\t- Emotional Drivers: Motivators like fear of error or desire for clarity.\n\t- Personal Values: User's core values or principles.\n\t- Career & Work Preferences: Interests, titles, domains related to work.\n\t- Productivity Style: User's work rhythm, focus preference, or task habits.\n\t- Demographic Information: Education level, fields of study, or similar data.\n\t- Geographic & Cultural Context: Physical location or cultural background.\n\t- Financial Profile: Any relevant information about financial behavior or context.\n\t- Health & Wellness: Physical/mental health indicators.\n\t- Education & Knowledge Level: Degrees, subjects, or demonstrated expertise.\n\t- Platform Behavior: Patterns in how the user interacts with the platform.\n\t- Tech Proficiency: Languages, tools, frameworks the user knows.\n\t- Hobbies & Interests: Non-work-related interests.\n\t- Social Identity: Group affiliations or demographics.\n\t- Media Consumption Habits: Types of media consumed (e.g., blogs, podcasts).\n\t- Life Goals & Milestones: Short- or long-term aspirations.\n\t- Relationship & Family Context: Any information about personal life.\n\t- Risk Tolerance: Comfort with uncertainty, experimentation, or failure.\n\t- Assistant Trust Level: Whether and when the user trusts assistant responses.\n\t- Time Usage Patterns: Frequency and habits of use.\n\t- Preferred Content Format: Formats preferred for answers (e.g., tables, bullet points).\n\t- Assistant Usage Patterns: Habits or styles in how the user engages with the assistant.\n\t- Language Preferences: Preferred tone and structure of assistant's language.\n\t- Motivation Triggers: Traits that drive engagement or satisfaction.\n\t- Behavior Under Stress: How the user reacts to failures or inaccurate responses.\n\n        To update the profile, you will output a JSON document containing a list of commands to be executed in sequence.\n\n        CRITICAL: You MUST use the command format below. Do NOT create nested objects or use any other format.\n\n        The following output will add a feature:\n        [\n            {\n                \"command\": \"add\",\n                \"tag\": \"Preferred Content Format\",\n                \"feature\": \"unicode_for_math\",\n                \"value\": true\n            }\n        ]\n        The following will delete all values associated with the feature:\n        [\n            {\n                \"command\": \"delete\",\n                \"tag\" : \"Language Preferences\",\n                \"feature\": \"format\"\n            }\n        ]\n        The following will update a feature:\n        [\n            {\n                \"command\": \"delete\",\n                \"tag\": \"Platform Behavior\",\n                \"feature\": \"prefers_detailed_responses\",\n                \"value\": true\n            },\n            {\n                \"command\": \"add\",\n                \"tag\" : \"Platform Behavior\",\n                \"feature\": \"prefers_detailed_response\",\n                \"value\": false\n            }\n        ]\n\n        Example Scenarios:\n        Query: \"Hi! My name is Katara\"\n        [\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"name\",\n                \"value\": \"Katara\"\n            }\n        ]\n        Query: \"I'm planning a dinner party for 8 people next weekend and want to impress my guests with something special. Can you suggest a menu that's elegant but not too difficult for a home cook to manage?\"\n        [\n            {\n                \"command\": \"add\",\n                \"tag\": \"Hobbies & Interests\",\n                \"feature\": \"home_cook\",\n                \"value\": \"User cooks fancy food\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Financial Profile\",\n                \"feature\": \"upper_class\",\n                \"value\": \"User entertains guests at dinner parties, suggesting affluence.\"\n            }\n        ]\n        Query: my boss (for the summer) is totally washed. he forgot how to all the basics but still thinks he does\n        [\n            {\n                \"command\": \"add\",\n                \"tag\": \"Psychological Profile\",\n                \"feature\": \"work_superior_frustration\",\n                \"value\": \"User is frustrated with their boss for perceived incompetence\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"summer_job\",\n                \"value\": \"User is working a temporary job for the summer\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Communication Style\",\n                \"feature\": \"informal_speech\",\n                \"value\": \"User speaks with all lower case letters and contemporary slang terms.\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"young_adult\",\n                \"value\": \"User is young, possibly still in college\"\n            }\n        ]\n        Query: Can you go through my inbox and flag any urgent emails from clients, then update the project status spreadsheet with the latest deliverable dates from those emails? Also send a quick message to my manager letting her know I'll have the budget report ready by end of day tomorrow.\n        [\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"traditional_office_job\",\n                \"value\": \"User does clerical work, reporting to a manager\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"client_facing_role\",\n                \"value\": \"User handles communication of deadlines to and from clients\"\n            },\n            {\n                \"command\": \"add\",\n                \"tag\": \"Demographic Information\",\n                \"feature\": \"autonomy_at_work\",\n                \"value\": \"User sets their own deadlines and subtasks.\"\n            }\n        }\n        Further Guidelines:\n        - Not everything you ought to record will be explicitly stated. Make inferences.\n        - If you are less confident about a particular entry, you should still include it, but make sure that the language you use (briefly) expresses this uncertainty in the value field\n        - Look at the text from as many distinct angles as you can find, remember you are the \"wide layer\".\n        - Keep only the key details (highest-entropy) in the feature name. The nuances go in the value field.\n        - Do not couple together distinct details. Just because the user associates together certain details, doesn't mean you should\n        - Do not create new tags which you don't see in the example profile. However, you can and should create new features.\n        - If a user asks for a summary of a report, code, or other content, that content may not necessarily be written by the user, and might not be relevant to the user's profile.\n        - Do not delete anything unless a user asks you to\n        - Only return the empty list [] if the query contains absolutely no personal information about the user (e.g., asking about the weather, requesting code without personal context, etc.). Names, basic demographics, preferences, and any personal details should ALWAYS be extracted.\n        - Listen to any additional instructions specific to the execution context provided underneath 'EXTRA EXTERNAL INSTRUCTIONS'\n        - First, think about what should go in the profile inside <think> </think> tags. Then output only a valid JSON.\n        - REMEMBER: Always use the command format with \"command\", \"tag\", \"feature\", and \"value\" keys. Never use nested objects or any other format.\n    "
        },
        {
            "role": "user",
            "content": "The old feature set is provided below:\n<OLD_PROFILE>\n{\"Demographic Information\": {\"name\": \"alice\"}, \"Hobbies & Interests\": {\"sandwiches\": \"User loves sandwiches.\", \"swimming\": \"User likes swimming.\"}}\n</OLD_PROFILE>\n\nThe history is provided below:\n<HISTORY>\nHello, my name is bob\n</HISTORY>\n"
        }
    ],
    "model": "qwen3-max",
    "response_format": {
        "type": "json_schema",
        "json_schema": {
            "schema": {
                "$defs": {
                    "SemanticCommand": {
                        "description": "Normalized instruction emitted by the LLM to mutate semantic features.",
                        "properties": {
                            "command": {
                                "$ref": "#/$defs/SemanticCommandType"
                            },
                            "feature": {
                                "title": "Feature",
                                "type": "string"
                            },
                            "tag": {
                                "title": "Tag",
                                "type": "string"
                            },
                            "value": {
                                "title": "Value",
                                "type": "string"
                            }
                        },
                        "required": [
                            "command",
                            "feature",
                            "tag",
                            "value"
                        ],
                        "title": "SemanticCommand",
                        "type": "object",
                        "additionalProperties": false
                    },
                    "SemanticCommandType": {
                        "description": "Semantic memory actions that can be applied to a feature.",
                        "enum": [
                            "add",
                            "delete"
                        ],
                        "title": "SemanticCommandType",
                        "type": "string"
                    }
                },
                "description": "Schema used to validate parsed feature-update commands returned by the LLM.",
                "properties": {
                    "commands": {
                        "items": {
                            "$ref": "#/$defs/SemanticCommand"
                        },
                        "title": "Commands",
                        "type": "array"
                    }
                },
                "title": "_SemanticFeatureUpdateRes",
                "type": "object",
                "additionalProperties": false,
                "required": [
                    "commands"
                ]
            },
            "name": "_SemanticFeatureUpdateRes",
            "strict": true
        }
    },
    "stream": true
}
curl --location 'https://dashscope.aliyuncs.com/compatible-mode/v1/chat/completions' \
    --header "Authorization: Bearer sk-xxxx" \
    --header 'Content-Type: application/json' \
    --data @request.json
Image

Expected behavior

The LLM does not get stuck in infinite loops.

Environment

  • OS: Linux

Additional context

No response

Activity

  1. added theissue type on Jan 26, 2026
  2. jealous commented on Jan 26, 2026

    @jealous
    Contributor

    I think I met the same issue.

  3. jealous commented on Jan 26, 2026

    @jealous
    Contributor

    Looks like a lot of items to consolidate.

    Image
  4. o-love commented on Jan 27, 2026

    @o-love
    Contributor

    There are a few issues happening in Cedric's environment.
    Which also seems to match up with The original issue.

    Three main things are happening when consolidating messages after ingestion:

    • Features aren't being properly grouped together.
    • The groups aren't being properly filtered based on number of entries.
    • There seems to be a lot of features with empty name and content, and odd names and content.

    All this causes thousands of calls of no value to the LLM provider.
    The expected behavior would be to create a few groups of a decent amount of features to then be consolidated into a single or a few features.
    So groups of 20 similar features to be consolidated into 1 or 2 features.
    Instead of the spam of groups of 1 or 2 features to be consolidated together.

    Alongside this, there seems to be a lot of empty or low quality features in Cedric's environment. I'm not sure how these got added to Cedric's environment but It would be useful to ensure that features generated by the LLM need a minimum quality. At the very least no empty fields.

    cc @jgong

  5. added
    priority: highIssue is urgent or highly impactful. Needs to be addressed as soon as possible.
    on Jan 27, 2026
  6. o-love commented on Jan 27, 2026

    @o-love
    Contributor

    Setting priority to High. Since so many calls to the llm will have a significant cost implication

  7. o-love commented on Jan 27, 2026

    @o-love
    Contributor

    To note, the default threshold for consolidation is 20 features with the same tag.
    The expected behavior is that after ingesting new memories into a category. We check if any tag in that category has more than 20 tags.
    If so we then call a LLM to consolidate the features in that tag.

  8. added 3 commits that reference this issue on Jan 27, 2026
  9. added a commit that references this issue on Jan 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

priority: highIssue is urgent or highly impactful. Needs to be addressed as soon as possible.

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions