Skip to content

[Bug]: Excessive background semantic/profile memory LLM calls #1453

Description

@wenhaocs

Describe the bug

MemMachine appears to trigger an unexpectedly large number of LLM calls for
semantic/profile memory ingestion after a relatively small amount of user
activity.

I used our internal product which integrates Memchine to ask around 40 questions over a small knowledge
set: 4 PDF files, each under 80 pages. I did not start any other jobs or run any
batch ingestion manually.

However, starting around midnight and continuing into the evening, the system
kept calling the LLM in the background and generated about $200 in model
costs.

A captured single LLM request shows that the call was not answering my PDF
question directly. It was a MemMachine profile/semantic memory extraction cal:

  • system prompt: "Your job is to handle memory extraction for a memory system..."
  • user prompt includes:
    • <OLD_PROFILE>{}</OLD_PROFILE>
    • <HISTORY>...previous assistant answer about Apple's gross margin...</HISTORY>

So the unexpected cost seems related to background semantic/profile memory
ingestion or repeated add_memory processing, not normal user-facing QA.

Steps to reproduce

Possible minimal reproduction path:

  1. Enable both episodic and semantic memory.
  2. Call add_memory repeatedly for the same user/project, passing full
    conversation context each time.
  3. Wait for semantic ingestion to run.
  4. Observe whether MemMachine repeatedly sends profile extraction prompts to the
    LLM even when there is no new user activity.

Expected behavior

After the pending memory ingestion backlog is processed, MemMachine should stop
calling the LLM while idle.

LLM calls for profile/semantic memory extraction should be bounded and
predictable. For a small document set and around 40 user questions, background
memory processing should not continue for many hours or generate very high
model costs.

Ideally MemMachine should also provide safeguards such as:

  • max background ingestion calls per project/user per time window
  • deduplication/idempotency for already-ingested messages
  • an option to disable automatic semantic/profile ingestion from MCP
  • cost/concurrency limits for background ingestion

Environment

MemMachine commit: af5a1a3

Additional context

one LLM call.txt

Activity

  1. added theissue type on Jun 15, 2026
  2. hegu-1 commented on Jul 17, 2026

    @hegu-1

    This failure mode suggests the background extractor lacks a stable work identity. Rate limits alone would cap the bill but not stop duplicate semantic/profile work from being scheduled.

    I would key each extraction job by something like (tenant, user, memory_kind, source_revision, extractor_version). Enqueue becomes idempotent; completion stores a receipt; retries reuse the same job id. If full conversation context is submitted repeatedly, the source revision/hash prevents reprocessing unchanged history.

    Add three independent controls:

    • per-user/project concurrency and cost budget;
    • a queue/status API showing pending/running/completed/failed jobs and estimated cost;
    • a hard pause/cancel switch that workers check between calls.

    Profile extraction should probably be delta-based: compare new source ids against the last completed watermark rather than send the full history each time. If the watermark is missing or moves backward, fail visibly instead of starting an unbounded replay.

    Regression tests should simulate duplicate add_memory, worker restart, retry after timeout, stale watermark, and queue cancellation. The acceptance invariant is simple: with no new source revision, the system reaches idle and performs zero further LLM calls.

    This also protects personal continuity as user-owned state rather than an invisible background process. The privacy/control principles I use are here: https://github.com/hegu-1/personal-ai-os/blob/main/docs/privacy.md

  3. github-actions commented on Sep 16, 2026

    @github-actions
    Contributor

    This issue has been automatically marked as stale because it has not had recent activity. It will be closed in 14 days if no further activity occurs. If this issue is still relevant, please comment to keep it open, or add the keep-open label if it should remain open indefinitely (e.g., a roadmap item awaiting a champion). Thank you for your contributions.

  4. github-actions commented on Sep 30, 2026

    @github-actions
    Contributor

    This issue was automatically closed due to inactivity. If it is still relevant, please reopen it or file a new issue referencing this one.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions