Skip to content

[coding agents 2/3] Serve memory_query and memory_expand over MCP at /v1/mcp - #1691

Draft
edwinyyyu wants to merge 57 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/event-memory-v1-mcp-main
Draft

edwinyyyu wants to merge 57 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/event-memory-v1-mcp-main

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Slice 2 of design/coding_agent_integration.md (#1579): the two recall tools a coding agent calls, served by the server over MCP at /v1/mcp. Stacked on #1690 (the v1 API), whose seam it uses; its diff shows #1690 and the stack below it until they merge. Rebuilt on 2026-09-28 with the stack; its four commits applied unchanged.

What it serves

A second FastMCP app, mounted at /v1/mcp beside the untouched v2 app at /mcp, with the tenant read from the X-MemMachine-Tenant header into a context variable of its own, so the two apps' contexts never mix. Tenants resolve through the same EpisodicSessionTenantEventMemories the routes use, against the one MemMachine instance the process holds, so a tenant reached through a tool and through a route is one memory. Both apps' lifespans are chained by the application's, since a mounted app's lifespan is not run by its mounter.

Tools, with no count or size parameter on either, as the design requires:

  • memory_query(cue, within=None, kinds=None, since=None, until=None): the tenant's EventMemory.query with the tenant's defaults (query_limit, expand_context), reranked through the tenant's reranker over rerank_candidates when it has one, each hit's window rendered with the session and segment markers, best first, one blank line between hits. within is a session id, kinds block kinds, since and until ISO 8601 timestamps that must carry an offset.
  • memory_expand(id, direction="around"): id is [segment:<hex32>] or the bare uuid; a [segments:a..b] run is refused with the instruction to give one end. The step is expand_before + expand_after from the defaults: around spends a quarter backward and the rest forward, earlier and later all of it one way. Each side renders with the markers, so its edges are the next handles; a side that ran out says so in one line.

The descriptions are the only prompting: what a good cue is (the context a memory was encoded in, the user's own wording, query the surroundings and expand when the target cannot be pinned, a lead is another query), how to read the markers, and that a window is an append read once. The rendered timestamp format is one constant, medium date and long time in the zone the timestamp carries.

Errors (no header, unknown tenant, no event backend, a bad id or timestamp, an anchor the tenant does not hold) are tool errors with a plain message.

Verified

28 tests through fastmcp's in-memory client with the v1 fake resolver: the tool surface (names, descriptions, schema properties exactly the five and two parameters), query rendering and ordering under AngleEmbedder, every filter reaching the memory, reranking, the three directions with their counts asserted against the defaults, every id form, every error, and one end-to-end streamable-HTTP handshake through the mounted app that reads the header. At this head, on #1663 fb38a72 (2026-10-09), which carries #1628's declared-only stores: ruff, ruff format, ty as CI runs it on packages/server and uv lock --check pass, and docs/openapi.json matches its generator; the episodic memory and v1 unit tests (519) pass against the in-memory store that now enforces the declared schema. Unit and integration suites run again when this PR comes up for review; they last passed on 2026-10-07 on #1663 e0dedbd: 2294 server tests, 136 PostgreSQL store integration tests and 258 client tests. docs/openapi.json is unaffected: the generator builds from the routers only, and regenerating it compares equal.

Choices to review

  • memory_expand's wire parameter is id through a pydantic alias, since ruff forbids a Python parameter named id and the repo forbids per-file ignores.
  • fastmcp's default stateful transport is kept, as in the v2 app; the header is read on the initialize request, which is how a static agent config sends it.
  • The tenant is resolved before arguments are parsed, so a bad argument sent to an unknown tenant reports the tenant.

Stack

Two stacks, one line of branches. Every PR but #1693 targets feat/horizontal-scaling, so a diff shows everything below it on that branch until that merges; #1693 is client-only, branches from main and targets it. Rebuilt on 2026-09-28: #1684 split into time bounds and sources (#1684) and sessions (#1715), the block kind moved to #1687, and each PR restacked in dependency order. Rebased on 2026-10-01 after #1713 merged, dropping the merge commit that carried it; on 2026-10-06 after #1733 was squash-merged into feat/horizontal-scaling; on 2026-10-07 after that branch took main's #1707, which makes episode uids UUIDs; and on 2026-10-09, after #1736 was squash-merged into feat/horizontal-scaling, onto #1663's head c0a2bbd, and the same day onto fb38a72, after #1628 moved beneath #1663 and #1813 was squash-merged into feat/horizontal-scaling. #1715 and everything above it stay deferred with the coding-agent features.

Event memory, on #1663:

# PR Base Content
1/7 #1684 #1663 time bounds, sources and expansion (port of #1597 without session)
2/7 #1686 #1684 the write transaction, and the rename to the event memory store (port of #1659)
3/7 #1685 #1686 eviction at ingest (port of #1617)
4/7 #1687 #1685 the block kind column, parameter and record key; context parts; kind-keyed tables (port of #1611)
5/7 #1692 #1687 the capture block kinds: tool_call, tool_result, injected, thinking
6/7 #1715 #1692 sessions: the session column, index and walk, session_ids (deferred)
7/7 #1688 #1715 session blocks and id markers in rendering (port of #1632)

Coding agents, slices of design/coding_agent_integration.md (#1579), on the event memory stack; 3/3 shares no code with the server, so its branch is on main, but it configures the endpoint 2/3 serves and writes the kinds #1692 registers, so it merges after both:

# PR Base Content
1/3 #1690 #1688 the v1 event-memory API: tenants, query, expand, events
2/3 #1691 (this PR) #1690 memory_query and memory_expand served at /v1/mcp
3/3 #1693 main Claude Code and Codex in the client: the MCP entry, the Stop hook, and capture

🤖 Generated with Claude Code

https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE

@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from 1e5a19b to f518a3f Compare September 17, 2026 21:49
@edwinyyyu edwinyyyu changed the title [coding agents 2/4] Serve memory_query and memory_expand over MCP at /v1/mcp [coding agents 2/5] Serve memory_query and memory_expand over MCP at /v1/mcp Sep 17, 2026

@edwinyyyu edwinyyyu left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

One dependency note (line comment).

from datetime import UTC, datetime, timedelta
from uuid import uuid4

import httpx2

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

httpx2 is declared by no pyproject in the workspace; it reaches this test only as fastmcp-slim's dependency, so a fastmcp-slim bump that drops it breaks the test rather than an import in src.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Declared in 1220ad4: httpx2>=2.0.0 in the root dev group, the floor being the first release with the ASGI transport, the client, Timeout and Auth the test names (checked in the 2.0.0 wheel); locked with uv 0.12.15, the release CI installs, so the lock gains one line.

@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from f518a3f to fd31354 Compare September 17, 2026 22:49
@edwinyyyu edwinyyyu changed the title [coding agents 2/5] Serve memory_query and memory_expand over MCP at /v1/mcp [coding agents 2/3] Serve memory_query and memory_expand over MCP at /v1/mcp Sep 17, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch 2 times, most recently from 1220ad4 to b84c1cd Compare September 18, 2026 00:09
@edwinyyyu
edwinyyyu marked this pull request as draft September 18, 2026 16:45
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from b84c1cd to 092da5c Compare September 18, 2026 19:50
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from 092da5c to c9710af Compare September 28, 2026 18:02
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from c9710af to 1974ff0 Compare October 2, 2026 00:18
@edwinyyyu edwinyyyu assigned edwinyyyu and unassigned edwinyyyu Oct 2, 2026
@edwinyyyu edwinyyyu added the poc Proof-of-concept implementation for a solution, feature, idea, etc. label Oct 2, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch 2 times, most recently from 6f6fb31 to a890a13 Compare October 7, 2026 20:11
@edwinyyyu
edwinyyyu changed the base branch from main to feat/horizontal-scaling October 9, 2026 00:51
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch 3 times, most recently from 4b90bde to 17192c6 Compare October 9, 2026 20:14
edwinyyyu and others added 29 commits October 9, 2026 17:30
…hine#1597, without session)

The part of MemMachine#1597 that is not session: events and segments carry a
nullable source id; query and expand take since/until and source_ids;
expand walks outward from a seed, which is an address located whatever
the filters say, and returns a Neighborhood that excludes it; query
answers QueryHits (score, seed, neighborhood); rendering is
render_segments with a DateTimeFormat; the segment store answers
get_segments and get_segment_neighborhoods in place of
get_segment_contexts; the v2 adapter lifts timestamp/created_at bounds
and producer_id conjuncts out of the filter into the typed parameters;
reserved property keys live in their own module and a caller cannot
write one; the timestamp column holds a UTC instant.

On this base:
- A neighborhood is MemMachine#1713's walk: the partition's order next to the
  seed, filters applied inside its window of 1,000 segments per side.
  The ordering index keeps main's shape.
- The vector stage gets MemMachine#1684's predicates on the reserved keys and,
  joined with AND, the conjuncts of the property filter the vector store
  declares, as MemMachine#1702's routing has it; the segment store still gets the
  whole filter. A record's declared properties come from its
  derivative's segment.
- Episode uids are UUIDs since MemMachine#1707: the search path reads each hit's
  `_episode_uid` back as a UUID, and the tests key their episodes by
  `_uid(name)` as main's do.
- Session and block kind are split out: the block kind column,
  parameter and record key go to the kinds PR, which is their first
  consumer, and session goes to its own PR.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The segment store gains a write transaction, `write()`, whose block the
caller may fill with work of its own: the event memory upserts its vector
records inside it, so segments commit only once the vector store has
acknowledged their records, and a failed upsert rolls them back. A forget
can then only see links whose records exist, and its delete is issued
after the upsert's acknowledgment, which closes the orphan race between
encode and forget rather than repairing after it.

An event is held at most once: a new event table, keyed by incarnation
and event uuid, is inserted first with ON CONFLICT DO NOTHING RETURNING,
and a batch naming a held event is rejected whole with
`SegmentStoreEventAlreadyStoredError` before anything is stored. Segments
carry a cascading foreign key to the event row; `delete_events` replaces
forget's by-event path, `delete_segments` stays for eviction and leaves
the event held, and `get_derivative_uuids_by_event_uuids` replaces the
two lookups whose only caller was forget. The purge reclaims event rows
after the segments, on the same budget.

The residue a crash can leave, a record acknowledged by the vector store
whose commit never happened, is repaired on retrieval: a hit without a
link is re-checked under `write(exclusive=True)`, which waits for every
write in flight, and a record still without a link is deleted after the
fence is released. An upsert that fails deletes the same ids before the
error propagates, since it may have been applied first.

On this base the change meets MemMachine#1661's final store: the writer's inserts
are the store's own, moved; the leaked-link reclaim MemMachine#1661 dropped stays
dropped; and the event rows purge the way the segments do, after them,
continuing from a cursor of their own on the queue entry,
events_purged_through, since a call that ends exactly on the last
segment leaves purged_through naming a segment.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Deleting the segment partition first waits for every write in flight,
whose records land before it commits, and blocks new ones, so the
vector partition's deletion that follows removes every record that
could ever have landed in it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Mechanical: the store now holds events, segments and links, so its name
follows its dependent. `SegmentStore` and every derived name become
`EventMemoryStore` and theirs, the module and test directories move with
them, the tables and indexes take the `event_memory_store_` prefix, and
the configuration key `segment_store` becomes `event_memory_store` in the
API spec, the client, the docs, the sample configs, the Helm chart and
the compose files. No behavior changes; tables are recreated. Applied by
the same substitutions to this tree rather than rebased.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Ported from the agentic_expansion branch, cosine only. EventMemory
takes an optional EvictionOptions: a cosine similarity threshold at or
above which eviction is considered, the number of stored derivatives
at or above it fetched per new derivative, and how many of a new
derivative and those stored ones to keep when there are more. None
keeps every derivative and issues no query.

At encode, after embedding, `_decide_eviction` settles the whole
decision before anything is written: the batch's predecessors of each
derivative are computed from the batch's own embeddings, earlier
indices only, so a batch evicts what serial ingestion would; one
batched neighbor query fetches the stored derivatives at or above the
threshold; the stored ones' timestamps come from their segments
through the store's lookup, since the vector store answers uuids and
scores only, and a neighbor whose segment is gone is not a member;
then, per derivative, the members over target_size are trimmed from
the temporal middle, the earliest target_size // 2 and the latest
remainder kept. The batch is sorted by timestamp first, so the
predecessor rule matches serial order.

The encode's write() block then carries both link writes: it adds the
surviving links and unlinks the displaced derivatives through
delete_derivatives, which the writer interface, the SQLAlchemy writer
and the fake writer gain, and upserts the surviving records inside the
same transaction, with the compensating delete an upsert failure
already had. The displaced records leave the vector store only after
that block commits: a delete that fails then leaves records no link
names, which read repair reclaims, rather than links naming records
that are gone. Skipped batch derivatives are never written, and a
displaced derivative's segment stays stored.

Tests: the branch's eviction tests on the new shapes, and
delete_derivatives through write() on both dialects, including an
unlink rolled back with its block.

On this base the surviving records carry their segment's declared
properties, as every record does since the routing change, and the
eviction's neighbor query and deletes go to the vector store partition.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Moved here from MemMachine#1684's port, where every block was text and the column
would have held one value with nothing dispatching on it; this PR's
kind-keyed segmenter and deriver tables are its first consumer. The
segment row carries its block's kind in a column of its own, filled by
the store from the block, since the codec's bytes are opaque to SQL;
query, expand, get_segments and get_segment_neighborhoods take
block_kinds; and the vector record carries the kind under its reserved
key, which the vector stage filters with an In predicate.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
…s (speedkick)

The second half of the event memory handoff, on top of the
source/expansion/eviction change: the content model and the processing
steps become registered families keyed by kind.

Blocks. `Block` is an ABC with `kind` as the discriminator every
registered family uses and `render(options) -> str | None`; the union
stays closed over `TextBlock`; the segment row's `block_kind` reads
`block.kind`.

Context. `Context = Mapping[str, ContextPart]`, keyed by part kind,
never None; `Author` is the first registered kind, an unregistered
kind decodes to `UnknownPart` and round-trips unchanged; `with_part` composes one. `ProducerContext`, `NullContext` and the
discriminated union go; the codec encodes a context as `{kind:
fields}`. The server puts the producer id in an `Author` part beside
`source_id`, so rendering keeps its `producer: text` shape.

Segmenter and deriver. `Segmenter` and `Deriver` are tables from block
kind to handler, built from handlers in order with a later handler
replacing an earlier one for its kind. `BlockSegmenter[B]` and `BlockDeriver[B]` are the one-kind
handler contracts, typed by the block class: `split(event, block) ->
list[Piece]` and `derive(segment, block) -> list[str]`. The table
builds every envelope; a handler decides pieces or texts and nothing
else. A kind with no handler passes through as one segment and derives
nothing; nothing raises on a kind. `TextSegmenter`, `WholeTextDeriver`
and `SentenceTextDeriver` keep their names as `text` handlers;
`PassthroughSegmenter` and its configuration name go: one segment per
block is the identity segmentation, and a default gets no name, so an
omitted `segmenter` means it.

Composition. A derivative is text (`Derivative.text`, the text to
embed, plus the segment's `block_kind` for the record), so every
derivative gets the same context processing. `format_header` is the
one composition point for the embedded text and the rendered header:
the timestamp, then the context parts `parts` names in that order
(default `("author",)`), then the content. Parts carry no order; a part
not listed contributes nothing; a new kind is placed by listing it.
`parts` is a parameter of the composers, the handler and `render`,
not a field of `DatetimeFormat`, which writes a timestamp and nothing
else. A handler owns the composition of what it embeds, its
`datetime_format` and `parts`, so a kind or a handler can embed under
its own; the Deriver table composes each derivative's text with them;
rendering for display takes the caller's per call.

Tests: table dispatch, override order, envelope copying; text handlers returning bare
content; composition order, unlisted parts, kind-name validation, and
the anchor over a derivative; context and block round-trips including
an unregistered part kind.

On this base events carry no session, so no segment, derivative or
record copies one; the test harness builds records on the vector store
partition; main's splitting-segmenter test compares the default
segmenter with a text handler; and the event sample configuration main
added documents the text handler as the other sample configurations do.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
RegisteredBlock as the field type left three type errors where a
handler's or the codec's Block was assigned to the field. The fields
are typed Block, decoded through the registered union by a
before-validator and serialized as the concrete kind; the handler
tests keep a typed block in hand instead of narrowing the field.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A context was a mapping the caller keyed by hand, and the rule that
the key equals the part's kind lived in prose. Context is a class now:
built from parts, read by kind, so there is no key to get wrong; two
parts of one kind are rejected at construction. It owns its wire form
through a Pydantic core schema, so a model field of the type accepts
an instance or `{kind: fields}` and serializes to `{kind: fields}`,
and the per-model validators, `with_part`, `encode_context` and
`decode_context` go.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The kind-keyed tables removed its config, its service-locator branch
and its test, and the module survived with no consumer. The long-term
memory's over-fetch comment named it too; it now says why a segmenter
can carry one episode several times. The expand_context comment main
added names it as well; it now says "no segmenter handler".

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A coding agent's transcript carries more than messages: the tool calls
the agent makes, what those tools return, and text that entered the
conversation without a user typing it. Each is a block kind of its own,
registered so the v1 events route accepts it and the timeline holds it.

`tool_call` carries the tool's name and its input. The input is a JSON
object rather than a mapping of property values: a tool's arguments
nest, and a block is content the event carries, not something the event
is filtered by. `tool_result` carries the tool's name, its output, and
whether the tool failed instead of returning; a result says nothing
about failure unless it says so. `injected` carries the text and what
put it there: a hook, a skill, a compaction, a reminder, a command, or
something else.

Each kind renders on one line, through the same `render` a timeline
reader already gets for a message: a call's input as compact JSON, a
failed result behind an `[error]` marker, injected text behind its
source. A window mixing messages and tool events reads as one timeline.

None of the three declares a segmenter or a deriver, so the tables'
fallbacks decide: one segment per block, however long, and no
derivative and no vector record. A tool event is on the timeline
and off the search surface, reached by expanding from a
message, which is the capture policy the coding-agent integration
settled on.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The kinds' round trip through the codec, their encoded form, and what
each renders: a call as its tool and its input on one line, a result as
its tool and its output with an `[error]` marker when it failed, and
injected text as its source and its text. An empty or overlong tool
name, an unlisted injection source and an unregistered kind are
rejected.

Segmentation and derivation are asserted where the tables decide them:
a long tool result is one segment while the text handler splits the
message beside it, and the text deriver derives nothing from any of the
three. End to end, an event carrying a message, a call and a result is
three segments and one vector record, the query reaches the message,
and expanding from it returns the two tool events.

The route's tests go with the v1 route, which sits above this change.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Capture holds everything the session log holds, never what a model
selected, so the agent's reasoning is a kind of its own rather than
something dropped on the way in: `{"kind": "thinking", "text": ...}`,
holding the reasoning as the transcript recorded it.

It is treated as the tool events are. It declares no segmenter and no
deriver, so reasoning is one segment however long and yields no
derivative and no vector record: it is on the timeline and off the
search surface, reached by expanding from a message. It renders as
`thinking: <text>`, one line of the same timeline a message and a tool
event render into.

The v1 route, which sits above this change, accepts it through the
registered union and names it with the other kinds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
An event belongs to a session, a conversation's own timeline, and a
walk never leaves it. Event, Segment and Derivative carry a required,
bounded session_id; the segment row holds it in a NOT NULL column, the
ordering index's second, after the incarnation
(event_memory_store_sg__in_se_ts_ev_ix_of), and
a neighborhood's walk pins the seed's session, on SQLite as a bound
parameter per seed and on PostgreSQL through the lateral join, inside
the same context window. query, expand and get_segments take
session_ids, an empty list keeping nothing and None every session; the
vector record carries the session under its reserved key, so the vector
stage evaluates session_ids too; expand raises LookupError for a seed
outside the named sessions, which it checks through get_segments. The v2
adapter has no conversation id to give, so every event it ingests is in
one reserved session, memmachine_default.

On this base the session goes onto the event memory store's names, the
write transaction's segment insert and the kind-keyed segmenter and
deriver tables, which copy it into every segment and derivative they
build.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`render` takes any segments, drops one given twice, and lays them out as
a block per session, the sessions in the order of their latest
timestamps with a blank line between, each block in the store's order,
one line per run of adjacent pieces of one event. Nothing in the text
says whether two lines are adjacent in the store, since a filtered walk
can omit a neighbor and no segment carries a position.

`ids` marks what a reader can name back: `"session"` heads each block
with `[session:"<id>"]`, the id JSON-quoted since a session id is any
string; `"segment"` starts each line with `[segment:<hex>]`, or
`[segments:<first>..<last>]` when the line holds more than one, so the
marker says which id opens the event and which closes it without a word
of prompting, and a one-segment event carries one id. Uuids are 32 hex
digits. With no ids and one session the text is what `render` produced
before, which `rerank` relies on.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A caller that speaks events rather than episodes needs the EventMemory
LongTermMemory built and the reranker it scores with; both were private,
so the only way in was the episode translation the event API has no use
for. The two accessors state what holds: the memory is None on the
declarative backend and once drop_session_partition has deleted the
collection and the partition, and the reranker is None when the tenant
has none configured.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The v1 API addresses a tenant by an opaque name and never sees a
lifecycle, so the routes ask one seam, TenantEventMemories, for the
tenant's EventMemory, its reranker and the defaults a request may omit.
What a name resolves to is the deployment's concern: replacing the
implementation moves every tenant to another substrate without touching
a route, which is what the tenant registry will do.

The default implementation gives each tenant one episodic-memory session
keyed `v1/tenants/<name>`, configured with the event backend and without
short-term memory. A v2 session key is an organization id and a project
id joined by a slash, and neither id may hold one, so a v2 key carries
exactly one slash and a v1 key carries at least two: no tenant name,
whatever it contains, addresses a v2 project, and the partition keys the
segment store derives differ with them.

Creation is idempotent and materializes the partition and the collection,
so a deployment that cannot serve the tenant says so at creation rather
than at the tenant's first search. A session whose configuration names no
event memory this deployment can build is reported as the component not
being enabled, which is what it is from a client's side.

Defaults live in one place, EpisodicMemoryDefaults, because the episodic
memory settings carry no limit, expansion or rerank width to take them
from.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Seven routes under /v1, beside the v2 router and changing nothing about
it: a tenant's creation, reading and deletion, search and expansion over
the tenant's episodic memory, and the ingestion and removal of events.
Search runs the vector query, reranks with the tenant's reranker when the
request asks and it has one, cuts to the limit, and renders each hit's
window with the session and segment markers a further request continues
from; expansion walks out from a segment uuid and renders each side the
same way.

Failures answer `{"error": {"code", "message"}}` under the design's
closed set of codes, mapped by the router's own route class so no other
router's errors change shape. A failure with no code of its own is
`internal`, whose traceback is logged and never answered with.

Three departures this deployment forces, each stated in the route
descriptions so a client written now keeps working when they lift: a
segment carries no ingestion position, because there is no event store;
ingest is synchronous, so `wait` is accepted and ignored and a repeated
event id stores a second copy rather than being rejected; and `filter` is
an expression in the server's filter grammar, which is the only filter
this tree parses. Naming a reranker is rejected rather than ignored: the
tenant's own reranker is the one the server scores with.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The routes are driven through the test client against a resolver whose
event memory is the real one over the in-memory segment store and vector
collection, so a test exercises the rendering, the filters and the
reranking rather than a mock's return value. Ordering is asserted only
under the angle embedder, since the plain fake embedder ties every score.

Covered: each route, the error mapping including the traceback that never
leaves the process, the id markers and the seed's place in a hit's
window, reranking on and off with its score floor, and the defaults
filling in the counts a request omits.

The default resolver is tested against a mocked MemMachine: the session
it configures, creation that finds the tenant already there, deletion of
both the memory and the row, and every resolution failure. Its namespace
is checked by the shape of the key it builds, which no v2 session key can
have.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The document is what the clients are generated from and what the API
reference renders, so the v1 routes belong in it with the descriptions
that state where this deployment departs from the API a client will
eventually meet. The generator loads the v1 router beside the v2 one,
names its three tags, and strips FastAPI's generated suffix from v1
operation ids as it already does for v2, so a generated client calls
`search_episodic_memory` rather than the path spelled out. No v2 path,
schema or operation id changes.

On this base the capture kinds are registered below the v1 route, so
the events route accepts them from the start and the document carries
their schemas in EventSpec's block union; regenerated with the tool.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`EventMemory.query` is the verb the memory uses, so the API uses it too:
the route is `POST /v1/tenants/{tenant}/episodic-memory/query`, its
function and operation id are `query_episodic_memory`, and the bodies
are `QueryRequest`, `QueryHitBody` and `QueryResponse`. The hit body
takes the `Body` suffix that `SegmentBody` and `TenantBody` carry,
because the memory's own `QueryHit` keeps its name and the two are used
side by side in the router.

The tenant's default is `query_limit`, and every docstring, description,
tag description, test name and comment that called the operation a
search now calls it a query. "Vector search" stays where it names the
vector store's technique rather than the operation: the stage that runs
before the reranker, and `vector_search_limit`, the parameter the memory
takes.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The event memory store holds an event once, so `encode_events` raises
`EventMemoryStoreEventAlreadyStoredError` when a batch names an event
the tenant already holds, and the batch stores nothing. The v1 error
contract gains the code `event_exists` for that, and the route class
answers the error with 409 and a message naming the uuids, so a client
retrying a batch learns which ids it may not reuse and that the rest
of the batch was not stored either.

The add-events description says so: a repeated id is a rejection of
the whole batch, not a second copy. The route declares the 409 and the
generated OpenAPI document carries both, and a router test reingests a
held id beside a fresh event and finds neither stored.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`EventSpec.blocks` decodes through the registered union, so the route
accepts a thinking block, a tool call, a tool result and injected text
with no change to its code, and a kind no registration names stays a
validation failure, answered 422 `invalid_request` like any other
malformed body.

What the route says about the field changes: the description names the
five kinds and which of them a query matches, so a client reading the
document knows that a message is the search surface and everything
else is reached by expanding from one. The kinds' schemas are in
docs/openapi.json since the route arrived; only the description
changes here, regenerated with docs/tools/generate_openapi.py.

The route side of the capture kinds, which register below the session
change while the route sits above it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A batch of one message event and four capture events (thinking, a tool
call, a tool result, injected text) is stored, only the message is a
hit, and expanding from it renders the four on their own lines; a
block missing a field its kind requires is 422 `invalid_request`
naming the field.

The route tests of the capture kinds, which register below the session
change while the route sits above it; their codec, segmenter, deriver
and memory tests stay with the kinds.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`memory_query` and `memory_expand` on a FastMCP instance of their own,
over the seam the v1 routes resolve a tenant through. The tenant is the
`X-MemMachine-Tenant` header of the MCP request, read by an ASGI
middleware into a context variable this module owns, so this app's tenant
and the v2 MCP app's ids never reach each other. The `MemMachine` the
tools resolve against is the one the process started, so a tenant reached
through a tool and the same tenant reached through a route are one
memory.

Neither tool takes a count. A query answers with the tenant's
`query_limit` hits, each rendered with the expansion the tenant
configures; a step of `memory_expand` is the tenant's `expand_before`
plus `expand_after`, spent a quarter backward and the rest forward for
`around`, and wholly on one side for `earlier` and `later`. The model
chooses a place and a direction rather than an amount, and the
descriptions carry the cue guidance and the marker grammar, which is the
only prompting the integration does.

Every failure a caller can cause -- no header, an unknown tenant, a
tenant with no event memory, a malformed timestamp or segment id, a range
marker where one end was wanted, an anchor the tenant does not hold -- is
a `ToolError` carrying one plain sentence.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The v2 MCP app keeps `/mcp` and its lifespan untouched; the v1 app is
mounted beside it. A mounted application's lifespan is not run by the
application that mounts it, so the server's lifespan now chains both MCP
apps' lifespans after its own resources, and each app's session manager
is started before the first request reaches it.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The tools are driven through fastmcp's in-memory client against the
FastMCP instance, with the tenant in the context variable and the v1
tests' fake resolver in place of a `MemMachine`: what a query renders,
the order of its hits, the filters reaching the memory, the reranker
fetching candidates and the hits then cut to the limit, the three
directions of an expansion with their counts asserted against the
tenant's defaults, every accepted form of a segment id, the refused run
of segments, and every failure a caller can cause.

One test runs the whole handshake against the mounted application over an
ASGI transport, which Starlette's `TestClient` cannot do for streamable
HTTP: the tenant is on the request and in no context variable of the
test, so a tool that answers at all read the header.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The HTTP test drives the mounted app through httpx2's ASGI transport,
and httpx2 reached it only as fastmcp-slim's dependency. It is now in
the dev group, with a floor at the first release that has the
transport, the client, Timeout and Auth the test names.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
@edwinyyyu
edwinyyyu force-pushed the feat/event-memory-v1-mcp-main branch from 27aacfa to 49d9dc0 Compare October 10, 2026 00:33

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

poc Proof-of-concept implementation for a solution, feature, idea, etc.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant