Repository navigation
Conversation
5e3d1b8 to
27ba848
Compare
edwinyyyu
left a comment
There was a problem hiding this comment.
- The 409
event_existspromise (line comment) has no path to be kept by the PRs named. events/deletereturnsNoneat 200, which FastAPI serializes as the JSON literalnull, and the OpenAPI 200 response advertisesapplication/json; the description says "no body". 204 would make that true.
| An event's `id` is the client's dedupe key: a client that retries a | ||
| batch supplies the ids it used, and a batch without them is stored | ||
| again. A tenant holds an event once: a batch naming an id the tenant | ||
| already holds is rejected whole with `event_exists`, and nothing in it |
There was a problem hiding this comment.
ErrorCode (errors.py) is the closed set tenant_not_found | component_not_enabled | invalid_request | internal, and #1686's EventMemoryStoreEventAlreadyStoredError is a plain Exception, so once #1686 is absorbed a repeated id is answered 500 internal by the except Exception arm, not 409 event_exists. Either the code and the mapping land here (or with #1686's absorption), or this text and the PR body should say 500.
There was a problem hiding this comment.
This describes the branch before its last commit, "Answer a batch naming a held event with 409 event_exists" (7c43319 after today's rebase): it adds event_exists to ErrorCode, maps EventMemoryStoreEventAlreadyStoredError onto 409 in the envelope route, and tests the route. The body now names the commit.
7c43319 to
18b71fb
Compare
18b71fb to
2dad8e8
Compare
2dad8e8 to
3fd3250
Compare
3fd3250 to
cdfa711
Compare
2cd4fda to
effd7ac
Compare
fb9a8a0 to
c09fb68
Compare
The test of MemMachine#1813, merged into this PR's base, in partition terms: the server defaults new collections to strict mode, the store's collection reads back with it off, and a filter on a partition's unindexed property is served. Co-Authored-By: Claude Opus 5.5 <[email protected]>
A partition stored every property of a record and filtered on any key, which made a caller's arbitrary keys part of the store's schema: the SQLite stores kept them in a JSON column and filtered with json_extract, Milvus the keys its schema did not declare in a JSON field, and a filter on a key the store never indexed scanned. Since EventMemory routes a filter on an undeclared key to the segment store, the vector store need not hold undeclared keys at all. A partition now stores the properties its store declares and no others. `upsert` raises UndeclaredPropertyKeyError before anything is sent for a record naming an undeclared key, and PropertyTypeMismatchError for a value of another type than its key declares; `query` raises UndeclaredPropertyKeyError for a filter naming an undeclared key and UnsupportedFilterError for a node outside the partition's `supported_filter_nodes`. Both SQLite stores keep one typed, indexed, nullable column per declared key on the records table (sql_columns.py); sqlite-vec 0.1.9 rejects NULL in a vec0 metadata column and a declared key is optional per record, so that store keeps the columns on the records table and hands the KNN a `rowid IN (SELECT ...)` allowlist, evaluating the filter during the search instead of after it. Qdrant drops the JSON copy and keeps a payload field per declared key; Milvus drops its JSON field and keeps its typed field per declared key. Datetimes are stored as microseconds since the epoch where a backend has no datetime type. Since every key a filter may name is now indexed, the Qdrant store creates its collection in strict mode (`unindexed_filtering_retrieve` and `_update` false, Qdrant Cloud's default): a filter on an unindexed key is refused by the server instead of scanned for. A leaf whose value is of another type than its key declares matches nothing, as on the SQL stores; the Qdrant compiler answers it with a filter no point satisfies, since the server would refuse the condition for the field's index, and Milvus no longer compares an int with a float key. Local mode does not record the setting, so a unit test checks the request and integration tests the server's answer. `declared_schema_contract.py` states the contract every backend's test module runs: which records a filtered search admits, over fixtures small enough that every backend searches them exactly, checked after each upsert so an approximate index fails on recall, by name, and not on the filter. On the registry-backed stores, the checks are the base handle's: `upsert` runs require_declared_properties in place of the type check it ran, and `query` runs require_supported_filter against the subclass's `supported_filter_nodes`, which each subclass now implements. A datetime column on the SQLite stores has a `tz_<key>` column beside it holding the UTC offset in seconds, written with the value and read by no filter, so a stored datetime is the value written, as on Milvus, Qdrant and the segment store; the microseconds column alone would keep only the instant. The Qdrant and Milvus design documents describe the declared-only properties and Qdrant's strict mode. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
A declared datetime on the SQLite stores had a `tz_<key>` column beside its microseconds column, holding the UTC offset in seconds that no filter reads. The vector stores keep a datetime property's instant only, as MemMachine#1736 now does for Milvus and MemMachine#1788 for Qdrant, and the segment store keeps the offset, so the column, `offset_column_name`, and the value written to it go: each declared property is one column, and a datetime stays microseconds since the epoch. The tests read a stored datetime back as its instant in UTC, and the roundtrip test checks that the stored instant equals the written one. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Every embedder MemMachine ships produces vectors meant to be compared by cosine -- OpenAI hard-coded it, Bedrock defaulted to it, SentenceTransformer only reported what the model declared -- while every layer that touched a score paid for the other three metrics in direction flags, threshold directions, and per-backend tables mapping the enum onto native metric names. `SimilarityMetric` is gone; scores are cosine similarities in [-1, 1] and the names say so: `QueryMatch.score` and `SearchMatch.score` become `cosine_similarity`, and `query(score_threshold=)` becomes `query(min_cosine_similarity=)`, which no longer needs a direction to be meaningful and is refused when not finite, as the threshold was. A store and its partitions no longer have a `similarity_metric`, the schema a partition is registered under no longer records one, and a search engine factory takes the dimensions alone. The Bedrock embedder's `similarity_metric` config key and semantic memory's `vector_similarity_metric` go with it, and the install and configuration docs drop them. The vector graph stores carried a metric per stored embedding, as a companion property beside every vector; that is gone and `Node.embeddings` holds plain vectors. NebulaGraph's `cosine()` cannot take `APPROXIMATE` and its vector indexes offer only L2 and IP. Cosine similarity between unit vectors is their inner product, so embeddings are normalized on the way in and compared with `inner_product()` against an IP index. The cosine half of MemMachine#1598 (`abf92a3a4` on speedkick), re-derived on the one-collection store: the rest of MemMachine#1598 (queries answer record UUIDs and scores, no `get`, semantic memory's `vector_uuid`) and all of MemMachine#1603 are in registry-backed base and its partition handle lose the metric, the stores' own threshold checks become `require_valid_min_cosine_similarity`, and the design documents describe cosine scoring and a schema without a metric. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
…hine#1597, without session) The part of MemMachine#1597 that is not session: events and segments carry a nullable source id; query and expand take since/until and source_ids; expand walks outward from a seed, which is an address located whatever the filters say, and returns a Neighborhood that excludes it; query answers QueryHits (score, seed, neighborhood); rendering is render_segments with a DateTimeFormat; the segment store answers get_segments and get_segment_neighborhoods in place of get_segment_contexts; the v2 adapter lifts timestamp/created_at bounds and producer_id conjuncts out of the filter into the typed parameters; reserved property keys live in their own module and a caller cannot write one; the timestamp column holds a UTC instant. On this base: - A neighborhood is MemMachine#1713's walk: the partition's order next to the seed, filters applied inside its window of 1,000 segments per side. The ordering index keeps main's shape. - The vector stage gets MemMachine#1684's predicates on the reserved keys and, joined with AND, the conjuncts of the property filter the vector store declares, as MemMachine#1702's routing has it; the segment store still gets the whole filter. A record's declared properties come from its derivative's segment. - Episode uids are UUIDs since MemMachine#1707: the search path reads each hit's `_episode_uid` back as a UUID, and the tests key their episodes by `_uid(name)` as main's do. - Session and block kind are split out: the block kind column, parameter and record key go to the kinds PR, which is their first consumer, and session goes to its own PR. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The segment store gains a write transaction, `write()`, whose block the caller may fill with work of its own: the event memory upserts its vector records inside it, so segments commit only once the vector store has acknowledged their records, and a failed upsert rolls them back. A forget can then only see links whose records exist, and its delete is issued after the upsert's acknowledgment, which closes the orphan race between encode and forget rather than repairing after it. An event is held at most once: a new event table, keyed by incarnation and event uuid, is inserted first with ON CONFLICT DO NOTHING RETURNING, and a batch naming a held event is rejected whole with `SegmentStoreEventAlreadyStoredError` before anything is stored. Segments carry a cascading foreign key to the event row; `delete_events` replaces forget's by-event path, `delete_segments` stays for eviction and leaves the event held, and `get_derivative_uuids_by_event_uuids` replaces the two lookups whose only caller was forget. The purge reclaims event rows after the segments, on the same budget. The residue a crash can leave, a record acknowledged by the vector store whose commit never happened, is repaired on retrieval: a hit without a link is re-checked under `write(exclusive=True)`, which waits for every write in flight, and a record still without a link is deleted after the fence is released. An upsert that fails deletes the same ids before the error propagates, since it may have been applied first. On this base the change meets MemMachine#1661's final store: the writer's inserts are the store's own, moved; the leaked-link reclaim MemMachine#1661 dropped stays dropped; and the event rows purge the way the segments do, after them, continuing from a cursor of their own on the queue entry, events_purged_through, since a call that ends exactly on the last segment leaves purged_through naming a segment. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Deleting the segment partition first waits for every write in flight, whose records land before it commits, and blocks new ones, so the vector partition's deletion that follows removes every record that could ever have landed in it. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Mechanical: the store now holds events, segments and links, so its name follows its dependent. `SegmentStore` and every derived name become `EventMemoryStore` and theirs, the module and test directories move with them, the tables and indexes take the `event_memory_store_` prefix, and the configuration key `segment_store` becomes `event_memory_store` in the API spec, the client, the docs, the sample configs, the Helm chart and the compose files. No behavior changes; tables are recreated. Applied by the same substitutions to this tree rather than rebased. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Ported from the agentic_expansion branch, cosine only. EventMemory takes an optional EvictionOptions: a cosine similarity threshold at or above which eviction is considered, the number of stored derivatives at or above it fetched per new derivative, and how many of a new derivative and those stored ones to keep when there are more. None keeps every derivative and issues no query. At encode, after embedding, `_decide_eviction` settles the whole decision before anything is written: the batch's predecessors of each derivative are computed from the batch's own embeddings, earlier indices only, so a batch evicts what serial ingestion would; one batched neighbor query fetches the stored derivatives at or above the threshold; the stored ones' timestamps come from their segments through the store's lookup, since the vector store answers uuids and scores only, and a neighbor whose segment is gone is not a member; then, per derivative, the members over target_size are trimmed from the temporal middle, the earliest target_size // 2 and the latest remainder kept. The batch is sorted by timestamp first, so the predecessor rule matches serial order. The encode's write() block then carries both link writes: it adds the surviving links and unlinks the displaced derivatives through delete_derivatives, which the writer interface, the SQLAlchemy writer and the fake writer gain, and upserts the surviving records inside the same transaction, with the compensating delete an upsert failure already had. The displaced records leave the vector store only after that block commits: a delete that fails then leaves records no link names, which read repair reclaims, rather than links naming records that are gone. Skipped batch derivatives are never written, and a displaced derivative's segment stays stored. Tests: the branch's eviction tests on the new shapes, and delete_derivatives through write() on both dialects, including an unlink rolled back with its block. On this base the surviving records carry their segment's declared properties, as every record does since the routing change, and the eviction's neighbor query and deletes go to the vector store partition. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Moved here from MemMachine#1684's port, where every block was text and the column would have held one value with nothing dispatching on it; this PR's kind-keyed segmenter and deriver tables are its first consumer. The segment row carries its block's kind in a column of its own, filled by the store from the block, since the codec's bytes are opaque to SQL; query, expand, get_segments and get_segment_neighborhoods take block_kinds; and the vector record carries the kind under its reserved key, which the vector stage filters with an In predicate. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
…s (speedkick)
The second half of the event memory handoff, on top of the
source/expansion/eviction change: the content model and the processing
steps become registered families keyed by kind.
Blocks. `Block` is an ABC with `kind` as the discriminator every
registered family uses and `render(options) -> str | None`; the union
stays closed over `TextBlock`; the segment row's `block_kind` reads
`block.kind`.
Context. `Context = Mapping[str, ContextPart]`, keyed by part kind,
never None; `Author` is the first registered kind, an unregistered
kind decodes to `UnknownPart` and round-trips unchanged; `with_part` composes one. `ProducerContext`, `NullContext` and the
discriminated union go; the codec encodes a context as `{kind:
fields}`. The server puts the producer id in an `Author` part beside
`source_id`, so rendering keeps its `producer: text` shape.
Segmenter and deriver. `Segmenter` and `Deriver` are tables from block
kind to handler, built from handlers in order with a later handler
replacing an earlier one for its kind. `BlockSegmenter[B]` and `BlockDeriver[B]` are the one-kind
handler contracts, typed by the block class: `split(event, block) ->
list[Piece]` and `derive(segment, block) -> list[str]`. The table
builds every envelope; a handler decides pieces or texts and nothing
else. A kind with no handler passes through as one segment and derives
nothing; nothing raises on a kind. `TextSegmenter`, `WholeTextDeriver`
and `SentenceTextDeriver` keep their names as `text` handlers;
`PassthroughSegmenter` and its configuration name go: one segment per
block is the identity segmentation, and a default gets no name, so an
omitted `segmenter` means it.
Composition. A derivative is text (`Derivative.text`, the text to
embed, plus the segment's `block_kind` for the record), so every
derivative gets the same context processing. `format_header` is the
one composition point for the embedded text and the rendered header:
the timestamp, then the context parts `parts` names in that order
(default `("author",)`), then the content. Parts carry no order; a part
not listed contributes nothing; a new kind is placed by listing it.
`parts` is a parameter of the composers, the handler and `render`,
not a field of `DatetimeFormat`, which writes a timestamp and nothing
else. A handler owns the composition of what it embeds, its
`datetime_format` and `parts`, so a kind or a handler can embed under
its own; the Deriver table composes each derivative's text with them;
rendering for display takes the caller's per call.
Tests: table dispatch, override order, envelope copying; text handlers returning bare
content; composition order, unlisted parts, kind-name validation, and
the anchor over a derivative; context and block round-trips including
an unregistered part kind.
On this base events carry no session, so no segment, derivative or
record copies one; the test harness builds records on the vector store
partition; main's splitting-segmenter test compares the default
segmenter with a text handler; and the event sample configuration main
added documents the text handler as the other sample configurations do.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
RegisteredBlock as the field type left three type errors where a handler's or the codec's Block was assigned to the field. The fields are typed Block, decoded through the registered union by a before-validator and serialized as the concrete kind; the handler tests keep a typed block in hand instead of narrowing the field. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A context was a mapping the caller keyed by hand, and the rule that
the key equals the part's kind lived in prose. Context is a class now:
built from parts, read by kind, so there is no key to get wrong; two
parts of one kind are rejected at construction. It owns its wire form
through a Pydantic core schema, so a model field of the type accepts
an instance or `{kind: fields}` and serializes to `{kind: fields}`,
and the per-model validators, `with_part`, `encode_context` and
`decode_context` go.
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The kind-keyed tables removed its config, its service-locator branch and its test, and the module survived with no consumer. The long-term memory's over-fetch comment named it too; it now says why a segmenter can carry one episode several times. The expand_context comment main added names it as well; it now says "no segmenter handler". Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A coding agent's transcript carries more than messages: the tool calls the agent makes, what those tools return, and text that entered the conversation without a user typing it. Each is a block kind of its own, registered so the v1 events route accepts it and the timeline holds it. `tool_call` carries the tool's name and its input. The input is a JSON object rather than a mapping of property values: a tool's arguments nest, and a block is content the event carries, not something the event is filtered by. `tool_result` carries the tool's name, its output, and whether the tool failed instead of returning; a result says nothing about failure unless it says so. `injected` carries the text and what put it there: a hook, a skill, a compaction, a reminder, a command, or something else. Each kind renders on one line, through the same `render` a timeline reader already gets for a message: a call's input as compact JSON, a failed result behind an `[error]` marker, injected text behind its source. A window mixing messages and tool events reads as one timeline. None of the three declares a segmenter or a deriver, so the tables' fallbacks decide: one segment per block, however long, and no derivative and no vector record. A tool event is on the timeline and off the search surface, reached by expanding from a message, which is the capture policy the coding-agent integration settled on. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The kinds' round trip through the codec, their encoded form, and what each renders: a call as its tool and its input on one line, a result as its tool and its output with an `[error]` marker when it failed, and injected text as its source and its text. An empty or overlong tool name, an unlisted injection source and an unregistered kind are rejected. Segmentation and derivation are asserted where the tables decide them: a long tool result is one segment while the text handler splits the message beside it, and the text deriver derives nothing from any of the three. End to end, an event carrying a message, a call and a result is three segments and one vector record, the query reaches the message, and expanding from it returns the two tool events. The route's tests go with the v1 route, which sits above this change. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Capture holds everything the session log holds, never what a model
selected, so the agent's reasoning is a kind of its own rather than
something dropped on the way in: `{"kind": "thinking", "text": ...}`,
holding the reasoning as the transcript recorded it.
It is treated as the tool events are. It declares no segmenter and no
deriver, so reasoning is one segment however long and yields no
derivative and no vector record: it is on the timeline and off the
search surface, reached by expanding from a message. It renders as
`thinking: <text>`, one line of the same timeline a message and a tool
event render into.
The v1 route, which sits above this change, accepts it through the
registered union and names it with the other kinds.
Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
An event belongs to a session, a conversation's own timeline, and a walk never leaves it. Event, Segment and Derivative carry a required, bounded session_id; the segment row holds it in a NOT NULL column, the ordering index's second, after the incarnation (event_memory_store_sg__in_se_ts_ev_ix_of), and a neighborhood's walk pins the seed's session, on SQLite as a bound parameter per seed and on PostgreSQL through the lateral join, inside the same context window. query, expand and get_segments take session_ids, an empty list keeping nothing and None every session; the vector record carries the session under its reserved key, so the vector stage evaluates session_ids too; expand raises LookupError for a seed outside the named sessions, which it checks through get_segments. The v2 adapter has no conversation id to give, so every event it ingests is in one reserved session, memmachine_default. On this base the session goes onto the event memory store's names, the write transaction's segment insert and the kind-keyed segmenter and deriver tables, which copy it into every segment and derivative they build. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`render` takes any segments, drops one given twice, and lays them out as a block per session, the sessions in the order of their latest timestamps with a blank line between, each block in the store's order, one line per run of adjacent pieces of one event. Nothing in the text says whether two lines are adjacent in the store, since a filtered walk can omit a neighbor and no segment carries a position. `ids` marks what a reader can name back: `"session"` heads each block with `[session:"<id>"]`, the id JSON-quoted since a session id is any string; `"segment"` starts each line with `[segment:<hex>]`, or `[segments:<first>..<last>]` when the line holds more than one, so the marker says which id opens the event and which closes it without a word of prompting, and a one-segment event carries one id. Uuids are 32 hex digits. With no ids and one session the text is what `render` produced before, which `rerank` relies on. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A caller that speaks events rather than episodes needs the EventMemory LongTermMemory built and the reranker it scores with; both were private, so the only way in was the episode translation the event API has no use for. The two accessors state what holds: the memory is None on the declarative backend and once drop_session_partition has deleted the collection and the partition, and the reranker is None when the tenant has none configured. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The v1 API addresses a tenant by an opaque name and never sees a lifecycle, so the routes ask one seam, TenantEventMemories, for the tenant's EventMemory, its reranker and the defaults a request may omit. What a name resolves to is the deployment's concern: replacing the implementation moves every tenant to another substrate without touching a route, which is what the tenant registry will do. The default implementation gives each tenant one episodic-memory session keyed `v1/tenants/<name>`, configured with the event backend and without short-term memory. A v2 session key is an organization id and a project id joined by a slash, and neither id may hold one, so a v2 key carries exactly one slash and a v1 key carries at least two: no tenant name, whatever it contains, addresses a v2 project, and the partition keys the segment store derives differ with them. Creation is idempotent and materializes the partition and the collection, so a deployment that cannot serve the tenant says so at creation rather than at the tenant's first search. A session whose configuration names no event memory this deployment can build is reported as the component not being enabled, which is what it is from a client's side. Defaults live in one place, EpisodicMemoryDefaults, because the episodic memory settings carry no limit, expansion or rerank width to take them from. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
Seven routes under /v1, beside the v2 router and changing nothing about
it: a tenant's creation, reading and deletion, search and expansion over
the tenant's episodic memory, and the ingestion and removal of events.
Search runs the vector query, reranks with the tenant's reranker when the
request asks and it has one, cuts to the limit, and renders each hit's
window with the session and segment markers a further request continues
from; expansion walks out from a segment uuid and renders each side the
same way.
Failures answer `{"error": {"code", "message"}}` under the design's
closed set of codes, mapped by the router's own route class so no other
router's errors change shape. A failure with no code of its own is
`internal`, whose traceback is logged and never answered with.
Three departures this deployment forces, each stated in the route
descriptions so a client written now keeps working when they lift: a
segment carries no ingestion position, because there is no event store;
ingest is synchronous, so `wait` is accepted and ignored and a repeated
event id stores a second copy rather than being rejected; and `filter` is
an expression in the server's filter grammar, which is the only filter
this tree parses. Naming a reranker is rejected rather than ignored: the
tenant's own reranker is the one the server scores with.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The routes are driven through the test client against a resolver whose event memory is the real one over the in-memory segment store and vector collection, so a test exercises the rendering, the filters and the reranking rather than a mock's return value. Ordering is asserted only under the angle embedder, since the plain fake embedder ties every score. Covered: each route, the error mapping including the traceback that never leaves the process, the id markers and the seed's place in a hit's window, reranking on and off with its score floor, and the defaults filling in the counts a request omits. The default resolver is tested against a mocked MemMachine: the session it configures, creation that finds the tenant already there, deletion of both the memory and the row, and every resolution failure. Its namespace is checked by the shape of the key it builds, which no v2 session key can have. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The document is what the clients are generated from and what the API reference renders, so the v1 routes belong in it with the descriptions that state where this deployment departs from the API a client will eventually meet. The generator loads the v1 router beside the v2 one, names its three tags, and strips FastAPI's generated suffix from v1 operation ids as it already does for v2, so a generated client calls `search_episodic_memory` rather than the path spelled out. No v2 path, schema or operation id changes. On this base the capture kinds are registered below the v1 route, so the events route accepts them from the start and the document carries their schemas in EventSpec's block union; regenerated with the tool. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`EventMemory.query` is the verb the memory uses, so the API uses it too:
the route is `POST /v1/tenants/{tenant}/episodic-memory/query`, its
function and operation id are `query_episodic_memory`, and the bodies
are `QueryRequest`, `QueryHitBody` and `QueryResponse`. The hit body
takes the `Body` suffix that `SegmentBody` and `TenantBody` carry,
because the memory's own `QueryHit` keeps its name and the two are used
side by side in the router.
The tenant's default is `query_limit`, and every docstring, description,
tag description, test name and comment that called the operation a
search now calls it a query. "Vector search" stays where it names the
vector store's technique rather than the operation: the stage that runs
before the reranker, and `vector_search_limit`, the parameter the memory
takes.
Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
The event memory store holds an event once, so `encode_events` raises `EventMemoryStoreEventAlreadyStoredError` when a batch names an event the tenant already holds, and the batch stores nothing. The v1 error contract gains the code `event_exists` for that, and the route class answers the error with 409 and a message naming the uuids, so a client retrying a batch learns which ids it may not reuse and that the rest of the batch was not stored either. The add-events description says so: a repeated id is a rejection of the whole batch, not a second copy. The route declares the 409 and the generated OpenAPI document carries both, and a router test reingests a held id beside a fresh event and finds neither stored. Co-Authored-By: Claude Fable 5.1 <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
`EventSpec.blocks` decodes through the registered union, so the route accepts a thinking block, a tool call, a tool result and injected text with no change to its code, and a kind no registration names stays a validation failure, answered 422 `invalid_request` like any other malformed body. What the route says about the field changes: the description names the five kinds and which of them a query matches, so a client reading the document knows that a message is the search surface and everything else is reached by expanding from one. The kinds' schemas are in docs/openapi.json since the route arrived; only the description changes here, regenerated with docs/tools/generate_openapi.py. The route side of the capture kinds, which register below the session change while the route sits above it. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
A batch of one message event and four capture events (thinking, a tool call, a tool result, injected text) is stored, only the message is a hit, and expanding from it renders the four on their own lines; a block missing a field its kind requires is 422 `invalid_request` naming the field. The route tests of the capture kinds, which register below the session change while the route sits above it; their codec, segmenter, deriver and memory tests stay with the kinds. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE
0de1ee6 to
ba8d3bf
Compare
Slice 1 of
design/coding_agent_integration.md(#1579): the tenant-scoped v1 HTTP API over the event memory, with the bodiesdesign/server_redesign.md("Server API") specifies, on today's server. Stacked on #1688, the top of the event memory stack (below). Rebuilt on 2026-09-28 with the stack: the last two commits are the route's side of the capture kinds (#1692), which register below the session change while the route sits above it, so the route names the kinds and tests them here. It goes through neither the v2 router norEpisode.One verb from
EventMemory.queryup: the operation is a query, its routequery, its hits query hits; "search" names only the vector store's technique.Routes (prefix
/v1)PUT /tenants/{tenant}201 on creation, 200 when it existed;GET200 / 404;DELETE204, idempotent. A tenant is an opaque name, at most 255 bytes.POST /tenants/{tenant}/episodic-memory/query:query,limit,min_cosine_similarity,expand_context,rerank(candidates,min_score, or null),since,until,session_ids,source_ids,block_kinds,filter,datetime_format,parts. Hits arescore,seed,segments(uuid,event_uuid,index,offset,timestampwith offset,session_id,source_id,context,block,properties) andtext, the window rendered withrender_segments(ids=("session", "segment")), so every hit carries the[session:"..."]and[segment:<hex>]markers a client continues from.POST .../episodic-memory/expand:anchor,before,after, the filters,datetime_format,parts; answersbeforeandafteras segment lists plusbefore_textandafter_text.POST /tenants/{tenant}/events: a list of events (idoptional, minted if absent;timestampoptional, server time if absent;session_idrequired;source_id;context;blocks;properties) →encode_events, 200{stored: [...]}. Ingest is synchronous, so?wait=is accepted and ignored.POST .../events/delete:{ids}→forget_events.Errors answer
{"error": {"code", "message"}}on the v1 routes only:tenant_not_found404,component_not_enabled404,event_exists409,invalid_request422,internal500 with the traceback logged and never returned. The 409 is the last commit: it adds the code toErrorCodeand mapsEventMemoryStoreEventAlreadyStoredErrorfrom #1686 onto it, with a route test; a review comment on this PR describes the tree before that commit.The seam
TenantEventMemories(create,exists,delete,resolve -> event memory, reranker, defaults) is all a route's logic depends on; the router's default dependency constructs the one implementation,EpisodicSessionTenantEventMemories, which makes a tenant one episodic-memory session under the keyv1/tenants/<name>, created with the event backend and no short-term memory, every resource the deployment's. It cannot collide with the v2 API: a v2 session key is<org>/<project>with ids that cannot contain a slash, so it holds exactly one, and a v1 key holds at least two. Delete drops the session and with it the partition and the collection.LongTermMemorygains two accessors,event_memoryandreranker, with no episode translation on this path.Departures from the redesign, stated in the route descriptions
positionon segments and no watermark route: there is no event store yet. Segments carryevent_uuidwhere the design writesevent_id.event_exists, naming the uuids; ids are the client's dedupe key.session_idis required on an event where the design has it optional:EventMemoryrequires one, and this path supplies no default the way the episode translation does.expandtakes a segment uuid as its anchor, not a segment or event uuid, and answersbefore_textandafter_textrather than onetextper side.filteris the existing string grammar; no JSON-tree parser exists on this tree and none was written.rerank.rerankeris rejected withinvalid_requestrather than ignored: naming a reranker would put the resource manager into a route, and the tenant's reranker is the deployment's.EpisodicMemoryDefaults(query limit 10, expansion 2, rerank candidates 40, expand 5 before and 5 after); no episodic-memory setting existed for them.Choices to review
An anchor the tenant does not hold is 422
invalid_request(the closed set has no segment-level 404);events/deleteanswers 200 with a JSONnullbody, since the route returns nothing and the redesign's head position does not exist here, where the tenant delete answers 204; a delete racing a resolve that still holds the manager's reference surfaces as 500 rather than a 409 outside the closed set.Verified
docs/openapi.jsonregenerated bydocs/tools/generate_openapi.py, which now loads the v1 router as well: five paths and their schemas added, the capture kinds' schemas inEventSpec's block union from the start, nothing of v2 changed; after every commit that touches the routes the document compares equal to the generator's output. The v1 suite holds 63 tests, 42 on the routes through the test client over the realEventMemoryon the in-memory fakes and 21 on the resolver. At this head, on #1663 fb38a72 (2026-10-09), which carries #1628's declared-only stores: ruff, ruff format, ty as CI runs it onpackages/serveranduv lock --checkpass, anddocs/openapi.jsonmatches its generator. Unit and integration suites run again when this PR comes up for review; they last passed on 2026-10-07 on #1663 e0dedbd: 2266 server tests, 136 PostgreSQL store integration tests and 258 client tests.Stack
Two stacks, one line of branches. Every PR but #1693 targets
feat/horizontal-scaling, so a diff shows everything below it on that branch until that merges; #1693 is client-only, branches frommainand targets it. Rebuilt on 2026-09-28: #1684 split into time bounds and sources (#1684) and sessions (#1715), the block kind moved to #1687, and each PR restacked in dependency order. Rebased on 2026-10-01 after #1713 merged, dropping the merge commit that carried it; on 2026-10-06 after #1733 was squash-merged intofeat/horizontal-scaling; on 2026-10-07 after that branch took main's #1707, which makes episode uids UUIDs; and on 2026-10-09, after #1736 was squash-merged intofeat/horizontal-scaling, onto #1663's head c0a2bbd, and the same day onto fb38a72, after #1628 moved beneath #1663 and #1813 was squash-merged intofeat/horizontal-scaling. #1715 and everything above it stay deferred with the coding-agent features.Event memory, on #1663:
tool_call,tool_result,injected,thinkingsession_ids(deferred)Coding agents, slices of
design/coding_agent_integration.md(#1579), on the event memory stack; 3/3 shares no code with the server, so its branch is onmain, but it configures the endpoint 2/3 serves and writes the kinds #1692 registers, so it merges after both:memory_queryandmemory_expandserved at/v1/mcpStophook, and capture🤖 Generated with Claude Code
https://claude.ai/code/session_01YBbQgZiCqeoLu83EkbEFHE