Skip to content

[vector store scale-out 2/6] Answer vector store queries with record UUIDs and scores, and refuse invalid inputs - #1733

Merged
malatewang merged 14 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-query-contract-speedkick
Oct 5, 2026
Merged

malatewang merged 14 commits into
MemMachine:feat/horizontal-scalingfrom
edwinyyyu:feat/vector-store-query-contract-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Summary

A vector store query now answers record UUIDs and scores. Callers resolve a hit through the store that owns the mapping, and nothing reads a record back from the vector store. Every store refuses inputs no collection can hold or search, and the contract states what a query sees of earlier writes.

  • Queries answer UUIDs and scores. A query returned each match's properties, and its vector on request. These were copies of what the callers' own stores hold, and only as fresh as the vector store's reads. Event memory now maps a hit to its segment through the segment store's derivative rows (new get_segment_uuids_by_derivative_uuids). Semantic memory maps a hit to its feature through a new vector_uuid column on the feature row, and its vector records carry no properties.

  • get, close_collection, return_vector and return_properties are removed.

  • Validation happens before anything is sent, in every store:

    • a declared property's type (a bool is not an int, and an int is not a float);
    • a record vector's or query vector's width;
    • a query vector's finiteness, and a score threshold's;
    • a query limit, which must be positive.

    Record refuses a coordinate that is not finite, and identifiers are matched whole, so a trailing newline no longer passes. A memory search whose top_k is not positive is refused before any work: SearchMemoriesSpec.top_k is positive, so request validation answers 422 naming it, and LongTermMemory.search_scored refuses such a limit before the query is embedded.

  • Consistency: an upsert or delete is durable once it returns, and queries may not reflect it right away. A store that guarantees more states it.

Breaking, with no migration:

  • The semantic feature table gains a vector_uuid column, which create_all does not add to an existing table.
  • The semantic collection declares no indexed properties, so a collection created before this no longer matches the configuration asked for.

Refused where main accepted:

  • A project's properties_schema declares types for its own keys, so an episode whose metadata holds another type for one is refused at the vector write, after the episode store, semantic memory and the event's segments have it. REST metadata values are all strings, so through REST every add carrying a key declared with any type but str is refused. On main such values were written as strings, and SQLite filters on the declared type skipped them. [user properties 2/2] Remove per-project filterable properties (port of #1606) #1670 and [user properties 1/2] Keep undeclared properties out of the vector store #1702, stacked above, take user properties out of the vector store, which closes this.
  • A memory search whose top_k is not positive is refused by request validation (422), before the query is embedded. On main it answered no matches on the SQLite stores and Milvus, and a 500 on Qdrant.

The Qdrant and Milvus stores keep main's layouts here. The PRs above this one move them onto a collection registry.

Why get is removed

  • Its one production caller was a bug. On main, only semantic memory's update_feature calls VectorStoreCollection.get. An upsert replaces a whole record, and the feature row holds no embedding, so update_feature read the stored vector back to write it again with fresh properties. That read-modify-write fails two ways (Semantic memory's update_feature reads its vector back to rewrite it: concurrent updates lose an embedding, and a missed read fails the update half-applied #1721):

    • An update of the text alone can write the old embedding back over a concurrent update's new one.
    • An update whose read misses the record fails after its row has committed.

    After commit 3, the vector record carries no properties, so an update writes the vector store only when it has a new embedding, and reads nothing. Event memory and semantic memory only query.

  • What get returns is a copy, not the record. The vector store is the authority for vectors alone. A record's content lives in its caller's own store: the feature row, or the segment store. A copy read back is only as fresh as the backend's reads, and once written back, a stale read becomes a lasting wrong write. Keeping get would keep that trap open for the next caller.

  • A get that is safe to write from costs too much to promise on every backend. It would have to reflect every write that returned before it, from any process:

    A weaker get is Semantic memory's update_feature reads its vector back to rewrite it: concurrent updates lose an embedding, and a missed read fails the update half-applied #1721 again. With query as the only read, the contract states one rule for every backend (commit 6): a write is durable once it returns, and queries may not reflect it right away.

  • What existed only for it goes too:

    • the return flags;
    • every store's record parsing and get;
    • VectorSearchEngine.get_vectors in both engines;
    • sqlite-vec's vector decoding.

    A backend added later implements upsert, query and delete only.

Why queries return neither properties nor vectors

  • Properties. A returned property is the vector store's copy of a value its caller owns.

    • A mutable one, such as semantic memory's copied feature columns or a caller's metadata, is only as fresh as the backend's reads. Acting on it is unsafe, and a caller that needs the value has to read its own store anyway.
    • An immutable one, such as event memory's _segment_uuid, cannot go stale. But it duplicates a mapping its owner already serves from an index, and it costs a reserved property name.

    Properties are still stored and filterable. A filter on a mutable property sees the copy too.

  • Vectors. No query asked for one: every query on main passes return_vector=False. And a backend does not give back the vector written:

    • Under cosine, Qdrant normalizes a vector when it stores it (CosineMetric::preprocess, v1.19.1), and so does hnswlib. Given [3, 4, 0], both return [0.6, 0.8, 0].
    • A quantizing engine keeps only an approximation: turbovec's 4-bit codes, or usearch at f16 or i8.

    Promising the written vector back would bind every engine to keep full-precision originals beside its index.

Commits

  1. Look up the segments that derivatives belong to in the segment store.
  2. Resolve event memory's search hits through the segment store.
  3. Resolve semantic search hits through a vector_uuid column on the feature row (fixes Semantic memory's update_feature reads its vector back to rewrite it: concurrent updates lose an embedding, and a missed read fails the update half-applied #1721).
  4. Answer a vector store query with record UUIDs and scores; remove get and close_collection.
  5. Refuse vector store inputs no collection can hold or search.
  6. State the vector store's consistency in its contract.
  7. Name the derivative-to-segment map segments_by_derivatives.
  8. Drop a None check no property reaches, and a filter branch no caller takes.
  9. Tidy the tests the review found out of place.
  10. Fold the declared-type check into the vector store utilities.
  11. Resolve semantic search hits in one statement.
  12. Refuse a query limit that is not positive.
  13. Refuse a search for no memories before it does any work.
  14. Show the positive top_k in the OpenAPI document.

Stack

21 open PRs: one independent PR, and the vector store tree of short parallel branches. Every PR's GitHub base is main, since the branches are in a fork and a pull request can target only this repository's branches; the on column gives the order the PRs build on each other instead. A stacked PR's diff on GitHub includes the PRs under it until they merge.

Independent of the vector store tree, directly on main:

# PR change on
— #1624 Make no memory request create a project main

The vector store tree. Each PR builds on the one in its on column; PRs on the same parent are parallel branches and do not depend on each other. #1631 is closed, superseded by #1733–#1736, which hold its changes split in four, with review changes since. #1670 and #1702 sit beneath #1627, whose code depends on them. Until the PRs under it merge, their changes show in a stacked PR's diff.

# PR change on
[vector store scale-out 1/6] #1671 (merged) Remove custom sharding from the Qdrant store (port of #1654) main
[vector store scale-out 2/6] #1733 (this PR) Answer vector store queries with record UUIDs and scores, and refuse invalid inputs main
[vector store scale-out 3/6] #1734 Arbitrate vector store collections in a SQL registry, with an incarnation per collection life and a purge #1733
[vector store scale-out 4/6] #1735 Move the Qdrant store onto the collection registry #1734
[vector store scale-out 5/6] #1736 Move the Milvus store onto the collection registry, against a Milvus server #1735
— #1631 (closed) Superseded by #1733–#1736, which hold its changes split in four, with review changes since —
[user properties 1/2] #1670 Remove per-project filterable properties (port of #1606) #1736
[user properties 2/2] #1702 Keep user properties out of the vector store #1670
[vector store scale-out 6/6] #1627 Make a vector store one collection, with string-keyed partitions #1702
[session storage 1/2] #1622 Create a session's storage with the session, never on a request #1627
[session storage 2/2] #1625 Remove open-or-create from both stores, and close from the segment store #1622
[search results] #1663 Score every vector search by cosine similarity, and name scores for it (port of #1598's cosine half) #1627
[sqlite store fixes 1/7] #1460 Publish vector index files atomically (but not durably) #1663
[sqlite store fixes 2/7] #1469 Never reuse a row id in SQLiteVectorStore #1460
[sqlite store fixes 3/7] #1672 Own the search engine's concurrency in the store, not in each engine (port of #1612) #1469
[sqlite store fixes 4/7] #1673 Serialize a partition's writes so the engine sees them in order (port of #1607) #1672
[sqlite store fixes 5/7] #1674 Refuse a pending row replay cannot honor, instead of dropping it (port of #1608) #1673
[sqlite store fixes 6/7] #1675 Take SQLite's write lock at BEGIN, not at the first write (port of #1609) #1674
[sqlite store fixes 7/7] #1676 Give every write a fresh row id, so a key names one version (port of #1610) #1675
[qdrant options] #1618 Let a deployment tune a Qdrant collection's HNSW, optimizers and quantization #1663
[declared schema 1/2] #1628 Make a vector store filter only on the properties it declares #1663
[declared schema 2/2] #1616 Close the filter union, and make negation the complement on every backend #1628

This PR is its 14 commits, a24e702d3, b24f332f8, 38dc7c267, a7ee319ef, 0c68316f4, 80179289c, bde9edddc, 697c74a3d, 26e36bc9a, 3ec6638b3, 03bc7885b, b05becefa, 8fe6f7b30, 90ff29912, directly on main. #1734 is stacked on it. Split from #1631 (closed).

Verification

At each commit:

  • ruff check, ruff format --check and ty check are clean, run as CI runs them (uv run --frozen --all-extras ty check --project packages/server).
  • The full server suite without integration tests passes at commits 1-5: 2000, 2000, 2005, 1980 and 2031 tests.
  • Commit 6 adds a docstring paragraph only. At it, ruff check is clean.
  • Commit 7 renames a local variable. At it, ruff check and ty check are clean, and the event memory tests pass (199).
  • At commits 8-12, ruff check, ruff format --check and ty check are clean, and the full server suite without integration tests passes: 2031, 2031, 2031, 2031 and 2039 tests.
  • At commits 13 and 14, the same, with 2043 tests each. Commit 13's tests: request validation refuses a top_k of 0 and -1; long-term memory refuses those limits before the query is embedded; the event backend's window test covers its lower bound with a negative expand_context instead of a limit of 0 on the in-memory vector fake.

At the head, in test containers: the integration tests pass against PostgreSQL and Qdrant 1.17.0, the image main tests with (462 passed, 6 skipped), at commit 14. They cover the vector store, the vector-store semantic storage, event memory, long-term memory and the resource manager. The 6 skipped are three Neo4j-specific semantic storage cases, which skip on the other backends, and three long-term memory tests, for want of a NebulaGraph server. The Milvus store's tests run against Milvus Lite in the suite above.

The regression test for identifiers with a trailing newline arrives with the collection lifecycle contract, in the Qdrant PR.

🤖 Generated with Claude Code

edwinyyyu and others added 6 commits October 1, 2026 12:40
Event memory resolves a search hit, a derivative's vector record, to the
derivative's segment through a `_segment_uuid` property it copies onto
every vector record. The segment store already holds that mapping, on the
derivative's own row.

`SegmentStorePartition.get_segment_uuids_by_derivative_uuids` reads it.
The SQLAlchemy partition answers with one query on the derivative table,
served by its primary key, and checks the partition's liveness in the same
statement, as its other reads do. A derivative the partition does not hold
is left out of the result. The in-memory partition the event memory tests
use implements it too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A derivative's vector record carried its segment's UUID as the
`_segment_uuid` property, and a search read it back from each match. That
copy is only as fresh as the vector store's reads, and the segment store
holds the same mapping on the derivative's row.

A search now asks the vector store for no properties and maps the matched
derivatives to their segments with one
`get_segment_uuids_by_derivative_uuids` call. A match whose derivative the
segment store no longer holds, because its segment was deleted and its
vector outlived it, is dropped. Vector records no longer carry
`_segment_uuid`, and the schema event memory expects of its collection no
longer declares it; a collection created with it keeps declaring it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…ure row

A feature's vector record was keyed by a UUID derived from the feature's
id, and carried eleven properties copied from the feature's row, among
them `feature_id`, which a search read back to resolve a hit to its row.
The caller's feature metadata was merged into the same properties.
- `update_feature` read the stored vector back through
  `VectorStoreCollection.get` to write the record again with fresh
  properties, since an upsert replaces a whole record. The value read is
  written back, so a stale read becomes a lasting wrong write: of two
  concurrent updates of one feature, one with a new embedding and one
  without, the second can write the old embedding over the new one; and a
  read that misses the record fails the update after its row was
  committed (MemMachine#1721).
- Nothing filtered on the copied properties; the row is their authority.

The feature row now carries `vector_uuid`, a fresh UUID that keys its
vector record, and the record carries no properties. A search resolves
its hits through that column, in the order the search returned them,
dropping a hit whose row is gone. `update_feature` writes the vector
store only when given a new embedding, and reads nothing back. The delete
paths take the UUIDs of exactly the rows they delete with RETURNING. The
collection declares no indexed properties.

Breaking: a feature table created before this has no `vector_uuid`
column, and the table is created with `create_all`, which adds none; and
a collection created with the old declared properties no longer matches
the configuration `open_or_create_collection` asks for. No migration is
included.

Fixes MemMachine#1721.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…and close_collection

A query returned each match's record: its properties, and its vector on
request. They were copies of what the callers' own stores hold, only as
fresh as the vector store's reads, and since the previous two commits no
caller reads them. `get` has had no caller since semantic memory stopped
reading its vectors back, and a `get` safe to write from would need every
backend to reflect every write that returned before it, from any process.
Every store's `close_collection` did nothing, and nothing called it.

- `QueryMatch` carries a score and a `record_uuid`, and `query` loses
  `return_vector` and `return_properties`. Properties are still stored
  and filtered on.
- `VectorStoreCollection.get` and `VectorStore.close_collection` go, with
  what only served them: the stores' record parsing,
  `VectorSearchEngine.get_vectors`, and sqlite-vec's vector decoding.
- `Record` is input-only, so its vector is required and its properties
  default to `{}`. The model rejects a missing vector at construction, so
  the stores' checks for one and their coalescing of `None` properties go.

Mechanical: event memory and semantic memory read `match.record_uuid` in
place of `match.record.uuid` and stop passing the flags, and the
in-memory test collection follows the interface.

Tests of the return flags, of `get`, and of values read back through a
query go. Tests that check what a store holds read the backend past the
store: the SQLite stores' records tables, a Qdrant scroll of the
collection's partition, a Milvus `get` on the client. The tests of a
store refusing a `None` vector become tests that the model refuses a
missing one, plus one that properties default to `{}` and one that a
record without properties is stored.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Each store refuses, with a ValueError and before any backend call:
- a record whose declared property holds a value of another type, since
  the collection indexes and filters the property as its declared type
  (`require_declared_types`: the value's type is the declared type, and a
  bool is not an int);
- a record vector or a query vector whose width is not the collection's
  dimensions;
- a query vector with a coordinate that is not finite, or a score
  threshold that is not finite.

The model checks a record's own vector for finiteness: `Record.vector` is
`list[FiniteFloat]`, so pydantic refuses a NaN or infinite coordinate when
the record is built, in the pass that already validates each coordinate.
A query vector is a plain sequence, so the stores check it. An embedding
endpoint can return NaN whatever the caller does, so both kinds of vector
are checked.

Qdrant and Milvus check a filter's property keys with the other inputs,
before the early return for no query vectors.

`validate_identifier` uses `fullmatch`: `$` also matches before a trailing
newline, so `"name\n"` passed as a namespace, a collection name or a
property key.

The ABC states these refusals in `upsert` and `query`. Tests: each store
refuses each input, and a record refuses a coordinate that is not finite.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
The contract said nothing about what a query sees of earlier writes.
`VectorStoreCollection` now states what every store keeps:

  An `upsert` or `delete` is durable once it returns; queries may not
  reflect it right away. A store that guarantees more states it.

Neither event memory nor semantic memory reads its own writes back
through the vector store: a search's hits resolve through the segment
store or the feature row, which hold the mapping. Milvus reads at the
consistency level its collections are configured with, and a replicated
Qdrant may answer a query from a replica that has not applied a write
yet, so a stronger promise would bind every store to read settings with
costs of their own.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
This was referenced Oct 1, 2026

@marvinyu-memverge marvinyu-memverge left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks Edwin. Re-reviewed at b05bece (the five commits since bde9edd), and this looks good to me.

  1. Declared-type refusal: agreed. Removing the cause in #1670 and #1702 is better than reordering the add around the check, and the breaking list now names the REST case, which is the part a release note needs. When the stack lands relative to a release is the maintainers' call.
  2. Nit: thanks. _features_by_vector_uuids resolves and filters in one select and keeps the search order, and an empty hit list runs no select.

Also checked in the new commits:

  • require_valid_limit runs in all four stores before the empty-input return. Milvus and both SQLite stores answered no matches for a non-positive limit before, and the body discloses the 422.
  • The None branch dropped from Qdrant's payload build was unreachable: Record.properties is dict[str, PropertyValue], and PropertyValue has no None.
  • Moving require_declared_types into utils.py is a pure move.

return_vector: bool = False,
return_properties: bool = True,
) -> list[QueryResult]:
require_valid_limit(limit)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Handle zero-result event searches before enforcing a positive vector limit. SearchMemoriesSpec.top_k accepts 0, and LongTermMemory._search_scored_event computes vector_search_limit=0 from num_episodes_limit=0. This query now raises ValueError instead of returning no results. The existing zero-limit event test uses an in-memory vector fake that does not enforce this new check.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks. Addressed in 8fe6f7b and 90ff299, differently from the suggestion: a search whose top_k is not positive is now refused before any work, rather than answered with no results.

  • SearchMemoriesSpec.top_k is gt=0, so request validation answers 422 naming top_k. The bound is in the OpenAPI document too, and the Python client validates it through the shared spec.
  • LongTermMemory.search_scored refuses a limit that is not positive before the query is embedded, for callers of the Python API.

Why refuse rather than return nothing: before this, the refusal came from the vector store after the query was embedded, so a search for no results still cost an embedding call and returned nothing useful. A pure top-k search usually requires k ≥ 1; Qdrant answers a limit of 0 with a 422, and Pinecone requires topK of at least 1. On main the outcome depended on the store: no matches on the SQLite stores and Milvus, a 500 on Qdrant.

The fake-only zero-limit case is gone. The event backend's window test now covers its lower bound with a negative expand_context. New tests check both refusals; the long-term-memory one uses an embedder that fails if it is called.

🤖 Written by Claude Code (Claude Opus 5.5) on behalf of @edwinyyyu.

edwinyyyu and others added 2 commits October 2, 2026 17:41
A search whose top_k is not positive asks for nothing, yet it reached
the vector store, which refuses a limit that is not positive, only
after event memory had embedded the query: a request for nothing still
cost an embedding call. SearchMemoriesSpec.top_k is now positive
(gt=0), so request validation refuses it with a 422 naming top_k, the
MCP tool included, and LongTermMemory.search_scored refuses a
num_episodes_limit that is not positive before the query is embedded,
for callers of the Python API. The event backend's window test no
longer runs a limit of 0 on the in-memory vector fake, which does not
check limits; its lower bound is now exercised with a negative
expand_context.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
docs/openapi.json, as docs/tools/generate_openapi.py writes it for the
search specification's top_k: exclusiveMinimum 0 and the description.
The generator's other differences under the locked FastAPI are left to
the regeneration that carries them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@edwinyyyu
edwinyyyu requested a review from malatewang October 3, 2026 03:09
@edwinyyyu
edwinyyyu changed the base branch from main to feat/horizontal-scaling October 5, 2026 16:49
@malatewang
malatewang merged commit f24c0a6 into MemMachine:feat/horizontal-scaling Oct 5, 2026
46 checks passed
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 7, 2026
…UUIDs and scores, and refuse invalid inputs (MemMachine#1733)

* Look up the segments that derivatives belong to in the segment store

Event memory resolves a search hit, a derivative's vector record, to the
derivative's segment through a `_segment_uuid` property it copies onto
every vector record. The segment store already holds that mapping, on the
derivative's own row.

`SegmentStorePartition.get_segment_uuids_by_derivative_uuids` reads it.
The SQLAlchemy partition answers with one query on the derivative table,
served by its primary key, and checks the partition's liveness in the same
statement, as its other reads do. A derivative the partition does not hold
is left out of the result. The in-memory partition the event memory tests
use implements it too.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Resolve event memory's search hits through the segment store

A derivative's vector record carried its segment's UUID as the
`_segment_uuid` property, and a search read it back from each match. That
copy is only as fresh as the vector store's reads, and the segment store
holds the same mapping on the derivative's row.

A search now asks the vector store for no properties and maps the matched
derivatives to their segments with one
`get_segment_uuids_by_derivative_uuids` call. A match whose derivative the
segment store no longer holds, because its segment was deleted and its
vector outlived it, is dropped. Vector records no longer carry
`_segment_uuid`, and the schema event memory expects of its collection no
longer declares it; a collection created with it keeps declaring it.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Resolve semantic search hits through a vector_uuid column on the feature row

A feature's vector record was keyed by a UUID derived from the feature's
id, and carried eleven properties copied from the feature's row, among
them `feature_id`, which a search read back to resolve a hit to its row.
The caller's feature metadata was merged into the same properties.
- `update_feature` read the stored vector back through
  `VectorStoreCollection.get` to write the record again with fresh
  properties, since an upsert replaces a whole record. The value read is
  written back, so a stale read becomes a lasting wrong write: of two
  concurrent updates of one feature, one with a new embedding and one
  without, the second can write the old embedding over the new one; and a
  read that misses the record fails the update after its row was
  committed (MemMachine#1721).
- Nothing filtered on the copied properties; the row is their authority.

The feature row now carries `vector_uuid`, a fresh UUID that keys its
vector record, and the record carries no properties. A search resolves
its hits through that column, in the order the search returned them,
dropping a hit whose row is gone. `update_feature` writes the vector
store only when given a new embedding, and reads nothing back. The delete
paths take the UUIDs of exactly the rows they delete with RETURNING. The
collection declares no indexed properties.

Breaking: a feature table created before this has no `vector_uuid`
column, and the table is created with `create_all`, which adds none; and
a collection created with the old declared properties no longer matches
the configuration `open_or_create_collection` asks for. No migration is
included.

Fixes MemMachine#1721.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Answer a vector store query with record UUIDs and scores; remove get and close_collection

A query returned each match's record: its properties, and its vector on
request. They were copies of what the callers' own stores hold, only as
fresh as the vector store's reads, and since the previous two commits no
caller reads them. `get` has had no caller since semantic memory stopped
reading its vectors back, and a `get` safe to write from would need every
backend to reflect every write that returned before it, from any process.
Every store's `close_collection` did nothing, and nothing called it.

- `QueryMatch` carries a score and a `record_uuid`, and `query` loses
  `return_vector` and `return_properties`. Properties are still stored
  and filtered on.
- `VectorStoreCollection.get` and `VectorStore.close_collection` go, with
  what only served them: the stores' record parsing,
  `VectorSearchEngine.get_vectors`, and sqlite-vec's vector decoding.
- `Record` is input-only, so its vector is required and its properties
  default to `{}`. The model rejects a missing vector at construction, so
  the stores' checks for one and their coalescing of `None` properties go.

Mechanical: event memory and semantic memory read `match.record_uuid` in
place of `match.record.uuid` and stop passing the flags, and the
in-memory test collection follows the interface.

Tests of the return flags, of `get`, and of values read back through a
query go. Tests that check what a store holds read the backend past the
store: the SQLite stores' records tables, a Qdrant scroll of the
collection's partition, a Milvus `get` on the client. The tests of a
store refusing a `None` vector become tests that the model refuses a
missing one, plus one that properties default to `{}` and one that a
record without properties is stored.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse vector store inputs no collection can hold or search

Each store refuses, with a ValueError and before any backend call:
- a record whose declared property holds a value of another type, since
  the collection indexes and filters the property as its declared type
  (`require_declared_types`: the value's type is the declared type, and a
  bool is not an int);
- a record vector or a query vector whose width is not the collection's
  dimensions;
- a query vector with a coordinate that is not finite, or a score
  threshold that is not finite.

The model checks a record's own vector for finiteness: `Record.vector` is
`list[FiniteFloat]`, so pydantic refuses a NaN or infinite coordinate when
the record is built, in the pass that already validates each coordinate.
A query vector is a plain sequence, so the stores check it. An embedding
endpoint can return NaN whatever the caller does, so both kinds of vector
are checked.

Qdrant and Milvus check a filter's property keys with the other inputs,
before the early return for no query vectors.

`validate_identifier` uses `fullmatch`: `$` also matches before a trailing
newline, so `"name\n"` passed as a namespace, a collection name or a
property key.

The ABC states these refusals in `upsert` and `query`. Tests: each store
refuses each input, and a record refuses a coordinate that is not finite.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* State the vector store's consistency in its contract

The contract said nothing about what a query sees of earlier writes.
`VectorStoreCollection` now states what every store keeps:

  An `upsert` or `delete` is durable once it returns; queries may not
  reflect it right away. A store that guarantees more states it.

Neither event memory nor semantic memory reads its own writes back
through the vector store: a search's hits resolve through the segment
store or the feature row, which hold the mapping. Milvus reads at the
consistency level its collections are configured with, and a replicated
Qdrant may answer a query from a replica that has not applied a write
yet, so a stronger promise would bind every store to read settings with
costs of their own.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Name the derivative-to-segment map segments_by_derivatives

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Drop a None check no property reaches, and a filter branch no caller takes

- The Qdrant store's payload skipped a None property, but a Record's
  properties are PropertyValues, which exclude None, so the model refuses
  one before the store sees it.
- Semantic storage's _apply_feature_filter took a Select or a Delete, but
  only selects reach it: delete_feature_set filters its DELETE ...
  RETURNING itself. It takes and answers a Select, and the cast at its
  caller goes.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Tidy the tests the review found out of place

- The SQLite stores' four input-validation tests sat inside TestFilters,
  under a "Delete" banner meant for TestDelete; they move into
  TestInputValidation, and the banner moves above TestDelete.
- Two event memory schemas still declared `_segment_uuid`, which event
  memory no longer reserves: the context test declares `_timestamp`
  alone, and the missing-base-field test declares nothing.
- The 40,000-feature test's comment says which SQLite the bind limit is
  the feature store's.

Tests only; the same tests run and pass.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Fold the declared-type check into the vector store utilities

declared_properties.py held one function beside utils.py, which holds
the other input checks every store runs (dimensions, query vector,
score threshold, identifiers). It imports only common.data_types,
which imports nothing back, so it moves into utils.py as is.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Resolve semantic search hits in one statement

Since queries answer record UUIDs alone, a semantic search mapped its
hits to feature ids in one select, then loaded the features by id with
the filter in another: two statements in two sessions, and the first
ran even when the search found nothing. One select on the vector_uuid
column with the filter now loads the features, ordered as the hits
were, and no select runs for a search with no hits. Results are
unchanged: the same features, in the same order, under the same
filter.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse a query limit that is not positive

The contract said nothing of a limit at or below zero, and the stores
disagreed: the SQLite stores and Milvus answered empty results, and
Qdrant passed the limit on for the server to reject. A limit that is
not positive is now a ValueError in every store, stated in the query
contract beside the other refused inputs, and checked whether or not
there are query vectors.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Refuse a search for no memories before it does any work

A search whose top_k is not positive asks for nothing, yet it reached
the vector store, which refuses a limit that is not positive, only
after event memory had embedded the query: a request for nothing still
cost an embedding call. SearchMemoriesSpec.top_k is now positive
(gt=0), so request validation refuses it with a 422 naming top_k, the
MCP tool included, and LongTermMemory.search_scored refuses a
num_episodes_limit that is not positive before the query is embedded,
for callers of the Python API. The event backend's window test no
longer runs a limit of 0 on the in-memory vector fake, which does not
check limits; its lower bound is now exercised with a negative
expand_context.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

* Show the positive top_k in the OpenAPI document

docs/openapi.json, as docs/tools/generate_openapi.py writes it for the
search specification's top_k: exclusiveMinimum 0 and the description.
The generator's other differences under the locked FastAPI are left to
the regeneration that carries them.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <[email protected]>
Rebased onto main. Main refuses a top_k that is not positive since MemMachine#1740,
with ge=1 and its own tests, so this commit keeps that bound and those tests
in place of the gt=0 bound and the test described above, and
docs/openapi.json shows minimum 1. vector_store_semantic_storage.py keeps
importing datetime, which the history methods MemMachine#1707 added use.
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Oct 10, 2026
…nd their partitions

The design documents came from MemMachine#1733-MemMachine#1736, where a store holds logical
collections addressed by namespace and name, each with its own
configuration. Here a store is one collection, with its dimensions, metric
and declared schema fixed at construction, and a partition is one tenant's
records in it, addressed by key. The documents say so:

- the collection registry document becomes the partition registry
  document: a registry belongs to one store and is addressed by partition
  key; its tables are `partition_registry_pt` and `partition_registry_gc`,
  and a tombstone needs no location, since the incarnation alone finds a
  dead partition's records in the store's native collection; a partition
  created under another schema is refused; `startup` creates the tables, as
  every store's startup creates its durable resources; a decision records
  that a store is one collection;
- the Qdrant and Milvus documents lay out one native collection per store,
  named by the vector store name (`sys_` and the name on Milvus), created at
  startup, where a partition's creation makes nothing in the backend;
- the overview, isolation, consistency and purge documents speak of
  partitions, `purge_deleted_partitions` and `settle(partition)`, and the
  overview describes the store and its partitions.

The measurements and their conditions are unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

horizontal scaling Wrong or unsafe when more than one server process serves the same backends (replicas or workers)

Projects

None yet

3 participants