Skip to content

[vector store alt 1/9] Restore per-project filterable properties (speedkick) - #1636

Closed
edwinyyyu wants to merge 1 commit into
MemMachine:speedkickfrom
edwinyyyu:alt/restore-project-properties-speedkick
Closed

edwinyyyu wants to merge 1 commit into
MemMachine:speedkickfrom
edwinyyyu:alt/restore-project-properties-speedkick

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Alternative to the [vector store N/13] chain (#1597, #1622..#1616): the two exclude each other. This chain keeps per-project properties_schema (user-defined metadata indexes) and conforms Qdrant to the incarnation lifecycle; SQLite and Milvus follow later. Every PR is a draft.

Purpose of the change

Reverts the second commit of #1606 ("Remove per-project filterable properties") and keeps its first, the OpenAPI regeneration under the locked FastAPI 0.141.1: docs/openapi.json is regenerated here with the same tool, so it differs from the pre-#1606 document only by the ValidationError fields regeneration added.

A project declares properties_schema, a set of caller property keys with types, on its long-term memory configuration; the event backend merges it into the vector store collection's indexed schema and rejects filters on any other m.<key>. Under this design the schema is fixed for the life of the project, a schema change is a delete and a create, and the vector store maps every distinct schema to one physical collection that all projects declaring it share. What that costs, per backend, is written where a tenant will see it: a warning in the configuration docs and the field descriptions on the server configuration, the project API, the memory-configuration API and the Python SDK; the section below is the same text.

Tenant-created physical resources: limits and cross-tenant effects

Every project (tenant) declares its own properties_schema, fixed for the life of the project (to change it, delete the project and create it again; its memories start empty). Filters may name any m.<key>: a declared key is indexed where the backend supports it, an undeclared key is evaluated without an index. On Qdrant and Milvus the vector store maps every distinct (namespace, vector dimensions, properties_schema) to one physical collection that all projects declaring that schema share, created the first time a project declares it and never dropped by this code. One tenant's choices therefore change what other tenants get. Verified 2026-09-15 against the vendor docs and, for the Qdrant planner, Qdrant's dev source (read_view/filtering.rs, read_view/dispatch.rs, sample_estimation.rs).

  • Physical collections grow with distinct schemas. A project declaring a schema nobody else uses creates a physical collection for the whole deployment. Qdrant Cloud allows 1000 collections per cluster by default and every collection carries its own resource overhead; Milvus allows 65,536 per instance. No reclaim of an empty physical collection is built.
  • Qdrant indexes every declared key as a payload index on the shared physical collection, next to ~10 system keys MemMachine indexes itself. Each index costs memory for every tenant's points, and Qdrant (>= 1.16) degrades with large index counts. The store disables strict mode on every collection it creates, so Qdrant Cloud's default caps (max_payload_index_count 100, rejection of unindexed filters) do not apply; collections created by an earlier MemMachine keep their own setting.
  • Qdrant evaluates an undeclared key per point, within the project. Every query carries the project's tenant condition on an is_tenant keyword index, which is the planner's primary clause, so the candidates are the project's own points, never the physical collection. An undeclared condition estimates as unknown (min 0), so the planner brute-forces the project's posting list with one payload read per point when the project's count in a segment is below full_scan_threshold (10,000 KB / vector bytes: ~1,700 points at 1536 dims), and otherwise samples 1,000 points to choose between that brute force and HNSW traversal with a payload read per visited node, where the missing payload_m links for the undeclared value fragment traversal under a restrictive condition (fewer results). Payload is in the cold tier (disk) by default. The cost lands on the querying project; other tenants feel it only as load on the shared node.
  • Milvus indexes nothing (stated, not mitigated). The store creates no scalar index, so declared and undeclared keys cost the same: all keys are written as dynamic fields ($meta, 65,536 bytes per row, all keys together) and again into the properties JSON field, and the filter, tenant term included, is evaluated per row over the project's hash partition (1 of 16 by default), which holds every tenant hashing there. Milvus allows 64 fields and 1,024 partitions per collection, VARCHAR <= 65,535.
  • SQLite stores share nothing. Each project gets its own tables (and, for the search-engine store, its own index file). sqlite_vec takes the limit nearest neighbors and applies the filter afterwards (k = min(limit, 4096)), so a selective filter returns fewer results than limit, declared or not; the search-engine store evaluates the filter per visited candidate with one SQL statement each and indexes no key.
  • Cross-tenant effects, plainly: (1) the physical-collection count grows with every distinct schema ever declared and is bounded by the backend; (2) on Qdrant, index memory and query load are shared per physical collection, while undeclared filters cost the querying tenant, not its co-tenants' data; (3) on Milvus every filter scans the shared hash partition per row; (4) a schema cannot change in place.

Stack

Slice 1 of 9, every PR targeting speedkick; merge bottom-up.

# PR change
1 #1636 (this PR) Restore per-project filterable properties
2 #1637 Create a session's storage when the session is created, never on a request
3 #1638 Make the scripts and examples that write to a project create it first
4 #1639 Make no memory request create a project
5 #1640 Remove open-or-create and close from both stores
6 #1641 Rename both stores' lookup to get_collection and get_partition
7 #1642 Remove custom sharding from the Qdrant store, and disable strict mode on its collections
8 #1643 Mint an incarnation per collection life on Qdrant; delete logically, reclaim by purge
9 #1644 Bound every request to a remote vector store by a configured timeout

This PR's own change is its commit c23c3959e; the rest of its diff is the slices under it, and drops out as they merge. Sits directly on speedkick; slice 2 is stacked on it.

Verification

ruff check, ruff format --check, ty check, pytest packages/server/server_tests (integration deselected): server 1924 passed, 3 skipped; client 255 passed; run at this slice's own commit on 2026-09-15.

🤖 Generated with Claude Code

Reverts the second commit of MemMachine#1606 (a8322a7), "Remove per-project
filterable properties", and keeps its first, the OpenAPI regeneration
under the locked FastAPI 0.141.1: docs/openapi.json is regenerated here
with the same tool, so it differs from the pre-MemMachine#1606 document only by the
ValidationError fields that regeneration added.

A project declares `properties_schema`, a set of caller property keys with
types, on its long-term memory configuration; the event backend merges it
into the vector store collection's indexed schema and rejects filters on
any other `m.<key>`. Under this design the schema is fixed for the life of
the project, a schema change is a delete and a create, and the vector
store maps every distinct schema to one physical collection that all
projects declaring it share. What that costs, per backend, is written
where a tenant will see it: a warning in the configuration docs and the
field descriptions on the server configuration, the project API, the
memory-configuration API and the Python SDK.

Alternative to the one-collection chain (MemMachine#1622..MemMachine#1616): the two exclude
each other.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Reopen if stack chosen.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant