Skip to content

VectorStore.create_collection takes a per-collection config, promising per-tenant schemas that shared-container multitenancy cannot support #1573

Description

@edwinyyyu

What happened

VectorStore.create_collection(namespace, name, config) takes a VectorStoreCollectionConfig per logical collection, so the interface promises that two collections in the same store may have different vector dimensions, similarity metrics, and indexed property schemas.

Per-collection dimensions and metric are bounded in a way schemas are not — they mean a native container per embedding model, so the container count follows model count rather than collection count. They are still not free, for the reason in the next section.

Per-collection indexed property schemas are not, and cannot be made so under shared-container multitenancy. A payload or scalar index is a structure per (container, field) that spans every record in the container, not only the declaring collection's. N collections declaring k distinct fields each gives up to N×k structures, index builds run over all co-located data, and nothing can be dropped without tracking which collections still need it. The cost grows with collection churn and never shrinks.

This is not a MemMachine limitation. Milvus documents the same tradeoff directly: partition-key tenancy, the only tier that scales to millions of tenants, requires a shared schema; per-tenant schemas are available only at collection level (65,536 by default) or database level (64 by default). Isolating index sets by giving each collection its own container is the collection-per-tenant model that Qdrant's guidance also warns against.

No caller currently exercises the capability — properties_schema is a server-global setting, so collections partition by config revision rather than per collection. The hazard is latent in the interface, not the implementation: the signature is the only surviving trace of an intent that the architecture cannot deliver at scale, and it is free to remove only until something starts relying on it.

The broader reason: runtime config makes the container set unknowable

The index-cost argument above applies only to schemas. This one applies to any per-collection configuration, dimensions and metric included.

A native container's properties are fixed when it is created. If callers supply configuration at collection-creation time, the set of containers a deployment needs becomes a function of what callers happen to ask for, discovered at runtime. A deployment cannot provision ahead what it cannot enumerate ahead, so container creation lands on the request path.

That can be made to appear atomic to the caller, and the current design does so: native-first ordering plus idempotent, content-addressed creation means a crashed create leaves an empty container that the next matching create adopts. The objection is not that it is impossible. It is what it costs on three axes.

Scalability. The number of native containers stops being a deployment decision and becomes a consequence of caller behaviour. Providers cap containers hard — 20 to 200 indexes on Pinecone, 1000 collections per Qdrant cloud cluster, ~10,000 practical on Milvus — so an unbounded caller-driven set turns a capacity limit into a runtime failure triggered by what someone asked for. Container creation is also slow enough to matter on that path: a Pinecone index takes minutes, so create latency depends on whether a container happens to exist yet.

Safety. Appearing atomic rests on a precondition that is nowhere written down: creation must be idempotent and content-addressed, so the orphan is adoptable rather than leaked. A backend that names containers per collection leaks one per crashed create, and nothing in the code says why that differs. Preconditions that are not stated are violated by the next implementation. Separately, config in the entry is a second source of truth for something that also lives in deployment configuration, and the two can drift apart silently — the same family as #1562.

Maintainability. The machinery that sustains the appearance is real and permanent: ordering rationale, crash-window reasoning, an adoptability precondition, and — for any backend where creation is not adoptable — an intent ledger with a sweeper and grace periods, purely to reclaim what a failed create left behind. All of it exists to manage a problem the interface introduces. Moving configuration to deployment configuration deletes that class of code instead of perfecting it: the container set becomes finite and provisionable, and creating a collection becomes one registry insert.

A smaller benefit follows. With no config in the entry there is nothing to compare, so get_or_register's "never compares entries" policy, the config-mismatch error, and the question of what config equality means across backends all leave the interface with it — questions the interface currently invites and cannot answer.

What is being removed is tenant-defined indexes, not indexes

The distinction that decides this is who defines an indexed property, not whether one exists.

Server-defined properties are declared once by the deployment and are therefore identical for every collection that server serves. That is precisely what makes container sharing work: co-located collections carry the same index set by construction, so there is no union to accumulate, no collection paying memory and build time for another's fields, and a new collection costs no migration. Every index the system needs today is of this kind.

Tenant-defined properties — one collection indexing invoice_id while another indexes patient_id — are the ones that break the sharing invariant, because the container must then carry the union of whatever its occupants happened to ask for. That union is what grows with churn and cannot be reclaimed.

Removing the config parameter removes the second, not the first. Indexed properties keep existing and keep being declared; they move to deployment configuration, where being identical across co-located collections becomes a property of the design rather than a coincidence of every caller passing the same thing. What disappears is a capability nothing currently uses, and one that no vector database offers under shared-container tenancy anyway.

Expected

The interface should promise what the architecture can support: one schema per native container, configured per deployment.

Suggested direction

Drop the config parameter from create_collection and move collection configuration into deployment configuration, alongside the decision about which native containers exist (#1572). Registry entries then record where a collection's records live, not how it is shaped.

Callers that need a property they did not declare are not blocked: both Qdrant and Milvus accept undeclared fields and filter them, so "filterable" and "efficiently filterable" are separate tiers and only the second needs declaring. Declaring a generous superset per deployment covers the rest.

This needs confirmation from whoever intended caller-defined indexes, since removing the parameter forecloses it.

Notes

Related to #1535, which covers indexed_properties_schema being honoured inconsistently across backends and forming part of collection identity. This issue is the narrower question of whether per-collection configuration should be in the interface at all.

Activity

  1. added theissue type on Sep 2, 2026
  2. changed the issue type fromtoon Sep 15, 2026
  3. self-assigned this
    on Sep 15, 2026
  4. edwinyyyu commented on Sep 15, 2026

    @edwinyyyu
    ContributorAuthor

    Folded into #1648, which carries this body unchanged under "Folded issues" alongside the seven other issues on the same decision (who declares indexed properties, whether native resources are created at runtime, how logical collections map to them), and records where the two candidate resolutions stand. Nothing here is dropped; follow-up on #1648.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions