Skip to content

VectorStoreCollection's undeclared-property mandate: refusable by Qdrant strict mode, unimplementable where attribute types are fixed, unbounded in key cardinality #1534

Description

@edwinyyyu

Summary

VectorStoreCollection's docstring requires implementations to support undeclared, mixed-type record properties. Three problems follow: a shipped backend has a supported configuration in which it cannot satisfy the clause, the mixed-type half is unimplementable on stores that fix attribute types, and nothing bounds how many distinct property keys a collection accumulates.

packages/server/src/memmachine_server/common/vector_store/vector_store.py:29-33:

Implementations must support storing, filtering on, and returning
record properties not declared in the configured indexed properties schema.

The schema exists to support indexing on fixed-type record properties.
Record properties not declared in the schema may have mixed-type values.

Problem 1: Qdrant can be configured to reject exactly what the clause mandates

Filtering on an undeclared property is, by definition, filtering on a non-indexed payload key. Qdrant's strict mode has a setting for refusing that: setting unindexed_filtering_retrieve to false "prevents retrieving points by filtering on a non indexed payload key which can be very slow".

To be precise about the default, because it matters: strict mode is "Available as of v1.13.0" in open-source Qdrant, and on Qdrant Cloud "strict mode is enabled by default for new collections" — but the same page notes that an unset restriction "does not have any effect. You need to explicitly set the restrictions you want to enforce." So this is not a claim that Qdrant Cloud rejects the mandate out of the box. It is that an operator protecting a shared cluster has a first-class, documented switch that makes the shipped Qdrant backend non-conformant, and the contract gives them no reason to think they shouldn't flip it.

The filtering docs show the same behaviour applied to a filter condition: without an index a condition "still returns correct results but is not accelerated (it is checked per point rather than served by the index)", and "when strict mode is enabled with unindexed_filtering_retrieve or unindexed_filtering_update set to false, a prefix condition is rejected unless the field has a prefix-enabled keyword index".

Problem 2: mixed types are unimplementable where attribute types are fixed

Qdrant Milvus sqlite / sqlite-vec turbopuffer
store undeclared JSON payload dynamic field ($meta) JSON column attribute, type inferred from first write
filter undeclared unaccelerated scan; rejectable via strict mode JSON expression over $meta yes yes, attributes indexed by default
mixed types for one key yes stored, untagged yes, type-tagged error

turbopuffer infers an attribute's type from its first occurrence, and "Changing the attribute type of an existing attribute is currently an error." Schema changes of that kind "cannot be done in-place"; the documented remedy is "exporting documents and upserting into a new namespace". A key that arrives as an int on one record and a str on the next is therefore a hard error there, so a conforming turbopuffer backend cannot be written — it will not be attempted, or it will be written non-conformant and diverge silently. (turbopuffer is not a backend in this repo; it is cited as a representative store this clause excludes.)

It already diverges, in-tree. The two shipped remote backends disagree today:

  • SQL stores write type-tagged values ({"t": …, "v": …}) and gate every comparison on the tag, so a filter never matches a value of another type — _compile_properties_json_leaf, common/filter/sql_filter_util.py:160-167.
  • Milvus writes untagged dynamic fields — _normalize_property_filter_value at milvus_vector_store.py:92-96 converts only datetime, returning every other value unchanged — and inherits whatever Milvus's own JSON comparison does.

Same filter, same data, two answers, and nothing says which is correct.

It also creates an undocumented semantic split. Because the SQL compiler must gate on a type tag, Comparison(field, "!=", v) compiles at sql_filter_util.py:167 to and_(type_check, col != v) — matching only records holding a value of v's type that differs — while Not(Comparison(field, "=", v)) compiles at :203 to ~and_(type_check, col == v), which additionally matches records that do not carry the field or carry it with another type. Both are reasonable predicates; they are not the same predicate; and the difference is discoverable only by reading the compiler. That subtlety exists solely because the contract permits mixed types.

Finally, the allowance imposes machinery no caller asked for: the type-tagged {"t", "v"} encoding and its per-leaf type_check exist to make cross-type comparison well defined for undeclared keys.

Problem 3: key cardinality is unbounded, and on several stores it is a shared budget

Because undeclared keys must work, nothing bounds how many distinct property keys a collection accumulates. Severity depends on how the backend stores keys.

On the shipped backends it is a cost paid by neighbours. Qdrant, Milvus and the SQL stores hold undeclared keys in a JSON payload, so there is no hard cap — but an unaccelerated filter is "checked per point", against a payload that grows without bound, and because the clause also requires returning undeclared properties that cost is paid on ordinary reads too. On a collection shared across tenants (Qdrant's is_tenant=True partition-key index at qdrant_vector_store.py:759-765, Milvus partition keys) one caller's key sprawl is paid for by everyone co-located with it.

The indexed side is explicitly capped. max_payload_index_count is a Qdrant strict-mode parameter that "caps the maximum number of payload index that can exist on a collection". Qdrant Cloud documents "The maximum number of payload indexes per collection is set to 100 (max_payload_index_count is set to 100)" and that "Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0)". Since indexed_properties_schema is compared for equality when reopening a collection, that is a hard ceiling on how much of the property space can ever be made efficient — and it appears nowhere in the contract.

On a store that maps keys into a schema it becomes exhaustion. Elasticsearch is the canonical precedent (not a backend here; cited for the failure mode): index.mapping.total_fields.limit has "default value is 1000", and "Beyond this limit, Elasticsearch returns the error Limit of total fields [X] has been exceeded". The opt-in escape, ignore_dynamic_beyond_limit (default false), is worse for a filter contract: "the index request will not fail. Instead, fields that would exceed the limit are not added to the mapping... The fields that were not added to the mapping will be added to the _ignored field." Writes succeed, filters on those fields match nothing, and nobody is told.

Suggested direction

1. Make the type rule a caller contract rather than an implementation mandate:

A property key holds values of a single type within a collection. An implementation may reject a write that changes a key's type.

This is what turbopuffer already does (type inferred from first occurrence, later change is an error), so it makes the strictest backend the reference rather than the non-conformant one. It also lets the SQL stores' type tagging be defence in depth rather than load-bearing, and removes the != / Not(=) subtlety along with its cause.

2. State a minimum key-cardinality guarantee — an implementation must support at least N distinct property keys per collection — so callers get a portable number instead of discovering a database limit in production.

3. Separate the bundled promises, since they are not equally portable:

  • storing undeclared properties — universal, keep as required
  • filtering undeclared properties — normally available, but the contract must accommodate a store configured to refuse it (Qdrant strict mode), and an implementation that can only offer it by creating an index should say so, because on a shared collection that cost lands on co-located tenants
  • returning undeclared properties — worth stating separately, since it makes payload size a per-read cost
  • mixed types — replaced by the single-type contract in (1)

Sources. Code read at 231ce171. External claims quoted from:
Qdrant strict mode ·
Qdrant filtering ·
Qdrant Cloud cluster configuration ·
turbopuffer schema ·
Elasticsearch mapping limit settings

Related: #1535.


🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.

Activity

  1. added theissue type on Aug 28, 2026
  2. changed the title [-]VectorStoreCollection requires undeclared mixed-type properties, which schema-first stores cannot implement[/-] [+]VectorStoreCollection's undeclared-property mandate is unimplementable on schema-first stores and leaves key cardinality unbounded[/+] on Aug 28, 2026
  3. changed the title [-]VectorStoreCollection's undeclared-property mandate is unimplementable on schema-first stores and leaves key cardinality unbounded[/-] [+]VectorStoreCollection's undeclared-property mandate: unsatisfiable under Qdrant strict mode, unimplementable on schema-first stores, unbounded in key cardinality[/+] on Aug 28, 2026
  4. changed the title [-]VectorStoreCollection's undeclared-property mandate: unsatisfiable under Qdrant strict mode, unimplementable on schema-first stores, unbounded in key cardinality[/-] [+]VectorStoreCollection's undeclared-property mandate: refusable by Qdrant strict mode, unimplementable where attribute types are fixed, unbounded in key cardinality[/+] on Aug 28, 2026
  5. edwinyyyu commented on Sep 15, 2026

    @edwinyyyu
    ContributorAuthor

    Folded into #1648, which carries this body unchanged under "Folded issues" alongside the seven other issues on the same decision (who declares indexed properties, whether native resources are created at runtime, how logical collections map to them), and records where the two candidate resolutions stand. Nothing here is dropped; follow-up on #1648.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions