Summary
VectorStoreCollection's docstring requires implementations to support undeclared, mixed-type record properties. Three problems follow: a shipped backend has a supported configuration in which it cannot satisfy the clause, the mixed-type half is unimplementable on stores that fix attribute types, and nothing bounds how many distinct property keys a collection accumulates.
packages/server/src/memmachine_server/common/vector_store/vector_store.py:29-33:
Implementations must support storing, filtering on, and returning
record properties not declared in the configured indexed properties schema.
The schema exists to support indexing on fixed-type record properties.
Record properties not declared in the schema may have mixed-type values.
Problem 1: Qdrant can be configured to reject exactly what the clause mandates
Filtering on an undeclared property is, by definition, filtering on a non-indexed payload key. Qdrant's strict mode has a setting for refusing that: setting unindexed_filtering_retrieve to false "prevents retrieving points by filtering on a non indexed payload key which can be very slow".
To be precise about the default, because it matters: strict mode is "Available as of v1.13.0" in open-source Qdrant, and on Qdrant Cloud "strict mode is enabled by default for new collections" — but the same page notes that an unset restriction "does not have any effect. You need to explicitly set the restrictions you want to enforce." So this is not a claim that Qdrant Cloud rejects the mandate out of the box. It is that an operator protecting a shared cluster has a first-class, documented switch that makes the shipped Qdrant backend non-conformant, and the contract gives them no reason to think they shouldn't flip it.
The filtering docs show the same behaviour applied to a filter condition: without an index a condition "still returns correct results but is not accelerated (it is checked per point rather than served by the index)", and "when strict mode is enabled with unindexed_filtering_retrieve or unindexed_filtering_update set to false, a prefix condition is rejected unless the field has a prefix-enabled keyword index".
Problem 2: mixed types are unimplementable where attribute types are fixed
|
Qdrant |
Milvus |
sqlite / sqlite-vec |
turbopuffer |
| store undeclared |
JSON payload |
dynamic field ($meta) |
JSON column |
attribute, type inferred from first write |
| filter undeclared |
unaccelerated scan; rejectable via strict mode |
JSON expression over $meta |
yes |
yes, attributes indexed by default |
| mixed types for one key |
yes |
stored, untagged |
yes, type-tagged |
error |
turbopuffer infers an attribute's type from its first occurrence, and "Changing the attribute type of an existing attribute is currently an error." Schema changes of that kind "cannot be done in-place"; the documented remedy is "exporting documents and upserting into a new namespace". A key that arrives as an int on one record and a str on the next is therefore a hard error there, so a conforming turbopuffer backend cannot be written — it will not be attempted, or it will be written non-conformant and diverge silently. (turbopuffer is not a backend in this repo; it is cited as a representative store this clause excludes.)
It already diverges, in-tree. The two shipped remote backends disagree today:
- SQL stores write type-tagged values (
{"t": …, "v": …}) and gate every comparison on the tag, so a filter never matches a value of another type — _compile_properties_json_leaf, common/filter/sql_filter_util.py:160-167.
- Milvus writes untagged dynamic fields —
_normalize_property_filter_value at milvus_vector_store.py:92-96 converts only datetime, returning every other value unchanged — and inherits whatever Milvus's own JSON comparison does.
Same filter, same data, two answers, and nothing says which is correct.
It also creates an undocumented semantic split. Because the SQL compiler must gate on a type tag, Comparison(field, "!=", v) compiles at sql_filter_util.py:167 to and_(type_check, col != v) — matching only records holding a value of v's type that differs — while Not(Comparison(field, "=", v)) compiles at :203 to ~and_(type_check, col == v), which additionally matches records that do not carry the field or carry it with another type. Both are reasonable predicates; they are not the same predicate; and the difference is discoverable only by reading the compiler. That subtlety exists solely because the contract permits mixed types.
Finally, the allowance imposes machinery no caller asked for: the type-tagged {"t", "v"} encoding and its per-leaf type_check exist to make cross-type comparison well defined for undeclared keys.
Problem 3: key cardinality is unbounded, and on several stores it is a shared budget
Because undeclared keys must work, nothing bounds how many distinct property keys a collection accumulates. Severity depends on how the backend stores keys.
On the shipped backends it is a cost paid by neighbours. Qdrant, Milvus and the SQL stores hold undeclared keys in a JSON payload, so there is no hard cap — but an unaccelerated filter is "checked per point", against a payload that grows without bound, and because the clause also requires returning undeclared properties that cost is paid on ordinary reads too. On a collection shared across tenants (Qdrant's is_tenant=True partition-key index at qdrant_vector_store.py:759-765, Milvus partition keys) one caller's key sprawl is paid for by everyone co-located with it.
The indexed side is explicitly capped. max_payload_index_count is a Qdrant strict-mode parameter that "caps the maximum number of payload index that can exist on a collection". Qdrant Cloud documents "The maximum number of payload indexes per collection is set to 100 (max_payload_index_count is set to 100)" and that "Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0)". Since indexed_properties_schema is compared for equality when reopening a collection, that is a hard ceiling on how much of the property space can ever be made efficient — and it appears nowhere in the contract.
On a store that maps keys into a schema it becomes exhaustion. Elasticsearch is the canonical precedent (not a backend here; cited for the failure mode): index.mapping.total_fields.limit has "default value is 1000", and "Beyond this limit, Elasticsearch returns the error Limit of total fields [X] has been exceeded". The opt-in escape, ignore_dynamic_beyond_limit (default false), is worse for a filter contract: "the index request will not fail. Instead, fields that would exceed the limit are not added to the mapping... The fields that were not added to the mapping will be added to the _ignored field." Writes succeed, filters on those fields match nothing, and nobody is told.
Suggested direction
1. Make the type rule a caller contract rather than an implementation mandate:
A property key holds values of a single type within a collection. An implementation may reject a write that changes a key's type.
This is what turbopuffer already does (type inferred from first occurrence, later change is an error), so it makes the strictest backend the reference rather than the non-conformant one. It also lets the SQL stores' type tagging be defence in depth rather than load-bearing, and removes the != / Not(=) subtlety along with its cause.
2. State a minimum key-cardinality guarantee — an implementation must support at least N distinct property keys per collection — so callers get a portable number instead of discovering a database limit in production.
3. Separate the bundled promises, since they are not equally portable:
- storing undeclared properties — universal, keep as required
- filtering undeclared properties — normally available, but the contract must accommodate a store configured to refuse it (Qdrant strict mode), and an implementation that can only offer it by creating an index should say so, because on a shared collection that cost lands on co-located tenants
- returning undeclared properties — worth stating separately, since it makes payload size a per-read cost
- mixed types — replaced by the single-type contract in (1)
Sources. Code read at 231ce171. External claims quoted from:
Qdrant strict mode ·
Qdrant filtering ·
Qdrant Cloud cluster configuration ·
turbopuffer schema ·
Elasticsearch mapping limit settings
Related: #1535.
🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.
Summary
VectorStoreCollection's docstring requires implementations to support undeclared, mixed-type record properties. Three problems follow: a shipped backend has a supported configuration in which it cannot satisfy the clause, the mixed-type half is unimplementable on stores that fix attribute types, and nothing bounds how many distinct property keys a collection accumulates.packages/server/src/memmachine_server/common/vector_store/vector_store.py:29-33:Problem 1: Qdrant can be configured to reject exactly what the clause mandates
Filtering on an undeclared property is, by definition, filtering on a non-indexed payload key. Qdrant's strict mode has a setting for refusing that: setting
unindexed_filtering_retrieveto false "prevents retrieving points by filtering on a non indexed payload key which can be very slow".To be precise about the default, because it matters: strict mode is "Available as of v1.13.0" in open-source Qdrant, and on Qdrant Cloud "strict mode is enabled by default for new collections" — but the same page notes that an unset restriction "does not have any effect. You need to explicitly set the restrictions you want to enforce." So this is not a claim that Qdrant Cloud rejects the mandate out of the box. It is that an operator protecting a shared cluster has a first-class, documented switch that makes the shipped Qdrant backend non-conformant, and the contract gives them no reason to think they shouldn't flip it.
The filtering docs show the same behaviour applied to a filter condition: without an index a condition "still returns correct results but is not accelerated (it is checked per point rather than served by the index)", and "when strict mode is enabled with
unindexed_filtering_retrieveorunindexed_filtering_updateset to false, a prefix condition is rejected unless the field has a prefix-enabled keyword index".Problem 2: mixed types are unimplementable where attribute types are fixed
$meta)$metaturbopuffer infers an attribute's type from its first occurrence, and "Changing the attribute type of an existing attribute is currently an error." Schema changes of that kind "cannot be done in-place"; the documented remedy is "exporting documents and upserting into a new namespace". A key that arrives as an
inton one record and astron the next is therefore a hard error there, so a conforming turbopuffer backend cannot be written — it will not be attempted, or it will be written non-conformant and diverge silently. (turbopuffer is not a backend in this repo; it is cited as a representative store this clause excludes.)It already diverges, in-tree. The two shipped remote backends disagree today:
{"t": …, "v": …}) and gate every comparison on the tag, so a filter never matches a value of another type —_compile_properties_json_leaf,common/filter/sql_filter_util.py:160-167._normalize_property_filter_valueatmilvus_vector_store.py:92-96converts onlydatetime, returning every other value unchanged — and inherits whatever Milvus's own JSON comparison does.Same filter, same data, two answers, and nothing says which is correct.
It also creates an undocumented semantic split. Because the SQL compiler must gate on a type tag,
Comparison(field, "!=", v)compiles atsql_filter_util.py:167toand_(type_check, col != v)— matching only records holding a value ofv's type that differs — whileNot(Comparison(field, "=", v))compiles at:203to~and_(type_check, col == v), which additionally matches records that do not carry the field or carry it with another type. Both are reasonable predicates; they are not the same predicate; and the difference is discoverable only by reading the compiler. That subtlety exists solely because the contract permits mixed types.Finally, the allowance imposes machinery no caller asked for: the type-tagged
{"t", "v"}encoding and its per-leaftype_checkexist to make cross-type comparison well defined for undeclared keys.Problem 3: key cardinality is unbounded, and on several stores it is a shared budget
Because undeclared keys must work, nothing bounds how many distinct property keys a collection accumulates. Severity depends on how the backend stores keys.
On the shipped backends it is a cost paid by neighbours. Qdrant, Milvus and the SQL stores hold undeclared keys in a JSON payload, so there is no hard cap — but an unaccelerated filter is "checked per point", against a payload that grows without bound, and because the clause also requires returning undeclared properties that cost is paid on ordinary reads too. On a collection shared across tenants (Qdrant's
is_tenant=Truepartition-key index atqdrant_vector_store.py:759-765, Milvus partition keys) one caller's key sprawl is paid for by everyone co-located with it.The indexed side is explicitly capped.
max_payload_index_countis a Qdrant strict-mode parameter that "caps the maximum number of payload index that can exist on a collection". Qdrant Cloud documents "The maximum number of payload indexes per collection is set to 100 (max_payload_index_countis set to 100)" and that "Larger numbers of payload indexes lead to performance degradation (starting with Qdrant v1.16.0)". Sinceindexed_properties_schemais compared for equality when reopening a collection, that is a hard ceiling on how much of the property space can ever be made efficient — and it appears nowhere in the contract.On a store that maps keys into a schema it becomes exhaustion. Elasticsearch is the canonical precedent (not a backend here; cited for the failure mode):
index.mapping.total_fields.limithas "default value is 1000", and "Beyond this limit, Elasticsearch returns the errorLimit of total fields [X] has been exceeded". The opt-in escape,ignore_dynamic_beyond_limit(defaultfalse), is worse for a filter contract: "the index request will not fail. Instead, fields that would exceed the limit are not added to the mapping... The fields that were not added to the mapping will be added to the_ignoredfield." Writes succeed, filters on those fields match nothing, and nobody is told.Suggested direction
1. Make the type rule a caller contract rather than an implementation mandate:
This is what turbopuffer already does (type inferred from first occurrence, later change is an error), so it makes the strictest backend the reference rather than the non-conformant one. It also lets the SQL stores' type tagging be defence in depth rather than load-bearing, and removes the
!=/Not(=)subtlety along with its cause.2. State a minimum key-cardinality guarantee — an implementation must support at least N distinct property keys per collection — so callers get a portable number instead of discovering a database limit in production.
3. Separate the bundled promises, since they are not equally portable:
Sources. Code read at
231ce171. External claims quoted from:Qdrant strict mode ·
Qdrant filtering ·
Qdrant Cloud cluster configuration ·
turbopuffer schema ·
Elasticsearch mapping limit settings
Related: #1535.
🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.