Describe the bug
VectorStoreSemanticStorage declares a semantic feature's free-form content as indexed, filterable vector-store properties.
_vector_properties (vector_store_semantic_storage.py:821) writes value -- the feature's content -- and every scalar entry of caller-supplied metadata into the vector record's properties, and semantic_manager.py:156 declares value in indexed_properties_schema. Declaring a property in that schema is what causes each backend to build an index for it: a keyword payload index per field in Qdrant (qdrant_vector_store.py:_create_native_collection), a JSON expression index per field in sqlite-vec (sqlite_vec_vector_store.py:_ensure_collection_tables).
value is not a filter dimension. It is unique per record, free-form, and is precisely what the record's vector already encodes -- the collection is storing and indexing each record's content inside its own index, in a form that can only ever answer exact-string equality. Caller-supplied metadata is unbounded and unenumerated in the same way; it is written into the payload for whatever keys a caller happens to pass.
Consequences, all on the hot paths:
- Index build and maintenance over high-cardinality free text, per backend, that no query can use selectively.
_vector_search_features (:635) queries with return_properties=True and a limit of _DEFAULT_VECTOR_QUERY_LIMIT = 10_000 (:54, :646), so every semantic search transfers each matched record's full property set -- content included -- for up to 10,000 records, in order to read one field, feature_id (:657).
- Every metadata-only update re-reads the stored vector and rewrites the whole record to keep the copy current (
:295-311), so content duplication costs a vector round trip on writes that never touched the vector.
The authority for all of this is already relational: VectorSemanticFeature (:83) holds value, json_metadata, and the rest, with composite indexes for the real lookup shapes.
Steps to reproduce
- Ingest semantic features through
VectorStoreSemanticStorage against any vector store backend.
- Inspect the created collection's schema.
value is present as an indexed property, alongside whatever keys callers passed in metadata.
- Issue any semantic search. Each of up to 10,000 matched records returns its full payload, including
value, so that feature_id can be read from it.
Expected behavior
Content does not become an indexed vector-store property. A record's payload carries what the vector store itself needs -- here, the feature_id pointer back to VectorSemanticFeature -- and content stays in the relational authority that already holds it.
Concretely: drop value and the merged caller metadata from _vector_properties, and drop value from indexed_properties_schema. feature_id is read back but never filtered on, and VectorStoreCollection already contracts to store and return properties that are not declared in the schema, so it needs no schema entry.
Additional context
Separate from this defect, and deliberately not folded into it: the remaining declared properties (set_id/set, semantic_category_id/category_name/category, tag_id/tag, feature/feature_name) are legitimate filter dimensions -- enumerated, repeated, and exactly what a pushed-down pre-filter would use. They are unused today only because filtering is routed to SQL: no call site passes property_filter to the collection, and _resolve_feature_field (:774) maps every caller-facing filter field, these included, onto a VectorSemanticFeature column or json_metadata[key]. Whether they stay in the payload is a routing decision about where filtering happens, not a category error, and should be decided on its own terms.
If pre-filtering is ever pushed down to the vector store, it should cover only attributes that are immutable after ingest. Post-filtering at the relational authority can drop candidates but cannot recover ones a stale pre-filter wrongly excluded, so a mutable attribute used as a pre-filter converts staleness into unrecoverable false negatives.
Note on landing this: indexed_properties_schema is part of VectorStoreCollectionConfig, and both remote stores derive native collection names as sha256(config.model_dump_json()), so changing the schema repoints the logical collection at a new, empty native collection and orphans the existing vectors. This needs a reindex, or it should land together with storing the resolved native collection name in the collection registry (#1524) so that config evolution stops re-deriving storage identity.
🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.
Describe the bug
VectorStoreSemanticStoragedeclares a semantic feature's free-form content as indexed, filterable vector-store properties._vector_properties(vector_store_semantic_storage.py:821) writesvalue-- the feature's content -- and every scalar entry of caller-suppliedmetadatainto the vector record's properties, andsemantic_manager.py:156declaresvalueinindexed_properties_schema. Declaring a property in that schema is what causes each backend to build an index for it: a keyword payload index per field in Qdrant (qdrant_vector_store.py:_create_native_collection), a JSON expression index per field in sqlite-vec (sqlite_vec_vector_store.py:_ensure_collection_tables).valueis not a filter dimension. It is unique per record, free-form, and is precisely what the record's vector already encodes -- the collection is storing and indexing each record's content inside its own index, in a form that can only ever answer exact-string equality. Caller-suppliedmetadatais unbounded and unenumerated in the same way; it is written into the payload for whatever keys a caller happens to pass.Consequences, all on the hot paths:
_vector_search_features(:635) queries withreturn_properties=Trueand a limit of_DEFAULT_VECTOR_QUERY_LIMIT = 10_000(:54,:646), so every semantic search transfers each matched record's full property set -- content included -- for up to 10,000 records, in order to read one field,feature_id(:657).:295-311), so content duplication costs a vector round trip on writes that never touched the vector.The authority for all of this is already relational:
VectorSemanticFeature(:83) holdsvalue,json_metadata, and the rest, with composite indexes for the real lookup shapes.Steps to reproduce
VectorStoreSemanticStorageagainst any vector store backend.valueis present as an indexed property, alongside whatever keys callers passed inmetadata.value, so thatfeature_idcan be read from it.Expected behavior
Content does not become an indexed vector-store property. A record's payload carries what the vector store itself needs -- here, the
feature_idpointer back toVectorSemanticFeature-- and content stays in the relational authority that already holds it.Concretely: drop
valueand the merged callermetadatafrom_vector_properties, and dropvaluefromindexed_properties_schema.feature_idis read back but never filtered on, andVectorStoreCollectionalready contracts to store and return properties that are not declared in the schema, so it needs no schema entry.Additional context
Separate from this defect, and deliberately not folded into it: the remaining declared properties (
set_id/set,semantic_category_id/category_name/category,tag_id/tag,feature/feature_name) are legitimate filter dimensions -- enumerated, repeated, and exactly what a pushed-down pre-filter would use. They are unused today only because filtering is routed to SQL: no call site passesproperty_filterto the collection, and_resolve_feature_field(:774) maps every caller-facing filter field, these included, onto aVectorSemanticFeaturecolumn orjson_metadata[key]. Whether they stay in the payload is a routing decision about where filtering happens, not a category error, and should be decided on its own terms.If pre-filtering is ever pushed down to the vector store, it should cover only attributes that are immutable after ingest. Post-filtering at the relational authority can drop candidates but cannot recover ones a stale pre-filter wrongly excluded, so a mutable attribute used as a pre-filter converts staleness into unrecoverable false negatives.
Note on landing this:
indexed_properties_schemais part ofVectorStoreCollectionConfig, and both remote stores derive native collection names assha256(config.model_dump_json()), so changing the schema repoints the logical collection at a new, empty native collection and orphans the existing vectors. This needs a reindex, or it should land together with storing the resolved native collection name in the collection registry (#1524) so that config evolution stops re-deriving storage identity.🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.