Skip to content

[Bug]: Semantic memory indexes free-form content as filterable vector-store properties #1528

Description

@edwinyyyu

Describe the bug

VectorStoreSemanticStorage declares a semantic feature's free-form content as indexed, filterable vector-store properties.

_vector_properties (vector_store_semantic_storage.py:821) writes value -- the feature's content -- and every scalar entry of caller-supplied metadata into the vector record's properties, and semantic_manager.py:156 declares value in indexed_properties_schema. Declaring a property in that schema is what causes each backend to build an index for it: a keyword payload index per field in Qdrant (qdrant_vector_store.py:_create_native_collection), a JSON expression index per field in sqlite-vec (sqlite_vec_vector_store.py:_ensure_collection_tables).

value is not a filter dimension. It is unique per record, free-form, and is precisely what the record's vector already encodes -- the collection is storing and indexing each record's content inside its own index, in a form that can only ever answer exact-string equality. Caller-supplied metadata is unbounded and unenumerated in the same way; it is written into the payload for whatever keys a caller happens to pass.

Consequences, all on the hot paths:

  • Index build and maintenance over high-cardinality free text, per backend, that no query can use selectively.
  • _vector_search_features (:635) queries with return_properties=True and a limit of _DEFAULT_VECTOR_QUERY_LIMIT = 10_000 (:54, :646), so every semantic search transfers each matched record's full property set -- content included -- for up to 10,000 records, in order to read one field, feature_id (:657).
  • Every metadata-only update re-reads the stored vector and rewrites the whole record to keep the copy current (:295-311), so content duplication costs a vector round trip on writes that never touched the vector.

The authority for all of this is already relational: VectorSemanticFeature (:83) holds value, json_metadata, and the rest, with composite indexes for the real lookup shapes.

Steps to reproduce

  1. Ingest semantic features through VectorStoreSemanticStorage against any vector store backend.
  2. Inspect the created collection's schema. value is present as an indexed property, alongside whatever keys callers passed in metadata.
  3. Issue any semantic search. Each of up to 10,000 matched records returns its full payload, including value, so that feature_id can be read from it.

Expected behavior

Content does not become an indexed vector-store property. A record's payload carries what the vector store itself needs -- here, the feature_id pointer back to VectorSemanticFeature -- and content stays in the relational authority that already holds it.

Concretely: drop value and the merged caller metadata from _vector_properties, and drop value from indexed_properties_schema. feature_id is read back but never filtered on, and VectorStoreCollection already contracts to store and return properties that are not declared in the schema, so it needs no schema entry.

Additional context

Separate from this defect, and deliberately not folded into it: the remaining declared properties (set_id/set, semantic_category_id/category_name/category, tag_id/tag, feature/feature_name) are legitimate filter dimensions -- enumerated, repeated, and exactly what a pushed-down pre-filter would use. They are unused today only because filtering is routed to SQL: no call site passes property_filter to the collection, and _resolve_feature_field (:774) maps every caller-facing filter field, these included, onto a VectorSemanticFeature column or json_metadata[key]. Whether they stay in the payload is a routing decision about where filtering happens, not a category error, and should be decided on its own terms.

If pre-filtering is ever pushed down to the vector store, it should cover only attributes that are immutable after ingest. Post-filtering at the relational authority can drop candidates but cannot recover ones a stale pre-filter wrongly excluded, so a mutable attribute used as a pre-filter converts staleness into unrecoverable false negatives.

Note on landing this: indexed_properties_schema is part of VectorStoreCollectionConfig, and both remote stores derive native collection names as sha256(config.model_dump_json()), so changing the schema repoints the logical collection at a new, empty native collection and orphans the existing vectors. This needs a reindex, or it should land together with storing the resolved native collection name in the collection registry (#1524) so that config evolution stops re-deriving storage identity.


🤖 Written by Claude Code (Opus 5) on behalf of @edwinyyyu.

Activity

  1. changed the title [-][Feat]: Semantic memory stores non-filterable payload in the vector store[/-] [+][Bug]: Semantic memory indexes free-form content as filterable vector-store properties[/+] on Aug 26, 2026
  2. added theissue type on Aug 26, 2026
  3. edwinyyyu commented on Sep 15, 2026

    @edwinyyyu
    ContributorAuthor

    Folded into #1648, which carries this body unchanged under "Folded issues" alongside the seven other issues on the same decision (who declares indexed properties, whether native resources are created at runtime, how logical collections map to them), and records where the two candidate resolutions stand. Nothing here is dropped; follow-up on #1648.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions