Skip to content

fix(vector-store): create payload indexes even when the collection exists (port of #1578 to main) - #1681

Merged
malatewang merged 2 commits into
MemMachine:mainfrom
wanghy73:port-main/qdrant-payload-indexes-1578
Sep 18, 2026
Merged

malatewang merged 2 commits into
MemMachine:mainfrom
wanghy73:port-main/qdrant-payload-indexes-1578

Conversation

@wanghy73

@wanghy73 wanghy73 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Port of #1578 from speedkick to main.

Cherry-pick of a2a1754f7eca6d976380cac1a5789859be251572, the squash commit #1578 merged as on
speedkick. It applied to main with no conflicts, and the diff is
byte-identical to the one reviewed on #1578 (checked with git patch-id), so the
review there still stands.

One extra commit rides along: c2d234b (#1592), the follow-up on speedkick that
reformats a single assert this PR adds. Without it the repo-wide
ruff format --check in CI fails. No functional change.

Re-checked on this branch, based on main:

  • The two integration tests were run against a real Qdrant in testcontainers
    and both pass — so the fix works on main, not just on speedkick. Note CI
    will not run them (addopts = ["-m", "not integration"]).
  • ruff check and ruff format --check clean repo-wide on the pinned ruff 0.16.8.
  • Server unit suite: 1869 passed, 3 skipped — identical to main.

The body below is #1578's, unchanged.


The defect

_create_native_collection wrapped create_collection and every
create_payload_index
in one try, then swallowed "already exists" for the whole
block. A collection that already existed therefore raised on the first call, took the
already-exists path, and was left with no payload indexes at all — despite the
docstring promising both were created idempotently.

Two creators are easy to arrive at. The guarding lock is keyed on the
AsyncQdrantClient object:

_name_locks: ClassVar[WeakKeyDictionary[AsyncQdrantClient, defaultdict[tuple[str, str], asyncio.Lock]]]

so it serialises callers inside one process and nothing across them. With
MEMMACHINE_WORKERS above 1 each worker has its own client and its own lock. A crash
between the two calls leaves the same state.

The change

The collection and the indexes now sit under separate guards, and each index is created
individually and tolerant of already-exists.

Evidence

Against Qdrant 1.19 in testcontainers.

Before — test_indexes_are_created_when_the_collection_already_exists fails, and not
with one index missing:

AssertionError: the tenant partition index is missing: a collection that already
existed never had its payload indexes created ... present: []
assert 'sys-partition_key' in set()

payload_schema is empty — none of the twelve.

After — both new tests pass, along with the rest of the suite: 293 unit, 142
integration
. ruff and ruff format clean.

Scope, checked rather than assumed

A filtered query on an unindexed collection returns only the matching tenant's points, so
what a missing sys-partition_key index costs is the multitenant storage layout and
query speed, not isolation
. I verified this directly rather than reasoning from
is_tenant=True.

Only the already-exists path reproduces. Two clients creating simultaneously both
succeed, which the second test pins.

What a reviewer needs to know

These tests are marked integration and need a real server — local-mode Qdrant ignores
payload indexes, so the defect is invisible there. CI will not exercise them, since
addopts = ["-m", "not integration"]. To run them:

pytest packages/server/server_tests/memmachine_server/common/vector_store/test_qdrant_vector_store.py \
  -m integration -k TestCollectionLifecycleAcrossWorkers

Docker required; the fixture starts its own Qdrant.

Not affected

The mm-tb6 benchmark collection was checked and holds all twelve payload indexes
including sys-partition_key, covering every point — no measurement ran against an
unindexed collection.

@edwinyyyu edwinyyyu left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Will be changing this anyway. The amount of effort to get this right at runtime instead of at setup is another reason schema operations/DDL does not belong at runtime/should not be dynamic.

wanghy73 and others added 2 commits September 18, 2026 00:35
…ists (MemMachine#1578)

_create_native_collection wrapped create_collection and every
create_payload_index in one try and swallowed "already exists" for the whole
block. A collection that already existed therefore raised on the first call,
took the already-exists path, and was left with no payload indexes at all -
despite the docstring promising both were created idempotently.

Two creators are easy to arrive at. The guarding lock is keyed on the
AsyncQdrantClient object, so it serialises callers inside one process and
nothing across them; with MEMMACHINE_WORKERS above 1 each worker has its own
client and its own lock. A crash between the two calls leaves the same state.

The collection and the indexes now sit under separate guards, and each index
is created individually and tolerant of already-exists.

Verified against Qdrant 1.19 in testcontainers. Before the change,
test_indexes_are_created_when_the_collection_already_exists fails with an
empty payload_schema - not a missing index, none of the twelve. After it, both
new tests pass, along with the rest of the vector store suite: 293 unit and
142 integration.

Scope, checked rather than assumed: a filtered query on an unindexed
collection returns only the matching tenant's points, so what a missing
sys-partition_key index costs is the multitenant storage layout and query
speed, not isolation. Note also that only the already-exists path reproduces;
two clients creating simultaneously both succeed, which the second test pins.

The tests need a real server and are marked integration - local-mode Qdrant
ignores payload indexes, so the defect is invisible there and CI, which runs
with `-m "not integration"`, will not exercise them.

Claude-Session: https://claude.ai/code/session_01Nr9kacmpFVTTfkZRw6esxP

Co-authored-by: Claude Opus 5 <[email protected]>
(cherry picked from commit a2a1754)
Signed-off-by: Haiyan Wang <[email protected]>
MemMachine#1592)

style: reformat an assert ruff 0.15.14 formats differently

`ruff format --check` fails on speedkick as it stands: the pinned ruff
(0.15.14, from the root pyproject the lint workflow reads its version
from) moves the message of a multi-line `assert` onto its own line, and
one assert in the Qdrant store tests predates that. No behavior change.

Every PR targeting speedkick inherits the red format check until this
lands, which is why it is on its own rather than folded into one of them.

Claude-Session: https://claude.ai/code/session_01NKmF9xNph9QH3ozNw3ZnJL

Co-authored-by: Claude Opus 5 (1M context) <[email protected]>
(cherry picked from commit c2d234b)
Signed-off-by: Haiyan Wang <[email protected]>
@wanghy73
wanghy73 force-pushed the port-main/qdrant-payload-indexes-1578 branch from 0afcc53 to 1bd6b31 Compare September 18, 2026 00:36
@malatewang
malatewang merged commit 73ca138 into MemMachine:main Sep 18, 2026
44 checks passed
This was referenced Sep 18, 2026
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard points the registry at its SQLite
database.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 21, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
edwinyyyu added a commit to edwinyyyu/MemMachine that referenced this pull request Sep 25, 2026
…y; delete logically, reclaim by purge

QdrantVectorStore keeps its catalog in SqlCollectionRegistry instead of the
per-namespace `__registry` collections, and the per-process creation locks
go with them. A logical collection is identified to callers by its
(namespace, name) and inside the store by the incarnation the registry mints
when it is created: every point carries the incarnation in the payload field
that carried the name (the `is_tenant` index stays), and a handle is bound to
one incarnation. A collection deleted and re-created under the same name
starts empty, its predecessor's points are never seen by it and never
reclaimed out from under it, and the old handle raises
VectorStoreCollectionHandleStaleError on every operation, `get` included.

Creation creates the native collection first (idempotent, as MemMachine#1681 left it)
and registers last, so a crash between the two leaves an empty native
collection the next creation of the same configuration adopts, never a row
whose points have nowhere to go; the registry's primary key arbitrates, so a
racing creator on any process gets AlreadyExists and open-or-create adopts
the winner's row. Deletion is one registry transaction: the collection is
unreachable when it commits and its points wait on the queue. The new
`purge_deleted_collections` does one round on the oldest tombstone due: it
looks for one point under the incarnation in the native collection the
tombstone names and, if there is one, deletes by filter in a single
server-side operation, `wait=True`; the registry keeps or removes the
tombstone by what the round found.

A handle checks liveness before an operation, to refuse a handle known to be
dead, and after it, so an operation completed under an incarnation that died
meanwhile raises instead of reporting success. No lock spans the remote call:
a write that lands under a dead incarnation is the purge's to reclaim, which
is what the tombstone's retention is for.

The VectorStore contract states this: the ABC's "at most one process per
collection" sentence becomes each store's own statement (QdrantVectorStore
serves any process sharing the backend and the registry database; the SQLite
stores keep their bound in their class docstrings), a handle's staleness is
part of the collection contract, and `purge_deleted_collections` is part of
the store contract, returning False on the stores whose deletion reclaims
physically (both SQLite stores, and Milvus until its own commit).

QdrantConf gains `registry_database`, the name of a relational database
under `resources.databases`, required, and `tombstone_retention_seconds`
(default 86400); `registry_replication_factor` goes with the registry
collections it sized. DatabaseManager hands the store the engine and the
backend's id, which scopes its rows in the tables every Qdrant store on one
relational database shares. The wizard and the sample configurations' Qdrant
blocks point the registry at their SQLite database, and the docs' parameter
table and databases page describe the two keys and which stores serve any
process.

Tests: the collection lifecycle contract (stale handles, empty re-creation,
open-or-create adopting the live incarnation, idempotent deletion, purge
reclaiming what deletion deferred and sparing live collections, a write
landing under a dead incarnation) runs against the local-mode client and,
as integration tests, the REST and gRPC clients; the cross-worker tests use
two clients over one registry and pin that a strict create is created once.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants