Skip to content

Feat: Add vector store collection registry with SQLAlchemy implementation (collection registry stack 1/5) - #1526

Closed
edwinyyyu wants to merge 1 commit into
MemMachine:mainfrom
edwinyyyu:feat/config-registry
Closed

edwinyyyu wants to merge 1 commit into
MemMachine:mainfrom
edwinyyyu:feat/config-registry

Conversation

@edwinyyyu

@edwinyyyu edwinyyyu commented Aug 26, 2026 •

Copy link
Copy Markdown
Contributor

Purpose of the change

Horizontal scalability: multiple MemMachine server processes should be able to work with the same collections and partitions on one backend (#1524). Strict sharding — each collection managed by exactly one process, as the VectorStore contract requires today — has always worked; the restriction itself is the limitation, and what prevents lifting it is collection metadata management, not the data path. This PR introduces the enabling primitive — a cross-process-safe collection registry (CollectionRegistry ABC + a SQLAlchemy implementation, in the vector_store package). Today several backends keep collection/partition metadata in ad-hoc per-backend bookkeeping whose create/open-or-create sequences are non-atomic read-check-write, guarded only by per-process asyncio locks. A registry backed by a store with real unique constraints makes entry creation an atomic compare-and-set across processes.

A follow-up PR switches QdrantVectorStore's internal Qdrant-based __registry bookkeeping to this registry.

Description

Contract (common/vector_store/collection_registry.py): a CollectionRegistry is a durable catalog of one vector store's logical collections — a mapping from (namespace, name) to an immutable CollectionRegistryEntry. The registry is deliberately specific to vector store collections rather than a generic key->config registry: the realistic reuse axis (other vector store backends — Milvus follows in this stack) is within the collection domain, a specific API can be widened later as an additive refactor whereas narrowing a shipped generic ABC is breaking, and speaking the domain removes key/adapter plumbing from the API entirely. Operations:

  • register(namespace, name, entry) — atomic; raises CollectionAlreadyRegisteredError on an existing collection regardless of stored-entry equality
  • get(namespace, name) — stored entry or None
  • get_or_register(namespace, name, entry) — atomic; returns (stored_entry, registered); never compares entries — config-equality policy belongs to the store, which translates conflicts into its own domain errors
  • deregister(namespace, name) — idempotent
  • startup() — idempotently prepares storage

The entry (CollectionRegistryEntry = {config, native_collection_name, partition_key}) stores resolved identity, one format shared by all backends: config is the ABC-level VectorStoreCollectionConfig, the native name is pinned at registration (so config-serialization changes can never silently repoint collections at new empty native collections), and the partition key carries a per-registration generation (so records written through handles held across a deregistration stay invisible and are never resurrected). Evolution policy: extend by adding optional fields with defaults — no version bump, no migration; backend-specific needs use the same mechanism, not per-backend entry types.

There is deliberately no update operation: entries are immutable once registered (the config hash is the native collection identity). Adding one later is purely additive.

Format evolution without standing version machinery. There is deliberately no stored format version and no version check. The evolution policy (add-optional-only, defaults reproducing prior behavior — stated on the entry model and guarded by a canned v1-row regression test) is what makes that safe: old rows classify exactly through defaults, so a future breaking change can introduce an entry format version field (default 1) at the moment it is first needed — every pre-existing row is correctly version 1 by construction — and ship its migration as a windowed migrate-then-deploy step. The accepted residual: an old binary started against migrated data fails with validation errors at first read rather than being refused at boot. (An earlier draft carried a declarations table with a startup version check and a redeclare primitive; it was removed because it insured against exactly the scenario the evolution policy already insures, at the cost of real API and documentation surface.)

SQLAlchemy implementation (sqlalchemy_collection_registry.py): each registry owns a dedicated table collection_registry_<name> (key VARCHAR(255) PK, entry JSON, JSONB on PostgreSQL), so registries sharing a database are isolated at the table level and a store can only reach its own registry (idempotent create_all in startup(), matching the segment store's schema-management convention). Storage keys are f"{namespace}/{name}" — an implementation detail; "/" is outside the identifier charset, so distinct pairs can never collide ("__" would be ambiguous). The primary key is the concurrency arbiter:

  • register = plain INSERT; IntegrityError -> CollectionAlreadyRegisteredError (the pattern SQLAlchemySegmentStore.create_partition already uses)
  • get_or_register = SELECT fast path, then native INSERT .. ON CONFLICT DO NOTHING; on conflict re-SELECT the winner — no exception-driven control flow
  • deregister = single DELETE

Entries round-trip through pydantic (dump_python(mode="json") / validate_python); get_or_register returns the serialization round trip of the input, so the registering store sees exactly what every later reader gets. Supported dialects: PostgreSQL and SQLite (the two relational providers the server offers). The 32-byte registry-name limit keeps collection_registry_<name> within PostgreSQL's 63-byte identifier limit. No new dependencies.

Adopters in this stack: Qdrant (next PR) and Milvus (final PR). The SQLite vector stores keep their own _CollectionRow bookkeeping — their state is process-local anyway, so the registry buys them nothing today.

Design doc: design/collection_registry.md — contributor-facing design notes live in the root design/ directory, outside the deployed docs site (docs/) and outside packaged sources.

Fixes/Closes

Fixes #1524

Type of change

  • New feature (non-breaking change which adds functionality)

How Has This Been Tested?

  • Unit Test
  • Integration Test

server_tests/memmachine_server/common/vector_store/test_sqlalchemy_collection_registry.py, parametrized over SQLite (file) and PostgreSQL (testcontainer, integration-marked): startup idempotence (twice, and a second instance on the same engine), entry round trips, duplicate registration (same and different entry), get_or_register canonicalization and no-comparison semantics, idempotent deregister, registry isolation on a shared engine, key-ambiguity regression (("a__b","c") vs ("a","b__c")), identifier/name/engine validation, the canned v1-row readability guard, and concurrency tests (two instances over one engine, asyncio.gather): concurrent register -> exactly one AlreadyRegistered; concurrent get_or_register with mismatched entries -> both callers observe the same stored entry, exactly one registered=True.

Test Results: 45 passed (SQLite + PostgreSQL) locally; ruff check, ruff format --check, and ty check packages clean (no new diagnostics).

Checklist

  • I have signed the commit(s) within this pull request
  • My code follows the style guidelines of this project (See STYLE_GUIDE.md)
  • I have performed a self-review of my own code
  • I have commented my code
  • My changes generate no new warnings
  • I have added unit tests that prove my fix is effective or that my feature works
  • New and existing unit tests pass locally with my changes

Further comments

Design notes for the reviewer:

  • Collection-specific, not generic: see the scope note above; design/collection_registry.md records the rejected generic draft and the reasoning, so the narrowing reads as deliberate rather than as a missed abstraction.
  • Flat JSON, not typed columns: every access is get-by-key, pydantic validates at the boundary, and typed columns would turn entry evolution into DDL migrations in a codebase whose schema management is create_all-on-startup.
  • Table-per-registry, not one shared table: simpler keys, database-level isolation, whole-registry cleanup is a DROP TABLE.
  • No registry management surface (list/delete registries): registries are constructed by wiring code from configuration; the database catalog already lists collection_registry_* tables for observability, and management operations are exactly the power stores must not have.

@edwinyyyu edwinyyyu changed the title Feat: Add generic config registry with SQLAlchemy implementation Feat: Add generic config registry with SQLAlchemy implementation (1/4) Aug 26, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/config-registry branch 3 times, most recently from 10e77c4 to 4f59b07 Compare August 27, 2026 17:27
@edwinyyyu edwinyyyu changed the title Feat: Add generic config registry with SQLAlchemy implementation (1/4) Feat: Add vector store collection registry with SQLAlchemy implementation (1/5) Aug 27, 2026
@edwinyyyu edwinyyyu changed the title Feat: Add vector store collection registry with SQLAlchemy implementation (1/5) Feat: Add vector store collection registry with SQLAlchemy implementation (collection registry stack 1/5) Aug 27, 2026
@edwinyyyu
edwinyyyu force-pushed the feat/config-registry branch 11 times, most recently from 078a5b9 to 9629018 Compare August 28, 2026 00:25
@edwinyyyu
edwinyyyu force-pushed the feat/config-registry branch 2 times, most recently from 8dfd87d to 7e2ccfa Compare August 31, 2026 22:00
@edwinyyyu

Copy link
Copy Markdown
Contributor Author

Superseded by a reworked stack. The premise this PR was built on has changed in four ways, each now filed separately:

The defects this PR did fix are still fixed by the replacement, and are now filed on their own so they do not get lost: #1562 (native name re-derived on open) and #1563 (stale handles resurrecting records).

Closing rather than force-pushing, so the review discussion here stays attached to the design it was about. Replacement PRs to follow.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feat]: Cross-process collection registry for vector store metadata

1 participant