Repository navigation
Add Qdrant to docker-compose and the Helm chart, and upgrade PostgreSQL to 18 - #1698
Conversation
The sample configs call Qdrant "the recommended default for server
deployments" and then tell the reader to start it by hand:
docker run -p 6333:6333 -p 6334:6334 qdrant/qdrant
which leaves it off memmachine-network, so the app cannot reach it by name.
Postgres and Neo4j both have proper services; Qdrant did not. This adds one,
pinned to v1.19.1, plus QDRANT_HOST/PORT/GRPC_PORT for the app, a depends_on
gate, a named volume and the matching block in the env sample.
The healthcheck is spelled as an explicit bash command rather than CMD-SHELL.
The image carries bash but no curl, wget or nc, and its /bin/sh is dash, which
has no /dev/tcp - so the usual CMD-SHELL curl idiom cannot work here.
Verified against v1.19.1 rather than assumed:
- `docker compose up -d qdrant` reaches healthy in ~9s.
- The check returns 0 against the live port and 1 against a dead one, so it
reports failure rather than always passing.
- Create collection, upsert a point, search it back, delete - all succeed.
- A collection survives `docker compose restart`, so the volume is wired.
- Another container on memmachine-network reaches http://qdrant:6333 by name,
and 6334 accepts connections.
Note this only provides the container. Using it still means pointing
configuration.yml at `backend: event` with `vector_store` wired to a qdrant
entry under resources.databases.
Signed-off-by: Haiyan Wang <[email protected]>
The chart deployed Postgres and Neo4j but had no Qdrant at all, so the vector store for the event-backed long-term memory had to be stood up by hand and wired in by editing the rendered config. This adds it as a first-class component, following the pattern the other two already use. - qdrant-deployment.yaml and qdrant-service.yaml (REST 6333, gRPC 6334), both gated on qdrant.enabled - qdrant-pvc mounted at /qdrant/storage - a wait-for-qdrant initContainer beside the existing two - an event_vector_store entry under resources.databases in the ConfigMap - qdrant.* values, and the README's component, PVC, template and values tables Probes are tcpSocket, matching neo4j and postgres. The image carries no curl or wget, so an httpGet-style check would have needed a different mechanism for no benefit here. The ConfigMap entry is deliberately NOT gated on qdrant.enabled, matching db_neo4j. `enabled: false` means "skip the in-cluster objects and point at an external host", so gating the config too would have broken exactly that case. I had it wrong first and caught it by rendering with enabled=false. Verified: - helm lint clean. - kubeconform -strict against Kubernetes 1.29 schemas: every object valid across default, qdrant.enabled=false, external host, apiKey/https, and all-dependencies-off. 17 objects by default, 14 with Qdrant off - the Deployment, Service and PVC drop out and nothing else changes. - The rendered event_vector_store block validates against the real QdrantConf model for default, apiKey, prefer_grpc and external-host renders. - api_key is emitted only when qdrant.apiKey is set. Not verified: no live cluster was reachable from here, so this is schema and render validation rather than a deployed rollout. Signed-off-by: Haiyan Wang <[email protected]>
Moves docker-compose and the Helm chart from pgvector/pgvector:pg16 to :pg18 (PostgreSQL 18.6, pgvector 0.8.6, up from 16.15 / pgvector 0.8.x). The tag bump alone is not enough. PG18 images store data in a major-version subdirectory - PGDATA is /var/lib/postgresql/18/docker, and the declared VOLUME moved from /var/lib/postgresql/data to /var/lib/postgresql - so the mount point has to move with it. Both the compose volume and the chart's postgres-pvc mountPath are updated accordingly. Checked what happens if it is not, rather than assuming: mounting the old /var/lib/postgresql/data under pg18 makes the image refuse to start and print an explanation, even on an empty volume. It does not start up and quietly write to an unmounted path, so a missed mount cannot silently lose data. The same is true of an existing PG16 volume: pg18 detects the old layout and exits rather than touching it. That makes this a breaking change for existing deployments, which now need a dump/restore. Documented in the chart README with the commands. Verified: - compose: postgres reaches healthy in ~9s; server_version 18.6, data_directory /var/lib/postgresql/18/docker, pgvector 0.8.6. - A table with a vector column survives `compose up --force-recreate`, so the moved mount really does persist. - Simulated the upgrade from a populated PG16 volume: pg18 exits with the layout error instead of starting, in both the old and new mount positions. - helm lint clean; kubeconform -strict against Kubernetes 1.29 valid across default, postgres.enabled=false and qdrant.enabled=false. Signed-off-by: Haiyan Wang <[email protected]>
Review feedback from Shu: only one store should be deployed, and the previous commits gave no way to choose. The chart pinned `episodic_memory.long_term_memory.vector_graph_store: db_neo4j` and deployed Qdrant beside it, so Qdrant ran unused and switching meant editing the rendered config. Compose had the same problem, starting both. The server treats the two as a discriminated union on `backend` - DeclarativeLongTermMemoryConf wants vector_graph_store, EventLongTermMemoryConf wants vector_store plus segment_store - so exactly one applies. Helm: new `episodicMemory.longTermMemory.backend` (declarative | event) renders the matching long_term_memory block, and each store's Deployment, Service and PVC is gated on the backend needing it as well as on its own `enabled` flag. `enabled` keeps its meaning of in-cluster versus external host. The wait-for initContainers are gated on the backend alone, since an external host still needs waiting for. Default stays declarative, so existing installs do not move. Compose: neo4j and qdrant move into `declarative` and `event` profiles, and the app's depends_on entries for them become `required: false` so the store this profile did not start is skipped rather than failing the project. The env sample sets COMPOSE_PROFILES=declarative, preserving today's default. Verified: - Helm renders exactly one store Deployment per backend; declarative gives neo4j + wait-for-neo4j, event gives qdrant + wait-for-qdrant. - Both rendered configuration.yml files load through the real Configuration.load_yml_file: declarative resolves to db_neo4j, event to event_vector_store + db_postgres. - helm lint clean; kubeconform -strict on Kubernetes 1.29 valid for declarative, event, and each with its store external (14/14, 14/14, 11/11). - compose: `--profile declarative` lists neo4j and not qdrant, `--profile event` the reverse, and COMPOSE_PROFILES from .env does the same. - Brought the event profile up for real: postgres and qdrant both healthy, no neo4j container started. Checked rather than assumed that profiles alone were not enough: with a plain depends_on, a profiled service that is not active fails the project with "depends on undefined service". required: false is what makes it work. Signed-off-by: Haiyan Wang <[email protected]>
|
Good catch — you are right, and the PR as it stood had no answer. It pinned The server treats the two as a discriminated union on Helm — one value picks the backend, and only the store it needs is deployed: helm upgrade --install memmachine . --set episodicMemory.longTermMemory.backend=event
The unused store’s Deployment, Service and PVC are skipped even if its Compose — the two stores are now in docker compose --profile event up # qdrant, no neo4jThe env sample sets Verified:
One thing worth recording, since it is not obvious: compose profiles alone were not enough. A plain |
Switching compose to Qdrant took two steps that were not symmetric: set the
profile, then hand-edit configuration.yml. Helm needs only a value because the
chart writes the config itself; compose mounts a file the user supplies. This
adds that file, so the switch is a copy rather than an edit.
cp sample_configs/configuration.event.yml configuration.yml
COMPOSE_PROFILES=event docker compose up -d
It is written for the compose topology specifically. Hostnames are service
names - postgres and qdrant - because inside a container `localhost` is that
container, where nothing is listening. The existing samples use localhost, which
is right for their case (Qdrant in Docker, MemMachine on the host) and wrong for
this one. Neo4j is absent; the event backend does not use it.
No credentials in the file. `password` and `api_key` use the $ENV_NAME syntax
and resolve from what compose already passes in, so a default `docker compose
up` needs no edits beyond setting OPENAI_API_KEY in .env.
Verified, rather than assumed, against the real loader:
- Configuration.load_yml_file parses it and resolves backend=event,
vector_store=event_vector_store, segment_store=profile_storage.
- qdrant -> qdrant:6333 (grpc 6334), postgres -> postgres:5432, no neo4j conf.
- password and api_key both resolve from the environment.
One thing this turned up: only `password` and `api_key` support $ENV_NAME
(PasswordMixin / ApiKeyMixin via WithValueFromEnv). `user` and `db_name` are
taken literally - an earlier draft used $POSTGRES_USER and the loader kept the
string verbatim. They are spelled out, with a comment saying they must match
.env.
Signed-off-by: Haiyan Wang <[email protected]>
| # `enabled` flag is true. | ||
| episodicMemory: | ||
| longTermMemory: | ||
| backend: declarative # declarative | event |
There was a problem hiding this comment.
As we discussed, switch event memory as default.
There was a problem hiding this comment.
Done in 46b3b01 — backend now defaults to event.
One thing to be aware of: releases installed before chart 0.2.0 ran Neo4j, so a plain helm upgrade would now switch them to event and Helm would remove the Neo4j Deployment, Service and neo4j-pvc. The README calls this out and says to pin --set episodicMemory.longTermMemory.backend=declarative when upgrading. Docker compose still defaults to declarative; happy to switch that too if we want them aligned.
There was a problem hiding this comment.
Follow-up: docker compose now defaults to event too (b71b8d9), so the two deployments match.
COMPOSE_PROFILES=eventinsample_configs/env.dockercompose, so a defaultdocker compose upruns Postgres + Qdrant and no Neo4j.memmachine-compose.shgenerates aconfiguration.ymlwired for the selected profile (vector_store+segment_storefor event), and points the Qdrant host at theqdrantservice.- An existing
.envwithoutCOMPOSE_PROFILESwould have started no vector store at all; the script now adds one, inferringdeclarativefrom an existing config that uses Neo4j, so current setups keep their backend.
Note: I rewrote the last two commits to add the DCO sign-off, so the Helm changes are now 8beccba (was 46b3b01).
| {{- else }} | ||
| backend: declarative | ||
| vector_graph_store: db_neo4j | ||
| {{- end }} |
There was a problem hiding this comment.
Check invalid value here?
There was a problem hiding this comment.
Done in 46b3b01. Every template now reads the backend through a memmachine.ltmBackend helper (templates/_helpers.tpl) that fails the render on anything other than declarative or event, e.g.:
episodicMemory.longTermMemory.backend must be "declarative" or "event", got "Event"
The else here is also now an explicit else if eq $backend "declarative", so nothing falls through silently.
| selector: | ||
| app: qdrant | ||
| ports: | ||
| - name: rest |
There was a problem hiding this comment.
use the port number from values.yaml
There was a problem hiding this comment.
Done in 46b3b01. The Service ports are qdrant.port / qdrant.grpcPort, and they target the container's named ports (rest / grpc), so changing the values moves the Service port without touching what Qdrant listens on. Rendered with 7000/7001: the Service, the configmap and the wait-for-qdrant initContainer all pick up the new ports.
| image: qdrant/qdrant:v1.19.1 | ||
| prefer_grpc: false # true = talk gRPC (grpcPort) instead of REST | ||
| https: false # true when fronted by TLS, or for Qdrant Cloud | ||
| # apiKey: <QDRANT_API_KEY> # required for Qdrant Cloud; omit for in-cluster |
There was a problem hiding this comment.
Should move the api key out of yaml?
There was a problem hiding this comment.
Done in 46b3b01. The key no longer lands in the ConfigMap — configuration.yml renders api_key: $QDRANT_API_KEY, which the server resolves from the env var. The key comes from a Secret:
qdrant.existingSecret: a Secret you create out-of-band with aQDRANT_API_KEYkey (preferred; wins if both are set), orqdrant.apiKey: the chart createsqdrant-secretfrom it.
With either set, the in-cluster Qdrant also enforces the key via QDRANT__SERVICE__API_KEY, so it isn't left open. Verified the rendered config loads through Configuration.load_yml_file with the key resolved from the env.
… key out of values Address review on MemMachine#1698: - episodicMemory.longTermMemory.backend now defaults to `event`. The README warns that releases installed before chart 0.2.0 must pin `declarative` on upgrade, or Helm removes the Neo4j Deployment, Service and neo4j-pvc. - Every template reads the backend through a `memmachine.ltmBackend` helper that fails the render on anything but `declarative` or `event`, rather than letting a typo fall through to the declarative branch. - The Qdrant Service takes its ports from qdrant.port / qdrant.grpcPort and targets the container's named ports. - The Qdrant API key is no longer rendered into the ConfigMap. It comes from qdrant.existingSecret, or from a chart-created qdrant-secret when qdrant.apiKey is set; MemMachine receives it as $QDRANT_API_KEY and the in-cluster Qdrant enforces it via QDRANT__SERVICE__API_KEY. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Signed-off-by: Haiyan Wang <[email protected]>
… key out of values Address review on MemMachine#1698: - episodicMemory.longTermMemory.backend now defaults to `event`. The README warns that releases installed before chart 0.2.0 must pin `declarative` on upgrade, or Helm removes the Neo4j Deployment, Service and neo4j-pvc. - Every template reads the backend through a `memmachine.ltmBackend` helper that fails the render on anything but `declarative` or `event`, rather than letting a typo fall through to the declarative branch. - The Qdrant Service takes its ports from qdrant.port / qdrant.grpcPort and targets the container's named ports. - The Qdrant API key is no longer rendered into the ConfigMap. It comes from qdrant.existingSecret, or from a chart-created qdrant-secret when qdrant.apiKey is set; MemMachine receives it as $QDRANT_API_KEY and the in-cluster Qdrant enforces it via QDRANT__SERVICE__API_KEY. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Signed-off-by: Haiyan Wang <[email protected]>
Match the Helm chart: COMPOSE_PROFILES in the env sample is now `event`, so a default `docker compose up` runs Qdrant and no Neo4j. memmachine-compose.sh keeps configuration.yml in step with the profile: - A generated config is wired for the backend COMPOSE_PROFILES selects. The CPU/GPU samples stay declarative-first (they also serve non-compose installs); for event the script comments out vector_graph_store and activates the commented backend/vector_store/segment_store lines. - Qdrant hosts are rewritten from localhost to the `qdrant` service, as postgres and neo4j already were. - An .env without COMPOSE_PROFILES would start no vector store at all. The script now adds one, inferring declarative from an existing config with an active vector_graph_store and event otherwise, so existing setups keep their backend. Verified: generated configs for CPU/GPU x all four providers x both backends load through Configuration.load_yml_file with the expected backend and qdrant host; the awk steps give identical output under gawk, mawk and busybox awk; the default env brings up postgres and qdrant healthy with no neo4j. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]> Signed-off-by: Haiyan Wang <[email protected]>
9743dcd to
b71b8d9
Compare
|
@malatewang @edwinyyyu could you take another look? @malatewang, all four of your comments are addressed, with a reply on each thread:
The thread replies cite 46b3b01. That is the same commit before a rebase, and it is 8beccba on the branch now. @edwinyyyu, a lot has changed since your approval. 60e46a6 deploys only the vector store the chosen backend uses, and 8beccba and b71b8d9 make All 44 checks pass, and the branch merges cleanly. |
…n registry MemMachine#1698 added a Qdrant event_vector_store to the Helm chart and to sample_configs/configuration.event.yml, the configuration docker-compose's event profile uses, without collection_registry, which this PR makes required; both failed configuration validation at boot. Each names the relational database it already has: db_postgres in the chart, profile_storage in the sample. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…ck) (MemMachine#1606) * Regenerate the OpenAPI document under the locked FastAPI `docs/openapi.json` predates the FastAPI release in `uv.lock` (0.141.1), whose `ValidationError` component carries `input` and `ctx`; regenerating the document with `docs/tools/generate_openapi.py` adds the two fields and changes nothing else. Separate from the API changes above it so their diffs of this file show only what they change. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn * Remove per-project filterable properties A project could declare `properties_schema`, a set of caller property keys with types, on its long-term memory configuration; the event backend merged it into the vector store collection's indexed schema and rejected filters on any other `m.<key>`. That let a tenant create database resources (indexes, columns) by naming them in a request, which is what forced per-collection native resources named by a hash of their schema on the backends that limit them. The option is removed from the server configuration, the project API and the memory-configuration API, the Python SDK, the sample configurations, the configuration docs and the OpenAPI document. A filter may name any `m.<key>`; the stores evaluate it on the properties they hold. What a store indexes is decided by the deployment, not per project. A breaking API change on `speedkick`. Co-Authored-By: Claude Opus 5 (1M context) <[email protected]> Claude-Session: https://claude.ai/code/session_01ESpWYTmCR7X3bJEpoA8SAn Rebased onto MemMachine#1631: the per-project schema also leaves MemMachine#1631's service locator, which creates the session's collection in a retry loop, and the commented option goes from the event sample configuration MemMachine#1698 added. --------- Co-authored-by: Claude Opus 5 (1M context) <[email protected]> (cherry picked from commit a8322a7)
Purpose of the change
Two deployment gaps, plus a version bump.
Qdrant had no container anywhere. The sample configs call it "the recommended default for server deployments (concurrent access, scaling)" and then tell the reader to start it by hand:
which leaves it off
memmachine-networkand outside the cluster, so the app cannot reach it by service name. Postgres and Neo4j are proper services in both compose and the Helm chart; Qdrant was in neither. The config schema has supported it all along (QdrantConf) — only the deployment plumbing was missing.PostgreSQL moves 16 → 18.
Description
1.
docker-compose.yml— aqdrantservice pinned tov1.19.1onmemmachine-network, ports 6333 (REST) / 6334 (gRPC), a named volume at/qdrant/storage, a healthcheck,QDRANT_HOST/QDRANT_PORT/QDRANT_GRPC_PORTfor the app, adepends_ongate, and the matching block insample_configs/env.dockercompose.The healthcheck is an explicit
bash -crather than the usualCMD-SHELL+ curl. The image carriesbashbut nocurl,wgetornc, and its/bin/shis dash, which has no/dev/tcp— so the normal idiom cannot work here.2. Helm chart —
qdrant-deployment.yamlandqdrant-service.yaml(both gated onqdrant.enabled),qdrant-pvcat/qdrant/storage, await-for-qdrantinitContainer beside the existing two, anevent_vector_storeentry underresources.databasesin the ConfigMap,qdrant.*values, and the README's component / PVC / template / values tables. Probes aretcpSocket, matching neo4j and postgres.The ConfigMap entry is deliberately not gated on
qdrant.enabled, matchingdb_neo4j.enabled: falsemeans "skip the in-cluster objects and point at an external host", so gating the config too would have broken exactly that case. I had it wrong first and caught it by rendering withenabled=false.3. PostgreSQL 16 → 18 —
pgvector/pgvector:pg18(PostgreSQL 18.6, pgvector 0.8.6, up from 16.15).The tag bump alone is not enough. PG18 images store data in a major-version subdirectory, so the mount point has to move with it:
PGDATA/var/lib/postgresql/data/var/lib/postgresql/18/dockerVOLUME/var/lib/postgresql/data/var/lib/postgresqlBoth the compose volume and the chart's
postgres-pvcmountPath are updated.Type of change
How Has This Been Tested?
Qdrant, compose — against v1.19.1:
docker compose configdocker compose up -d qdrant01— it genuinely failsscore: 1.0compose restarthttp://qdrant:6333from another container on the networkThe negative test mattered: a healthcheck that always returns 0 would be worse than none.
Qdrant, Helm —
helm lintclean, andkubeconform -strictagainst Kubernetes 1.29 schemas:Exactly three objects drop out when Qdrant is disabled, and nothing else changes. The rendered
event_vector_storeblock was also fed through the realQdrantConfmodel for the default,apiKey,prefer_grpcand external-host renders — valid in each.api_keyis emitted only whenqdrant.apiKeyis set.PostgreSQL 18:
server_version18.6,data_directory/var/lib/postgresql/18/docker, pgvector 0.8.6.vectorcolumn survivescompose up --force-recreate, so the moved mount really does persist.Further comments
On the PostgreSQL upgrade and existing data. A PG16 data directory cannot be read by PG18 — normal for any PostgreSQL major upgrade. I simulated it: populate a PG16 volume, then point pg18 at it, and the container exits with the layout error instead of starting, in both the old and new mount positions. Nothing is lost or overwritten, but nothing starts either until the data is migrated.
This was raised and the answer was that there is no existing data to preserve — these are dev deployments — so no migration is needed here. The dump/restore commands are documented in the chart README for anyone who does have data.
Scope. This adds the container and makes it addressable. It does not switch anything over: using Qdrant still means pointing
configuration.ymlatbackend: eventwithvector_storewired to the qdrant entry. Neo4j serves the declarative backend, Qdrant the event one.Not verified. No Kubernetes cluster was reachable from here, so the Helm side is render and schema validation, not a deployed rollout.