Skip to content

deploy: update prod to v1.7.2 - #1216

Merged
rdimitrov merged 1 commit into
mainfrom
promote-prod-v1.7.2
Apr 27, 2026
Merged

rdimitrov merged 1 commit into
mainfrom
promote-prod-v1.7.2

Conversation

@rdimitrov

Copy link
Copy Markdown
Member

Summary

Promotes v1.7.2 (PR #1215) to production:

  • Row-constructor cursor pagination — fixes the 760ms list query that was the root cause of today's pool-exhaustion alerts. Local benchmark shows /v0/servers paginated requests dropping from ~30ms (cold cache, 100K rows) to ~0.05ms.
  • Per-phase publish slog — publish complete events now include validate / lock / remotes / version_checks / unmark / db_create timings so the next slow publish is self-diagnostic.
  • pg_stat_statements — aggregate query visibility we didn't have during today's incident.

Deployment caveat

The CNPG cluster spec change for pg_stat_statements triggers an in-pod postgres restart on the next prod Pulumi run. Staging took ~45–70s of registry unavailability during the equivalent restart. The DB retry-with-backoff in v1.7.1 covers most of it, but registry pods may bounce once before they reconnect (one staging pod hit Failed to connect after 8 attempts then recovered on its next restart).

Time the merge for a low-traffic UTC window — the alert history suggests very early UTC is quietest.

Post-merge action

Once the prod Pulumi run completes, run once on the existing cluster:

PATH=/opt/homebrew/share/google-cloud-sdk/bin:$PATH
kubectl exec -i registry-pg-1 -c postgres \
  --context gke_mcp-registry-prod_us-central1-b_mcp-registry-prod \
  -- psql -U postgres -c "CREATE EXTENSION IF NOT EXISTS pg_stat_statements"

Verify:

kubectl exec -i registry-pg-1 -c postgres \
  --context gke_mcp-registry-prod_us-central1-b_mcp-registry-prod \
  -- psql -U postgres -d app -c "
    SELECT round(mean_exec_time::numeric, 2) AS mean_ms, calls,
           substring(query, 1, 80) AS q
    FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 10;"

Test plan

  • v1.7.2 release built and pushed (ghcr.io/modelcontextprotocol/registry:1.7.2)
  • Staging deployed cleanly; pg_stat_statements is collecting on staging
  • Prod Pulumi run applies cleanly; brief PG restart
  • Run CREATE EXTENSION on prod
  • Confirm publish complete events contain per-phase timings on the next publish

🤖 Generated with Claude Code

Promotes #1215 to production: row-constructor cursor pagination fix,
per-phase publish slog, and pg_stat_statements enablement.

Note: the CNPG spec change in v1.7.2 triggers an in-pod postgres
restart on the next prod Pulumi run. Staging took ~45-70s of registry
unavailability during the equivalent restart — DB retry budget covered
most of it but exited on attempt 8/8 once before pod recovery. Same
will happen on prod when this PR is merged. Time the merge for a
low-traffic UTC window. Run `CREATE EXTENSION IF NOT EXISTS
pg_stat_statements` on prod once after the deploy (existing cluster;
fresh clusters get it via postInitApplicationSQL).

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
@rdimitrov
rdimitrov merged commit a4d9d82 into main Apr 27, 2026
5 checks passed
@rdimitrov
rdimitrov deleted the promote-prod-v1.7.2 branch April 27, 2026 19:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant