Skip to content

Alert when one worker replica stops dequeuing while jobs wait #154

Description

@HMarzban

Problem

Worker health cannot see a single parked replica. getStoreQueueOldestWaitingAgeMs reads the shared wait list (apps/hocuspocus.server/src/lib/queue.ts:366-371), so a replica that has stopped dequeuing still reports healthy while its sibling drains the queue.

Store jobs carry document saves. A parked replica therefore delays persistence, and nothing pages.

What to do

Add a Grafana alert rule in scripts/observability/grafana/provisioning/alerting/rules-workers.yml. Fire when completed jobs_total stays flat on one instance for five minutes while queue_jobs{state="waiting"} is above zero.

Blocked by

Needs per-replica labels. Confirm those first, in the sibling issue about count by (instance).

Acceptance

  • Stopping one worker container while the other drains raises the alert within five minutes.
  • A normal single-replica restart does not raise it.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions