You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Measure cursor traffic per edit in production after the cursor guards deploy #353
Every open editor used to send its cursor again after each remote edit, even when the cursor had not moved. On a busy pad on 2026-09-26, production sent many awareness frames per document update. Commit e84934acd adds two client guards that stop the unchanged and echoed cursor frames. That commit ships in the production deploy of 204f8ef62. This issue measures the same ratio in production before and after that deploy, and records how they compare.
Acceptance criteria
The baseline ratio for the 2026-09-26 busy window is saved before Prometheus drops it. Prometheus keeps 15 days or 4 GB, whichever runs out first. Save it by 2026-10-10.
The same query runs over a production window after the deploy. That window has at least 10 open editors and at least 3 people typing.
For each window, the maintainer keeps the start and end time, the ratio, and the peak ws_active_connections in private notes, not on this issue.
A comment on this issue says whether the after ratio per connection (ratio divided by peak ws_active_connections) is at most half the before value. It gives no absolute numbers.
If the after value is more than half, the comment names the next check to run. One example: confirm that the open tabs run the new webapp build.
Blocked by
None — can start now.
Agent brief
Type: HITL — only the maintainer can read production Grafana at https://grafana.docs.plus and run a live multi-editor session. The maintainer runs the queries and posts the result as a comment on this issue. The after window needs a real session with several editors: plan one, or wait for the next busy pad.
Category: enhancement
Current behavior:
apps/webapp/src/hooks/guardProviderAwareness.ts skips a cursor write that equals the current local value. It also stops the provider from sending back awareness it just received.
useYdocAndProvider.ts calls guardProviderAwareness on every provider (apps/webapp/src/hooks/useYdocAndProvider.ts:244).
On the server, onAwarenessUpdate counts ws_messages_total{type="awareness"} and ws_awareness_updates_total (apps/hocuspocus.server/src/hocuspocus.server.ts:239-242). onChange records each applied update in the ydoc_update_bytes histogram (:230-232).
A local benchmark on 2026-09-27 suggests that this counter misses part of the cursor traffic (unverified; the benchmark is not in the repo). So compare each window only with the saved baseline from the same query.
Prometheus scrapes every hocuspocus-server replica every 30 s (scripts/observability/prometheus/prometheus.yml). Retention is 15d and 4GB (docker-compose.observability.yml, --storage.tsdb.retention.*).
A deploy reloads open tabs once, when the new service worker takes control (apps/webapp/src/hooks/useServiceWorker.ts, handleControllerChange).
Desired behavior: Two ratios from one query, before and after. This issue records only how they compare.
The query, run in Grafana Explore against Prometheus. Set the time range to the window.
Use this exact query for both windows. Whether the 2026-09-26 figure came from this exact query is unverified. So compare against the saved baseline.
To find the baseline window, graph sum(ws_active_connections) on 2026-09-26. Pick the busy period with the highest peak.
Where to start:apps/webapp/src/hooks/guardProviderAwareness.ts (guardProviderAwareness). apps/hocuspocus.server/src/lib/metrics.ts (wsMessagesTotal, ydocUpdateBytes, wsActiveConnections). apps/hocuspocus.server/src/hocuspocus.server.ts (onAwarenessUpdate, onChange). Line numbers are hints as of 2026-09-28; the agent searches by symbol.
Rules that apply:apps/hocuspocus.server/CLAUDE.md §Production And Docker Compose. This issue also sets two limits: read production numbers from Grafana only, and run no script inside a production container.
Verify:
Run gh run list --workflow prod.docs.plus.yml --limit 5. Confirm the run for 204f8ef62 shows success. The after window must start after that run ends.
Find the baseline window as above. Run the query over it. Save the value and a Grafana snapshot.
Run a session: at least 10 editors with the pad focused, and at least 3 typing for 5 minutes.
Run the query over that window. Post the comment the acceptance criteria describe.
Out of scope
Fixing the ws_messages_total help text, or adding a socket-level frame counter. File those as their own issues if the numbers need them.
Deleting the guards. That waits for fixed upstream releases, as the comment in guardProviderAwareness.ts says.
Parent
#328.
What to build
Every open editor used to send its cursor again after each remote edit, even when the cursor had not moved. On a busy pad on 2026-09-26, production sent many awareness frames per document update. Commit
e84934acdadds two client guards that stop the unchanged and echoed cursor frames. That commit ships in the production deploy of204f8ef62. This issue measures the same ratio in production before and after that deploy, and records how they compare.Acceptance criteria
ws_active_connectionsin private notes, not on this issue.ws_active_connections) is at most half the before value. It gives no absolute numbers.Blocked by
None — can start now.
Agent brief
Type: HITL — only the maintainer can read production Grafana at
https://grafana.docs.plusand run a live multi-editor session. The maintainer runs the queries and posts the result as a comment on this issue. The after window needs a real session with several editors: plan one, or wait for the next busy pad.Category: enhancement
Current behavior:
apps/webapp/src/hooks/guardProviderAwareness.tsskips acursorwrite that equals the current local value. It also stops the provider from sending back awareness it just received.useYdocAndProvider.tscallsguardProviderAwarenesson every provider (apps/webapp/src/hooks/useYdocAndProvider.ts:244).onAwarenessUpdatecountsws_messages_total{type="awareness"}andws_awareness_updates_total(apps/hocuspocus.server/src/hocuspocus.server.ts:239-242).onChangerecords each applied update in theydoc_update_byteshistogram (:230-232).hocuspocus-serverreplica every 30 s (scripts/observability/prometheus/prometheus.yml). Retention is15dand4GB(docker-compose.observability.yml,--storage.tsdb.retention.*).apps/webapp/src/hooks/useServiceWorker.ts,handleControllerChange).Desired behavior: Two ratios from one query, before and after. This issue records only how they compare.
The query, run in Grafana Explore against Prometheus. Set the time range to the window.
Use this exact query for both windows. Whether the 2026-09-26 figure came from this exact query is unverified. So compare against the saved baseline.
To find the baseline window, graph
sum(ws_active_connections)on 2026-09-26. Pick the busy period with the highest peak.Where to start:
apps/webapp/src/hooks/guardProviderAwareness.ts(guardProviderAwareness).apps/hocuspocus.server/src/lib/metrics.ts(wsMessagesTotal,ydocUpdateBytes,wsActiveConnections).apps/hocuspocus.server/src/hocuspocus.server.ts(onAwarenessUpdate,onChange). Line numbers are hints as of 2026-09-28; the agent searches by symbol.Rules that apply:
apps/hocuspocus.server/CLAUDE.md§Production And Docker Compose. This issue also sets two limits: read production numbers from Grafana only, and run no script inside a production container.Verify:
gh run list --workflow prod.docs.plus.yml --limit 5. Confirm the run for204f8ef62showssuccess. The after window must start after that run ends.Out of scope
ws_messages_totalhelp text, or adding a socket-level frame counter. File those as their own issues if the numbers need them.guardProviderAwareness.tssays.