Skip to content

Measure cursor traffic per edit in production after the cursor guards deploy #353

Description

@HMarzban

Parent

#328.

What to build

Every open editor used to send its cursor again after each remote edit, even when the cursor had not moved. On a busy pad on 2026-09-26, production sent many awareness frames per document update. Commit e84934acd adds two client guards that stop the unchanged and echoed cursor frames. That commit ships in the production deploy of 204f8ef62. This issue measures the same ratio in production before and after that deploy, and records how they compare.

Acceptance criteria

  • The baseline ratio for the 2026-09-26 busy window is saved before Prometheus drops it. Prometheus keeps 15 days or 4 GB, whichever runs out first. Save it by 2026-10-10.
  • The same query runs over a production window after the deploy. That window has at least 10 open editors and at least 3 people typing.
  • For each window, the maintainer keeps the start and end time, the ratio, and the peak ws_active_connections in private notes, not on this issue.
  • A comment on this issue says whether the after ratio per connection (ratio divided by peak ws_active_connections) is at most half the before value. It gives no absolute numbers.
  • If the after value is more than half, the comment names the next check to run. One example: confirm that the open tabs run the new webapp build.

Blocked by

None — can start now.

Agent brief

Type: HITL — only the maintainer can read production Grafana at https://grafana.docs.plus and run a live multi-editor session. The maintainer runs the queries and posts the result as a comment on this issue. The after window needs a real session with several editors: plan one, or wait for the next busy pad.

Category: enhancement

Current behavior:

  • apps/webapp/src/hooks/guardProviderAwareness.ts skips a cursor write that equals the current local value. It also stops the provider from sending back awareness it just received.
  • useYdocAndProvider.ts calls guardProviderAwareness on every provider (apps/webapp/src/hooks/useYdocAndProvider.ts:244).
  • The upstream cursor bug is public at Cursor plugin re-sends an unchanged cursor after every remote change (since 3.0.6) ueberdosis/y-tiptap#55.
  • On the server, onAwarenessUpdate counts ws_messages_total{type="awareness"} and ws_awareness_updates_total (apps/hocuspocus.server/src/hocuspocus.server.ts:239-242). onChange records each applied update in the ydoc_update_bytes histogram (:230-232).
  • A local benchmark on 2026-09-27 suggests that this counter misses part of the cursor traffic (unverified; the benchmark is not in the repo). So compare each window only with the saved baseline from the same query.
  • Prometheus scrapes every hocuspocus-server replica every 30 s (scripts/observability/prometheus/prometheus.yml). Retention is 15d and 4GB (docker-compose.observability.yml, --storage.tsdb.retention.*).
  • A deploy reloads open tabs once, when the new service worker takes control (apps/webapp/src/hooks/useServiceWorker.ts, handleControllerChange).

Desired behavior: Two ratios from one query, before and after. This issue records only how they compare.

The query, run in Grafana Explore against Prometheus. Set the time range to the window.

sum(increase(ws_messages_total{type="awareness"}[$__range]))
/
sum(increase(ydoc_update_bytes_count[$__range]))

Use this exact query for both windows. Whether the 2026-09-26 figure came from this exact query is unverified. So compare against the saved baseline.

To find the baseline window, graph sum(ws_active_connections) on 2026-09-26. Pick the busy period with the highest peak.

Where to start: apps/webapp/src/hooks/guardProviderAwareness.ts (guardProviderAwareness). apps/hocuspocus.server/src/lib/metrics.ts (wsMessagesTotal, ydocUpdateBytes, wsActiveConnections). apps/hocuspocus.server/src/hocuspocus.server.ts (onAwarenessUpdate, onChange). Line numbers are hints as of 2026-09-28; the agent searches by symbol.

Rules that apply: apps/hocuspocus.server/CLAUDE.md §Production And Docker Compose. This issue also sets two limits: read production numbers from Grafana only, and run no script inside a production container.

Verify:

  1. Run gh run list --workflow prod.docs.plus.yml --limit 5. Confirm the run for 204f8ef62 shows success. The after window must start after that run ends.
  2. Find the baseline window as above. Run the query over it. Save the value and a Grafana snapshot.
  3. Run a session: at least 10 editors with the pad focused, and at least 3 typing for 5 minutes.
  4. Run the query over that window. Post the comment the acceptance criteria describe.

Out of scope

  • Fixing the ws_messages_total help text, or adding a socket-level frame counter. File those as their own issues if the numbers need them.
  • Deleting the guards. That waits for fixed upstream releases, as the comment in guardProviderAwareness.ts says.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions