Skip to content

Revisit in-page WebMCP tools for browser agents in January 2027 #462

Description

@HMarzban

Summary

WebMCP is a draft browser API. A page registers tools, and an AI agent running inside the user's browser tab calls them, in the user's own session. docs.plus already has an MCP connector for claude.ai, ChatGPT and Claude Code. This RFC asks whether the pad should also register a small set of read-only tools for browser agents.

Parent: #230.

Maintainer ruling, 2026-10-07

  • Wait. Do not build the in-page tools now.
  • Revisit in January 2027. Whoever picks this up then: re-check the spec, the Chrome and Edge origin trials, and whether Gemini in Chrome or another agent calls WebMCP. If a major agent ships support, build the four-tool read-only slice described above. Post the findings here first.

The slice is under "Proposed first slice" below.

Build trigger

Build only when the January check finds that Gemini in Chrome or another major agent ships WebMCP support. The ChatGPT desktop browser was already confirmed (see Baseline) when the maintainer ruled "Wait". If its reach has grown past Work and Codex, say so in the findings.

Baseline (checked 2026-10-07)

Spec and API

  • Status: a W3C Web Machine Learning Community Group draft, not on the W3C Standards Track. Latest draft: 2 October 2026 (spec, repo).
  • It does not use the MCP wire protocol. It shares the vocabulary only (tools, input schemas).
  • The API changed during 2026:
    • navigator.modelContext moved to document.modelContext.
    • provideContext, clearContext and unregisterTool were removed. A tool is unregistered by aborting the AbortSignal passed to registerTool.
    • registerTool now returns a Promise. A duplicate name rejects with InvalidStateError.
  • Tool annotations include readOnlyHint, untrustedContentHint and consequentialHint.

Browsers

Agents that call it today

  • Confirmed: the ChatGPT desktop app's built-in browser ("Site tools"). It is available to ChatGPT Work and Codex, and supports only a subset of the API (learn.chatgpt.com/docs/webmcp).
  • Experimental: Brave Leo.
  • A developer bridge: chrome-devtools-mcp with --categoryExperimentalWebmcp=true exposes list_webmcp_tools and execute_webmcp_tool to Claude Code and other MCP clients.
  • Not confirmed: Gemini in Chrome ("will soon support", Chrome at I/O, 2026-05-19), Claude in Chrome, Edge Copilot.

Parity (AGENTS.md: Google Docs, Word, Notion)

  • None of the three has announced WebMCP tools.
  • Google Docs ships a remote MCP server instead (docsmcp.googleapis.com), which is the path docs.plus already took.

Proposed first slice (build only if the trigger is met)

Four tools. They read or navigate only, so they pass every existing permission gate unchanged.

Tool Same as the MCP connector tool? Existing code it calls
get_outline Same name and purpose. No slug, no rev. A walk over the top-level children of editor.state.doc only, like the connector's buildOutline (apps/hocuspocus.server/src/modules/mcp/domain/outline.ts:16). Keep every heading with a toc-id, the Title included. Skip a repeated toc-id, as buildOutline does. Then buildNestedToc (apps/webapp/src/components/toc/utils.ts:7).
read_document Same name and purpose; one section only, with section_id Resolve section_id through the same top-level walk, which gives the heading pos and child index. The private computeOwnRange (apps/webapp/src/components/TipTap/extensions/shared/match-section.ts:86) ends a section at the next heading of any level. That is the rule of the server's sectionAt (apps/hocuspocus.server/src/modules/document-content/domain/sections.ts:49-53). Export it and add it to shared/index.ts; do not write a second copy. Serialize with editor.markdown.serialize({ type: 'doc', content: doc.slice(from, to).content.toJSON() }), with from and to from computeOwnRange (@tiptap/markdown, loaded at TipTap/TipTap.tsx:21, :143).
get_selection In-page only resolveHeadingIdForDocPos (services/commentAnchor.ts:16), doc.textBetween, useFocusedHeadingStore.getState()
go_to_heading In-page only tocActions.navigateToHeading(id, { openChat: false }) (components/toc/hooks/tocActions.ts:58), the same call as the tick rail (components/toc/TocTickRail.tsx:746). It reads the live editor when it runs.

Do not use collectHeadings or findAllSections for the walk:

  • collectHeadings walks every depth. It would list headings inside a blockquote or a table cell, which own no section. It also carries the hyperlink picker's kind: 'heading' contract.
  • findAllSections skips the Title (child 0).

Where the code goes

  • apps/webapp/src/components/pages/document/hooks/useWebMcpTools.ts, one file. It checks for document.modelContext, registers the four tools with one AbortController, and aborts it in the effect cleanup. The cleanup also covers React StrictMode's double mount, where a duplicate name would reject.
  • No import() split. Every helper the tools call already ships in the pad bundle. A lazy chunk would hold only the four tool objects (read from code, not measured). Registering in the same tick also avoids a cleanup that runs before a lazy chunk resolves.
  • Mount: components/pages/document/layouts/PadEditorLifecycle.tsx, next to useEditorAndProvider. DocumentLayouts.tsx:31 mounts it only when !isHistory, above both the desktop and the mobile layout. /editor renders EditorPlayground, so it registers nothing.

Rules for every tool

  • Read the editor when the tool runs (useStore.getState().settings.editor.instance, check isDestroyed), never when it registers.
  • While settings.editor.providerSyncing is true, or with no editor, return { error: 'not-ready' }. Before sync, the editor holds only the local copy.
  • Mark read tools readOnlyHint: true, untrustedContentHint: true, because document text is user content.
  • Keep each result small. Chrome's WebMCP guidance advises about 1,500 characters per output (secure tools; not measured).
  • read_document returns { markdown, truncated } and cuts markdown at 1,500 characters. An unknown section_id returns { error: 'not-found' }.
  • Never call .focus(), because on iOS it raises the keyboard.

Out of scope

  • Write tools and chat tools. The MCP connector covers both. A write tool needs its own issue.
  • The declarative form API (still a TODO in the spec).

Acceptance criteria

January 2027 revisit:

  • A comment on this issue updates each Baseline line, with the new check date.
  • That comment says build or wait, and names the agent that meets the build trigger, if any. It answers the Open decisions.

Only if the trigger is met:

  • On a pad, the agent lists get_outline, read_document, get_selection and go_to_heading.
  • On /, on /editor and in History view, it lists no tools.
  • get_outline lists only top-level headings, and read_document resolves every id it lists.
  • While settings.editor.providerSyncing is true, every tool returns { error: 'not-ready' }, not stale text.
  • In dev (StrictMode), the console shows no InvalidStateError from registerTool.
  • go_to_heading on a phone does not raise the keyboard.

Verify (if built)

  1. In Chrome, turn on chrome://flags/#enable-webmcp-testing.
  2. Add the bridge to Claude Code: claude mcp add chrome-devtools -- bunx chrome-devtools-mcp@latest --categoryExperimentalWebmcp=true.
  3. Run list_webmcp_tools on a pad: four tools. On /, on /editor, and in History view: no tools.
  4. Run each tool with execute_webmcp_tool, and check the results against "Rules for every tool".

Open decisions (answer in the January findings comment)

  1. If built: join the Chrome origin trial on docs.plus? Recommended: yes, only when the slice ships, and record the token's end date. Unknown: whether the ChatGPT desktop browser needs the token at all.
  2. Reuse the connector tool names get_outline and read_document? Recommended: yes. The purpose matches. The output does not: the in-page tools return no rev and no [n] block numbers. Media also differs. The connector shows placeholders (redactMedia). The editor serializer writes media Markdown (extensions/extension-hypermultimedia/src/markdown/typedMediaMarkdown.ts). An agent that holds both must not mix their results.
  3. Does the parity rule (AGENTS.md:148) bind an integration surface? Recommended: no. The rule targets interaction behaviour. WebMCP, like the MCP connector, is an integration surface, and none of the three products ships WebMCP tools.

Activity

  1. HMarzban commented on Oct 7, 2026

    @HMarzban
    CollaboratorAuthor

    Maintainer ruling, 2026-10-07

    • Wait. Do not build the in-page tools now.
    • Revisit in January 2027. Whoever picks this up then: re-check the spec, the Chrome and Edge origin trials, and whether Gemini in Chrome or another agent calls WebMCP. If a major agent ships support, build the four-tool read-only slice described above. Post the findings here first.
  2. HMarzban commented on Oct 7, 2026

    @HMarzban
    CollaboratorAuthor

    Maintainer ruling update, 2026-10-07

    If the slice is built later:

    1. Join the Chrome origin trial on docs.plus when it ships, and record the token's end date.
    2. Reuse the connector tool names get_outline and read_document.
    3. The parity rule in AGENTS.md (follow Google Docs, Word and Notion) does not bind WebMCP. It is an integration surface, not an editor behaviour.

    A one-time check is scheduled for 2027-01-11. It will post its findings here and ask for a build-or-wait ruling.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    EditorTiptap & ProsemirrorIdeaenhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions