Skip to content

RFC: carry media node options through Markdown #244

Description

@HMarzban

Verdict

Give every media node a way to carry its options through Markdown, so a person or an agent can write a sized, floated, captioned embed by hand.

Two shapes, both optional:

  • Caption rides the standard CommonMark title slot: ![image](url "My caption").
  • Layout rides an option block after the closing parenthesis: ![youtube](url){width=800 float=right}.

Together:

![youtube](https://www.youtube.com/watch?v=dQw4w9WgXcQ "How docs.plus works"){width=800 float=right}

Emit only what differs from the default. A node nobody styled still exports as ![youtube](url). Most documents will not change shape at all.

Two things are decided here, and the grammar is the smaller one.

The caption also needs its HTML leg back. A user captions a video, copies it, and the caption is gone. Eight of nine nodes never emit a <figcaption>, so every HTML path drops it — clipboard copy included. That is one call per node and no new syntax, and it is the first slice.

The Markdown lane leaves the shared embed degradation, while DOCX and ODT keep it. Without that split the grammar changes nothing anyone would see.

What is broken today

Measured, not read. Each number comes from driving the built package or the product path, not from reading the source.

Media options do not survive an export. Set every node to width=800 height=450 float=right margin=… display=inline-block justifyContent=center plus a caption, then serialize. Four of sixty attribute slots survive. Only video and audio keep anything, and only width and height.

The product export never reaches the media hooks at all. importMarkdown builds the right node. exportMarkdown writes a bare link:

Input Node built on import Markdown written on export
![video](…/a.mp4) video [https://…/a.mp4](https://…/a.mp4)
![youtube](…) youtube [https://…](https://…)
![spotify](…) spotify [https://…](https://…)
![audio](…/a.mp3 width=450 height=120) audio, both set [https://…/a.mp3](https://…/a.mp3)

vimeo, soundcloud, x and loom behave the same way.

This is deliberate, and the comment says why. portableJson.ts:10-13:

Every embed is an iframe or a player that DOCX, Markdown and ODT cannot express. The src is the only part a reader can still follow.

toPortableJson rewrites all eight EMBED_NODE_TYPES into a hyperlink paragraph before serialization. So the eight renderMarkdown hooks in @docs.plus/extension-hypermultimedia are unreachable from the product export. They run in the clean-room Cypress suite and nowhere else. image is excluded on purpose and does export as a picture.

The syntax that exists today is not portable Markdown. typedMediaMarkdown.ts:79 writes size inside the destination parentheses: ![video](url width=640 height=360). CommonMark reads what follows the destination as a title, and a title must be quoted. Stock marked returns that input as literal text, not an image. It works inside docs.plus only because each node registers an inline tokenizer that runs before the standard image rule.

The caption is the <figcaption> a user types on the node. It is editable in the node view, it lives in the caption attribute, and that attribute is the source of truth. It persists through collaboration and JSON. All nine nodes carry it (captionAttribute(), caption.ts:116-131).

It survives almost no serialization. Measured on all nine nodes:

Leg Result
Stored attribute kept on all nine
HTML <figcaption> kept on image only
Markdown lost on all nine, image included
HTML written, then re-parsed kept on image only

Three lines explain it:

  • captionAttribute().renderHTML returns {} (caption.ts:130), so the caption never becomes an HTML attribute. It can only ride a <figure> wrapper.
  • wrapRenderWithCaption (caption.ts:132-143) has exactly one caller: image.ts:178.
  • parseHTML reads a caption only when element.tagName === 'FIGURE' (caption.ts:125-128), so parsing is symmetric with rendering.

The HTML leg is disclosed in the published 2.0.0 CHANGELOG — "HTML serialization carries the caption for image only — which includes clipboard copy/paste and the toolbar Copy action". It is stated as a limit, with no reason recorded. So a user who captions a video and copies it loses the caption, and that is a more everyday path than an export.

image also carries an unused title attribute. ![alt](url "A caption") parses the title and discards it.

![video](src) does not round-trip. It imports, picks up the schema defaults width 640 and height 480, then exports as ![video](src width=640 height=480).

Decision

Caption gets both of its missing legs

The caption is the editable <figcaption> on the node. Fixing it is two independent changes, and neither needs a new grammar.

HTML. Call wrapRenderWithCaption from the other eight nodes, not from image alone. parseHTML already reads a <figure> on every node, so the read side needs nothing. This alone fixes clipboard copy and the toolbar Copy action.

The open question is the wrapper's effect on an iframe embed. image.ts:178 moves the block layout style onto the <figure> when a caption is present, and each embed node pins its own height. Check soundcloud and spotify first — both floor their height and neither scales with the shell.

Markdown. ![image](url "My caption") maps to the caption attribute, both directions.

Valid CommonMark. Every other Markdown tool already renders the title slot as a caption or a tooltip. The attribute already exists on all nine nodes, so nothing new is stored.

Layout uses the option block

{key=value key=value} immediately after the closing parenthesis. Whitespace-separated. Quotes allowed around a value that holds a space.

The vocabulary is the one the code already defines at embedKit.ts:37-45:

width · height · display · float · clear · margin · justifyContent

An unknown key is dropped, not stored. A value that fails its type check is dropped, and the node keeps its default.

Why this shape, on evidence

An earlier draft of this RFC claimed the brace block "follows the Pandoc convention, so a model already writes it". That was asserted, not checked. It is now checked, against vendor documentation and against GitHub's live renderer.

There is no single industry convention. The leaders disagree:

Convention Platforms
Brace attribute block GitLab (native), Pandoc, Quarto, Material for MkDocs, MyST, kramdown, markdown-it-attrs
Raw HTML <img> only GitHub, Docusaurus, MDX, Next.js, Astro, Joplin, Typora, GitBook, Mintlify, Google's style guide
Pipe inside the alt text Obsidian
JSX component MDX, Astro
Nothing at all Notion, Bear, Craft, Microsoft Learn

CommonMark itself defines no attribute syntax. Its generic-attributes proposal was never adopted.

GitLab already ships this exact shape, which is the strongest single precedent:

![GitLab logo](img/markdown_logo_v17_11.png "Title Text"){width=100 height=100px}

A title slot and a brace block, together, on the same image, in a shipping product.

Measured through GitHub's /markdown API, which is the largest renderer we must degrade well on:

Input GitHub output
![video](url width=800 height=600) — what we ship today ![video](<a href=…>…</a> width=800 height=600) — the image construct collapses and the reader sees literal ![video](
![image](url){width=800 float=right} image renders; {width=800 float=right} prints as trailing text
![youtube](url "My caption"){width=800 float=right} image renders, alt="youtube" and title="My caption" both intact; braces print as text
![youtube|800](url) — Obsidian form renders cleanly with no stray text, but the pipe is swallowed into alt

So the shape we ship today is the one that breaks. The proposal degrades to visible-but-harmless text, and keeps the caption.

Three honest costs, stated plainly:

  1. The brace family is not one uniform key set. kramdown needs {:width="800"}, with a colon and quoted values. MyST writes w=100px, not width=. Only GitLab and Pandoc accept the bare {width=800} form.
  2. Only width and height mean anything outside docs.plus. Pandoc turns float=right into data-float="right". The other five keys are ours alone.
  3. The Obsidian pipe degrades better than anything else. It renders with zero stray characters. It loses because Obsidian documents only |WIDTH and |WIDTHxHEIGHT, so it cannot carry float, margin, a caption, or later player options.

Implementation cost does not favour any candidate. @tiptap/markdown 3.22.3 runs on marked and documents no image attribute convention at all. Every candidate needs a custom inline tokenizer. So the choice is on merit, not on effort.

Accept it from foreign Markdown

The block is read on import from any Markdown, not only from our own output. That is the point of the task. An agent writes the file, docs.plus reads it.

The isValid*Url validators keep running on src, unchanged, whatever the block carries.

Emit only what differs

An attribute equal to the node default is not written. A styled node writes only the keys the user changed.

This is what keeps an exported document readable, and it is the reason the grammar is safe to add. It also needs one explicit rule, because today's behaviour is ambiguous: compare against the schema default, not against "was it set". So a video at 640×480 writes nothing, and ![video](src) starts round-tripping unchanged.

Split the Markdown lane out of embed degradation

toPortableJson serves three lanes:

  • markdownExport.ts:30
  • odtExport.ts:343
  • documentHtml.ts:44, which feeds DOCX

Markdown leaves. DOCX and ODT stay. Markdown can express an embed once it has a syntax; a .docx and a .odt still cannot hold an iframe.

Without this split the grammar changes nothing a user or an agent would ever see. This is the reviewable part of the proposal.

Learning curve

The target is close to zero for an end user. Two facts decide whether that is reachable.

No convention is universal, so no syntax is already known to everyone. A GitHub user has never seen a brace block. A GitLab or Pandoc user has. The honest claim is not "everyone knows this"; it is "more vendors publish this than any other form, and it explains itself when read".

The end user almost never writes it. They use the toolbar. Markdown reaches them at three moments: they read an exported file, they hand-write a size once in a while, or an agent writes a file for them. So the curve is set by three properties, in this order:

  1. The common case needs no syntax. ![youtube](url) stays exactly that. This is what "emit only what differs" buys, and it is the largest single contributor.
  2. The keys explain themselves. width=800 float=right needs no lookup. |800 and ?w=800 both need one.
  3. A wrong guess degrades safely. Measured above: a foreign renderer still shows the media and prints the unknown keys as text.

Documentation is part of this work, not a follow-up. Nothing in docs/ teaches the Markdown media syntax today. docs/api/ holds README.md, authentication.md, quickstart.md and websocket.md, and none of them mentions Markdown. So the ![youtube](url) form we already ship is documented only in an npm README that an end user never opens. A syntax with a good curve and no page still costs a lookup that fails.

Alternatives considered

Query parameters on the src — ![youtube](url?w=800&float=right). No new grammar at all. Rejected: it corrupts the URL, it breaks the isValid*Url validators, and the parameters travel to the third-party host.

Inline HTML for styled media — emit <img width=… style="float:right"> when a value differs. Renders correctly in GitHub and Pandoc. Rejected: it stops being Markdown, many renderers strip it, and no HTML element renders a YouTube embed anyway.

Keep the current in-parenthesis syntax and extend it — cheapest to build. Rejected on measurement: GitHub's renderer returns it as literal text, so every file we write stays broken outside docs.plus. That defeats the goal.

The Obsidian pipe — ![youtube|800](url). It degrades better than every other candidate, with zero stray characters on GitHub. Rejected: Obsidian documents only |WIDTH and |WIDTHxHEIGHT, so it cannot carry float, margin, a caption, or player options. It also collides with the node type already in the alt slot.

Constraints the design must respect

Registration order decides inline tokens. In @tiptap/markdown 3.22.3, block tokens try the next handler when one yields nothing (MarkdownManager.ts:405). Inline and mark tokens take markdownHandlers[0] with no fallback (:247, :727). A losing extension cannot yield by returning nothing. image is an inline token, so the media hooks must win by order.

A backslash escape is destroyed on import. a \* literal star imports as a literal star. Both characters vanish. Any option value that holds a Markdown character is unsafe until this is fixed, so it lands in the same work.

marked keeps its defaults. No call site passes markedOptions, so breaks stays false.

One shared MarkdownManager. Constructing one mutates a process-global marked singleton and its tokenizers accumulate. Both lanes share a lazy instance for that reason (markdownExport.ts:21-27).

The import cap is 64 KB. MAX_MARKDOWN_CHARS is 64 * 1024, and the route answers 413 above it. An option block makes every media line longer. If the cap moves, the user-facing copy at apps/webapp/src/api/documents/conversionErrors.ts:9-11 moves in the same edit.

A URL-valued option needs the same gate as src. video.poster is the case that exists today. Route it through isSafeMediaSrc, not through the plain value parser.

Two media-import defects this work should fix

Both are attribute defects on the Markdown import path, so they belong here rather than in a separate issue.

  • Spotify import ignores the canonical URL and the default height. An imported track renders 352 px tall instead of 152 px, so the same track looks different depending on whether it arrived by paste or by import.
  • X import stores an un-normalized src, unlike every other X write path. Nothing breaks visually, because the read side normalizes again, but two identical embeds hold different stored values.

Suggested slices

  1. Caption in HTML for the other eight nodes. One call to wrapRenderWithCaption per node. Fixes clipboard copy and the toolbar Copy action. No new syntax, no schema change, and the read side already works.
  2. Caption through the Markdown title slot, all nine nodes, both directions.
  3. Split the Markdown lane out of toPortableJson, so the eight embed hooks become reachable. No new syntax; the existing hooks start running.
  4. Option block for the seven layout keys, emit-only-what-differs, in the shared createTypedMediaMarkdownHooks factory and in the image node.
  5. Per-node player options under the same grammar — controls, autoplay, loop, muted, preload, poster, and the x set. Same parser, larger vocabulary.
  6. A documentation page under docs/ that teaches the syntax. Nothing there teaches it today, so this is not optional polish.

Slice 1 is the smallest change with the largest user-visible win, and it does not touch Markdown at all. Slices 1, 2 and 3 are independent of each other.

Breaking change

The five packages are published at 2.0.0.

Slice 3 changes what renderMarkdown writes for audio and video: ![video](url width=800) becomes ![video](url){width=800}. An external consumer that reads our Markdown with its own parser breaks. Nothing else does — import keeps accepting the old shape.

That reads as a minor bump with a migration note, not a major, because the old input still parses. Worth a ruling before slice 3 lands.

Open questions

  1. justifyContent or justify-content on the wire? Six of seven keys are identical in both spellings. Only this one differs. Kebab-case reads as CSS and is easier to guess; camelCase matches the stored attribute exactly. Recommendation: write kebab-case, accept both on import.
  2. Does image join the same tokenizer? It uses the standard image token today, not the hm_* family. Giving it an option block means taking that token by registration order.
  3. Minor or major version for slice 3? See above.

Not this issue

Related work found by the same audit, filed separately: the pad turns a pasted Markdown link into a link mark instead of a hyperlink mark; a select-all Markdown paste can degrade to literal text; superscript and subscript are dropped on export and API.md does not say so; each webapp editor re-registers twelve tokenizers into the global marked.

How this was measured

  • All five clean-room suites green, exit 0: 426 Cypress specs and 22 Jest tests. extension-hypermultimedia 171/171, including markdown/markdown-round-trip.cy.ts at 26/26.
  • Attribute loss measured by driving the built dist/index.js through a real Editor with StarterKit and @tiptap/markdown 3.22.3.
  • Export behaviour measured by driving importMarkdown and exportMarkdown from apps/hocuspocus.server/src/modules/document-conversion/domain/ directly.
  • 38 Markdown constructs round-tripped through that same product path: 21 unchanged, 9 normalized acceptably, 3 lossy.
  • Caption measured on all nine nodes across four legs: stored attribute, getHTML(), Markdown serialize, and HTML written then re-parsed.
  • Convention survey across 30+ platforms, primary vendor documentation only, then reproduced independently through GitHub's POST /markdown API for the four candidate shapes.

Activity

  1. HMarzban commented on Sep 6, 2026

    @HMarzban
    CollaboratorAuthor

    Scope correction after the maintainer pointed at README section Caption.

    The caption in this RFC is the editable figcaption on the media node, stored in the caption attribute. That attribute is the source of truth and persists through collaboration and JSON.

    I first wrote that Markdown drops it. Measured on all nine nodes, it is worse than that:

    • stored attribute: kept on all nine
    • HTML figcaption: kept on image only
    • Markdown: lost on all nine, image included
    • HTML written then re-parsed: kept on image only

    Mechanism: captionAttribute().renderHTML returns an empty object (caption.ts:130), so the caption can only ride a figure wrapper, and wrapRenderWithCaption (caption.ts:132-143) has one caller, image.ts:178.

    The HTML leg is disclosed in the published 2.0.0 CHANGELOG as a limit, with no reason recorded. It hits clipboard copy and the toolbar Copy action, so it is more everyday than an export.

    The body now carries this, and the HTML leg is slice 1 because it needs one call per node and no new syntax.

  2. HMarzban commented on Sep 6, 2026

    @HMarzban
    CollaboratorAuthor

    Convention check, because the first draft asserted one without a source.

    I wrote that the brace block follows the Pandoc convention so a model already writes it. I checked nothing. It is now checked against vendor documentation for 30+ platforms, and reproduced through GitHub's POST /markdown API.

    There is no single industry convention. GitHub is HTML-only. GitLab ships a native brace block. Obsidian uses a pipe in the alt text. Docusaurus, MDX, Next.js and Astro use JSX. Notion, Bear and Craft support nothing.

    The strongest precedent is GitLab, which already ships the exact hybrid this RFC proposes: a title slot and a brace block on one image.

    Measured on GitHub's own renderer:

    • what we ship today, the in-parens form, comes back as literal text with the image construct collapsed
    • the brace form renders the image and prints the braces as trailing text
    • the hybrid renders the image and keeps alt and title intact
    • the Obsidian pipe renders with zero stray text, but cannot carry float, margin or a caption

    Three costs are now stated in the body. The brace family is not one uniform key set, since kramdown needs a colon and quoted values and MyST writes w= instead of width=. Only width and height mean anything outside docs.plus. And implementation cost is equal across every candidate, because @tiptap/markdown documents no image attribute convention at all, so all of them need a custom tokenizer.

    A Learning curve section is now in the body. It says plainly that no syntax is already known to everyone, and that the curve is carried mainly by the common case needing no syntax at all.

    One gap the survey exposed: nothing under docs/ teaches the Markdown media syntax today. A documentation page is now slice 6, not a follow-up.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions