Skip to content

fix(muyajs): allow CJK ideographs as flanking boundary for emphasis - #4355

Merged
Jocs merged 3 commits into
developfrom
fix/4307-cjk-emphasis-flanking
Jun 3, 2026
Merged

Jocs merged 3 commits into
developfrom
fix/4307-cjk-emphasis-flanking

Conversation

@Jocs

@Jocs Jocs commented Jun 3, 2026 •

Copy link
Copy Markdown
Member

Summary

Closes #4307.

**…** adjacent to a CJK ideograph + inner punctuation (e.g. 中文**"加粗"**中文, 日本語**(強調)**日本語, 한국어**[강조]**한국어) fails to render as strong in WYSIWYG. The same input has been broken since at least v0.17.1 (the parser code was byte-identical until this PR), and GitHub/GFM renders the same way.

⚠️ Standards compliance — this PR ships a non-standard extension

This is not a CommonMark conformance fix. It is a deliberate divergence from CommonMark spec 6.2 ("Emphasis and strong emphasis", https://spec.commonmark.org/0.31.2/#emphasis-and-strong-emphasis):

  • Spec 6.2 only counts Unicode whitespace and Unicode punctuation as flanking boundaries.
  • CJK Unified Ideographs are Unicode category Lo — neither whitespace nor punctuation.
  • Under a literal reading of the spec, 中文**"加粗"**中文 therefore MUST NOT open a strong run. GitHub/GFM is strictly spec-conformant and behaves the same way.

CJK scripts do not use spaces between words, so the strict rule effectively denies emphasis to any CJK paragraph that wraps the ** run with punctuation (quotes, parentheses, brackets, 、 …). This PR adopts the well-known CJK-friendly emphasis extension already shipped by Typora, VSCode markdownlint, Joplin, and most CJK-oriented Markdown tools: treat CJK ideographs as boundary-equivalent in clause (2b) of the flanking check.

Why this is safe:

  • The widening is additive only (||): it never rejects emphasis CommonMark accepts, it only accepts emphasis CommonMark rejects on the CJK-boundary edge case.
  • All ASCII / English inputs continue to parse byte-identically — verified by 3 regression cases in the new vitest and by the full 556-test unit suite remaining green.
  • Documented in the source inline (packages/muyajs/lib/parser/utils.js) with explicit "NON-STANDARD EXTENSION" header so future readers / spec auditors can see the divergence at a glance.

Implementation

  • packages/muyajs/lib/parser/utils.js — introduce CJK_REG (CJK Unified Ideographs BMP + Ext-A + Ext-B via surrogate pairs + Compatibility Ideographs, Hiragana, Katakana, Halfwidth Katakana, Hangul Syllables) and extend the (2b) clauses of canOpenEmphasis / canCloseEmphasis with a three-way ||.
  • packages/desktop/src/types/muya.d.ts — small ambient shim for muya/lib/parser so the new vitest's named-import passes vue-tsc, consistent with the existing muya bridges.

Tests

  • packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts — token-level vitest: 3 CJK + inner-punctuation cases (fail on develop, pass with this PR) + 3 regression cases (already pass on develop, still pass).
  • packages/desktop/test/e2e/strong-cjk.spec.ts — DOM-level Playwright spec: confirms <strong> is emitted in the live WYSIWYG render.

Test plan

  • pnpm test:unit — 556/556 pass (10 files)
  • pnpm -C packages/desktop exec playwright test test/e2e/ — 75 pass / 2 skip / 0 fail (incl. 2 new specs)
  • pnpm run lint — 0 errors
  • pnpm run typecheck — pass
  • Verified red-first: new specs fail on develop before the parser change, then pass after.
  • Manually verified the same inputs still fail in v0.17.1 (parser file is byte-identical to pre-fix develop), confirming this is a pre-existing bug, not a 0.19 regression.

🤖 Generated with Claude Code

Closes #4307.

Strict CommonMark 6.2 flanking rules treat a CJK Unified Ideograph as
neither Unicode whitespace nor Unicode punctuation, so `中文**"加粗"**中文`
(and the symmetric Japanese / Korean variants) failed clause (2b) of
canOpenEmphasis / canCloseEmphasis and never opened a strong run. Mirror
the well-known CJK-friendly emphasis extension shipped by Typora, VSCode
markdownlint, etc.: treat CJK Unified Ideographs (BMP + Ext-A + Ext-B
via surrogate pairs + Compatibility Ideographs), Hiragana, Katakana,
Halfwidth Katakana and Hangul Syllables as boundary-equivalent in
clause (2b).

Adds:
- packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts —
  token-level vitest covering three CJK + inner-punctuation cases plus
  three regression cases that already worked on develop.
- packages/desktop/test/e2e/strong-cjk.spec.ts — Playwright spec
  confirming `<strong>` is emitted in the live WYSIWYG DOM.
- muya/lib/parser shim in src/types/muya.d.ts so the new vitest's
  named-import passes vue-tsc.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
@github-actions

github-actions Bot commented Jun 3, 2026 •

Copy link
Copy Markdown

Website preview ready: https://pr-4355-marktext-website.ransixi.workers.dev

Built from f6a53aa55d9694660eb95de2013a9cf469430302 · Worker: marktext-website · Version: e59a4310-61ac-4a67-8858-934db2be7ac7 · Alias: pr-4355

@github-actions

github-actions Bot commented Jun 3, 2026 •

Copy link
Copy Markdown

Build artifacts for PR #4355:

Run: https://github.com/marktext/marktext/actions/runs/26866178020

Artifact Size Link
marktext-windows-arm64 270.6 MB Download
marktext-linux 592.9 MB Download
marktext-macos-x64 271.7 MB Download
marktext-windows-x64 271.9 MB Download
marktext-macos-arm64 261.4 MB Download

Strengthen the inline comments in parser/utils.js so it is unambiguous
that the CJK_REG widening of the flanking check is a *deliberate
divergence* from CommonMark spec 6.2, not a spec compliance fix. Note
that GitHub/GFM is strictly conformant by rejecting the same inputs,
and that the extension is additive only (it never rejects emphasis
CommonMark accepts).

No behavioral change.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR updates the legacy MuyaJS Markdown emphasis parsing so **...** can be recognized as strong emphasis when adjacent to CJK text plus inner punctuation (a deliberate, documented non-CommonMark extension), addressing issue #4307 in the WYSIWYG editor pipeline.

Changes:

  • Extend the emphasis “flanking” boundary checks in canOpenEmphasis / canCloseEmphasis to treat CJK characters as boundary-equivalent (additive widening).
  • Add a unit-level tokenizer regression suite covering CJK + inner-punctuation cases and existing “sanity” cases.
  • Add an Electron Playwright e2e regression asserting <strong> is emitted in the rendered editor DOM and add a small TS ambient module shim for muya/lib/parser.

Reviewed changes

Copilot reviewed 3 out of 4 changed files in this pull request and generated 1 comment.

File Description
packages/muyajs/lib/parser/utils.js Widens emphasis boundary rules to treat CJK characters as valid flanking boundaries (non-standard extension).
packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts Adds token-level unit tests ensuring strong is produced in CJK + inner-punctuation scenarios.
packages/desktop/test/e2e/strong-cjk.spec.ts Adds DOM-level e2e coverage verifying <strong> appears in WYSIWYG rendering for the reported cases.
packages/desktop/src/types/muya.d.ts Adds ambient declarations so named imports from muya/lib/parser typecheck in the new unit test.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +151 to 155
// Non-standard CJK widening (see CJK_REG block above) — additive only:
// CJK ideographs are accepted as "preceded by" boundary, on top of the
// CommonMark-defined whitespace/punctuation set.
if (PUNCTUATION_REG.test(followedChar) && !(UNICODE_WHITESPACE_REG.test(precededChar) || PUNCTUATION_REG.test(precededChar) || CJK_REG.test(precededChar))) {
return false

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — verified and addressed in f6a53aa.

charAt / src[i] only return one UTF-16 code unit, so the surrogate-pair branch of CJK_REG (and, I now realize, the long list of surrogate-pair alternations in the existing CommonMark PUNCTUATION_REG) was effectively dead. Both are now hit because every flanking-boundary read goes through new lastCodePointChar / codePointCharAt helpers that return 1 or 2 code units as appropriate.

Added a vitest case 𠀀𠀁**"加粗"**𠀀𠀁 (CJK Ext-B, U+20000/U+20001) to lock the surrogate-pair behavior down. Full suite still green (557 unit / 75 e2e).

Address Copilot inline review on PR #4355: CJK_REG contains a
surrogate-pair branch ([\uD840-\uD87F][\uDC00-\uDFFF]) intended to
match non-BMP CJK ideographs (CJK Ext-B onward), but the existing
flanking-boundary readers used charAt / src[i] — single UTF-16 code
unit access — so the regex branch could never fire in practice.

Add lastCodePointChar / codePointCharAt helpers that read a full
Unicode code point (1 or 2 UTF-16 code units) and route every flanking
read in canOpenEmphasis / canCloseEmphasis through them.

Side benefit: PUNCTUATION_REG also contains numerous surrogate-pair
alternations (\uD800[\uDD00-\uDD02\uDF9F\uDFD0] and similar) for
non-BMP Unicode punctuation that were dead for the same reason. Those
spec-defined cases are now correctly tested too.

Adds an extra vitest case using 𠀀 / 𠀁 (CJK Ext-B, U+20000/U+20001)
to lock in the surrogate-pair behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
@Jocs
Jocs merged commit 7ca908b into develop Jun 3, 2026
13 checks passed
@Jocs
Jocs deleted the fix/4307-cjk-emphasis-flanking branch June 3, 2026 05:59
codeforhome added a commit to codeforhome/marktext that referenced this pull request Jun 3, 2026
Merges 37 upstream commits into the fork, including the pnpm monorepo
conversion (marktext#4302), muya TS rewrite migration (marktext#4314), website package,
and fixes through CJK emphasis boundary (marktext#4355).

Conflict resolutions:
- package.json: adopt upstream monorepo root shell; drop fork's
  root-level deps block; move 4 fork-added deps (diff, @types/diff,
  markdown-it, @types/markdown-it) to packages/desktop/package.json
- build.yml: keep both fork's workflow_dispatch+dynamic-matrix and
  upstream's paths-ignore: packages/muya/**
- wsl.ts, comparisonPane.vue, wsl-paths.spec.ts: accepted git's rename
  detection — files live under packages/desktop/ in the monorepo layout
- docs/dev/FORK_PLAN.md: kept at repo-root docs/dev/ (private fork doc,
  not published website content)
- pnpm-lock.yaml: regenerated via pnpm install
thimbleberrysystems pushed a commit to thimbleberrysystems/WordBird that referenced this pull request Jun 21, 2026
…arktext#4355)

* fix(muyajs): allow CJK ideographs as flanking boundary for emphasis

Closes marktext#4307.

Strict CommonMark 6.2 flanking rules treat a CJK Unified Ideograph as
neither Unicode whitespace nor Unicode punctuation, so `中文**"加粗"**中文`
(and the symmetric Japanese / Korean variants) failed clause (2b) of
canOpenEmphasis / canCloseEmphasis and never opened a strong run. Mirror
the well-known CJK-friendly emphasis extension shipped by Typora, VSCode
markdownlint, etc.: treat CJK Unified Ideographs (BMP + Ext-A + Ext-B
via surrogate pairs + Compatibility Ideographs), Hiragana, Katakana,
Halfwidth Katakana and Hangul Syllables as boundary-equivalent in
clause (2b).

Adds:
- packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts —
  token-level vitest covering three CJK + inner-punctuation cases plus
  three regression cases that already worked on develop.
- packages/desktop/test/e2e/strong-cjk.spec.ts — Playwright spec
  confirming `<strong>` is emitted in the live WYSIWYG DOM.
- muya/lib/parser shim in src/types/muya.d.ts so the new vitest's
  named-import passes vue-tsc.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

* docs(muyajs): mark CJK emphasis widening as a non-standard extension

Strengthen the inline comments in parser/utils.js so it is unambiguous
that the CJK_REG widening of the flanking check is a *deliberate
divergence* from CommonMark spec 6.2, not a spec compliance fix. Note
that GitHub/GFM is strictly conformant by rejecting the same inputs,
and that the extension is additive only (it never rejects emphasis
CommonMark accepts).

No behavioral change.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

* fix(muyajs): make flanking-boundary char extraction code-point aware

Address Copilot inline review on PR marktext#4355: CJK_REG contains a
surrogate-pair branch ([\uD840-\uD87F][\uDC00-\uDFFF]) intended to
match non-BMP CJK ideographs (CJK Ext-B onward), but the existing
flanking-boundary readers used charAt / src[i] — single UTF-16 code
unit access — so the regex branch could never fire in practice.

Add lastCodePointChar / codePointCharAt helpers that read a full
Unicode code point (1 or 2 UTF-16 code units) and route every flanking
read in canOpenEmphasis / canCloseEmphasis through them.

Side benefit: PUNCTUATION_REG also contains numerous surrogate-pair
alternations (\uD800[\uDD00-\uDD02\uDF9F\uDFD0] and similar) for
non-BMP Unicode punctuation that were dead for the same reason. Those
spec-defined cases are now correctly tested too.

Adds an extra vitest case using 𠀀 / 𠀁 (CJK Ext-B, U+20000/U+20001)
to lock in the surrogate-pair behavior.

Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] 似乎不能正常处理**加粗语法的配对

2 participants