Repository navigation
fix(muyajs): allow CJK ideographs as flanking boundary for emphasis - #4355
Conversation
Closes #4307. Strict CommonMark 6.2 flanking rules treat a CJK Unified Ideograph as neither Unicode whitespace nor Unicode punctuation, so `中文**"加粗"**中文` (and the symmetric Japanese / Korean variants) failed clause (2b) of canOpenEmphasis / canCloseEmphasis and never opened a strong run. Mirror the well-known CJK-friendly emphasis extension shipped by Typora, VSCode markdownlint, etc.: treat CJK Unified Ideographs (BMP + Ext-A + Ext-B via surrogate pairs + Compatibility Ideographs), Hiragana, Katakana, Halfwidth Katakana and Hangul Syllables as boundary-equivalent in clause (2b). Adds: - packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts — token-level vitest covering three CJK + inner-punctuation cases plus three regression cases that already worked on develop. - packages/desktop/test/e2e/strong-cjk.spec.ts — Playwright spec confirming `<strong>` is emitted in the live WYSIWYG DOM. - muya/lib/parser shim in src/types/muya.d.ts so the new vitest's named-import passes vue-tsc. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
|
Website preview ready: https://pr-4355-marktext-website.ransixi.workers.dev Built from |
|
Build artifacts for PR #4355: Run: https://github.com/marktext/marktext/actions/runs/26866178020
|
Strengthen the inline comments in parser/utils.js so it is unambiguous that the CJK_REG widening of the flanking check is a *deliberate divergence* from CommonMark spec 6.2, not a spec compliance fix. Note that GitHub/GFM is strictly conformant by rejecting the same inputs, and that the extension is additive only (it never rejects emphasis CommonMark accepts). No behavioral change. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
There was a problem hiding this comment.
Pull request overview
This PR updates the legacy MuyaJS Markdown emphasis parsing so **...** can be recognized as strong emphasis when adjacent to CJK text plus inner punctuation (a deliberate, documented non-CommonMark extension), addressing issue #4307 in the WYSIWYG editor pipeline.
Changes:
- Extend the emphasis “flanking” boundary checks in
canOpenEmphasis/canCloseEmphasisto treat CJK characters as boundary-equivalent (additive widening). - Add a unit-level tokenizer regression suite covering CJK + inner-punctuation cases and existing “sanity” cases.
- Add an Electron Playwright e2e regression asserting
<strong>is emitted in the rendered editor DOM and add a small TS ambient module shim formuya/lib/parser.
Reviewed changes
Copilot reviewed 3 out of 4 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| packages/muyajs/lib/parser/utils.js | Widens emphasis boundary rules to treat CJK characters as valid flanking boundaries (non-standard extension). |
| packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts | Adds token-level unit tests ensuring strong is produced in CJK + inner-punctuation scenarios. |
| packages/desktop/test/e2e/strong-cjk.spec.ts | Adds DOM-level e2e coverage verifying <strong> appears in WYSIWYG rendering for the reported cases. |
| packages/desktop/src/types/muya.d.ts | Adds ambient declarations so named imports from muya/lib/parser typecheck in the new unit test. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| // Non-standard CJK widening (see CJK_REG block above) — additive only: | ||
| // CJK ideographs are accepted as "preceded by" boundary, on top of the | ||
| // CommonMark-defined whitespace/punctuation set. | ||
| if (PUNCTUATION_REG.test(followedChar) && !(UNICODE_WHITESPACE_REG.test(precededChar) || PUNCTUATION_REG.test(precededChar) || CJK_REG.test(precededChar))) { | ||
| return false |
There was a problem hiding this comment.
Good catch — verified and addressed in f6a53aa.
charAt / src[i] only return one UTF-16 code unit, so the surrogate-pair branch of CJK_REG (and, I now realize, the long list of surrogate-pair alternations in the existing CommonMark PUNCTUATION_REG) was effectively dead. Both are now hit because every flanking-boundary read goes through new lastCodePointChar / codePointCharAt helpers that return 1 or 2 code units as appropriate.
Added a vitest case 𠀀𠀁**"加粗"**𠀀𠀁 (CJK Ext-B, U+20000/U+20001) to lock the surrogate-pair behavior down. Full suite still green (557 unit / 75 e2e).
Address Copilot inline review on PR #4355: CJK_REG contains a surrogate-pair branch ([\uD840-\uD87F][\uDC00-\uDFFF]) intended to match non-BMP CJK ideographs (CJK Ext-B onward), but the existing flanking-boundary readers used charAt / src[i] — single UTF-16 code unit access — so the regex branch could never fire in practice. Add lastCodePointChar / codePointCharAt helpers that read a full Unicode code point (1 or 2 UTF-16 code units) and route every flanking read in canOpenEmphasis / canCloseEmphasis through them. Side benefit: PUNCTUATION_REG also contains numerous surrogate-pair alternations (\uD800[\uDD00-\uDD02\uDF9F\uDFD0] and similar) for non-BMP Unicode punctuation that were dead for the same reason. Those spec-defined cases are now correctly tested too. Adds an extra vitest case using 𠀀 / 𠀁 (CJK Ext-B, U+20000/U+20001) to lock in the surrogate-pair behavior. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]>
Merges 37 upstream commits into the fork, including the pnpm monorepo conversion (marktext#4302), muya TS rewrite migration (marktext#4314), website package, and fixes through CJK emphasis boundary (marktext#4355). Conflict resolutions: - package.json: adopt upstream monorepo root shell; drop fork's root-level deps block; move 4 fork-added deps (diff, @types/diff, markdown-it, @types/markdown-it) to packages/desktop/package.json - build.yml: keep both fork's workflow_dispatch+dynamic-matrix and upstream's paths-ignore: packages/muya/** - wsl.ts, comparisonPane.vue, wsl-paths.spec.ts: accepted git's rename detection — files live under packages/desktop/ in the monorepo layout - docs/dev/FORK_PLAN.md: kept at repo-root docs/dev/ (private fork doc, not published website content) - pnpm-lock.yaml: regenerated via pnpm install
…arktext#4355) * fix(muyajs): allow CJK ideographs as flanking boundary for emphasis Closes marktext#4307. Strict CommonMark 6.2 flanking rules treat a CJK Unified Ideograph as neither Unicode whitespace nor Unicode punctuation, so `中文**"加粗"**中文` (and the symmetric Japanese / Korean variants) failed clause (2b) of canOpenEmphasis / canCloseEmphasis and never opened a strong run. Mirror the well-known CJK-friendly emphasis extension shipped by Typora, VSCode markdownlint, etc.: treat CJK Unified Ideographs (BMP + Ext-A + Ext-B via surrogate pairs + Compatibility Ideographs), Hiragana, Katakana, Halfwidth Katakana and Hangul Syllables as boundary-equivalent in clause (2b). Adds: - packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts — token-level vitest covering three CJK + inner-punctuation cases plus three regression cases that already worked on develop. - packages/desktop/test/e2e/strong-cjk.spec.ts — Playwright spec confirming `<strong>` is emitted in the live WYSIWYG DOM. - muya/lib/parser shim in src/types/muya.d.ts so the new vitest's named-import passes vue-tsc. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * docs(muyajs): mark CJK emphasis widening as a non-standard extension Strengthen the inline comments in parser/utils.js so it is unambiguous that the CJK_REG widening of the flanking check is a *deliberate divergence* from CommonMark spec 6.2, not a spec compliance fix. Note that GitHub/GFM is strictly conformant by rejecting the same inputs, and that the extension is additive only (it never rejects emphasis CommonMark accepts). No behavioral change. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> * fix(muyajs): make flanking-boundary char extraction code-point aware Address Copilot inline review on PR marktext#4355: CJK_REG contains a surrogate-pair branch ([\uD840-\uD87F][\uDC00-\uDFFF]) intended to match non-BMP CJK ideographs (CJK Ext-B onward), but the existing flanking-boundary readers used charAt / src[i] — single UTF-16 code unit access — so the regex branch could never fire in practice. Add lastCodePointChar / codePointCharAt helpers that read a full Unicode code point (1 or 2 UTF-16 code units) and route every flanking read in canOpenEmphasis / canCloseEmphasis through them. Side benefit: PUNCTUATION_REG also contains numerous surrogate-pair alternations (\uD800[\uDD00-\uDD02\uDF9F\uDFD0] and similar) for non-BMP Unicode punctuation that were dead for the same reason. Those spec-defined cases are now correctly tested too. Adds an extra vitest case using 𠀀 / 𠀁 (CJK Ext-B, U+20000/U+20001) to lock in the surrogate-pair behavior. Co-Authored-By: Claude Opus 4.7 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 4.7 (1M context) <[email protected]>
Summary
Closes #4307.
**…**adjacent to a CJK ideograph + inner punctuation (e.g.中文**"加粗"**中文,日本語**(強調)**日本語,한국어**[강조]**한국어) fails to render as strong in WYSIWYG. The same input has been broken since at least v0.17.1 (the parser code was byte-identical until this PR), and GitHub/GFM renders the same way.This is not a CommonMark conformance fix. It is a deliberate divergence from CommonMark spec 6.2 ("Emphasis and strong emphasis", https://spec.commonmark.org/0.31.2/#emphasis-and-strong-emphasis):
Lo— neither whitespace nor punctuation.中文**"加粗"**中文therefore MUST NOT open a strong run. GitHub/GFM is strictly spec-conformant and behaves the same way.CJK scripts do not use spaces between words, so the strict rule effectively denies emphasis to any CJK paragraph that wraps the
**run with punctuation (quotes, parentheses, brackets, 、 …). This PR adopts the well-known CJK-friendly emphasis extension already shipped by Typora, VSCode markdownlint, Joplin, and most CJK-oriented Markdown tools: treat CJK ideographs as boundary-equivalent in clause (2b) of the flanking check.Why this is safe:
||): it never rejects emphasis CommonMark accepts, it only accepts emphasis CommonMark rejects on the CJK-boundary edge case.packages/muyajs/lib/parser/utils.js) with explicit "NON-STANDARD EXTENSION" header so future readers / spec auditors can see the divergence at a glance.Implementation
packages/muyajs/lib/parser/utils.js— introduceCJK_REG(CJK Unified Ideographs BMP + Ext-A + Ext-B via surrogate pairs + Compatibility Ideographs, Hiragana, Katakana, Halfwidth Katakana, Hangul Syllables) and extend the (2b) clauses ofcanOpenEmphasis/canCloseEmphasiswith a three-way||.packages/desktop/src/types/muya.d.ts— small ambient shim formuya/lib/parserso the new vitest's named-import passesvue-tsc, consistent with the existing muya bridges.Tests
packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts— token-level vitest: 3 CJK + inner-punctuation cases (fail ondevelop, pass with this PR) + 3 regression cases (already pass ondevelop, still pass).packages/desktop/test/e2e/strong-cjk.spec.ts— DOM-level Playwright spec: confirms<strong>is emitted in the live WYSIWYG render.Test plan
pnpm test:unit— 556/556 pass (10 files)pnpm -C packages/desktop exec playwright test test/e2e/— 75 pass / 2 skip / 0 fail (incl. 2 new specs)pnpm run lint— 0 errorspnpm run typecheck— passdevelopbefore the parser change, then pass after.🤖 Generated with Claude Code