Skip to content

test(muya): port MarkText regression cases (#4341/#4307/#4190) - #4399

Merged
Jocs merged 2 commits into
developfrom
test/muya-port-marktext-regressions
Jun 8, 2026
Merged

Jocs merged 2 commits into
developfrom
test/muya-port-marktext-regressions

Conversation

@Jocs

@Jocs Jocs commented Jun 8, 2026

Copy link
Copy Markdown
Member

What

Ports three MarkText-specific regression cases from the legacy packages/muyajs desktop specs into @muyajs/core's own test suite, to lock in behavioral fidelity for the migration off legacy muyajs. Tests only — no engine changes.

Legacy sources read:

New muya specs (all under packages/muya/src/state/__tests__/, picked up by pnpm test):

  • nestedMixedLists.spec.ts
  • strongCjkFlanking.spec.ts
  • tableNormalization.spec.ts

Fidelity verification — per case

#4341 nested mixed lists (ul-in-ol / ol-in-ul) — ✅ PASS (muya matches legacy)

The MarkdownToState state tree nests the differing-type list (bullet-list inside the second order-list item, and vice-versa) under the correct list-item — it does not collapse into a paragraph (the legacy failure mode). md → state → md round-trips identically and is stable on a second pass. Ported as structural + round-trip assertions. The legacy engine used ContentState/ExportMarkdown; muya uses MarkdownToState/StateToMarkdown, so the indentation is regenerated by the serializer rather than hard-coded.

#4190 table normalization (ragged body rows) — ✅ PASS (muya matches legacy)

StateToMarkdown.serializeTable clamps every body row to the header column count (an extra cell is silently dropped) and tolerates a body row with fewer cells than the header — never throwing in either case, exactly matching legacy ExportMarkdown.normalizeTable. The legacy spec hand-built a malformed table > thead/tbody > tr > th/td block tree; the muya equivalent hand-builds a malformed ITableState because a GFM markdown round trip can never produce a ragged table (the block parser already pads/truncates rows), so a hand-built state is the only way to exercise the serializer's own column-clamp guard.

#4307 CJK strong flanking (**"加粗"** against a CJK boundary) — ⚠️ GAP (documented)

@muyajs/core does NOT reproduce the legacy behavior. muya tokenises inline markdown with marked@16, which implements the CommonMark emphasis "flanking" rule literally. CommonMark classifies CJK ideographs and Hangul as "other" (neither whitespace nor punctuation). For a left-flanking ** run, clause (2b) requires the char before the run to be whitespace/punctuation whenever the char after it is punctuation. In 例子例子**"加粗"**例子例子 the char after ** is " (punctuation) and the char before it is 子 (CJK → "other"), so the run is not left-flanking and marked emits a literal **.

Verified the gap is in marked itself (not muya glue): marked.parseInline('例子例子**"加粗"**例子例子') returns the literal **, while 中文**加粗**中文 (CJK on the inner side) bolds correctly.

Legacy MarkText shipped its own inline tokenizer whose canOpen/canCloseEmphasis flanking helpers treat CJK as punctuation; marked has no such patch. Fixing it requires either patching the dependency or shipping a custom inline-emphasis tokenizer extension — out of scope for a tests-only fidelity PR.

Handling: the four CJK cases assert the correct (legacy) expected behavior under it.fails, so:

  • the suite stays green while the gap exists, and
  • the moment the engine starts recognising these (a marked upgrade or a flanking patch) the it.fails flips red, forcing promotion to a plain it — fidelity can only go up.

Sanity cases that already work (中文**加粗**中文, before **"normal"** after, before**normal**after) are plain it so a future fix can't regress them.

Follow-up: close the CJK-flanking gap in @muyajs/core (custom inline emphasis tokenizer or a marked patch that treats CJK as punctuation for flanking), then promote the four it.fails cases to it.

Verification

pnpm -C packages/muya lint            # 0 errors (pre-existing complexity warnings only)
pnpm -C packages/muya lint:types      # pass
pnpm -C packages/muya check-circular   # pass
pnpm -C packages/muya test            # 451 passed | 4 expected fail (455)
pnpm -C packages/muya test:spec       # 1347 passed

The 4 "expected fail" are the documented #4307 CJK gap cases.

🤖 Generated with Claude Code

Port three MarkText-specific regression cases from the legacy
`packages/muyajs` desktop specs into `@muyajs/core`'s own suite to lock
in behavioral fidelity for the migration off legacy muyajs. Tests only —
no engine changes.

- #4341 nested mixed lists (ul-in-ol / ol-in-ul): PASS. The state tree
  from MarkdownToState nests the differing-type list under the correct
  list-item (no paragraph collapse), and md -> state -> md round-trips
  identically. Ported as structural + round-trip assertions.
- #4190 table normalization (body row with more/fewer cells than the
  header): PASS. StateToMarkdown.serializeTable clamps each row to the
  header column count (extra cell dropped) and never throws. The legacy
  spec hand-built a malformed block tree for ExportMarkdown.normalizeTable;
  the muya equivalent hand-builds a malformed ITableState because a GFM
  round trip can never produce a ragged table.
- #4307 CJK strong flanking (`**"加粗"**` against a CJK boundary):
  documented engine GAP. marked@16 implements the CommonMark flanking
  rule literally and classifies CJK ideographs / Hangul as "other"
  (neither whitespace nor punctuation), so `**` adjacent to a CJK char
  with punctuation-bounded inner content does not open/close emphasis.
  Legacy muyajs shipped a custom tokenizer that treats CJK as punctuation
  for flanking; marked does not. The four CJK cases assert the CORRECT
  (legacy) behavior under `it.fails`, so the suite stays green while the
  gap exists and flips red the moment the gap closes. Sanity cases that
  already work are plain `it` so a future fix can't regress them.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
Copilot AI review requested due to automatic review settings June 8, 2026 07:28

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Ports MarkText legacy regression coverage into @muyajs/core’s state-layer unit tests to ensure behavioral fidelity during the migration off legacy packages/muyajs (tests only; no engine changes).

Changes:

  • Add nested mixed-list regression tests ensuring list structure is preserved and md → state → md round-trips stably (#4341).
  • Add table export normalization regression tests for ragged rows (extra/fewer cells) to ensure serialization doesn’t throw and clamps columns (#4190).
  • Add CJK strong-emphasis flanking regression cases, documenting the current marked gap via it.fails while keeping the suite green (#4307).

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.

File Description
packages/muya/src/state/tests/nestedMixedLists.spec.ts Adds regression coverage for mixed nested lists, including structure assertions and round-trip stability.
packages/muya/src/state/tests/strongCjkFlanking.spec.ts Adds strong-emphasis CJK boundary cases; records current engine gap with expected-fail tests plus sanity cases that must keep passing.
packages/muya/src/state/tests/tableNormalization.spec.ts Adds table serialization normalization tests for ragged rows (but overlaps existing coverage in the suite).

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +55 to +56
describe('table normalization on export (#4190)', () => {
it('exports a well-formed table without error', () => {

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — you're right, this was duplicate coverage. I verified tableNormalization.spec.ts against the existing serializeTable — row width mismatch suite in stateToMarkdown.spec.ts (#4222/#4190): both exercise the exact same three scenarios (well-formed table, body row with extra cells → dropped/no-crash, body row with fewer cells → no-crash). The only delta here was slightly stricter pipe-count assertions on the extra-cell case, but those verify the same "extra cell dropped" outcome the existing suite already asserts via .not.toContain('| 3') / .not.toContain('| 4'), so there was no genuinely-new behavior to preserve.

Deleted tableNormalization.spec.ts in 9e8a00b and kept stateToMarkdown.spec.ts as the single source of truth for table-width normalization. The other two ported specs (nestedMixedLists.spec.ts #4341, strongCjkFlanking.spec.ts #4307) cover different regressions and were left in place. Full suite stays green (lint / lint:types / check-circular / test / test:spec).

The serializeTable ragged-row coverage in tableNormalization.spec.ts
duplicated the existing serializeTable — row width mismatch suite in
stateToMarkdown.spec.ts (both #4222/#4190): same well-formed,
extra-cell-dropped, and short-row scenarios. Keeping two suites for one
behavior risks them drifting apart, so remove the redundant file and
keep the single existing suite as the source of truth.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
@Jocs
Jocs merged commit baf81fe into develop Jun 8, 2026
9 checks passed
@Jocs
Jocs deleted the test/muya-port-marktext-regressions branch June 8, 2026 08:03
Jocs added a commit that referenced this pull request Jun 8, 2026
…4401)

* fix(muya): treat CJK as punctuation for strong/em flanking (#4307)

Strong/em delimited with `**` directly against a CJK character whose inner
content is punctuation-bounded did not bold in muya, e.g.
`例子例子**"加粗"**例子例子`, `日本語**(強調)**日本語`, `한국어**[강조]**한국어`,
and the non-BMP `𠀀𠀁**"加粗"**𠀀𠀁`. The legacy muyajs engine bolded these.

CommonMark §6.2 classifies CJK ideographs / Hangul / Kana as "other" (Lo),
neither whitespace nor punctuation, so a `**` run wrapped by CJK with
punctuation-bounded content is not left/right-flanking and stays literal.
muya has TWO inline-tokenization paths and both carry the same CJK-as-
punctuation widening the legacy engine shipped:

- Static / export path (marked@16): a `cjkEmStrong` tokenizer override that
  rebuilds marked's emStrong flanking regexes with CJK folded into the
  punctuation class and removed from the alphanumeric class, registered in
  getHighlightHtml and getClipboardHtml. Faithful copy of marked's emStrong
  body; rules are swapped in/out per-call so the shared tokenizer rules are
  never left mutated.
- Live editor path (inlineRenderer): CJK widening added to
  canOpen/canCloseEmphasis in inlineRenderer/utils.ts, plus full code-point
  reading so the non-BMP CJK Ext-B surrogate-pair branch is live.

CJK ranges (matching legacy CJK_REG): Hiragana+Katakana U+3040–U+30FF, CJK
Ext-A U+3400–U+4DBF, CJK Unified U+4E00–U+9FFF, CJK Compatibility
U+F900–U+FAFF, Hangul Syllables U+AC00–U+D7AF, Halfwidth Katakana
U+FF66–U+FF9D, and CJK Ext-B U+20000–U+2A6DF (non-BMP).

The widening is additive — it never bolds anything CommonMark accepts as
non-emphasis — so the CommonMark 0.31 + GFM conformance suites are unchanged
(1347/1347, no unexpected passes). Promotes the #4399 `it.fails` placeholders
to passing `it` cases and adds live-editor-path coverage plus negative cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

* test(muya): assert negative CJK flanking cases emit neither strong nor em

The static/export-path negative cases in strongCjkFlanking.spec.ts only
checked for <strong> via rendersStrong. Since #4307 widens the
emphasis/strong flanking logic, a regression could surface as unexpected
<em> output while still passing a <strong>-only assertion. Add a rendersEm
helper mirroring rendersStrong and assert NEITHER tag is produced for the
negative cases.

The live editor path already covers this: tokenizesEmphasis returns true
for either a strong or em token, so its negative .toBe(false) already
rejects both. Only the static path needed strengthening.

Addresses Copilot review on #4401.

Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants