Repository navigation
test(muya): port MarkText regression cases (#4341/#4307/#4190) - #4399
Conversation
Port three MarkText-specific regression cases from the legacy `packages/muyajs` desktop specs into `@muyajs/core`'s own suite to lock in behavioral fidelity for the migration off legacy muyajs. Tests only — no engine changes. - #4341 nested mixed lists (ul-in-ol / ol-in-ul): PASS. The state tree from MarkdownToState nests the differing-type list under the correct list-item (no paragraph collapse), and md -> state -> md round-trips identically. Ported as structural + round-trip assertions. - #4190 table normalization (body row with more/fewer cells than the header): PASS. StateToMarkdown.serializeTable clamps each row to the header column count (extra cell dropped) and never throws. The legacy spec hand-built a malformed block tree for ExportMarkdown.normalizeTable; the muya equivalent hand-builds a malformed ITableState because a GFM round trip can never produce a ragged table. - #4307 CJK strong flanking (`**"加粗"**` against a CJK boundary): documented engine GAP. marked@16 implements the CommonMark flanking rule literally and classifies CJK ideographs / Hangul as "other" (neither whitespace nor punctuation), so `**` adjacent to a CJK char with punctuation-bounded inner content does not open/close emphasis. Legacy muyajs shipped a custom tokenizer that treats CJK as punctuation for flanking; marked does not. The four CJK cases assert the CORRECT (legacy) behavior under `it.fails`, so the suite stays green while the gap exists and flips red the moment the gap closes. Sanity cases that already work are plain `it` so a future fix can't regress them. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
There was a problem hiding this comment.
Pull request overview
Ports MarkText legacy regression coverage into @muyajs/core’s state-layer unit tests to ensure behavioral fidelity during the migration off legacy packages/muyajs (tests only; no engine changes).
Changes:
- Add nested mixed-list regression tests ensuring list structure is preserved and
md → state → mdround-trips stably (#4341). - Add table export normalization regression tests for ragged rows (extra/fewer cells) to ensure serialization doesn’t throw and clamps columns (#4190).
- Add CJK strong-emphasis flanking regression cases, documenting the current
markedgap viait.failswhile keeping the suite green (#4307).
Reviewed changes
Copilot reviewed 3 out of 3 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
| packages/muya/src/state/tests/nestedMixedLists.spec.ts | Adds regression coverage for mixed nested lists, including structure assertions and round-trip stability. |
| packages/muya/src/state/tests/strongCjkFlanking.spec.ts | Adds strong-emphasis CJK boundary cases; records current engine gap with expected-fail tests plus sanity cases that must keep passing. |
| packages/muya/src/state/tests/tableNormalization.spec.ts | Adds table serialization normalization tests for ragged rows (but overlaps existing coverage in the suite). |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| describe('table normalization on export (#4190)', () => { | ||
| it('exports a well-formed table without error', () => { |
There was a problem hiding this comment.
Good catch — you're right, this was duplicate coverage. I verified tableNormalization.spec.ts against the existing serializeTable — row width mismatch suite in stateToMarkdown.spec.ts (#4222/#4190): both exercise the exact same three scenarios (well-formed table, body row with extra cells → dropped/no-crash, body row with fewer cells → no-crash). The only delta here was slightly stricter pipe-count assertions on the extra-cell case, but those verify the same "extra cell dropped" outcome the existing suite already asserts via .not.toContain('| 3') / .not.toContain('| 4'), so there was no genuinely-new behavior to preserve.
Deleted tableNormalization.spec.ts in 9e8a00b and kept stateToMarkdown.spec.ts as the single source of truth for table-width normalization. The other two ported specs (nestedMixedLists.spec.ts #4341, strongCjkFlanking.spec.ts #4307) cover different regressions and were left in place. Full suite stays green (lint / lint:types / check-circular / test / test:spec).
The serializeTable ragged-row coverage in tableNormalization.spec.ts duplicated the existing serializeTable — row width mismatch suite in stateToMarkdown.spec.ts (both #4222/#4190): same well-formed, extra-cell-dropped, and short-row scenarios. Keeping two suites for one behavior risks them drifting apart, so remove the redundant file and keep the single existing suite as the source of truth. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]>
…4401) * fix(muya): treat CJK as punctuation for strong/em flanking (#4307) Strong/em delimited with `**` directly against a CJK character whose inner content is punctuation-bounded did not bold in muya, e.g. `例子例子**"加粗"**例子例子`, `日本語**(強調)**日本語`, `한국어**[강조]**한국어`, and the non-BMP `𠀀𠀁**"加粗"**𠀀𠀁`. The legacy muyajs engine bolded these. CommonMark §6.2 classifies CJK ideographs / Hangul / Kana as "other" (Lo), neither whitespace nor punctuation, so a `**` run wrapped by CJK with punctuation-bounded content is not left/right-flanking and stays literal. muya has TWO inline-tokenization paths and both carry the same CJK-as- punctuation widening the legacy engine shipped: - Static / export path (marked@16): a `cjkEmStrong` tokenizer override that rebuilds marked's emStrong flanking regexes with CJK folded into the punctuation class and removed from the alphanumeric class, registered in getHighlightHtml and getClipboardHtml. Faithful copy of marked's emStrong body; rules are swapped in/out per-call so the shared tokenizer rules are never left mutated. - Live editor path (inlineRenderer): CJK widening added to canOpen/canCloseEmphasis in inlineRenderer/utils.ts, plus full code-point reading so the non-BMP CJK Ext-B surrogate-pair branch is live. CJK ranges (matching legacy CJK_REG): Hiragana+Katakana U+3040–U+30FF, CJK Ext-A U+3400–U+4DBF, CJK Unified U+4E00–U+9FFF, CJK Compatibility U+F900–U+FAFF, Hangul Syllables U+AC00–U+D7AF, Halfwidth Katakana U+FF66–U+FF9D, and CJK Ext-B U+20000–U+2A6DF (non-BMP). The widening is additive — it never bolds anything CommonMark accepts as non-emphasis — so the CommonMark 0.31 + GFM conformance suites are unchanged (1347/1347, no unexpected passes). Promotes the #4399 `it.fails` placeholders to passing `it` cases and adds live-editor-path coverage plus negative cases. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> * test(muya): assert negative CJK flanking cases emit neither strong nor em The static/export-path negative cases in strongCjkFlanking.spec.ts only checked for <strong> via rendersStrong. Since #4307 widens the emphasis/strong flanking logic, a regression could surface as unexpected <em> output while still passing a <strong>-only assertion. Add a rendersEm helper mirroring rendersStrong and assert NEITHER tag is produced for the negative cases. The live editor path already covers this: tokenizesEmphasis returns true for either a strong or em token, so its negative .toBe(false) already rejects both. Only the static path needed strengthening. Addresses Copilot review on #4401. Co-Authored-By: Claude Opus 4.8 (1M context) <[email protected]> --------- Co-authored-by: Claude Opus 4.8 (1M context) <[email protected]>
What
Ports three MarkText-specific regression cases from the legacy
packages/muyajsdesktop specs into@muyajs/core's own test suite, to lock in behavioral fidelity for the migration off legacy muyajs. Tests only — no engine changes.Legacy sources read:
packages/desktop/test/unit/specs/markdown-nested-mixed-lists.spec.ts([Bug] Regression: nested mixed lists no longer work #4341)packages/desktop/test/unit/specs/markdown-strong-cjk.spec.ts([Bug] 似乎不能正常处理**加粗语法的配对 #4307)packages/desktop/test/unit/specs/export-markdown.spec.ts—ExportMarkdown.normalizeTablecases (Unexpected error: Cannot read properties of undefined (reading 'width') #4190)New muya specs (all under
packages/muya/src/state/__tests__/, picked up bypnpm test):nestedMixedLists.spec.tsstrongCjkFlanking.spec.tstableNormalization.spec.tsFidelity verification — per case
#4341 nested mixed lists (ul-in-ol / ol-in-ul) — ✅ PASS (muya matches legacy)
The
MarkdownToStatestate tree nests the differing-type list (bullet-listinside the secondorder-listitem, and vice-versa) under the correctlist-item— it does not collapse into a paragraph (the legacy failure mode).md → state → mdround-trips identically and is stable on a second pass. Ported as structural + round-trip assertions. The legacy engine usedContentState/ExportMarkdown; muya usesMarkdownToState/StateToMarkdown, so the indentation is regenerated by the serializer rather than hard-coded.#4190 table normalization (ragged body rows) — ✅ PASS (muya matches legacy)
StateToMarkdown.serializeTableclamps every body row to the header column count (an extra cell is silently dropped) and tolerates a body row with fewer cells than the header — never throwing in either case, exactly matching legacyExportMarkdown.normalizeTable. The legacy spec hand-built a malformedtable > thead/tbody > tr > th/tdblock tree; the muya equivalent hand-builds a malformedITableStatebecause a GFM markdown round trip can never produce a ragged table (the block parser already pads/truncates rows), so a hand-built state is the only way to exercise the serializer's own column-clamp guard.#4307 CJK strong flanking (⚠️ GAP (documented)
**"加粗"**against a CJK boundary) —@muyajs/core does NOT reproduce the legacy behavior. muya tokenises inline markdown with
marked@16, which implements the CommonMark emphasis "flanking" rule literally. CommonMark classifies CJK ideographs and Hangul as "other" (neither whitespace nor punctuation). For a left-flanking**run, clause (2b) requires the char before the run to be whitespace/punctuation whenever the char after it is punctuation. In例子例子**"加粗"**例子例子the char after**is"(punctuation) and the char before it is子(CJK → "other"), so the run is not left-flanking and marked emits a literal**.Verified the gap is in
markeditself (not muya glue):marked.parseInline('例子例子**"加粗"**例子例子')returns the literal**, while中文**加粗**中文(CJK on the inner side) bolds correctly.Legacy MarkText shipped its own inline tokenizer whose
canOpen/canCloseEmphasisflanking helpers treat CJK as punctuation; marked has no such patch. Fixing it requires either patching the dependency or shipping a custom inline-emphasis tokenizer extension — out of scope for a tests-only fidelity PR.Handling: the four CJK cases assert the correct (legacy) expected behavior under
it.fails, so:it.failsflips red, forcing promotion to a plainit— fidelity can only go up.Sanity cases that already work (
中文**加粗**中文,before **"normal"** after,before**normal**after) are plainitso a future fix can't regress them.Follow-up: close the CJK-flanking gap in
@muyajs/core(custom inline emphasis tokenizer or a marked patch that treats CJK as punctuation for flanking), then promote the fourit.failscases toit.Verification
The 4 "expected fail" are the documented #4307 CJK gap cases.
🤖 Generated with Claude Code