Skip to content

Support reading comment text/content from Nihon Kohden EEG files #13633

Description

@eulerleibniz

Describe the new feature or enhancement

Description

Nihon Kohden EEG files contain comments/annotations with textual and image content, but currently this information is not accessible when reading the data with MNE.

I previously raised this question on the MNE Discourse forum, where the limitation was discussed:
https://mne.discourse.group/t/how-to-read-the-nihon-kohden-eeg-files-comments-content-of-the-comment/11680

At the moment, MNE can detect the presence/timing of comments but does not expose the actual text content of those comments to the user.

Why this matters

The comment text in Nihon Kohden recordings often contains clinically and experimentally relevant metadata, such as:

  • Seizure/Event types and full descriptions added by experts
  • Manual annotations added during recording added by nurses

Losing this information during import makes downstream analysis specially in case of machine learning incomplete and requires users to rely on external vendor software.

Current behavior

Given a code like this:

import mne 
print(mne.__version__)
raw = mne.io.read_raw_nihon("FJ00231Z.EEG", preload=False) # OR preload=True
for ann in raw.annotations:
    print(ann)

This will output the annotations as expected, but it does not read the comments content. Here is a sample output of the code:

1.11.0 # MNE version
Loading FJ00231Z.EEG
Found 21E file, reading channel names.
Reading header from Path\To\File\EEG2100\FJ00231Z.EEG # EEG 2100 Device
Found PNT file, reading metadata.
Found LOG file, reading events.

OrderedDict({'onset': np.float64(11768.0), 'duration': np.float64(0.0), 'description': np.str_('eye close'), 'orig_time': datetime.datetime(2025, 12, 27, 9, 49, 31, tzinfo=datetime.timezone.utc), 'extras': {}})

OrderedDict({'onset': np.float64(13307.568), 'duration': np.float64(0.0), 'description': np.str_('P_COMMENT'), 'orig_time': datetime.datetime(2025, 12, 27, 9, 49, 31, tzinfo=datetime.timezone.utc), 'extras': {}})

Those annotations with 'description': np.str_('P_COMMENT') contain comments such as this image and are not currently readable:

Image
  • Nihon Kohden EEG files can be read
  • Annotations can be read
  • Comments timing are available
  • Comments are all shown as P_COMMENT and the text/content is not accessible

Expected behavior

At least the Comment text should be parsed and exposed, ideally as:

  • Annotations.description, or
  • In the extra field of the annotation ideally as a dict

Describe your proposed implementation

Basic code snippet

Here is the code i use currently to read the comments from a given .CMT file. It should be integrated into mne.io.read_raw_nihon code but the code is not complete so it can wait.

import re
import string
from dataclasses import dataclass


@dataclass
class Comment:
    timestamp: int
    text: str


TS_RE = re.compile(rb"(\d{20})") # The timing of each annotation which is 20 digit long

PRINTABLE = set(bytes(string.printable, "ascii")) # The .CMT file contains many NULL and control characters


def clean_bytes(b: bytes) -> str:
    # keep printable ASCII, including space and newline - NOT TESTED WITH COMMENTS INCLUDING IMAGE LINKS
    cleaned_byte = bytes(c if c in PRINTABLE else ord(" ") for c in b)
    cleaned_str = cleaned_byte.decode("ascii", errors="ignore").strip()
    cleaned_str = cleaned_str[10:].lstrip().lstrip("\t\n\x0b\r\x0c") # at least 10 control characters are included before the text
    return cleaned_str


def parse_cmt(path: str):
    data = open(path, "rb").read()
    matches = list(TS_RE.finditer(data))
    records = []
    for i, m in enumerate(matches):
        ts = int(m.group(1).decode("ascii"))
        start = m.end()
        end = matches[i + 1].start() if i + 1 < len(matches) else len(data)

        raw_text = data[start:end]
        text = clean_bytes(raw_text)

        if text:
            records.append(Comment(timestamp=ts, text=text))

    return records


records = parse_cmt("ROOT_PATH/NKT/EEG2100/FJ00231Z.CMT")
for r in records:
    print("📝")
    print(r.timestamp)
    print(r.text)

Limitations:

  • Tested only on EEG 2100 devices
  • Its just a workaround for starter
  • Nihon Kohden Comments can containt color code, background transparency, and even a reference image. None of these are read here as i've never seen experts really use them.

Describe possible alternatives

No alternatives currently

Additional context

No response

Activity

  1. agramfort commented on Feb 4, 2026

    @agramfort
    SponsorMember
  2. wmvanvliet commented on Feb 4, 2026

    @wmvanvliet
    Contributor

    This looks like a pretty straightforward addition to our nihon reader to me, as long as we keep the comments text-only. We could replace the currently unhelpful P_COMMENT string with the actual text of the comment.

    But yes, we need an example .CMT file, also to add to our testing data for unit tests.

  3. eulerleibniz commented on Feb 4, 2026

    @eulerleibniz
    Author

    I just picked one of my files, removed its personal health information, and added multiple types of comments to it. The comments contain single line, multi-line, colored background, transparent background, image reference, and current view reference. Here i will share the files for you. I tested to see if my code can read the comments correctly and it worked for me - It can read the text only of course.

    NOTE: I noticed that the timestamp in the CMT file has microsecond accuracy, but the timestamp in LOG file and mne annotations has millisecond accuracy. So keep that in mind in case you want to cross check them.

    Files.zip

  4. wmvanvliet commented on Feb 5, 2026

    @wmvanvliet
    Contributor

    Your implementation above is a little brittle. I've done some digging with a hex editor and found the following for the structure of the .CMT files:

    File header: 1024 bytes (0x400)
        Version string: 16 bytes ASCII
        Filename: 48 bytes ASCII
        Creation time: 32 bytes ASCII 
        Unknown: 45 bytes
        Version string: 16 bytes ASCII
        Unknown: 867 bytes
        
    Comments header: 41 bytes
        Unknown: 1 byte
        Version string: 16 bytes ASCII
        Unknown: 4 bytes
        Number of comments blocks: 4 bytes (32 bit unsigned integer?)
        Unknown: 16 bytes
        
    Comment block: 560 bytes
        Unknown: 4 bytes
        Timestamp: 32 bytes ASCII
        Appearance (colors and such): 24 bytes
        Comment text: 384 bytes
        Image filename: 64 bytes
        Unknown: 52 bytes
    

    Based on this, you should be able to write the code more like the rest of the nihon.py file: seek to the proper byte offset in the file and read a certain number of bytes.

  5. eulerleibniz commented on Feb 7, 2026

    @eulerleibniz
    Author

    Thanks for the tip @wmvanvliet .I didn't know about hex editors usage. Here is the final code i wrote and i will use for myself. I can integrate it to nihon code if you think its ok:

    from __future__ import annotations
    
    import pathlib
    from dataclasses import dataclass
    from datetime import datetime
    from typing import Literal, Optional
    
    # Each comment block is fixed-size: 20 (ts) + 24 (unk) + 512 (payload) + 4 (trail)
    CMT_COMMENT_BLOCK_SIZE = 560
    
    CommentType = Literal["text_only", "specific_reference", "current_view"]
    
    
    # -----------------------------
    # Helpers
    # -----------------------------
    
    
    def _decode_cstr(b: bytes) -> str:
        """Decode a fixed-width, null-padded byte field to a Python string."""
        return b.split(b"\x00", 1)[0].decode("utf-8", errors="replace").strip()
    
    
    def _hex(b: bytes) -> str:
        """Hex-encode bytes (lowercase). Useful while reverse-engineering."""
        return b.hex()
    
    
    def _parse_nk_datetime(raw: str) -> Optional[datetime]:
        """
        Best-effort parse for Nihon Kohden-style timestamps.
    
        Observed formats (digits only):
          - 14 digits: YYYYMMDDhhmmss
          - 15 digits: YYYYMMDDhhmmssT      (T = tenths of a second; 0-9)
          - 20 digits: YYYYMMDDhhmmssffffff (microseconds)
    
        Returns None if parsing fails.
        """
        s = raw.strip()
        if not s or not s.isdigit():
            return None
    
        try:
            if len(s) == 14:
                return datetime.strptime(s, "%Y%m%d%H%M%S")
    
            if len(s) == 15:
                # Interpret last digit as tenths-of-second.
                base = datetime.strptime(s[:14], "%Y%m%d%H%M%S")
                tenths = int(s[14])  # 0..9
                return base.replace(microsecond=tenths * 100_000)
    
            if len(s) == 20:
                return datetime.strptime(s, "%Y%m%d%H%M%S%f")
    
        except ValueError:
            return None
    
        return None
    
    
    # -----------------------------
    # Data models
    # -----------------------------
    
    
    @dataclass(frozen=True)
    class CmtFileHeader:
        """
        Parsed header fields from a Nihon Kohden .CMT file.
    
        Offsets are absolute from the start of the file (byte 0).
        """
    
        device_version: str  # [0:16]
        ref_file_1: str  # [16:32]
        ref_file_2: str  # [32:48]
        ref_file_3: str  # [48:64]
    
        # Keep both raw + parsed for debugging / reverse-engineering
        creation_time_raw: str  # [64:79]
        creation_time: Optional[datetime]
    
        version_info_1: str  # [129:145]
        version_info_2: str  # [150:166]
        version_info_3: str  # [1025:1041]
        unknown_1041_1069_hex: str  # [1041:1069]
    
        @classmethod
        def from_bytes(cls, data: bytes) -> "CmtFileHeader":
            """
            Parse the .CMT header from raw file bytes.
    
            Known header layout (absolute offsets):
              - device_version        @ [0:16]
              - ref_file_1            @ [16:32]
              - ref_file_2            @ [32:48]
              - ref_file_3            @ [48:64]
              - creation_time_raw     @ [64:79]   (parsed via _parse_nk_datetime)
              - version_info_1        @ [129:145]
              - version_info_2        @ [150:166]
              - version_info_3        @ [1025:1041]
              - unknown region (hex)  @ [1041:1069]
            """
            creation_raw = _decode_cstr(data[64:79])
            return cls(
                device_version=_decode_cstr(data[0:16]),
                ref_file_1=_decode_cstr(data[16:32]),
                ref_file_2=_decode_cstr(data[32:48]),
                ref_file_3=_decode_cstr(data[48:64]),
                creation_time_raw=creation_raw,
                creation_time=_parse_nk_datetime(creation_raw),
                version_info_1=_decode_cstr(data[129:145]),
                version_info_2=_decode_cstr(data[150:166]),
                version_info_3=_decode_cstr(data[1025:1041]),
                unknown_1041_1069_hex=_hex(data[1041:1069]),
            )
    
    
    @dataclass(frozen=True)
    class CmtComment:
        """
        A parsed comment block from a Nihon Kohden .CMT file.
    
        Canonical fields:
          - type:
              * text_only:
                  Free-text comment with no associated file.
              * current_view:
                  Reference to an image representing the *current screen/view*
                  at the time the comment was written (typically a PNG).
              * specific_reference:
                  Reference to an arbitrary file (e.g. BMP, PNG, etc.).
                  NOTE: Only the filename is stored in the CMT file.
                  The parent directory/path is resolved elsewhere by the system.
          - timestamp:
              Parsed datetime (best effort; may include sub-second precision)
          - text:
              Decoded comment text
          - reference_name:
              Optional filename associated with the comment (no directory)
    
        We keep timestamp_raw for auditing and reverse-engineering.
        """
    
        index: int
        timestamp_raw: str
        timestamp: Optional[datetime]
        type: CommentType
        text: str
        reference_name: Optional[str] = None
    
        @classmethod
        def from_block(cls, block: bytes, *, index: int) -> "CmtComment":
            """
            Parse one fixed-size comment block (560 bytes).
    
            Block layout (relative offsets within the 560-byte block):
              - timestamp20           @ [0:20]
              - unknown/control       @ [20:44]     (kept for future RE)
              - payload               @ [44:556]
              - trailer/control       @ [556:560]   (kept for future RE)
    
            Payload layout (512 bytes):
              - text/base             @ [0:384]
              - slot64                @ [384:448]
              - ctrl64                @ [448:512]   (TYPE CLASSIFIER)
    
            ctrl64 is a control/pattern region that is mostly NULL bytes with a few
            fixed markers. These markers were identified empirically and are used
            to classify comment type:
    
              - specific_reference:
                  ctrl64[0]  == '`'
                  ctrl64[4]  == '8'
                  ctrl64[8]  == '@'
    
              - current_view:
                  ctrl64[4]  == 'q'
                  ctrl64[12] == '}'
            """
            TS_LEN = 20
            UNK_LEN = 24
            PAYLOAD_LEN = 512
            TRAIL_LEN = 4
    
            TEXT_LEN = 384
            SLOT_LEN = 64
            CTRL_LEN = 64
    
            if len(block) != CMT_COMMENT_BLOCK_SIZE:
                raise ValueError(f"Expected {CMT_COMMENT_BLOCK_SIZE} bytes, got {len(block)}")
    
            # ---- block-level slices ----
            ts_b = block[0:TS_LEN]  # [0:20]
            unk_b = block[TS_LEN : TS_LEN + UNK_LEN]  # [20:44] (future RE)
            payload_b = block[TS_LEN + UNK_LEN : TS_LEN + UNK_LEN + PAYLOAD_LEN]  # [44:556]
            trail_b = block[-TRAIL_LEN:]  # [556:560] (future RE)
    
            # Timestamp is ASCII digits (usually), but tolerate nulls/spaces.
            timestamp_raw = ts_b.decode("ascii", errors="replace").strip("\x00").strip()
            timestamp = _parse_nk_datetime(timestamp_raw)
    
            # ---- classify first via ctrl64 (last 64 bytes of payload) ----
            ctrl64 = payload_b[TEXT_LEN + SLOT_LEN : TEXT_LEN + SLOT_LEN + CTRL_LEN]  # [448:512]
    
            is_specific_ref = ctrl64[0] == ord("`") and ctrl64[4] == ord("8") and ctrl64[8] == ord("@")
            is_current_view = ctrl64[4] == ord("q") and ctrl64[12] == ord("}")
    
            # Future reverse-engineering breadcrumbs (not stored):
            # unknown24_hex = _hex(unk_b)
            # trailer4_hex  = _hex(trail_b)
            # ctrl64_hex    = _hex(ctrl64)
    
            # ---- extract depending on type ----
            if is_current_view:
                text384 = payload_b[0:TEXT_LEN]  # payload[0:384]
                slot64 = payload_b[TEXT_LEN : TEXT_LEN + SLOT_LEN]  # payload[384:448]
    
                # Future RE hooks (not stored):
                # - many files use reference like "YYYYMMDDhhmmss.png" (seconds resolution)
                # - you can parse first 14 digits into a datetime:
                #     view_name = _decode_cstr(slot64)
                #     dt14 = view_name[:14]
                #     view_dt = _parse_nk_datetime(dt14)  # would parse as seconds
    
                return cls(
                    index=index,
                    timestamp_raw=timestamp_raw,
                    timestamp=timestamp,
                    type="current_view",
                    text=_decode_cstr(text384),
                    reference_name=_decode_cstr(slot64),
                )
    
            if is_specific_ref:
                text384 = payload_b[0:TEXT_LEN]  # payload[0:384]
                slot64 = payload_b[TEXT_LEN : TEXT_LEN + SLOT_LEN]  # payload[384:448]
                return cls(
                    index=index,
                    timestamp_raw=timestamp_raw,
                    timestamp=timestamp,
                    type="specific_reference",
                    text=_decode_cstr(text384),
                    reference_name=_decode_cstr(slot64),
                )
    
            # text_only: entire payload is text
            return cls(
                index=index,
                timestamp_raw=timestamp_raw,
                timestamp=timestamp,
                type="text_only",
                text=_decode_cstr(payload_b),
                reference_name=None,
            )
    
    
    @dataclass(frozen=True)
    class CmtFile:
        path: str
        size_bytes: int
        header: CmtFileHeader
        comment_start_offset: int
        comments: list[CmtComment]
    
    
    # -----------------------------
    # Main reader
    # -----------------------------
    
    
    def read_cmt(path: str | pathlib.Path, comment_start_offset: int = 1069) -> CmtFile:
        """
        Read a Nihon Kohden EEG-1100A .CMT file (based on the confirmed layout).
    
        comment_start_offset defaults to 1069 based on empirical observation for EEG-1100A
        files (this may vary for other devices / software versions).
        """
        path = str(path)
        with open(path, "rb") as f:
            data = f.read()
    
        header = CmtFileHeader.from_bytes(data)
    
        n_blocks = max(0, (len(data) - comment_start_offset) // CMT_COMMENT_BLOCK_SIZE)
    
        comments: list[CmtComment] = []
        for i in range(n_blocks):
            base = comment_start_offset + i * CMT_COMMENT_BLOCK_SIZE
            block = data[base : base + CMT_COMMENT_BLOCK_SIZE]
            comments.append(CmtComment.from_block(block, index=i))
    
        return CmtFile(
            path=path,
            size_bytes=len(data),
            header=header,
            comment_start_offset=comment_start_offset,
            comments=comments,
        )
    
    
    # -----------------------------
    # Example usage
    # -----------------------------
    if __name__ == "__main__":
        cmt = read_cmt("FJ00231Z.CMT")
    
        print("\n📁 CMT FILE SUMMARY")
        print("────────────────────────")
        print("Device        :", cmt.header.device_version)
        print("Created (raw) :", cmt.header.creation_time_raw)
        print("Created (dt)  :", cmt.header.creation_time)
        print("Comments      :", len(cmt.comments))
    
        print("\n📝 COMMENTS")
        print("────────────────────────")
    
        for c in cmt.comments:
            ref = c.reference_name if c.reference_name else "—"
    
            print(
                f"\n[{c.index:02d}] {c.type.upper()}\n"
                f"Timestamp : {c.timestamp_raw} ({c.timestamp})\n"
                f"Reference : {ref}\n"
                f"Text      ⬇️:\n",
                f"{c.text}",
            )
  6. wmvanvliet commented on Feb 9, 2026

    @wmvanvliet
    Contributor

    I've created an initial PR for this (#13642). @eulerleibniz If you could make an example file available which we could include in our testing dataset (this one should also include the .EEG file). We can add unit tests for this functionality and complete the PR.

  7. eulerleibniz commented on Feb 9, 2026

    @eulerleibniz
    Author

    @wmvanvliet at the moment i don't have access to any software that allows me to crop the egg files. That means i would need to upload a gigantic 1.3GB file, not to mention it violates privacy policy terms of my access to this data. Honestly i don't think you need the .eeg file, but i do my best to get a tiny sample if possible.

  8. wmvanvliet commented on Feb 9, 2026

    @wmvanvliet
    Contributor

    We probably cannot use files recorded from actual people who are not you. But would it be possible to record a few seconds from the machine without anyone connected to it? We do need an .EEG to test the syncing of the comments to the rest of the data.

  9. eulerleibniz commented on Feb 12, 2026

    @eulerleibniz
    Author

    I've tried asking for some data and they said that's not possible :(. So for the time being I cannot provide any sample data @wmvanvliet

  10. wmvanvliet commented on Feb 23, 2026

    @wmvanvliet
    Contributor

    Then we'll table this issue until such time as we can get our hands on some sample data.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions