Skip to content

Modified UTF-8 (CESU-8) corruption in JNI string marshaling (GetJNIString/NewJNIString) #6

Description

@Thrameos

Summary

GetJNIString()/NewJNIString() in native/jni_util.cpp used env->GetStringUTFChars()/env->NewStringUTF() -- JNI's modified UTF-8 (CESU-8) encoding, not standard UTF-8:

  • A supplementary-plane character (emoji, many CJK extension characters, etc.) is encoded as two 3-byte CESU-8 surrogate sequences instead of one proper 4-byte UTF-8 sequence.
  • An embedded NUL is encoded as an overlong 2-byte sequence (0xC0 0x80) instead of a real 0x00 byte.

The result was assigned straight into a CefString, which CEF/Chromium treats as standard UTF-8 -- so any string containing a supplementary-plane character or an embedded NUL crossing the JNI boundary in either direction gets silently corrupted.

Real-world symptom

Reported by a user embedding JCEF: content containing an emoji/astral character came back mangled, and their own application code hung because it was waiting for an exact string match that the corruption prevented from ever occurring (an application-level hang caused by the corruption, not a hang inside JCEF itself).

Reproduction (added in #1)

ModifiedUtf8Test.java isolates both directions so a bug in one can't be masked by the other:

  • nativeToJavaAstralCharacterSurvivesTitleChange: JS generates U+1F600 (😀) purely natively (String.fromCodePoint, no Java input), read back via onTitleChange (NewJNIString). Failed pre-fix: expected 😀, got ð (a single garbage character -- NewStringUTF() only understands 1-3-byte modified-UTF-8 sequences, and CEF's real 4-byte UTF-8 sequence for U+1F600 isn't valid modified UTF-8).
  • javaToNativeAstralCharacterSurvivesExecuteJavaScript: a Java string containing 😀 passed to executeJavaScript() (GetJNIString), verified inside JS so the readback can't be corrupted by NewJNIString too. Failed pre-fix: expected PASS, got FAIL:fffd (0xFFFD = Unicode replacement character -- GetStringUTFChars() encoded the Java surrogate pair as two lone-surrogate CESU-8 sequences, which CEF's real UTF-8 decoder rejected as invalid).

Fix (in #1)

Goes through JNI's non-UTF string functions (NewString()/GetStringChars()) instead, copying CefString's native UTF-16 buffer directly (jchar/char16_t are both exactly 16 bits) -- no transcoding, no ambiguity, and cheaper than either the buggy version or a UTF-8-round-trip fix would be.

Prior art (context, not an excuse to skip fixing it here)

JetBrains/jcef independently hit and fixed this exact bug in 2022 (commit 5d6a595, JBR-4468: fixed string conversion, essentially the same UTF-16-direct approach) but never filed an issue or PR against this upstream project -- confirmed via issue search, zero hits for UTF/surrogate/CESU/emoji/GetStringUTFChars/NewStringUTF prior to this report.

Environment

  • CEF 146.0.10+g8219561+chromium-146.0.7680.179, linux64

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions