Repository navigation
Compose serialization into byte[] to avoid per-call Unsafe overhead - #12
Conversation
Since JDK 24 (JEP 498), every sun.misc.Unsafe memory-access call runs a
beforeMemoryAccess() check containing a volatile load. The generated
writeTo() used Unsafe.putByte/putInt/... per element, paying that cost
for every field write on modern JDKs — measured at +30% on nested-heavy
messages (BaseCommand serialize: 20.2 ops/us default vs 26.2 with
--sun-misc-unsafe-memory-access=allow on JDK 26).
Serialization now composes into a plain byte[], whose stores compile to
raw memory accesses on every JDK:
- Heap buffers are written in place through their backing array.
- Direct, composite and other buffers are composed in a scratch array
cached on the (typically pooled) message instance and transferred
with a single bulk writeBytes().
- Nested messages write via _writeTo(byte[], int) into the same array:
no per-child ensureWritable/memoryAddress/writerIndex round-trips.
- ASCII strings copy straight from String.value with System.arraycopy;
parsed string/bytes fields pass through with the indexed
getBytes(idx, array, i, len); toByteArray() writes directly into the
result array.
The write hot loop no longer uses sun.misc.Unsafe at all. writeTo() to
a CompositeByteBuf now works (previously threw
UnsupportedOperationException from array()), and serialization is fully
functional when Unsafe is unavailable.
Serialize throughput (JMH, Apple M-series, ops/us):
JDK 26 before / after JDK 17 before / after
Pulsar BaseCommand 20.2 / 25.3 (+25%) 25.5 / 25.8
AddressBook 21.5 / 26.3 (+22%) 19.2 / 24.5 (+27%)
Pulsar MessageMetadata 14.4 / 14.2 13.4 / 14.0
AddressBook fill+ser 14.6 / 15.1 13.2 / 14.1
Tradeoffs: messages under ~15 bytes serialized to non-heap buffers pay
~2-3ns for the bulk transfer; non-ASCII strings allocate one temporary
array during encoding (previously streamed via reserveAndWriteUtf8).
There was a problem hiding this comment.
Pull request overview
This PR reworks generated protobuf serialization to compose directly into a byte[] (with an index cursor) and then bulk-write to the target ByteBuf, avoiding per-field sun.misc.Unsafe memory-access overhead introduced by newer JDKs and improving behavior for non-array/non-addressable buffers.
Changes:
- Replace per-element Unsafe writes with array-based raw write helpers and a generated
_writeTo(byte[], int)fast path for nested messages. - Update all generated field serializers (numbers/strings/bytes/messages/maps) to write into the shared array cursor.
- Add a regression test covering heap/direct/composite write targets and re-serialization after parse.
Reviewed changes
Copilot reviewed 13 out of 13 changed files in this pull request and generated 1 comment.
Show a summary per file
| File | Description |
|---|---|
| tests/src/test/java/io/streamnative/lightproto/tests/WriteTargetTypesTest.java | Adds coverage for consistent writeTo() output across heap/direct/composite buffers and parsed-message round-trip. |
| code-generator/src/main/resources/io/streamnative/lightproto/generator/LightProtoCodec.java | Replaces Unsafe hot-loop raw writes with array-based raw writers and introduces scratch-buffer support. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoMessage.java | Generates array-composition writeTo() and the new cross-package _writeTo(byte[], int) method plus per-instance scratch caching. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoField.java | Updates tag writing to advance an array cursor (_i) instead of an Unsafe address. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoNumberField.java | Updates numeric field serialization to write into the array cursor. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoBytesField.java | Updates bytes field serialization to copy into the array cursor and advance it. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoStringField.java | Updates string field serialization to write UTF-8 bytes into the array cursor (including parsed-buffer passthrough). |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoRepeatedNumberField.java | Updates packed/unpacked repeated number serialization to the array cursor. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoRepeatedBytesField.java | Updates repeated bytes serialization to copy into the array cursor. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoRepeatedStringField.java | Updates repeated string serialization to the array cursor and parsed-buffer passthrough. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoMessageField.java | Updates nested-message serialization to call _writeTo(byte[], int) instead of per-child ByteBuf writes. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoRepeatedMessageField.java | Updates repeated message serialization to call _writeTo(byte[], int) for each element. |
| code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoMapField.java | Updates map-entry serialization to write into the array cursor and avoids _i loop-variable collision. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| /** Returns current if it can hold size bytes, otherwise a larger replacement. */ | ||
| static byte[] scratchFor(byte[] current, int size) { | ||
| if (current != null && current.length >= size) { | ||
| return current; | ||
| } |
There was a problem hiding this comment.
Good catch on the near-cap band — fixed in b060c40: growth is now clamped to SCRATCH_RETAIN_MAX for retainable sizes (so a ~0.5–1 MiB message settles on a retained cap-sized array instead of re-allocating an unretainable 2× array on every write), and beyond-cap outliers allocate exactly since they're never retained. Added testScratchFor covering reuse, doubling, the near-cap clamp, and the exact-size outlier path. The 64-byte floor for a null scratch is intentional: it's retained on first use and avoids growth churn for tiny messages, so the size==0 case allocates once per instance at most.
Review feedback: doubling growth in scratchFor() could produce an array larger than SCRATCH_RETAIN_MAX for messages just under the cap. Such an array is never retained on the instance, so every writeTo() of a ~0.5-1 MiB message to a non-heap buffer would re-allocate ~2x the message size. Growth is now clamped to the retain cap for retainable sizes, and beyond-cap outliers allocate exactly (they are never retained, so amortization is pointless). Covered by testScratchFor.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 14 out of 14 changed files in this pull request and generated no new comments.
Comments suppressed due to low confidence (1)
code-generator/src/main/java/io/streamnative/lightproto/generator/LightProtoMessage.java:240
- writeTo(ByteBuf) now computes size / ensures capacity / allocates scratch before validating required fields (the required-field check only happens inside _writeTo). If required fields are missing, this can unnecessarily grow the target buffer or allocate scratch before throwing, and it changes behavior vs the previous implementation which validated first. Consider calling checkRequiredFields() at the start of writeTo() (while keeping the check in _writeTo for callers like toByteArray() / nested writes).
w.format(" @Override public int writeTo(io.netty.buffer.ByteBuf _b) {\n");
w.format(" int _serializedSize = getSerializedSize();\n");
w.format(" if (_b.hasArray()) {\n");
…20) * Add large-message serialize benchmark and non-array target identity sweep LargeMessageBenchmark serializes the Pulsar topic-list shape from 600 B to 4.6 MB, a varint-dense repeated-int64 message at 2 KB and 8 KB, and a 2 MB bytes payload into pooled direct buffers. NonArrayTargetIdentityTest checks that writeTo() to direct, offset and multi-component composite targets is byte-identical to the heap-array path for sizes swept byte by byte across every plausible internal boundary, for built and parsed messages and across repeated writes. * Write direct buffers in place through their NIO view above 512 bytes Since #12, writeTo() to a non-array buffer stages the whole message in a heap byte[] scratch and bulk-copies it. Messages above SCRATCH_RETAIN_MAX (1 MiB) never retain that scratch, so every write allocated a fresh full-size array — multi-MB G1-humongous allocations that OOMed Pulsar's proxy back-pressure test (apache/pulsar#26256, together with the clear() retention fixed in #19). Below the cap the copy itself was still paid. A single-region direct buffer exposes its memory as a java.nio.ByteBuffer through ByteBuf.internalNioBuffer(). Absolute puts on a DirectByteBuffer compile to a bounds check plus a jdk.internal.misc.Unsafe store — which, unlike sun.misc.Unsafe, carries no JDK 24+ deprecation check — so the message can be written in place: no scratch array and no bulk copy, at any size. writeTo() now dispatches heap buffers in place through the backing array (unchanged), single-region direct buffers larger than NIO_WRITE_MIN (512 bytes) through the NIO view, and everything else (small messages; composites and other buffers without a single NIO region) through the scratch path as before. Above the threshold no direct-buffer write touches the scratch, so it only grows past 512 bytes for composite targets. The threshold exists because the view's per-put cost is a fixed tax per message while the copy it saves grows with size. Interleaved JMH on pooled direct buffers (JDK 21/26): the view is 15-19% slower on the ~70-byte varint-dense MessageMetadata, at parity on BaseCommand, 20% faster at 600 bytes, and 35-40% faster from 6 KB to 100 KB; on the 2 MB / 4.6 MB cases it removes the per-write allocation (-70% / -55%) and matches the per-field ByteBuf-API write-through of #18, which it replaces. The field emitters are parameterized over the write sink (WriteSink.ARRAY / WriteSink.NIO): one emitter produces both _writeTo(byte[], int) and _writeTo(ByteBuffer, int), differing only in the sink variable and in how bulk data is copied out of a ByteBuf; every raw writer in LightProtoCodec is overloaded for both sinks. NonArrayTargetIdentityTest sweeps sizes byte by byte across every boundary (64 B .. 1 MiB) on direct, offset and multi-component composite targets, for built and parsed messages, repeated strings incl. non-ASCII, bytes payloads, nested trees and the Pulsar BaseCommand shape. NioWriteTest checks the routing flips exactly at NIO_WRITE_MIN, that composites keep the scratch path, and (via ThreadMXBean.getThreadAllocatedBytes) that 5 writes of a 4.6 MB topic list or a 5 MB payload allocate less than a quarter of one message. LargeMessageBenchmark covers 600 B .. 4.6 MB topic lists, varint-dense 2 KB / 8 KB messages and a 2 MB payload.
Its last readers were the Unsafe raw writers that streamnative#12 replaced with the byte[] writers. The NIO writers added later take the byte order from the buffer itself, with nb.order().
* Remove the unused ByteBuf writers from LightProtoCodec Generated code serializes through the raw byte[] and NIO writers, so LightProtoCodec's ByteBuf writers (writeVarInt, writeVarInt64, writeSignedVarInt*, writeFixedInt*, writeFloat, writeDouble and writeString) were only called by tests. The codec is copied into every generated package, along with these 78 unused lines. LightProtoCodecTest and SegmentedByteBufTest now encode with the raw writers that generated code uses. Their assertions are unchanged, so the protobuf-java CodedInputStream checks now cover the production writers, which had no direct unit tests. * Remove the unused LITTLE_ENDIAN constant from LightProtoCodec Its last readers were the Unsafe raw writers that #12 replaced with the byte[] writers. The NIO writers added later take the byte order from the buffer itself, with nb.order().
Motivation
Since JDK 24 (JEP 498), every
sun.misc.Unsafememory-access call runs abeforeMemoryAccess()check containing a volatile load (visible with-XX:+PrintInlining:getByte/putByte → beforeMemoryAccess → isMemoryAccessWarned → getBooleanVolatile). The generatedwriteTo()usedUnsafe.putByte/putInt/...per element, paying that cost for every field write on modern JDKs.Sizing the effect on JDK 26 (serialize, ops/µs): BaseCommand 20.2 default vs 26.2 with
--sun-misc-unsafe-memory-access=allow(+30%), concentrated in nested-heavy messages with dense streams of small writes; JDK 17 (pre-JEP 498) agrees at 25.5.Changes
Serialization now composes into a plain
byte[], whose stores compile to raw memory accesses on every JDK — nosun.misc.Unsafeanywhere in the write hot loop:ensureWritable, which may replace it).writeBytes(). Scratch arrays above 1 MiB are not retained._writeTo(byte[], int)into the same array — eliminating the per-childensureWritable/memoryAddress/writerIndexround-trips the old path did for every sub-message.String.valueviaSystem.arraycopy; parsed string/bytes fields pass through with indexedgetBytes(idx, array, i, len);toByteArray()writes directly into the result array.Also fixed along the way
writeTo()to aCompositeByteBuf(or any buffer without a backing array or memory address) now works — it previously threwUnsupportedOperationExceptionfromarray().Unsafeis unavailable (the ASCII string fast path falls back togetBytes(ISO_8859_1)).Results
Serialize throughput (JMH, Apple M-series, ops/µs):
BaseCommand serialize vs protobuf-java goes from 1.2× to 1.5×. The JDK 26 and JDK 17 columns now match: the JEP 498 performance cliff is gone, and the gains on AddressBook/BaseCommand come from the removed per-child buffer round-trips, which are JDK-independent.
Zero steady-state allocation confirmed with
-prof gc(≈10⁻⁴ B/op).Tradeoffs
SimpleBenchmark.lightProtoSerialize, which serializes a bare 8-byte Point due to a pre-existing benchmark bug).reserveAndWriteUtf8); ASCII strings — the overwhelmingly common case — got faster.Verification
WriteTargetTypesTest: byte-identical output across heap (withensureWritablegrowth), pooled heap witharrayOffset, pooled direct, and composite targets, plus re-serialization of parsed messages (lazy passthrough writes).