Repository navigation
Investigate Android interop overhead during .NET MAUI startup, especially PeekPeer #12764
Description
Activity
- addedneeds-triageIssues that need to be assigned.Issues that need to be assigned.and removedneeds-triageIssues that need to be assigned.Issues that need to be assigned.
on Sep 11, 2026 - addedArea: PerformanceIssues with performance.Issues with performance.
on Sep 11, 2026 I added benchmark-only coverage and ran the focused matrix on a Samsung Galaxy
S23 (SM-S911B), Android 16/API 36, arm64, Release, CoreCLR, and the trimmable
type map. The device was at 38.1 C when recorded. No shipping runtime behavior
was changed.Component costs
Operation Mean Allocation Current GetIdentityHashCode90.54 ns 0 B Legacy JNI call with pre-cached class/method 99.51 ns 32 B IsSameObject48.24 ns 0 B WeakReference.TryGetTarget17.04 ns 0 B Weak reference + IsSameObject53.58 ns 0 B NewGlobalRef183.96 ns 0 B This rules out a simple “cache the
System.identityHashCodemethod ID” win. The
currentJniSystem.IdentityHashCodeimplementation already caches its type and
method and uses the generated JNI function-table path without allocating. The
legacy pre-cached call is about 10% slower and allocates aJValue[].There is no portable way to derive stable Java object identity from the raw
jobjectvalue: equivalent local/global references may have different handle
values, reference values can be recycled, and the VM may move objects.End-to-end lookup locality
The current cached
GetObject<T>/registered-peer path costs about 187–213 ns
per lookup in the looped benchmark. A benchmark-only weak recent-peer cache
checksIsSameObjectbefore performing the identity-hash and global-registry
lookup.Pattern Current registry Recent-peer fast path Effect One repeated peer 191 ns 42.7 ns 78% faster Two peers alternating every call 187 ns 265 ns 42% slower Two peers, runs of eight 197 ns 71.0 ns 64% faster Eight peers alternating every call 187 ns 282 ns 51% slower Eight peers, runs of eight 188 ns 70.2 ns 63% faster 64 peers alternating every call 213 ns 256 ns 20% slower 64 peers, runs of eight 203 ns 72.1 ns 65% faster The run-of-eight result is consistent across working-set sizes because each
group pays one miss/current-registry lookup followed by seven cheap
IsSameObjecthits. This makes callback locality the key unknown. Before
shipping this approach, we should instrument real MAUI startup to record
consecutive peer reuse or replay the callback peer sequence from a trace.Callback comparison
Operation Mean Direct managed primitive call 1.93 ns JNI -> Java -> managed primitive callback 365.5 ns JNI -> Java -> managed callback with a string 542.3 ns A roughly 190 ns peer lookup is large relative to the 365 ns primitive
roundtrip, so avoiding it on local repeated-peer callbacks could materially
reduce callback overhead.Forced GC/bridge-pressure stress
The stress benchmark performs 65,536 cached peer lookups while a dedicated
thread repeatedly forces blocking full collections. Every measured batch
verifies that the benchmark's worker completed a collection. It intentionally
includes GC suspension, bridge processing, scheduling, and any final wait for a
collection; it is not a direct benchmark of GC-bridge code.Forced pressure Registered peers Mean per lookup Max No 64 619 ns 1.73 us No 512 524 ns 756 ns Yes 64 21.6 us 36.2 us Yes 512 18.7 us 23.6 us The 30–40x degradation under forced collection is much larger than the normal
registry cost, while increasing the registered peer count from 64 to 512 did
not make the pressured case slower. This supports investigating
GC/safepoint/bridge interaction, but it does not show that the peer dictionary
ors_instancesLockblocks on the bridge. The bridge uses a separate context
registry; runtime suspension and JNI/reference access remain plausible causes.Recommended next steps
- Add low-overhead startup counters for
PeekPeercalls, hits/misses,
consecutive same-peer reuse, identity-hash calls,IsSameObjectcalls, and
time/call counts overlapping bridge generations. - Prototype a thread-local weak recent-peer +
IsSameObjectfast path and
validate its hit rate against MAUI startup before altering the global
registry. - Correlate lookup latency with GC start/end and bridge start/end events using
the original.nettraceor explicit EventSource counters. - Only investigate a more specialized identity-hash JNI helper if additional
raw function-table benchmarking shows removable exception-checking or
wrapper cost; the existing public path itself is already about 90 ns and
allocation-free on this device.
The benchmark changes are currently on the investigation branch and include
component, locality, callback, and forced-GC groups plus more reliable
BenchmarkDotNet reporting.- Add low-overhead startup counters for
Follow-up BenchmarkDotNet result on the Galaxy S23:
Operation Mean Allocated ExceptionCheck()35.93 ns 0 B ExceptionOccurred()71.88 ns 0 B IsSameObject()71.04 ns 0 B System.identityHashCode()through the current wrapper122.42 ns 0 B This hotter-device run makes the decomposition especially clear: the identity-hash wrapper is approximately one Java/JNI call plus one
ExceptionOccurred()call.System.identityHashCode(Object)declares no exception and explicitly returns zero fornull, so a specialized internal path which omits the generic post-call exception lookup is worth benchmarking/implementing. The genericCallStaticIntMethodwrapper checks because arbitrary Java static methods can throw; that policy is unnecessarily conservative for this known runtime method.Implemented and benchmarked a specialized
System.identityHashCodeJNI path. It caches the bootstrapjava/lang/SystemjclassandjmethodIDas raw per-JniRuntimeIntPtrfields during runtime construction, uses UTF-8 member lookup, and callsCallStaticIntMethodAdirectly without the genericExceptionOccurred()roundtrip. TheJniTyperemains tracked by its owning runtime; sequential Java.Interop proxy runtimes get their own valid handles. A null regression test verifiesidentityHashCode(null) == 0and no pending exception.Galaxy S23 results (15 iterations):
Method Mean Allocated Optimized GetIdentityHashCode63.71 ns 0 B Generic call with exception check 98.11 ns 0 B ExceptionOccurred()alone51.27 ns 0 B IsSameObject()48.64 ns 0 B This is about 35% faster for identity hashing. In the looped locality benchmark, the normal repeated cached peer lookup improved from about 191 ns to 158 ns (~17%).
Final
PeekPeerrerun after inlining initialization and using the per-runtime raw JNI cache (Galaxy S23, 15 iterations):Scenario Mean Allocated Current registered-peer lookup, repeated peer 158.0 ns 0 B Current lookup, working sets 2-64 156-159 ns in stable cells 0 B Benchmark-only recent-peer IsSameObjecthit40.7-42.3 ns 0 B Recent-peer path, runs of eight 64.9-66.2 ns 0 B Recent-peer miss on every alternating callback 215-220 ns 0 B The original pre-optimization repeated-peer result was about 191 ns, so the no-exception-check identity hash path improves the normal registry lookup by roughly 17%. The recent-peer optimization still depends strongly on actual callback locality.
MAUI startup A/B on Samsung A16
I tested this change against the exact runtime source used by the installed
Android runtime pack.Build
- Device: Samsung Galaxy A16 (
SM-A165F), Android 16 - App:
dotnet new maui --sample-content - App configuration: Release,
android-arm64, CoreCLR, trimmable type map - Installed runtime:
11.0.0-rc.2.26455.110 - VMR commit:
52ecb082fd3889636b5793bcaa9f4ca7cb9deb71 - Corresponding
dotnet/runtimecommit:
459f6b60db0a6fd1ed05aedd4ab4c669b9b5bada - Patched runtime: the four commits from #131952 applied to that exact commit
- Runtime build command for both variants:
./build.sh clr.runtime -os android -arch arm64 -c Release -rebuild - Base
libcoreclr.so:
77eb5c08ef0992bd5f85a4d2b6172b1c427b43f5b3d44e8d707a1c00621b6572 - Patched
libcoreclr.so:
ffaa7f11beef1e006e48e555aef2e4487d03f1cb14f6dfa313080982960364b0
The two APKs were cloned from one base APK and re-signed after replacing
lib/arm64-v8a/libcoreclr.so. Excluding signatures, the only differing APK
entry waslibcoreclr.so.Measurement
- ART compilation:
cmd package compile -m speed -f - Cold launch:
am start -S -W - Metric:
TotalTime - App data cleared after each install
- Three warmup cold launches per block
- Twelve measured cold launches per block
- Two counterbalanced passes, eight install blocks each
- 96 launches per variant
- Device temperature during measured runs: 28.1-28.5 C
Results
Variant Mean Median StdDev Base 2,598.23 ms 2,594.0 ms 42.42 ms Selective weak wait 2,571.54 ms 2,569.5 ms 33.74 ms Difference -26.69 ms (-1.03%) -24.5 ms - Bootstrap 95% CI for the mean difference: -37.71 to -16.09 ms
- Two-sided permutation p-value: < 0.00001
- 10% trimmed-mean difference: -24.15 ms
- Pass 1: -33.12 ms (-1.27%)
- Reverse-order pass 2: -20.25 ms (-0.78%)
Bridge confirmation
A separate diagnostic startup with GC logging enabled recorded:
386bridge SCCs23cross-references- callback at
14:44:29.887 - cleanup completion at
14:44:29.929
That is an approximately 42 ms accepted bridge round during startup. The
observed ~27 ms first-display improvement is consistent with removing
unnecessary UI-thread weak-reference waits during part of that round; this PR
does not make the bridge itself complete faster.Conclusion
On this peer-heavy MAUI sample, selective weak-reference waiting produces a
repeatable, statistically detectable improvement of about 20-30 ms, or
roughly 1% of cold startup time.These local CoreCLR builds do not have the official runtime pack's PGO/BOLT
optimization, so the absolute startup values should not be compared with
shipping builds. Both sides use identical local build settings, making the
relative A/B result the meaningful value.- Device: Samsung Galaxy A16 (
Android framework version
net11.0-android (Preview)
Affected platform version
Current .NET for Android trimmable type-map path. The analyzed capture was an optimized, trimmable .NET MAUI startup build using a custom MIBC. Exact deployed .NET MAUI, AndroidX, Syncfusion, device, and Android versions were not recorded with the Speedscope file and should be captured in follow-up measurements.
Description
A sampled startup profile of an optimized, trimmable .NET MAUI application suggests that Android interop contributes meaningfully to first-layout/startup overhead. The most interesting signal is
Microsoft.Android.Runtime.JavaMarshalRegisteredPeers.PeekPeer(), which appears beneath Java-to-managed callbacks and property mapping while MAUI is creating and binding native views.The profile's nested
OnMeasurestacks are structurally legitimate: MAUI enters Java to measure a child, Java synchronously calls back into managed layout/adapter code, and RecyclerView may create and bind item views during measurement. Counting nested measurement frames only once gives approximately 396 ms attributed beneathOnMeasure.Within that interval:
TemplatedItemViewHolder.BindElement.SetHandlerBind; do not add these valuesJavaMarshalRegisteredPeers.PeekPeerTrimmableTypeMapValueManager.CreatePeerThread.PollGCThe largest outer measurement region was approximately 102 ms. About 89 ms was beneath item binding, while approximately 48 ms included
PeekPeer.PeekPeer()currently performs several potentially relevant operations:System.identityHashCode()through JNI for every incoming reference.IsSameObject()to disambiguate hash collisions.Source:
android/src/Mono.Android/Microsoft.Android.Runtime/JavaMarshalRegisteredPeers.cs
Lines 195 to 218 in 26e3d80
One approximately 31 ms
PeekPeerinterval markedUNMANAGED_CODE_TIMEoverlaps GC-bridge processing on another thread. This makes GC interaction, suspension, or lock/coordination effects worth investigating. It does not prove that the dictionary lookup, identity hash, or JNI transition itself consumed 31 ms: the Speedscope output is reconstructed from samples, andCPU_TIME/UNMANAGED_CODE_TIMEare not precise scheduler or method-duration measurements.The goal of this issue is to quantify the interop contribution accurately and identify changes that reduce startup overhead without weakening Java peer identity or GC-bridge correctness.
Suggested investigation areas:
GetPeer()/PeekPeer()to measure call count, hit/miss rate, identity-hash cost, lock wait/hold time, bucket sizes, weak-reference access, andIsSameObject()calls during startup..nettraceplus Android scheduler/native tracing.System.identityHashCode()andIsSameObject()calls, while preserving identity across local/global references and hash collisions.Relevant MAUI behavior:
PlatformInterop.measureAndGetWidthAndHeight()synchronously callsview.measure(), so its inclusive duration contains all descendant Java and managed callback work rather than just JNI overhead: https://github.com/dotnet/maui/blob/b96aa036b89fe41fe1ce6cae63a3f2d1550e5682/src/Core/AndroidNative/maui/src/main/java/com/microsoft/maui/PlatformInterop.java#L437-L442Steps to Reproduce
dotnet-trace/EventPipe and retain the original.nettracein addition to exporting Speedscope JSON.ContentViewGroup.OnMeasure,LayoutViewGroup.OnMeasure,TemplatedItemViewHolder.Bind,Java.Lang.Object.GetObject,JniValueManager.GetPeer, andJavaMarshalRegisteredPeers.PeekPeer.Did you find any workaround?
No runtime workaround has been established. Reducing initial item-template complexity or avoiding excess item realization may reduce the number of interop operations, but that does not address the underlying peer-lookup cost and has not been validated as a general workaround.
Relevant log output
Important profiling limitation: TraceEvent assigns elapsed time until the next relevant observation to the previous sampled stack. These values identify areas to investigate but should not be interpreted as instrumented method timings. The original
.nettraceand correlated native/scheduler tracing are needed for causal attribution.