Skip to content

Investigate Android interop overhead during .NET MAUI startup, especially PeekPeer #12764

Description

@simonrozsival

Android framework version

net11.0-android (Preview)

Affected platform version

Current .NET for Android trimmable type-map path. The analyzed capture was an optimized, trimmable .NET MAUI startup build using a custom MIBC. Exact deployed .NET MAUI, AndroidX, Syncfusion, device, and Android versions were not recorded with the Speedscope file and should be captured in follow-up measurements.

Description

A sampled startup profile of an optimized, trimmable .NET MAUI application suggests that Android interop contributes meaningfully to first-layout/startup overhead. The most interesting signal is Microsoft.Android.Runtime.JavaMarshalRegisteredPeers.PeekPeer(), which appears beneath Java-to-managed callbacks and property mapping while MAUI is creating and binding native views.

The profile's nested OnMeasure stacks are structurally legitimate: MAUI enters Java to measure a child, Java synchronously calls back into managed layout/adapter code, and RecyclerView may create and bind item views during measurement. Counting nested measurement frames only once gives approximately 396 ms attributed beneath OnMeasure.

Within that interval:

Stack attribution Approximate time Notes
TemplatedItemViewHolder.Bind 250 ms About 63% of measurement-attributed time; includes template creation, binding, handler creation, and native realization
Element.SetHandler 204 ms Overlaps Bind; do not add these values
JavaMarshalRegisteredPeers.PeekPeer 63 ms Existing-peer lookup path
TrimmableTypeMapValueManager.CreatePeer 1.6 ms New peer resolution through the trimmable type map was small specifically within measurement
Thread.PollGC 33 ms Sample attribution associated with GC/safepoint activity, not necessarily a 33 ms pause

The largest outer measurement region was approximately 102 ms. About 89 ms was beneath item binding, while approximately 48 ms included PeekPeer.

PeekPeer() currently performs several potentially relevant operations:

  1. Calls Java System.identityHashCode() through JNI for every incoming reference.
  2. Acquires the process-wide registered-peer lock.
  3. Looks up the identity-hash bucket.
  4. Reads weak-reference targets.
  5. Uses JNI IsSameObject() to disambiguate hash collisions.

Source:

public static IJavaPeerable? PeekPeer (JniObjectReference reference)
{
if (!reference.IsValid)
return null;
int key = JniEnvironment.References.GetIdentityHashCode (reference);
lock (s_instancesLock) {
if (!RegisteredInstances.TryGetValue (key, out RegisteredPeerBucket peers))
return null;
for (int i = peers.Count - 1; i >= 0; i--) {
if (peers [i].Target is IJavaPeerable peer
&& JniEnvironment.Types.IsSameObject (reference, peer.PeerReference))
{
return peer;
}
}
if (peers.Count == 0)
RegisteredInstances.Remove (key);
}
return null;
}

One approximately 31 ms PeekPeer interval marked UNMANAGED_CODE_TIME overlaps GC-bridge processing on another thread. This makes GC interaction, suspension, or lock/coordination effects worth investigating. It does not prove that the dictionary lookup, identity hash, or JNI transition itself consumed 31 ms: the Speedscope output is reconstructed from samples, and CPU_TIME/UNMANAGED_CODE_TIME are not precise scheduler or method-duration measurements.

The goal of this issue is to quantify the interop contribution accurately and identify changes that reduce startup overhead without weakening Java peer identity or GC-bridge correctness.

Suggested investigation areas:

  1. Add focused instrumentation or counters around GetPeer()/PeekPeer() to measure call count, hit/miss rate, identity-hash cost, lock wait/hold time, bucket sizes, weak-reference access, and IsSameObject() calls during startup.
  2. Correlate peer lookup with CLR GC, Java GC, and GC-bridge phases using the original .nettrace plus Android scheduler/native tracing.
  3. Determine whether repeated Java-to-managed callbacks or MAUI property mappings perform avoidable peer lookups for the same object during initial handler creation and layout.
  4. Evaluate whether the peer registry can reduce or avoid repeated JNI System.identityHashCode() and IsSameObject() calls, while preserving identity across local/global references and hash collisions.
  5. Evaluate contention-reduction approaches for the process-wide peer registry, such as partitioning or a read-optimized representation, while accounting for peer reconciliation and bridge collection.
  6. Compare the trimmable type-map path with other type-map/runtime configurations using the same application and startup scenario.
  7. Separate the cost of the JNI transition from native/ART work, GC waits, runtime suspension, and Speedscope sample-gap attribution before selecting an optimization.

Relevant MAUI behavior: PlatformInterop.measureAndGetWidthAndHeight() synchronously calls view.measure(), so its inclusive duration contains all descendant Java and managed callback work rather than just JNI overhead: https://github.com/dotnet/maui/blob/b96aa036b89fe41fe1ce6cae63a3f2d1550e5682/src/Core/AndroidNative/maui/src/main/java/com/microsoft/maui/PlatformInterop.java#L437-L442

Steps to Reproduce

  1. Build an optimized, trimmable .NET MAUI Android application using the trimmable type-map path. The investigated build also used a custom MIBC.
  2. Use a startup page containing nested MAUI layouts and a RecyclerView-backed items control with non-trivial item templates. The analyzed application also used Syncfusion controls.
  3. Collect startup tracing with dotnet-trace/EventPipe and retain the original .nettrace in addition to exporting Speedscope JSON.
  4. Inspect the first-layout call stacks beneath ContentViewGroup.OnMeasure, LayoutViewGroup.OnMeasure, TemplatedItemViewHolder.Bind, Java.Lang.Object.GetObject, JniValueManager.GetPeer, and JavaMarshalRegisteredPeers.PeekPeer.
  5. Correlate those samples with GC and GC-bridge events and, ideally, Perfetto or simpleperf data.
  6. Repeat with targeted instrumentation and alternate type-map/runtime configurations to quantify call frequency and actual operation latency.

Did you find any workaround?

No runtime workaround has been established. Reducing initial item-template complexity or avoiding excess item realization may reduce the number of interop operations, but that does not address the underlying peer-lookup cost and has not been validated as a general workaround.

Relevant log output

OnMeasure union (nested frames counted once):       ~395.749 ms
TemplatedItemViewHolder.Bind under OnMeasure:       ~249.674 ms
Element.SetHandler under OnMeasure:                 ~203.691 ms (overlaps Bind)
JavaMarshalRegisteredPeers.PeekPeer under measure:   ~63.105 ms
TrimmableTypeMapValueManager.CreatePeer:               ~1.639 ms

Largest outer OnMeasure region:                     ~101.656 ms
  TemplatedItemViewHolder.Bind:                       ~89.284 ms
  JavaMarshalRegisteredPeers.PeekPeer:                ~47.865 ms

Suspicious overlap:
  main thread PeekPeer -> UNMANAGED_CODE_TIME:       1162.593-1193.455 ms
  GC bridge BridgeProcessingStarted:                1162.669-1191.913 ms
  ProcessCollectedContexts begins:                  1191.913 ms

Important profiling limitation: TraceEvent assigns elapsed time until the next relevant observation to the previous sampled stack. These values identify areas to investigate but should not be interpreted as instrumented method timings. The original .nettrace and correlated native/scheduler tracing are needed for causal attribution.

Activity

  1. added
    needs-triageIssues that need to be assigned.
    and removed
    needs-triageIssues that need to be assigned.
    on Sep 11, 2026
  2. added this to the .NET 12 milestone on Sep 11, 2026
  3. simonrozsival commented on Sep 11, 2026

    @simonrozsival
    MemberAuthor

    I added benchmark-only coverage and ran the focused matrix on a Samsung Galaxy
    S23 (SM-S911B), Android 16/API 36, arm64, Release, CoreCLR, and the trimmable
    type map. The device was at 38.1 C when recorded. No shipping runtime behavior
    was changed.

    Component costs

    Operation Mean Allocation
    Current GetIdentityHashCode 90.54 ns 0 B
    Legacy JNI call with pre-cached class/method 99.51 ns 32 B
    IsSameObject 48.24 ns 0 B
    WeakReference.TryGetTarget 17.04 ns 0 B
    Weak reference + IsSameObject 53.58 ns 0 B
    NewGlobalRef 183.96 ns 0 B

    This rules out a simple “cache the System.identityHashCode method ID” win. The
    current JniSystem.IdentityHashCode implementation already caches its type and
    method and uses the generated JNI function-table path without allocating. The
    legacy pre-cached call is about 10% slower and allocates a JValue[].

    There is no portable way to derive stable Java object identity from the raw
    jobject value: equivalent local/global references may have different handle
    values, reference values can be recycled, and the VM may move objects.

    End-to-end lookup locality

    The current cached GetObject<T>/registered-peer path costs about 187–213 ns
    per lookup in the looped benchmark. A benchmark-only weak recent-peer cache
    checks IsSameObject before performing the identity-hash and global-registry
    lookup.

    Pattern Current registry Recent-peer fast path Effect
    One repeated peer 191 ns 42.7 ns 78% faster
    Two peers alternating every call 187 ns 265 ns 42% slower
    Two peers, runs of eight 197 ns 71.0 ns 64% faster
    Eight peers alternating every call 187 ns 282 ns 51% slower
    Eight peers, runs of eight 188 ns 70.2 ns 63% faster
    64 peers alternating every call 213 ns 256 ns 20% slower
    64 peers, runs of eight 203 ns 72.1 ns 65% faster

    The run-of-eight result is consistent across working-set sizes because each
    group pays one miss/current-registry lookup followed by seven cheap
    IsSameObject hits. This makes callback locality the key unknown. Before
    shipping this approach, we should instrument real MAUI startup to record
    consecutive peer reuse or replay the callback peer sequence from a trace.

    Callback comparison

    Operation Mean
    Direct managed primitive call 1.93 ns
    JNI -> Java -> managed primitive callback 365.5 ns
    JNI -> Java -> managed callback with a string 542.3 ns

    A roughly 190 ns peer lookup is large relative to the 365 ns primitive
    roundtrip, so avoiding it on local repeated-peer callbacks could materially
    reduce callback overhead.

    Forced GC/bridge-pressure stress

    The stress benchmark performs 65,536 cached peer lookups while a dedicated
    thread repeatedly forces blocking full collections. Every measured batch
    verifies that the benchmark's worker completed a collection. It intentionally
    includes GC suspension, bridge processing, scheduling, and any final wait for a
    collection; it is not a direct benchmark of GC-bridge code.

    Forced pressure Registered peers Mean per lookup Max
    No 64 619 ns 1.73 us
    No 512 524 ns 756 ns
    Yes 64 21.6 us 36.2 us
    Yes 512 18.7 us 23.6 us

    The 30–40x degradation under forced collection is much larger than the normal
    registry cost, while increasing the registered peer count from 64 to 512 did
    not make the pressured case slower. This supports investigating
    GC/safepoint/bridge interaction, but it does not show that the peer dictionary
    or s_instancesLock blocks on the bridge. The bridge uses a separate context
    registry; runtime suspension and JNI/reference access remain plausible causes.

    Recommended next steps

    1. Add low-overhead startup counters for PeekPeer calls, hits/misses,
      consecutive same-peer reuse, identity-hash calls, IsSameObject calls, and
      time/call counts overlapping bridge generations.
    2. Prototype a thread-local weak recent-peer + IsSameObject fast path and
      validate its hit rate against MAUI startup before altering the global
      registry.
    3. Correlate lookup latency with GC start/end and bridge start/end events using
      the original .nettrace or explicit EventSource counters.
    4. Only investigate a more specialized identity-hash JNI helper if additional
      raw function-table benchmarking shows removable exception-checking or
      wrapper cost; the existing public path itself is already about 90 ns and
      allocation-free on this device.

    The benchmark changes are currently on the investigation branch and include
    component, locality, callback, and forced-GC groups plus more reliable
    BenchmarkDotNet reporting.

  4. simonrozsival commented on Sep 11, 2026

    @simonrozsival
    MemberAuthor

    Follow-up BenchmarkDotNet result on the Galaxy S23:

    Operation Mean Allocated
    ExceptionCheck() 35.93 ns 0 B
    ExceptionOccurred() 71.88 ns 0 B
    IsSameObject() 71.04 ns 0 B
    System.identityHashCode() through the current wrapper 122.42 ns 0 B

    This hotter-device run makes the decomposition especially clear: the identity-hash wrapper is approximately one Java/JNI call plus one ExceptionOccurred() call. System.identityHashCode(Object) declares no exception and explicitly returns zero for null, so a specialized internal path which omits the generic post-call exception lookup is worth benchmarking/implementing. The generic CallStaticIntMethod wrapper checks because arbitrary Java static methods can throw; that policy is unnecessarily conservative for this known runtime method.

  5. simonrozsival commented on Sep 11, 2026

    @simonrozsival
    MemberAuthor

    Implemented and benchmarked a specialized System.identityHashCode JNI path. It caches the bootstrap java/lang/System jclass and jmethodID as raw per-JniRuntime IntPtr fields during runtime construction, uses UTF-8 member lookup, and calls CallStaticIntMethodA directly without the generic ExceptionOccurred() roundtrip. The JniType remains tracked by its owning runtime; sequential Java.Interop proxy runtimes get their own valid handles. A null regression test verifies identityHashCode(null) == 0 and no pending exception.

    Galaxy S23 results (15 iterations):

    Method Mean Allocated
    Optimized GetIdentityHashCode 63.71 ns 0 B
    Generic call with exception check 98.11 ns 0 B
    ExceptionOccurred() alone 51.27 ns 0 B
    IsSameObject() 48.64 ns 0 B

    This is about 35% faster for identity hashing. In the looped locality benchmark, the normal repeated cached peer lookup improved from about 191 ns to 158 ns (~17%).

  6. simonrozsival commented on Sep 11, 2026

    @simonrozsival
    MemberAuthor

    Final PeekPeer rerun after inlining initialization and using the per-runtime raw JNI cache (Galaxy S23, 15 iterations):

    Scenario Mean Allocated
    Current registered-peer lookup, repeated peer 158.0 ns 0 B
    Current lookup, working sets 2-64 156-159 ns in stable cells 0 B
    Benchmark-only recent-peer IsSameObject hit 40.7-42.3 ns 0 B
    Recent-peer path, runs of eight 64.9-66.2 ns 0 B
    Recent-peer miss on every alternating callback 215-220 ns 0 B

    The original pre-optimization repeated-peer result was about 191 ns, so the no-exception-check identity hash path improves the normal registry lookup by roughly 17%. The recent-peer optimization still depends strongly on actual callback locality.

  7. simonrozsival commented on Sep 12, 2026

    @simonrozsival
    MemberAuthor

    MAUI startup A/B on Samsung A16

    I tested this change against the exact runtime source used by the installed
    Android runtime pack.

    Build

    • Device: Samsung Galaxy A16 (SM-A165F), Android 16
    • App: dotnet new maui --sample-content
    • App configuration: Release, android-arm64, CoreCLR, trimmable type map
    • Installed runtime: 11.0.0-rc.2.26455.110
    • VMR commit: 52ecb082fd3889636b5793bcaa9f4ca7cb9deb71
    • Corresponding dotnet/runtime commit:
      459f6b60db0a6fd1ed05aedd4ab4c669b9b5bada
    • Patched runtime: the four commits from #131952 applied to that exact commit
    • Runtime build command for both variants:
      ./build.sh clr.runtime -os android -arch arm64 -c Release -rebuild
    • Base libcoreclr.so:
      77eb5c08ef0992bd5f85a4d2b6172b1c427b43f5b3d44e8d707a1c00621b6572
    • Patched libcoreclr.so:
      ffaa7f11beef1e006e48e555aef2e4487d03f1cb14f6dfa313080982960364b0

    The two APKs were cloned from one base APK and re-signed after replacing
    lib/arm64-v8a/libcoreclr.so. Excluding signatures, the only differing APK
    entry was libcoreclr.so
    .

    Measurement

    • ART compilation: cmd package compile -m speed -f
    • Cold launch: am start -S -W
    • Metric: TotalTime
    • App data cleared after each install
    • Three warmup cold launches per block
    • Twelve measured cold launches per block
    • Two counterbalanced passes, eight install blocks each
    • 96 launches per variant
    • Device temperature during measured runs: 28.1-28.5 C

    Results

    Variant Mean Median StdDev
    Base 2,598.23 ms 2,594.0 ms 42.42 ms
    Selective weak wait 2,571.54 ms 2,569.5 ms 33.74 ms
    Difference -26.69 ms (-1.03%) -24.5 ms
    • Bootstrap 95% CI for the mean difference: -37.71 to -16.09 ms
    • Two-sided permutation p-value: < 0.00001
    • 10% trimmed-mean difference: -24.15 ms
    • Pass 1: -33.12 ms (-1.27%)
    • Reverse-order pass 2: -20.25 ms (-0.78%)

    Bridge confirmation

    A separate diagnostic startup with GC logging enabled recorded:

    • 386 bridge SCCs
    • 23 cross-references
    • callback at 14:44:29.887
    • cleanup completion at 14:44:29.929

    That is an approximately 42 ms accepted bridge round during startup. The
    observed ~27 ms first-display improvement is consistent with removing
    unnecessary UI-thread weak-reference waits during part of that round; this PR
    does not make the bridge itself complete faster.

    Conclusion

    On this peer-heavy MAUI sample, selective weak-reference waiting produces a
    repeatable, statistically detectable improvement of about 20-30 ms, or
    roughly 1% of cold startup time.

    These local CoreCLR builds do not have the official runtime pack's PGO/BOLT
    optimization, so the absolute startup values should not be compared with
    shipping builds. Both sides use identical local build settings, making the
    relative A/B result the meaningful value.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions