Skip to content

[CoreCLR/NativeAOT] Replace LLVM IR compilation with binary data packaging #10784

Description

@simonrozsival

Goal

Simplify the app build pipeline by eliminating LLVM IR code generation, llc compilation, and ld linking for CoreCLR and NativeAOT builds. This reduces build tool dependencies, build complexity, and long-term maintenance cost.

This must not come at the expense of a significant measurable startup performance regression. We need to measure the actual impact on real devices before and after.

Summary

Every .NET for Android app build generates 5-7 LLVM IR (.ll) files per ABI, compiles them with llc, and links them with ld into libxamarin-app.so. This shared library is almost entirely read-only data — configuration structs, lookup tables, and pre-allocated buffers. The LLVM IR pipeline is a heavyweight code generation + compilation step for what is fundamentally a data packaging problem.

We already have a simpler mechanism for packaging data into the APK: DSOWrapperGenerator wraps arbitrary binary files in a minimal ELF .so using llvm-objcopy, places them in lib/{abi}/, and the runtime mmaps them directly from the APK. This is how assembly stores (assemblies.blob) work today.

Proposal: Replace the LLVM IR → llc → ld pipeline with direct binary serialization → llvm-objcopy for all configuration data. The existing LlvmIrComposer subclasses already compute all the values in C# — we just change the output stage from "emit LLVM IR text" to "write raw bytes matching the C struct layout."

Scope: CoreCLR and NativeAOT only. MonoVM continues using LLVM IR until deprecated.

Dependency: Trimmable TypeMap (in progress) eliminates typemaps and marshal_methods. This proposal handles the remaining .ll files.

What's in libxamarin-app.so today

File Contents Replacement
typemaps.*.ll Java↔.NET type mapping tables Trimmable TypeMap (in progress)
marshal_methods.*.ll Marshal method init stub Trimmable TypeMap (in progress)
environment.*.ll ApplicationConfig struct, runtime properties, DSO cache, env vars Binary blob (this proposal)
compressed_assemblies.*.ll Decompression descriptors + zero-init buffer Binary blob + dynamic alloc
jni_remap.*.ll JNI type/method remapping (Intune MAM) Binary blob or managed Dictionary
pinvoke_preserve.*.ll P/Invoke symbol preservation + find_pinvoke() (CoreCLR unified linking only) Linker flags (this proposal)
jni_init_funcs.*.ll NativeAOT: JNI_OnLoad dispatch Generated C# (this proposal)

Proposed approach

1. Binary config blob in ELF wrapper (replaces environment, compressed_assemblies, jni_remap)

Build time:

The existing LlvmIrComposer subclasses (e.g., ApplicationConfigNativeAssemblyGeneratorCLR, CompressedAssembliesNativeAssemblyGenerator) already have a two-stage pipeline:

  1. Compose — compute all values, populate StructureInstance<ApplicationConfig>, List<StructureInstance<DSOCacheEntry>>, etc.
  2. Generate — serialize to LLVM IR text via LlvmIrGenerator

We replace stage 2: instead of LlvmIrGenerator emitting text, a new BinaryBlobWriter serializes StructureInstance<T> objects directly to bytes using the existing StructureInfo metadata (field offsets, sizes, alignment, padding). Same data, same layout, no compilation step.

The output config.bin has a simple section-based format:

[Header: magic, version, section_count, section_offsets[], section_sizes[]]
[Section 0: ApplicationConfig]           // binary-compatible with C struct
[Section 1: runtime property names]      // null-terminated string table
[Section 2: runtime property values]     // null-terminated string table
[Section 3: DSOCacheEntry[]]             // array of C structs
[Section 4: DSO name string data]
[Section 5: DSOApkEntry[]]              // template entries (fd filled at runtime)
[Section 6: compressed assembly descriptors]

Then wrap and package:

DSOWrapperGenerator.WrapIt(config.bin) → lib/{abi}/libruntime-config.so

This uses llvm-objcopy --add-section payload=config.bin — the same tool already used for assembly stores. No llc, no ld.

Runtime:

The zip scan (which already runs to find assembly stores) discovers libruntime-config.so, mmaps it from the APK, and get_wrapper_dso_payload_pointer_and_size() returns a direct pointer to the payload.

// Same infrastructure as assembly store loading
auto [data, size] = get_wrapper_dso_payload_pointer_and_size(mmap_info, "libruntime-config.so");

// Parse header, cast section pointers directly to C structs
auto header = static_cast<const ConfigBlobHeader*>(data);
auto base = static_cast<const uint8_t*>(data);
application_config = reinterpret_cast<const ApplicationConfig*>(base + header->sections[0].offset);
dso_cache = reinterpret_cast<const DSOCacheEntry*>(base + header->sections[3].offset);

No parsing, no copying, no deserialization. The MSBuild task writes bytes matching the C struct memory layout. The C++ code casts pointers into the mmap'd region. Data is demand-paged from the APK by the kernel — same mechanism as today.

Changes to libmonodroid.so: Replace extern declarations (currently resolved by libxamarin-app.so at load time) with static pointer globals initialized from the mmap'd blob. This follows the same pattern already used for assembly_store data.

2. Dynamic allocation (replaces BSS pre-allocated buffers)

The zero-initialized buffers in libxamarin-app.so (assembly store slots, decompression buffer) use BSS sections, which the kernel backs with mmap(MAP_ANONYMOUS) + demand paging. Allocating with new[] uses the same kernel mechanism for large allocations. Replace:

// Before: LLVM IR pre-allocates in BSS
extern uint8_t uncompressed_assemblies_data_buffer[];
extern AssemblyStoreSingleAssemblyRuntimeData assembly_store_bundled_assemblies[];

// After: allocate at startup (size from config blob)
auto buffer = new uint8_t[config->total_uncompressed_size]();
auto assemblies = new AssemblyStoreSingleAssemblyRuntimeData[config->assembly_count]();

3. Environment variables → Java Os.setenv()

Generate Java code calling Os.setenv() before initInternal(), following the pattern NativeAOT already uses (NativeAotEnvironmentVars.java).

4. pinvoke_preserve.*.ll → Linker flags + dlsym (CoreCLR unified linking only)

This is the one file with actual executable code: find_pinvoke() maps (library_hash, entrypoint_hash) → function pointer via nested switch statements. It serves two purposes:

Linker symbol preservation — references to symbols like @SystemNative_Bind prevent --gc-sections from stripping them. Replace with --undefined=<symbol> linker flags. PinvokeScanner already produces the symbol list, and NativeLinker.cs already supports --export-dynamic-symbol — the infrastructure is in place. (There's even a TODO in dynamic.cc:88 where the team considered this approach.)

Runtime P/Invoke resolution — replace with dlsym(RTLD_DEFAULT, entrypoint_name), which already exists as a fallback in dynamic.cc. P/Invoke results are cached by CoreCLR — each entrypoint is resolved once. Performance impact to be measured.

5. jni_init_funcs.*.ll → Generated C# (NativeAOT only)

Replace the LLVM IR function pointer array with generated C# using [DllImport("__Internal")]:

static class JniInitFunctions
{
    [DllImport("__Internal")]
    static extern int JNI_OnLoad_SystemNative (IntPtr vm, IntPtr reserved);

    [DllImport("__Internal")]
    static extern int JNI_OnLoad_CryptoNative (IntPtr vm, IntPtr reserved);

    public static void CallAll (IntPtr vm)
    {
        JNI_OnLoad_SystemNative (vm, IntPtr.Zero);
        JNI_OnLoad_CryptoNative (vm, IntPtr.Zero);
    }
}

NativeAOT compiles [DllImport("__Internal")] to direct native call instructions — zero overhead, compile-time symbol resolution, missing symbol = link error (not runtime crash). The DirectPInvoke infrastructure already exists in Microsoft.Android.Sdk.NativeAOT.targets.

Performance

The primary goal is long-term maintainability and build simplification. However, this must not come at the expense of a significant measurable startup regression.

The proposed approach uses the same mmap-from-APK mechanism as today — config data is still accessed via direct pointer dereferences into memory-mapped regions. The main differences are: (1) eliminating dlopen("libxamarin-app.so") and its symbol resolution overhead, (2) replacing BSS pre-allocated buffers with dynamic new[], and (3) replacing find_pinvoke() with dlsym for unified linking.

We need to measure the actual performance impact on real devices (high-end and low-end) with representative apps before and after. A feature flag should allow A/B comparison.

Build time

Current Proposed
5-7 llc invocations per ABI (LLVM IR compilation) Eliminated
1 ld invocation per ABI (native linking) Eliminated
— 1 llvm-objcopy per ABI (already used for assembly stores)

Work items

Phase 1: Binary blob infrastructure

  • BinaryBlobWriter: serialize StructureInstance<T> to raw bytes using StructureInfo metadata
  • Define blob header format (magic, version, section table)
  • MSBuild task: reuse existing LlvmIrComposer.Compose() → BinaryBlobWriter → DSOWrapperGenerator.WrapIt()
  • C++ init_runtime_config(): mmap blob from APK, parse header, set global pointers
  • Convert extern declarations in xamarin-app.hh to static pointer globals (gated on MonoVM compat)

Phase 2: Migrate data (incremental, per section)

  • ApplicationConfig struct
  • Runtime properties (name/value string tables for coreclr_initialize())
  • DSO cache + APK entries + name data
  • Compressed assembly descriptors
  • Pre-allocated buffers → dynamic new[]
  • Environment variables → Java Os.setenv()
  • JNI remapping tables

Phase 3: Executable code replacements

  • pinvoke_preserve.ll → --undefined linker flags + dlsym(RTLD_DEFAULT)
  • jni_init_funcs.ll (NativeAOT) → generated C# with [DllImport("__Internal")]
  • NativeAOT environment.ll → generated Java or C#

Phase 4: Cleanup

  • Remove System.loadLibrary("xamarin-app") for CoreCLR/NativeAOT
  • Gate LLVM IR generators to MonoVM-only
  • Remove libxamarin-app.so from CoreCLR/NativeAOT APK
  • Gate llc/ld to MonoVM builds only

Risks and mitigations

Risk Mitigation
Startup regression Benchmark on real devices before/after. Feature flag for A/B. Old path remains until validated.
Struct layout drift (MSBuild writer vs C++ reader) Reuse existing StructureInfo metadata for binary layout. Version header enables forward compat.
MonoVM compatibility All changes gated behind runtime check. MonoVM path unchanged.
Desktop designer application_dso_stub.cc remains.

Activity

  1. added
    Area: NativeAOTIssues that only occur when using NativeAOT.
    Area: CoreCLRIssues that only occur when using CoreCLR.
    and removed
    needs-triageIssues that need to be assigned.
    on Feb 9, 2026
  2. added this to the .NET 11 milestone on Feb 9, 2026
  3. simonrozsival commented on Feb 9, 2026

    @simonrozsival
    MemberAuthor

    /cc @grendello what are your thoughts?

  4. grendello commented on Feb 9, 2026

    @grendello
    Contributor

    I think it may complicate things more than expected. One aspect, very important, of using LLVM tools to build libxamarin-app.so is that they take care of every platform requirement - page alignment, efficient padding of structures etc. With your proposal it's your blob generator that has to take care of all that. If you want to cast structs to point into the blob, you have to ensure that the data is aligned on the proper page boundary so that CPU data caching isn't badly affected - that effectively means you probably want to align the blob bits to 16k, which will waste some storage space. You don't have to do it, but then the performance of accessing the data will suffer (on some devices more, on some devices less). You also need to keep in mind the (unlikely, but not improbable) need to use different bit order on different platforms (big endian vs little endian).

    Furthermore, some data in what we generate today (esp. DSO Cache) needs to store pointers - that means it has to be writable, and if you mmap something from the APK it will be read-only, since it's backed by read-only storage (application cannot write to its own APK files after they are installed). Supporting this would require splitting up structures, so that the writable portions are in libmonodroid.so while the R/O data remains in the blob. It would probably require more dynamic allocation, which may slow things a bit.

    Additionally, mmap is very expensive and might not be faster than fixing up the handful of native symbols in libxamarin-app.so (it's not dlopen-ed, libmonodroid.so references it so the library is loaded by the system linker).
    Likewise, new will likely be slower than BSS section init. Maintenance will be likely easier, since there's less to learn (no need to learn LLVM IR)

    With regards to performance, I can't tell off hand, but I think it might be slower. You need to measure it, there's no other way.

  5. simonrozsival commented on Feb 9, 2026

    @simonrozsival
    MemberAuthor

    This is what copilot tells me:

    Alignment and padding — The existing StructureInfo metadata in the C# code already computes field offsets, alignment, and padding (it has to, to emit correct LLVM IR). A binary writer would use the same metadata. Section-level alignment within the blob is straightforward — pad each section to its required alignment in the header. The 16KB page alignment is handled by
    llvm-objcopy --set-section-alignment payload=0x4000, same as assembly stores. This is not a new problem — DSOWrapperGenerator + get_wrapper_dso_payload_pointer_and_size() already solve it for assembly store blobs.

    Endianness — Every Android target in .NET for Android is little-endian. The existing LLVM IR generators also assume LE. This is a theoretical concern with zero practical risk today.

    Writable data in mmap'd region — This is a valid concern. Structures like DSOApkEntry (fd/offset filled during zip scan), DSOCacheEntry (dlopen handle, to_load flag), and assembly store slots are written to at runtime. An mmap'd blob from the APK is read-only. The solution is to split: read-only fields (hashes, name indices, config scalars, string tables) stay in the
    mmap'd blob, writable companion arrays are allocated with new[] at startup. This adds some complexity but it's a well-understood pattern — and the proposal already lists pre-allocated buffers as "dynamic new[] replacements." We should make the read-only vs writable split more explicit in the proposal.

    mmap cost — The config data would be embedded in the same ELF wrapper as assembly stores (or placed as a sibling that's part of the same zip scan). The zip scan already mmaps the assembly store blob — the config data would just be at a different offset within the same mmap'd region. There is no additional mmap call. The total mmap'd size grows slightly but demand
    paging means only touched pages are loaded.

    new vs BSS — Fair point for small allocations. BSS is part of the initial mmap of the .so — zero additional syscalls. new for small arrays goes through malloc. For large buffers both end up as mmap(MAP_ANONYMOUS), but for small arrays like assembly_store_bundled_assemblies[N] there's real malloc overhead. Needs measurement.

    Bottom line — Fully agree that performance needs to be measured, not assumed. The primary motivation is build simplification and maintainability, and we should not ship this if it causes a significant measurable startup regression. Assembly stores already prove the ELF-wrapper + mmap pattern works for data packaging — we're extending that pattern, not inventing
    something new.

  6. grendello commented on Feb 10, 2026

    @grendello
    Contributor

    This is what copilot tells me:

    Alignment and padding — The existing StructureInfo metadata in the C# code already computes field offsets, alignment, and padding (it has to, to emit correct LLVM IR). A binary writer would use the same metadata. Section-level alignment within the blob is straightforward — pad each section to its required alignment in the header. The 16KB page alignment is handled by
    llvm-objcopy --set-section-alignment payload=0x4000, same as assembly stores. This is not a new problem — DSOWrapperGenerator + get_wrapper_dso_payload_pointer_and_size() already solve it for assembly store blobs.

    Fair, this could work, but it actually increases code complexity. While now the generated info is used, effectively, by the 3rd party assembler and linker, you would need to handle it in the binary writer - more code to maintain. Not much of an improvement, IMO.

    Endianness — Every Android target in .NET for Android is little-endian. The existing LLVM IR generators also assume LE. This is a theoretical concern with zero practical risk today.

    It's a theoretical problem until you acknowledge that Arm can do both LE and BE, and what is done today in all targets is irrelevant. Also, it affects the practicality of the solution you're designing. Yes, LLVM IR assumes LE, but to switch it to a BE target is a minimal change in text format of the input file - the rest is handled by the LLVM tools, which also don't need to change as they already support it. With the theoretical binary writer, you have to handle all that in your own code, again - increases complexity of the locally maintained code. One of the key reasons why LLVM IR was chosen (before it we generated native assembly directly) was precisely the fact that it provides an abstraction layer over the target platform and allows us to emit effectively the same code/data for all platforms (the difference is mostly in LLVM IR attributes, and since LLVM IR doesn't support include files and preprocessor, we had to output the data to separate files).

    Writable data in mmap'd region — This is a valid concern. Structures like DSOApkEntry (fd/offset filled during zip scan), DSOCacheEntry (dlopen handle, to_load flag), and assembly store slots are written to at runtime. An mmap'd blob from the APK is read-only. The solution is to split: read-only fields (hashes, name indices, config scalars, string tables) stay in the
    mmap'd blob, writable companion arrays are allocated with new[] at startup. This adds some complexity but it's a well-understood pattern — and the proposal already lists pre-allocated buffers as "dynamic new[] replacements." We should make the read-only vs writable split more explicit in the proposal.

    Adding more new[] calls counters the idea of doing everything possible statically at the build time and it will decrease performance at startup. Also, splitting up the structures removes the benefits of data cache locality, thus possibly also decreasing startup performance.

    mmap cost — The config data would be embedded in the same ELF wrapper as assembly stores (or placed as a sibling that's part of the same zip scan). The zip scan already mmaps the assembly store blob — the config data would just be at a different offset within the same mmap'd region. There is no additional mmap call. The total mmap'd size grows slightly but demand

    Adding the new data to the assembly blob is, IMO, a mistake - it mixes responsibilities, concerns and purposes. It's bad design, simply put. Also, the assembly blob isn't always present - it's not there for Debug builds (assemblies are synced to the device with FastDev and exist as discrete files on the device's filesystem).

    paging means only touched pages are loaded.

    Not true in this case. Assembly blob will be paged in pretty quickly, as the application will need to read the assemblies (that wouldn't change compared to what we have now). It's even more true for the structures that would be added to the end of the blob - they are needed during the startup. The difference with today's setup is that they would have to mmap-ed instead of loaded by the system linker (which does it more efficiently).

    new vs BSS — Fair point for small allocations. BSS is part of the initial mmap of the .so — zero additional syscalls. new for small arrays goes through malloc. For large buffers both end up as mmap(MAP_ANONYMOUS), but for small arrays like assembly_store_bundled_assemblies[N] there's real malloc overhead. Needs measurement.

    The assumption of mmap(MAP_ANONYMOUS) is used is irrelevant and, possibly, incorrect. To know what's going on behind the scenes would require looking at the system linker source code. Besides, it's an implementation detail that's irrelevant here as there's no guarantee that the same mode of allocation is used on all the devices. What's important, however, is the fact that the system linker does more efficiently than local code would. Definitely needs measuring instead of guessing what the system does.

    Bottom line — Fully agree that performance needs to be measured, not assumed. The primary motivation is build simplification and maintainability, and we should not ship this if it causes a significant measurable startup regression. Assembly stores already prove the ELF-wrapper + mmap pattern works for data packaging — we're extending that pattern, not inventing
    something new.

    As mentioned above, you would shift a lot of maintenance responsibility (runtime implementation, data format, data ordering, direct target platform dependencies and requirements) to your own code - that decreases maintainability. The solution would also increase the complexity of code involved in the startup sequence. Measurement is a must, however it requires you to develop the full solution (optimized) and not a POC, which is a lot of work - but it's the only way to see if the idea works.

  7. simonrozsival commented on Sep 28, 2026

    @simonrozsival
    MemberAuthor

    Superseded by #12940, now the single tracking issue for removing app-build LLVM IR from CoreCLR and NativeAOT. The consolidated plan incorporates the design review here: ABI layout and byte order, writable versus APK-backed read-only state, FastDev without an assembly store, and startup allocation/mapping costs. Closing as superseded, not as implemented; follow-up PRs and decisions belong on #12940.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Area: CoreCLRIssues that only occur when using CoreCLR.Area: NativeAOTIssues that only occur when using NativeAOT.

    Type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions