Repository navigation
[CoreCLR/NativeAOT] Replace LLVM IR compilation with binary data packaging #10784
Description
Activity
- addedneeds-triageIssues that need to be assigned.Issues that need to be assigned.
on Feb 9, 2026 - addedArea: NativeAOTIssues that only occur when using NativeAOT.Issues that only occur when using NativeAOT.Area: CoreCLRIssues that only occur when using CoreCLR.Issues that only occur when using CoreCLR.and removedneeds-triageIssues that need to be assigned.Issues that need to be assigned.
on Feb 9, 2026 /cc @grendello what are your thoughts?
I think it may complicate things more than expected. One aspect, very important, of using LLVM tools to build
libxamarin-app.sois that they take care of every platform requirement - page alignment, efficient padding of structures etc. With your proposal it's your blob generator that has to take care of all that. If you want to cast structs to point into the blob, you have to ensure that the data is aligned on the proper page boundary so that CPU data caching isn't badly affected - that effectively means you probably want to align the blob bits to 16k, which will waste some storage space. You don't have to do it, but then the performance of accessing the data will suffer (on some devices more, on some devices less). You also need to keep in mind the (unlikely, but not improbable) need to use different bit order on different platforms (big endian vs little endian).Furthermore, some data in what we generate today (esp. DSO Cache) needs to store pointers - that means it has to be writable, and if you mmap something from the APK it will be read-only, since it's backed by read-only storage (application cannot write to its own APK files after they are installed). Supporting this would require splitting up structures, so that the writable portions are in
libmonodroid.sowhile the R/O data remains in the blob. It would probably require more dynamic allocation, which may slow things a bit.Additionally,
mmapis very expensive and might not be faster than fixing up the handful of native symbols inlibxamarin-app.so(it's notdlopen-ed,libmonodroid.soreferences it so the library is loaded by the system linker).
Likewise,newwill likely be slower than BSS section init. Maintenance will be likely easier, since there's less to learn (no need to learn LLVM IR)With regards to performance, I can't tell off hand, but I think it might be slower. You need to measure it, there's no other way.
This is what copilot tells me:
Alignment and padding — The existing StructureInfo metadata in the C# code already computes field offsets, alignment, and padding (it has to, to emit correct LLVM IR). A binary writer would use the same metadata. Section-level alignment within the blob is straightforward — pad each section to its required alignment in the header. The 16KB page alignment is handled by
llvm-objcopy --set-section-alignment payload=0x4000, same as assembly stores. This is not a new problem — DSOWrapperGenerator + get_wrapper_dso_payload_pointer_and_size() already solve it for assembly store blobs.Endianness — Every Android target in .NET for Android is little-endian. The existing LLVM IR generators also assume LE. This is a theoretical concern with zero practical risk today.
Writable data in mmap'd region — This is a valid concern. Structures like DSOApkEntry (fd/offset filled during zip scan), DSOCacheEntry (dlopen handle, to_load flag), and assembly store slots are written to at runtime. An mmap'd blob from the APK is read-only. The solution is to split: read-only fields (hashes, name indices, config scalars, string tables) stay in the
mmap'd blob, writable companion arrays are allocated with new[] at startup. This adds some complexity but it's a well-understood pattern — and the proposal already lists pre-allocated buffers as "dynamic new[] replacements." We should make the read-only vs writable split more explicit in the proposal.mmap cost — The config data would be embedded in the same ELF wrapper as assembly stores (or placed as a sibling that's part of the same zip scan). The zip scan already mmaps the assembly store blob — the config data would just be at a different offset within the same mmap'd region. There is no additional mmap call. The total mmap'd size grows slightly but demand
paging means only touched pages are loaded.new vs BSS — Fair point for small allocations. BSS is part of the initial mmap of the .so — zero additional syscalls. new for small arrays goes through malloc. For large buffers both end up as mmap(MAP_ANONYMOUS), but for small arrays like assembly_store_bundled_assemblies[N] there's real malloc overhead. Needs measurement.
Bottom line — Fully agree that performance needs to be measured, not assumed. The primary motivation is build simplification and maintainability, and we should not ship this if it causes a significant measurable startup regression. Assembly stores already prove the ELF-wrapper + mmap pattern works for data packaging — we're extending that pattern, not inventing
something new.This is what copilot tells me:
Alignment and padding — The existing StructureInfo metadata in the C# code already computes field offsets, alignment, and padding (it has to, to emit correct LLVM IR). A binary writer would use the same metadata. Section-level alignment within the blob is straightforward — pad each section to its required alignment in the header. The 16KB page alignment is handled by
llvm-objcopy --set-section-alignment payload=0x4000, same as assembly stores. This is not a new problem — DSOWrapperGenerator + get_wrapper_dso_payload_pointer_and_size() already solve it for assembly store blobs.Fair, this could work, but it actually increases code complexity. While now the generated info is used, effectively, by the 3rd party assembler and linker, you would need to handle it in the binary writer - more code to maintain. Not much of an improvement, IMO.
Endianness — Every Android target in .NET for Android is little-endian. The existing LLVM IR generators also assume LE. This is a theoretical concern with zero practical risk today.
It's a theoretical problem until you acknowledge that Arm can do both LE and BE, and what is done today in all targets is irrelevant. Also, it affects the practicality of the solution you're designing. Yes, LLVM IR assumes LE, but to switch it to a BE target is a minimal change in text format of the input file - the rest is handled by the LLVM tools, which also don't need to change as they already support it. With the theoretical binary writer, you have to handle all that in your own code, again - increases complexity of the locally maintained code. One of the key reasons why LLVM IR was chosen (before it we generated native assembly directly) was precisely the fact that it provides an abstraction layer over the target platform and allows us to emit effectively the same code/data for all platforms (the difference is mostly in LLVM IR attributes, and since LLVM IR doesn't support include files and preprocessor, we had to output the data to separate files).
Writable data in mmap'd region — This is a valid concern. Structures like DSOApkEntry (fd/offset filled during zip scan), DSOCacheEntry (dlopen handle, to_load flag), and assembly store slots are written to at runtime. An mmap'd blob from the APK is read-only. The solution is to split: read-only fields (hashes, name indices, config scalars, string tables) stay in the
mmap'd blob, writable companion arrays are allocated with new[] at startup. This adds some complexity but it's a well-understood pattern — and the proposal already lists pre-allocated buffers as "dynamic new[] replacements." We should make the read-only vs writable split more explicit in the proposal.Adding more
new[]calls counters the idea of doing everything possible statically at the build time and it will decrease performance at startup. Also, splitting up the structures removes the benefits of data cache locality, thus possibly also decreasing startup performance.mmap cost — The config data would be embedded in the same ELF wrapper as assembly stores (or placed as a sibling that's part of the same zip scan). The zip scan already mmaps the assembly store blob — the config data would just be at a different offset within the same mmap'd region. There is no additional mmap call. The total mmap'd size grows slightly but demand
Adding the new data to the assembly blob is, IMO, a mistake - it mixes responsibilities, concerns and purposes. It's bad design, simply put. Also, the assembly blob isn't always present - it's not there for Debug builds (assemblies are synced to the device with FastDev and exist as discrete files on the device's filesystem).
paging means only touched pages are loaded.
Not true in this case. Assembly blob will be paged in pretty quickly, as the application will need to read the assemblies (that wouldn't change compared to what we have now). It's even more true for the structures that would be added to the end of the blob - they are needed during the startup. The difference with today's setup is that they would have to
mmap-ed instead of loaded by the system linker (which does it more efficiently).new vs BSS — Fair point for small allocations. BSS is part of the initial mmap of the .so — zero additional syscalls. new for small arrays goes through malloc. For large buffers both end up as mmap(MAP_ANONYMOUS), but for small arrays like assembly_store_bundled_assemblies[N] there's real malloc overhead. Needs measurement.
The assumption of
mmap(MAP_ANONYMOUS)is used is irrelevant and, possibly, incorrect. To know what's going on behind the scenes would require looking at the system linker source code. Besides, it's an implementation detail that's irrelevant here as there's no guarantee that the same mode of allocation is used on all the devices. What's important, however, is the fact that the system linker does more efficiently than local code would. Definitely needs measuring instead of guessing what the system does.Bottom line — Fully agree that performance needs to be measured, not assumed. The primary motivation is build simplification and maintainability, and we should not ship this if it causes a significant measurable startup regression. Assembly stores already prove the ELF-wrapper + mmap pattern works for data packaging — we're extending that pattern, not inventing
something new.As mentioned above, you would shift a lot of maintenance responsibility (runtime implementation, data format, data ordering, direct target platform dependencies and requirements) to your own code - that decreases maintainability. The solution would also increase the complexity of code involved in the startup sequence. Measurement is a must, however it requires you to develop the full solution (optimized) and not a POC, which is a lot of work - but it's the only way to see if the idea works.
Superseded by #12940, now the single tracking issue for removing app-build LLVM IR from CoreCLR and NativeAOT. The consolidated plan incorporates the design review here: ABI layout and byte order, writable versus APK-backed read-only state, FastDev without an assembly store, and startup allocation/mapping costs. Closing as superseded, not as implemented; follow-up PRs and decisions belong on #12940.
Goal
Simplify the app build pipeline by eliminating LLVM IR code generation,
llccompilation, andldlinking for CoreCLR and NativeAOT builds. This reduces build tool dependencies, build complexity, and long-term maintenance cost.This must not come at the expense of a significant measurable startup performance regression. We need to measure the actual impact on real devices before and after.
Summary
Every .NET for Android app build generates 5-7 LLVM IR (
.ll) files per ABI, compiles them withllc, and links them withldintolibxamarin-app.so. This shared library is almost entirely read-only data — configuration structs, lookup tables, and pre-allocated buffers. The LLVM IR pipeline is a heavyweight code generation + compilation step for what is fundamentally a data packaging problem.We already have a simpler mechanism for packaging data into the APK:
DSOWrapperGeneratorwraps arbitrary binary files in a minimal ELF.sousingllvm-objcopy, places them inlib/{abi}/, and the runtime mmaps them directly from the APK. This is how assembly stores (assemblies.blob) work today.Proposal: Replace the LLVM IR →
llc→ldpipeline with direct binary serialization →llvm-objcopyfor all configuration data. The existingLlvmIrComposersubclasses already compute all the values in C# — we just change the output stage from "emit LLVM IR text" to "write raw bytes matching the C struct layout."Scope: CoreCLR and NativeAOT only. MonoVM continues using LLVM IR until deprecated.
Dependency: Trimmable TypeMap (in progress) eliminates
typemapsandmarshal_methods. This proposal handles the remaining.llfiles.What's in
libxamarin-app.sotodaytypemaps.*.llmarshal_methods.*.llenvironment.*.llApplicationConfigstruct, runtime properties, DSO cache, env varscompressed_assemblies.*.lljni_remap.*.llpinvoke_preserve.*.llfind_pinvoke()(CoreCLR unified linking only)jni_init_funcs.*.llProposed approach
1. Binary config blob in ELF wrapper (replaces
environment,compressed_assemblies,jni_remap)Build time:
The existing
LlvmIrComposersubclasses (e.g.,ApplicationConfigNativeAssemblyGeneratorCLR,CompressedAssembliesNativeAssemblyGenerator) already have a two-stage pipeline:StructureInstance<ApplicationConfig>,List<StructureInstance<DSOCacheEntry>>, etc.LlvmIrGeneratorWe replace stage 2: instead of
LlvmIrGeneratoremitting text, a newBinaryBlobWriterserializesStructureInstance<T>objects directly to bytes using the existingStructureInfometadata (field offsets, sizes, alignment, padding). Same data, same layout, no compilation step.The output
config.binhas a simple section-based format:Then wrap and package:
This uses
llvm-objcopy --add-section payload=config.bin— the same tool already used for assembly stores. Nollc, nold.Runtime:
The zip scan (which already runs to find assembly stores) discovers
libruntime-config.so, mmaps it from the APK, andget_wrapper_dso_payload_pointer_and_size()returns a direct pointer to the payload.No parsing, no copying, no deserialization. The MSBuild task writes bytes matching the C struct memory layout. The C++ code casts pointers into the mmap'd region. Data is demand-paged from the APK by the kernel — same mechanism as today.
Changes to
libmonodroid.so: Replaceexterndeclarations (currently resolved bylibxamarin-app.soat load time) with static pointer globals initialized from the mmap'd blob. This follows the same pattern already used forassembly_storedata.2. Dynamic allocation (replaces BSS pre-allocated buffers)
The zero-initialized buffers in
libxamarin-app.so(assembly store slots, decompression buffer) use BSS sections, which the kernel backs withmmap(MAP_ANONYMOUS)+ demand paging. Allocating withnew[]uses the same kernel mechanism for large allocations. Replace:3. Environment variables → Java
Os.setenv()Generate Java code calling
Os.setenv()beforeinitInternal(), following the pattern NativeAOT already uses (NativeAotEnvironmentVars.java).4.
pinvoke_preserve.*.ll→ Linker flags +dlsym(CoreCLR unified linking only)This is the one file with actual executable code:
find_pinvoke()maps(library_hash, entrypoint_hash)→ function pointer via nested switch statements. It serves two purposes:Linker symbol preservation — references to symbols like
@SystemNative_Bindprevent--gc-sectionsfrom stripping them. Replace with--undefined=<symbol>linker flags.PinvokeScanneralready produces the symbol list, andNativeLinker.csalready supports--export-dynamic-symbol— the infrastructure is in place. (There's even a TODO indynamic.cc:88where the team considered this approach.)Runtime P/Invoke resolution — replace with
dlsym(RTLD_DEFAULT, entrypoint_name), which already exists as a fallback indynamic.cc. P/Invoke results are cached by CoreCLR — each entrypoint is resolved once. Performance impact to be measured.5.
jni_init_funcs.*.ll→ Generated C# (NativeAOT only)Replace the LLVM IR function pointer array with generated C# using
[DllImport("__Internal")]:NativeAOT compiles
[DllImport("__Internal")]to direct native call instructions — zero overhead, compile-time symbol resolution, missing symbol = link error (not runtime crash). TheDirectPInvokeinfrastructure already exists inMicrosoft.Android.Sdk.NativeAOT.targets.Performance
The primary goal is long-term maintainability and build simplification. However, this must not come at the expense of a significant measurable startup regression.
The proposed approach uses the same mmap-from-APK mechanism as today — config data is still accessed via direct pointer dereferences into memory-mapped regions. The main differences are: (1) eliminating
dlopen("libxamarin-app.so")and its symbol resolution overhead, (2) replacing BSS pre-allocated buffers with dynamicnew[], and (3) replacingfind_pinvoke()withdlsymfor unified linking.We need to measure the actual performance impact on real devices (high-end and low-end) with representative apps before and after. A feature flag should allow A/B comparison.
Build time
llcinvocations per ABI (LLVM IR compilation)ldinvocation per ABI (native linking)llvm-objcopyper ABI (already used for assembly stores)Work items
Phase 1: Binary blob infrastructure
BinaryBlobWriter: serializeStructureInstance<T>to raw bytes usingStructureInfometadataLlvmIrComposer.Compose()→BinaryBlobWriter→DSOWrapperGenerator.WrapIt()init_runtime_config(): mmap blob from APK, parse header, set global pointersexterndeclarations inxamarin-app.hhto static pointer globals (gated on MonoVM compat)Phase 2: Migrate data (incremental, per section)
ApplicationConfigstructcoreclr_initialize())new[]Os.setenv()Phase 3: Executable code replacements
pinvoke_preserve.ll→--undefinedlinker flags +dlsym(RTLD_DEFAULT)jni_init_funcs.ll(NativeAOT) → generated C# with[DllImport("__Internal")]environment.ll→ generated Java or C#Phase 4: Cleanup
System.loadLibrary("xamarin-app")for CoreCLR/NativeAOTlibxamarin-app.sofrom CoreCLR/NativeAOT APKllc/ldto MonoVM builds onlyRisks and mitigations
StructureInfometadata for binary layout. Version header enables forward compat.application_dso_stub.ccremains.