Repository navigation
JIT: inefficient codegen for calls returning 16-byte structs on Linux x64 / arm64 #8571
Copy link
Copy link
Closed
Labels
arch-x64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issuePerformance related issue
Milestone
Description
Activity
- changed the title
[-]JIT: inefficient codegen for calls returning 16-byte structs on Linux x64[/-][+]JIT: inefficient codegen for calls returning 16-byte structs on Linux x64 / arm64[/+]on Mar 25, 2020 This is also applicable for ARM64 as well. When profiling this method on windows ARM4, we generate the following. The ctor call is inlined.
G_M43267_IG01: A9BE7BFD stp fp, lr, [sp,#-32]! 910003FD mov fp, sp F9000BBF str xzr, [fp,#16] // [V02 tmp1] ;; bbWeight=1 PerfScore 2.50 G_M43267_IG02: F9400400 ldr x0, [x0,#8] B40001C0 cbz x0, G_M43267_IG06 ;; bbWeight=1 PerfScore 4.00 G_M43267_IG03: B9400801 ldr w1, [x0,#8] 2A0103E1 mov w1, w1 F100283F cmp x1, #10 54000143 blo G_M43267_IG06 ;; bbWeight=0.50 PerfScore 2.50 G_M43267_IG04: 91004000 add x0, x0, #16 52800141 mov w1, #10 910043A2 add x2, fp, #16 // [V02 tmp1] F9000040 str x0, [x2] B9000841 str w1, [x2,#8] F9400BA0 ldr x0, [fp,#16] // [V02 tmp1] F9400FA1 ldr x1, [fp,#24] // [V02 tmp1+0x08] ;; bbWeight=1 PerfScore 7.50 G_M43267_IG05: A8C27BFD ldp fp, lr, [sp],#32 D65F03C0 ret lr ;; bbWeight=1 PerfScore 2.00 G_M43267_IG06: 97E65C22 bl System.ThrowHelper:ThrowArgumentOutOfRangeException() D43E0000 bkpt ;; bbWeight=0 PerfScore 0.00Here
G_M43267_IG04doesn't need to be so long and can be:add x0, x0, #16 mov w1, #10 # At this point, x0 and x1 has the struct that we are trying to returnThis issue seems to pertain to an older version of binarytrees (perhaps https://github.com/dotnet/coreclr/blob/414ab4ee1a6f31ae63f166de2b9d4d0af640574f/tests/src/JIT/Performance/CodeQuality/BenchmarksGame/binarytrees/binarytrees.csharp.cs). For that version, after #36862, we are largely keeping the return values in registers.
- ghost locked as resolved and limited conversation to collaborators
on Dec 21, 2020
Metadata
Metadata
Assignees
Labels
arch-x64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIenhancementProduct code improvement that does NOT require public API changes/additionsProduct code improvement that does NOT require public API changes/additionsoptimizationtenet-performancePerformance related issuePerformance related issue
From the binarytrees performance benchmark, initial call to
bottomUpTreefromBench(other calls to this method have similar issues)bottomUpTreehas similar issues at its recursive call sites, and also does some redundant zeroing of temp structs that were zeroed in the prolog:Note this latter bit of code could simply be something like
category:cq
theme:structs
skill-level:expert
cost:large