Repository navigation
[Arm64] addressing mode inefficiencies in Guid:op_Equality(Guid,Guid):bool #35622
Copy link
Copy link
Closed
Labels
arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIoptimization
Milestone
Description
Activity
- addedarea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
on Apr 29, 2020 - addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Apr 29, 2020 Related: #35635
Changing the Guid:op_Equality implementation to:
public static bool operator ==(Guid a, Guid b) => Unsafe.As<int, long>(ref a._a) == Unsafe.As<int, long>(ref b._a) && Unsafe.As<byte, long>(ref a._d) == Unsafe.As<byte, long>(ref b._d);fixes the odd address base recalculations and only leaves the unnecessary stack usage.
arm64 assembly with new Guid implementation
G_M51749_IG01: A9BD7BFD stp fp, lr, [sp,#-48]! 910003FD mov fp, sp F90013A0 str x0, [fp,#32] F90017A1 str x1, [fp,#40] F9000BA2 str x2, [fp,#16] F9000FA3 str x3, [fp,#24] ;; bbWeight=1 PerfScore 5.50 G_M51749_IG02: F94013A0 ldr x0, [fp,#32] F9400BA1 ldr x1, [fp,#16] EB01001F cmp x0, x1 540000E1 bne G_M51749_IG05 ;; bbWeight=1 PerfScore 5.50 G_M51749_IG03: F94017A0 ldr x0, [fp,#40] F9400FA1 ldr x1, [fp,#24] EB01001F cmp x0, x1 9A9F17E0 cset x0, eq ;; bbWeight=0.50 PerfScore 2.50 G_M51749_IG04: A8C37BFD ldp fp, lr, [sp],#48 D65F03C0 ret lr ;; bbWeight=0.50 PerfScore 1.00 G_M51749_IG05: 52800000 mov w0, #0 ;; bbWeight=0.50 PerfScore 0.25 G_M51749_IG06: A8C37BFD ldp fp, lr, [sp],#48 D65F03C0 ret lr
- added a commit that references this issue
on Apr 30, 2020 - removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on May 4, 2020 - addedJitUntriagedCLR JIT issues needing additional triageCLR JIT issues needing additional triage
on Oct 28, 2020 - removedJitUntriagedCLR JIT issues needing additional triageCLR JIT issues needing additional triage
on Nov 14, 2020 - addedneeds-further-triageIssue has been initially triaged, but needs deeper consideration or reconsiderationIssue has been initially triaged, but needs deeper consideration or reconsideration
on Mar 23, 2021 - removedneeds-further-triageIssue has been initially triaged, but needs deeper consideration or reconsiderationIssue has been initially triaged, but needs deeper consideration or reconsideration
on Apr 8, 2021 The latest code is much better:
; Assembly listing for method System.Guid:EqualsCore(byref,byref):bool ; Emitting BLENDED_CODE for generic ARM64 CPU - Windows ; optimized code ; fp based frame ; partially interruptible ; No PGO data ; invoked as altjit ; Final local variable assignments ; ; V00 arg0 [V00,T00] ( 3, 3 ) byref -> x0 single-def ; V01 arg1 [V01,T01] ( 3, 3 ) byref -> x1 single-def ;* V02 loc0 [V02 ] ( 0, 0 ) byref -> zero-ref ;* V03 loc1 [V03 ] ( 0, 0 ) byref -> zero-ref ;# V04 OutArgs [V04 ] ( 1, 1 ) lclBlk ( 0) [sp+00H] "OutgoingArgSpace" ; V05 rat0 [V05,T02] ( 3, 6 ) simd16 -> d16 HFA(simd16) "ReplaceWithLclVar is creating a new local variable" ; ; Lcl frame size = 0 G_M26697_IG01: stp fp, lr, [sp, #-0x10]! mov fp, sp ;; size=8 bbWeight=1 PerfScore 1.50 G_M26697_IG02: ld1 {v16.16b}, [x0] ld1 {v17.16b}, [x1] cmeq v16.16b, v16.16b, v17.16b uminp v16.4s, v16.4s, v16.4s umov x0, v16.d[0] cmn x0, #1 cset x0, eq ;; size=28 bbWeight=1 PerfScore 10.00 G_M26697_IG03: ldp fp, lr, [sp], #0x10 ret lr
- ghost locked as resolved and limited conversation to collaborators
on Oct 30, 2022
Metadata
Metadata
Assignees
Labels
arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIoptimization
The arm64 generated code for
Guid::op_Equality()could be better by (1) incorporating thefpaddress calculation into theldraddressing modes, and (2) not using stack at all.The code:
This code itself is weird, comparing 4
intvalues instead of comparing field-by-field of oneint, twoshort, and eightbyte. It should compare 2longon 64-bit.x64 code is pretty direct translation of this C# code.
arm64 first pushes the 2 16-byte struct-in-register-pair arguments to stack, then reloads each 4-byte element one at a time to compare. The base address of the stack locals are computed over and over, instead of being folded into the subsequent addressing modes that add the offset.
x64 assembly
arm64 assembly
Possible arm64 assembly after fixing address calculations
The JIT shouldn't need to put the argument structs on the stack at all. In which case we could generate code like the following (also assuming we can compare full registers, and not 4 bytes at a time).
Possible arm64 assembly fully optimized
category:cq
theme:optimization
skill-level:intermediate
cost:medium