Repository navigation
ARM64: loop array indexing inefficiencies #34810
Description
Activity
- addedarea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIuntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Apr 10, 2020 A case of a simple optimization where a STR or LDR immediately followed by an address variable addition could be transformed to subsume the addition into the STR/LDR instruction was discussed here in the context of the intrinsics.
It looks like what you are suggesting would be a sequence of transformations, loop induction variable strength reduction, where we either wouldn't maintain
iseparately, or would maintain bothiand<array base> + 16 + i * 4in the loop.cc @AndyAyersMS
Just verified that gcc seems to do that optimization but clang doesn't.
https://godbolt.org/z/Wp9Xhu- removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Apr 13, 2020 Just verified that gcc seems to do that optimization but clang doesn't.
https://godbolt.org/z/Wp9XhuClang uses a more complicated addressing mode but also equally valid. (your example is missing an
-O1).str w8, [x9, x8, lsl #2] add x8, x8, #1 // =1 cmp x8, #10 // =10 b.ne .LBB0_1Of course simpler addressing modes are always preferred :)
- changed the title
[-]ARM64: Use post-index addressing mode to access array elements[/-][+]ARM64: loop array indexing inefficiencies[/+]on Apr 23, 2020 Note that we have to be careful with ref/byref creation and reporting. E.g., hoisting
<array base> + 16out of the loop to create a pointer to the array element base would create a byref pointer that needs to be reported. Note the comment infgMorphArrayIndex:// Be careful to only create the byref pointer when the full index expression is added to the array reference. // We don't want to create a partial byref address expression that doesn't include the full index offset: // a byref must point within the containing object. It is dangerous (especially when optimizations come into // play) to create a "partial" byref that doesn't point exactly to the correct object; there is risk that // the partial byref will not point within the object, and thus not get updated correctly during a GC. // This is mostly a risk in fully-interruptible code regions.The PR where this comment was introduced: dotnet/coreclr#17524
Right, if we have an address computation where the full computation tree has a mixture of positive and negative adjustments to the address, we need to be careful not to reassociate too broadly; all the intermediate results must be addresses within the bounds of the parent object.
Given that ARM doesn't have base + scaled index + offset addressing mode, it seems like we really need to be able to hoist
<object base> [ref] + <array first element offset> [native int]out of a loop as a byref.- addedJitUntriagedCLR JIT issues needing additional triageCLR JIT issues needing additional triage
on Oct 28, 2020 6 remaining items
- removedneeds-further-triageIssue has been initially triaged, but needs deeper consideration or reconsiderationIssue has been initially triaged, but needs deeper consideration or reconsideration
on Jun 7, 2021 I think it worth moving this to 7.0 as I'd expect noticeable perf improvements from it:

I tried to implement it via https://github.com/dotnet/runtime/pull/60085/files and even emitted something similar but it needs more work.
I made some progress on this and re-assigning to myself if you don't mind
Reacted by Tamar Christina@EgorBo you said that you completed this work. Can you link your PR and close this issue?
Yes, I believe this can be closed via a series of PRs for addressing modes, mainly
and follow ups:
- ghost locked as resolved and limited conversation to collaborators
on Mar 26, 2022
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsDone
The line
arr[i] = 1generates the following code to calculate the address of element to save the value.vs. how x64 generates:
The ARM64 pattern can be optimized to use post-index addressing mode using:
category:cq
theme:optimization
skill-level:intermediate
cost:medium