Skip to content

R2R: Cache target of indirect cell address to optimize redundant cell address loading #38890

Description

@kunalspathak

Today we don't CSE loading the target of indirect cell address because during CSE we don't have that information in the IR. It happens in later phase like lower.

Consider the following code pattern:

        ...
        9000000B          adrp    x11, [RELOC #0x1de756537c0]
        9100016B          add     x11, x11, #0
        F9400160          ldr     x0, [x11]
        D63F0000          blr     x0
        ...
        9000000B          adrp    x11, [RELOC #0x1de756537c0]
        9100016B          add     x11, x11, #0
        F9400160          ldr     x0, [x11]
        D63F0000          blr     x0
        AA0003EF          mov     x15, x0
       ...

If we can optimize it using peephole or more some ambitious final instructions scanner phase to something like this to:

        ...
        9000000B          adrp    x11, [RELOC #0x1de756537c0]
        9100016B          add     x11, x11, #0
        F9400160          ldr     x0, [x11]
                          mov     xR, x0     ; store x0 in some register xR
        D63F0000          blr     x0
        ...
                          mov     x0, xR     ; retrieve xR into x0
        D63F0000          blr     x0
        AA0003EF          mov     x15, x0
       ...

With this, we can get an improvement of 8 bytes + 1 elimination of memory access.
I wrote an analyzer asm to find out how many addresses are CSE candidates and the number is huge. From what I noticed, it would by little over 2MB of size reduction.

Processed 191816 methods. Found 29246 methods containing 259123 groups.

Details: cse-candidates.txt

category:cq
theme:cse
skill-level:expert
cost:large
impact:large

Activity

  1. added
    area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
    untriagedNew issue has not been triaged by the area owner
    on Jul 7, 2020
  2. removed
    untriagedNew issue has not been triaged by the area owner
    on Jul 8, 2020
  3. added this to the Future milestone on Jul 8, 2020
  4. jakobbotsch commented on Sep 28, 2021

    @jakobbotsch
    Member

    Note that the contents of the indirection cell is not invariant. In the example here the second time the call happens the indirection cell will be pointing to the R2R code address and later to tier-1 code, while the first time around it might be pointing to an import thunk.

    With that said it might still be ok to do this optimization as the import thunk/delay load helper should be idempotent so calling it a second time in rare cases should be ok (though with loops we should consider avoiding it).

  5. EgorBo commented on Jul 20, 2023

    @EgorBo
    Member

    Together with @jakobbotsch we've just checked that it seems that we already do the right thing here, e.g.:

    image

    The address is hoisted, the ldr is not (as expected)

  6. ghost locked as resolved and limited conversation to collaborators on Aug 19, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    arch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMItenet-performancePerformance related issue

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions