Repository navigation
Use cmeq, cmge, cmgt (zero) when one of the operands is Vector64/128<T>.Zero #33972
Copy link
Copy link
Closed
Labels
Priority:3Work that is nice to haveWork that is nice to havearch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIoptimization
Milestone
Description
Activity
- addedarea-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMI
on Mar 23, 2020 - addeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Mar 23, 2020 runtime/src/libraries/System.Collections/src/System/Collections/BitArray.cs
Lines 183 to 195 in c74407f
// JIT does not support code hoisting for SIMD yet // However comparison against zero can be replaced to cmeq against zero (vceqzq_s8) // See dotnet/runtime#33972 for details Vector128<byte> zero = Vector128<byte>.Zero; fixed (bool* ptr = values) { for (; (i + Vector128ByteCount * 2u) <= (uint)values.Length; i += Vector128ByteCount * 2u) { // Same logic as SSE2 path, however we lack MoveMask (equivalent) instruction // As a workaround, mask out the relevant bit after comparison // and combine by ORing all of them together (In this case, adding all of them does the same thing) Vector128<byte> lowerVector = AdvSimd.LoadVector128((byte*)ptr + i); Vector128<byte> lowerIsFalse = AdvSimd.CompareEqual(lowerVector, zero); - removeduntriagedNew issue has not been triaged by the area ownerNew issue has not been triaged by the area owner
on Apr 4, 2020 - addedneeds-further-triageIssue has been initially triaged, but needs deeper consideration or reconsiderationIssue has been initially triaged, but needs deeper consideration or reconsideration
on Mar 23, 2021 - removedneeds-further-triageIssue has been initially triaged, but needs deeper consideration or reconsiderationIssue has been initially triaged, but needs deeper consideration or reconsideration
on Jun 7, 2021 I believe this item was addressed fully.
@TIHan can you please confirm and close the issue?Yes, I believe it was. There is more opportunity with other instructions like 'cmle' and 'cmlt', but based on the title of this issue, we got them covered.
- ghost locked as resolved and limited conversation to collaborators
on Apr 15, 2022
Metadata
Metadata
Assignees
Labels
Priority:3Work that is nice to haveWork that is nice to havearch-arm64area-CodeGen-coreclrCLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMICLR JIT compiler in src/coreclr/src/jit and related components such as SuperPMIoptimization
For example,
The code as in #33749 (comment) BitArray:.ctor
should be optimized down to
This applies to all the intrinsics that are mapped to cmeq, cmge, cmgt, cmle, cmlt, fcmeq, fcmge, fcmgt, fcmle and fcmlt instructions
category:cq
theme:hardware-intrinsics
skill-level:intermediate
cost:small