Skip to content

Enable Avx2 optimizations on Porter-Duff operations. #2340

Description

@JimBobSquarePants

Discussed in #2328

Originally posted by JimBobSquarePants January 26, 2023
Given that #1433 is not achievable within a reasonable timescale for v3 and might well be better implemented in a pattern using static virtual members in interfaces I'd like to propose that we enhance the existing Vector4 based Porter-Duff implementations with Vector256<float> equivalents.

This will give us a 2X performance improvement on supported devices over the existing implementation which we badly need if we are ever able to release a v1 of ImageSharp.Drawing.

We should be able to implement each separable method in the same fashion as the existing methods. The pixel blender implementations can call these.

For example. The existing Over method...

public static Vector4 Over(Vector4 destination, Vector4 source, Vector4 blend)
{
    // calculate weights
    float blendW = destination.W * source.W;
    float dstW = destination.W - blendW;
    float srcW = source.W - blendW;

    // calculate final alpha
    float alpha = dstW + source.W;

    // calculate final color
    Vector4 color = (destination * dstW) + (source * srcW) + (blend * blendW);

    // unpremultiply
    color /= MathF.Max(alpha, 0.001F);
    color.W = alpha;

    return color;
}

...Could be implemented in the following manner.

public static Vector256<float> Over(Vector256<float> destination, Vector256<float> source, Vector256<float> blend)
{
    const int BlendAlphaControl = 0b_10_00_10_00;
    const int ShuffleAlphaControl = 0b_11_11_11_11;

    // calculate weights
    var sW = Avx.Shuffle(source, source, ShuffleAlphaControl);
    var dstW = Avx.Shuffle(destination, destination, ShuffleAlphaControl);
    var blendW = Avx.Multiply(sW, dstW);

    dstW = Avx.Subtract(dstW, blendW);
    var srcW = Avx.Subtract(sW, blendW);

    // calculate final alpha
    Vector256<float> alpha = Avx2.Add(dstW, sW);

    // calculate final color
    Vector256<float> color = Avx2.Add(Avx2.Add(Avx2.Multiply(destination, dstW), Avx2.Multiply(source, srcW)), Avx2.Multiply(blend, blendW));

    // unpremultiply
    Vector256<float> alphaEpsilon = Avx2.Max(alpha, Vector256.Create(0.001f));
    color = Avx2.Divide(color, alphaEpsilon);
    return Avx.Blend(color, alpha, BlendAlphaControl);
}

For someone with a good grasp of the SIMD APIs this should not be a huge undertaking.

Thoughts @antonfirsov @br3aker ??

Activity

  1. added this to the 3.0.0 milestone on Feb 2, 2023
  2. antonfirsov commented on Feb 7, 2023

    @antonfirsov
    Member

    The method Vector256<float> Over(Vector256<float> destination, Vector256<float> source, Vector256<float> blend) would transform 2 color/pixel instances at the same time. This is an AVX2 implementation detail that shouldn't be exposed on the public API. We should go for whatever optimizations that make the span-based bulk methods faster, this what perf critical code should use anyways. This means that the AVX2 code above would be likely most efficient if it lives in the loop iterating through the spans.

    Thinking strategically, I have doubts if we need this in 3.0. Assuming the main goal is to optimize Drawing, 2x speedup is good but not dramatic enough to close the perf gap between IS.Drawing and competitors. Don't have a good grasp how much work would be to implement the optimizations for most important blenders, but if it's a lot, I'm not sure if it's worth to delay 3.0 with this work. Also note that the optimization would only kick in in the first Drawing version which is based on ImageSharp 3.0.

  3. JimBobSquarePants commented on Feb 8, 2023

    @JimBobSquarePants
    MemberAuthor

    Yep. We'd never expose the detail. Rather create Vector256 versions of all the internal PorterDuffFunctions methods currently implemented using Vector4. We'd then call those methods in each blender

    protected override void BlendFunction(Span<Vector4> destination, ReadOnlySpan<Vector4> background, ReadOnlySpan<Vector4> source, float amount)
    {
    amount = Numerics.Clamp(amount, 0, 1);
    for (int i = 0; i < destination.Length; i++)
    {
    destination[i] = PorterDuffFunctions.NormalSrc(background[i], source[i], amount);
    }
    }

    I don't think the workload is too great though I could be wrong. I had a look through and I believe I'd be capable of most of them though others could do them much, much faster.

    I'm really concerned about Drawing and the performance issues we have there (along with the complete lack of progress - It's been a beta now for 5 years!) and really want to be able to deliver a V1 asap. A 2X speedup would be the minimum I would expect for a V1 which should probably target ImageSharp unless someone has a good argument otherwise.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions