Skip to content

Optimize pixel blending with integer arithmetics #1433

Description

@antonfirsov

Problem

The current float and Vector4 -based pixel blender API does not give us too much space for introducing the rasterization perf improvements needed for SixLabors/ImageSharp.Drawing#102. With our current bulk API, the maximum we can do is to process 2 pixels in one AVX batch, since we can fit only 8 float-s into one AVX register. This means that the expected speedup for blending is around or below 2x. With this we would keep lagging behind Skia and GDI significantly.

Idea

We should explore API-s and implementations working with UInt16-based fixed point arithmetics. This is technically very similar to approach taken by the libjpeg decoder SIMD pipelines which we eventually also want to adapt. In theory, UInt16-based bulk processing should reduce the time spent in pixel blenders by ~4x (or more) when AVX2 is present.

This will require API additions similar to the following:

public abstract class PixelBlender<TPixel> 
{
    public void Blend<TPixelSrc>(
            Configuration configuration,
            Span<TPixel> destination,
            ReadOnlySpan<TPixel> background,
            ReadOnlySpan<TPixelSrc> source,
            // 'amount' is scaled to 0-255. Could be byte, but with UInt16 we will avoid some unnecessary conversions
            ReadOnlySpan<UInt16> amount); 

    protected virtual BlendFunction(
            Configuration configuration,
            Span<Rgba32> destination,
            ReadOnlySpan<Rgba32> background,
            ReadOnlySpan<Rgba32> source,
            ReadOnlySpan<UInt16> amount);
}


public static class PorterDuffFunctions
{
    /*public*/ static Vector4 NormalSrcOver(Span<Rgba32> destination,
            ReadOnlySpan<Rgba32> background,
            ReadOnlySpan<Rgba32> source,
            ReadOnlySpan<UInt16> opacity);
}

Update:
In the first API variant there was a type ScaledUInt16Vector4, but after thinking it through, I realized it is unnecessary. We should work with Rgba32 to maximize perf.

Activity

  1. JimBobSquarePants commented on Nov 21, 2020

    @JimBobSquarePants
    Member

    Agreed. This seems to be the smartest approach while keeping work on the CPU. Other libraries defer to the GPU for perf

  2. JimBobSquarePants commented on Sep 9, 2021

    @JimBobSquarePants
    Member

    Pinning this as It's super important for ImageSharp.Drawing

  3. JimBobSquarePants commented on Sep 29, 2021

    @JimBobSquarePants
    Member

    @br3aker I know it's not jpeg but is this something you'd be interested in? Seems like something you could churn out whereas it would take me ages.

  4. antonfirsov commented on Oct 3, 2021

    @antonfirsov
    MemberAuthor

    This is important in long term, but improving Jpeg decoder performance has higher priority at the moment IMO, so if @br3aker is more interested in Jpeg he's doing the right thing :)

    We need to close the 40% gap to vips and MagicScaler in thumbnail making. (See #1775).

  5. JimBobSquarePants commented on Oct 3, 2021

    @JimBobSquarePants
    Member

    Yep. Let's get that gap closed! We still need someone to look at this so hoping someone from the community can step up while I focus on the Fonts/Drawing libraries themselves.

  6. br3aker commented on Oct 5, 2021

    @br3aker
    Contributor

    Oops, totally missed this conversation, I'm sorry.

    We can work something out after jpeg decoder :)

  7. added this to the milestone on Nov 27, 2021
  8. JimBobSquarePants commented on Jan 18, 2022

    @JimBobSquarePants
    Member

    I think this is our only outstanding feature that is blocking V2 now. I don't think we can release ImageSharp.Drawing without it.

  9. 17 remaining items

  10. JimBobSquarePants commented on Jan 26, 2023

    @JimBobSquarePants
    Member

    Awesome. We should be able to incrementally implement versions of all the separable methods. I'll spin up an issue.

  11. tannergooding commented on Feb 16, 2023

    @tannergooding
    Contributor

    Worth noting that while some operations are just as fast for integral. Others, like multiplication, can be ~2x slower and so the exact perf benefit may then vary a bit depending exactly on what's needed for the blend operation. -- SIMD float multiplication is ~4 cycles, SIMD int32 multiplication is ~10 cycles, SIMD int16 multiplication is ~5 cycles, but only does the lower or upper elements so you need 2 instructions

  12. saucecontrol commented on Feb 16, 2023

    @saucecontrol
    Contributor

    I would add to that, FMA changes the equation as well, because in olden times you were looking at 4 cycles for float mul + 4 cycles for float add on the low end, compared to a best case of 5 + 1 for int16. With FMA, obviously we get that float add for free. A lot of the existing code in the wild was written pre-FMA, when the integer math hoops were more worth jumping through.

  13. JimBobSquarePants commented on Feb 17, 2023

    @JimBobSquarePants
    Member

    So are we saying here that it's actually best to stick with Avx and Vector256<float>?

  14. antonfirsov commented on Feb 17, 2023

    @antonfirsov
    MemberAuthor

    Yeah, naively I thought that we can squeeze out much more from integers, but feels like #2340 is more worthy to chase now, there is good potential to reach better than 2x speedup with FMA.

  15. JimBobSquarePants commented on Feb 17, 2023

    @JimBobSquarePants
    Member

    I guess that's good news then as it's a LOT less work.

    I'll need a lot of help to do this though, I can do bit and pieces but will struggle with certain parts.

  16. saucecontrol commented on Feb 17, 2023

    @saucecontrol
    Contributor

    So are we saying here that it's actually best to stick with Avx and Vector256<float>?

    I hate to drop an "it depends" on you but...

    I'd say the general rule of thumb is that the more complex the math, the more likely it is Vector256<float> will work out better. Balance that with the fact that if you're starting and ending with integer data, you pay an extra 4+4 cycles for the int->float->int in addition to only be able to fit half as many pixels in each vector. Obviously if you go float in one part, it's usually better to stay there, pipelining steps if possible.

    On the other hand, integer domain makes it easier to incorporate bit twiddling hacks, plus there are some integer-only instructions that do crazy amounts of work in ~5 cycles, like our old pals mpsadbw and pmaddubsw. If you can take advantage of those tricks, obviously it could sway performance heavily in favor of integer math.

    The porter-duff ops are simple enough they will probably be faster in integer domain, but taking into account the FMA factor, that win is probably not going to be big enough to justify added complexity if you're more comfortable in float land. You can always re-visit when you target a newer runtime with the simplified xplat intrinsics.

  17. JimBobSquarePants commented on Feb 17, 2023

    @JimBobSquarePants
    Member

    So an integer based version of PorterDuffFunctions would be something I'd use for Rgba32 compatible pixel formats. We can use the existing bulk shuffle operations to translate them for working against and call them using overloads of PixelOperations.GetPixelBlender.

    Unfortunately, implementing those would be outside my ability to deliver at present.

  18. JimBobSquarePants commented on Mar 17, 2023

    @JimBobSquarePants
    Member

    I don't think we need this anymore. The new blender is very fast.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions