Repository navigation
Optimize pixel blending with integer arithmetics #1433
Description
Activity
Agreed. This seems to be the smartest approach while keeping work on the CPU. Other libraries defer to the GPU for perf
- pinned this issue
on Sep 9, 2021 Pinning this as It's super important for ImageSharp.Drawing
@br3aker I know it's not jpeg but is this something you'd be interested in? Seems like something you could churn out whereas it would take me ages.
Yep. Let's get that gap closed! We still need someone to look at this so hoping someone from the community can step up while I focus on the Fonts/Drawing libraries themselves.
Reacted by Anton FirszovOops, totally missed this conversation, I'm sorry.
We can work something out after jpeg decoder :)
Reacted by James Jackson-SouthI think this is our only outstanding feature that is blocking V2 now. I don't think we can release ImageSharp.Drawing without it.
17 remaining items
Awesome. We should be able to incrementally implement versions of all the separable methods. I'll spin up an issue.
- unpinned this issue
on Feb 2, 2023 - pinned this issue
on Feb 13, 2023 Worth noting that while some operations are just as fast for
integral. Others, like multiplication, can be ~2x slower and so the exact perf benefit may then vary a bit depending exactly on what's needed for the blend operation. -- SIMDfloatmultiplication is ~4 cycles, SIMDint32multiplication is ~10 cycles, SIMDint16multiplication is ~5 cycles, but only does the lower or upper elements so you need 2 instructionsI would add to that, FMA changes the equation as well, because in olden times you were looking at 4 cycles for float mul + 4 cycles for float add on the low end, compared to a best case of 5 + 1 for int16. With FMA, obviously we get that float add for free. A lot of the existing code in the wild was written pre-FMA, when the integer math hoops were more worth jumping through.
So are we saying here that it's actually best to stick with
AvxandVector256<float>?Yeah, naively I thought that we can squeeze out much more from integers, but feels like #2340 is more worthy to chase now, there is good potential to reach better than 2x speedup with FMA.
Reacted by James Jackson-SouthI guess that's good news then as it's a LOT less work.
I'll need a lot of help to do this though, I can do bit and pieces but will struggle with certain parts.
So are we saying here that it's actually best to stick with
AvxandVector256<float>?I hate to drop an "it depends" on you but...
I'd say the general rule of thumb is that the more complex the math, the more likely it is
Vector256<float>will work out better. Balance that with the fact that if you're starting and ending with integer data, you pay an extra 4+4 cycles for the int->float->int in addition to only be able to fit half as many pixels in each vector. Obviously if you go float in one part, it's usually better to stay there, pipelining steps if possible.On the other hand, integer domain makes it easier to incorporate bit twiddling hacks, plus there are some integer-only instructions that do crazy amounts of work in ~5 cycles, like our old pals
mpsadbwandpmaddubsw. If you can take advantage of those tricks, obviously it could sway performance heavily in favor of integer math.The porter-duff ops are simple enough they will probably be faster in integer domain, but taking into account the FMA factor, that win is probably not going to be big enough to justify added complexity if you're more comfortable in float land. You can always re-visit when you target a newer runtime with the simplified xplat intrinsics.
So an integer based version of
PorterDuffFunctionswould be something I'd use forRgba32compatible pixel formats. We can use the existing bulk shuffle operations to translate them for working against and call them using overloads ofPixelOperations.GetPixelBlender.Unfortunately, implementing those would be outside my ability to deliver at present.
- unpinned this issue
on Feb 20, 2023 I don't think we need this anymore. The new blender is very fast.
Reacted by Anton Firszov
Problem
The current
floatandVector4-based pixel blender API does not give us too much space for introducing the rasterization perf improvements needed for SixLabors/ImageSharp.Drawing#102. With our current bulk API, the maximum we can do is to process 2 pixels in one AVX batch, since we can fit only 8float-s into one AVX register. This means that the expected speedup for blending is around or below 2x. With this we would keep lagging behind Skia and GDI significantly.Idea
We should explore API-s and implementations working with
UInt16-based fixed point arithmetics. This is technically very similar to approach taken by the libjpeg decoder SIMD pipelines which we eventually also want to adapt. In theory,UInt16-based bulk processing should reduce the time spent in pixel blenders by ~4x (or more) when AVX2 is present.This will require API additions similar to the following:
Update:
In the first API variant there was a type
ScaledUInt16Vector4, but after thinking it through, I realized it is unnecessary. We should work withRgba32to maximize perf.