Repository navigation
Enable Avx2 optimizations on Porter-Duff operations. #2340
Description
Activity
- pinned this issue
on Feb 2, 2023 The method
Vector256<float> Over(Vector256<float> destination, Vector256<float> source, Vector256<float> blend)would transform 2 color/pixel instances at the same time. This is an AVX2 implementation detail that shouldn't be exposed on the public API. We should go for whatever optimizations that make the span-based bulk methods faster, this what perf critical code should use anyways. This means that the AVX2 code above would be likely most efficient if it lives in the loop iterating through the spans.Thinking strategically, I have doubts if we need this in 3.0. Assuming the main goal is to optimize Drawing, 2x speedup is good but not dramatic enough to close the perf gap between IS.Drawing and competitors. Don't have a good grasp how much work would be to implement the optimizations for most important blenders, but if it's a lot, I'm not sure if it's worth to delay 3.0 with this work. Also note that the optimization would only kick in in the first Drawing version which is based on ImageSharp 3.0.
Reacted by James Jackson-SouthYep. We'd never expose the detail. Rather create Vector256 versions of all the internal
PorterDuffFunctionsmethods currently implemented usingVector4. We'd then call those methods in each blenderImageSharp/src/ImageSharp/PixelFormats/PixelBlenders/DefaultPixelBlenders.Generated.cs
Lines 43 to 50 in 37f5cc5
protected override void BlendFunction(Span<Vector4> destination, ReadOnlySpan<Vector4> background, ReadOnlySpan<Vector4> source, float amount) { amount = Numerics.Clamp(amount, 0, 1); for (int i = 0; i < destination.Length; i++) { destination[i] = PorterDuffFunctions.NormalSrc(background[i], source[i], amount); } } I don't think the workload is too great though I could be wrong. I had a look through and I believe I'd be capable of most of them though others could do them much, much faster.
I'm really concerned about Drawing and the performance issues we have there (along with the complete lack of progress - It's been a beta now for 5 years!) and really want to be able to deliver a V1 asap. A 2X speedup would be the minimum I would expect for a V1 which should probably target ImageSharp unless someone has a good argument otherwise.
- unpinned this issue
on Feb 20, 2023
Discussed in #2328
Originally posted by JimBobSquarePants January 26, 2023
Given that #1433 is not achievable within a reasonable timescale for v3 and might well be better implemented in a pattern using static virtual members in interfaces I'd like to propose that we enhance the existing
Vector4based Porter-Duff implementations withVector256<float>equivalents.This will give us a 2X performance improvement on supported devices over the existing implementation which we badly need if we are ever able to release a v1 of ImageSharp.Drawing.
We should be able to implement each separable method in the same fashion as the existing methods. The pixel blender implementations can call these.
For example. The existing
Overmethod......Could be implemented in the following manner.
For someone with a good grasp of the SIMD APIs this should not be a huge undertaking.
Thoughts @antonfirsov @br3aker ??