Repository navigation
Analysis & Discussion: Jpeg & Resize processing pipelines, improvement opportunities #1064
Description
Activity
Thanks for taking the time writing all of the above, it's very informative. 🤯
Focusing on Jpeg decoding optimization for now I would advocate for being as radical as possible.
I propose redefining the entire JpegPostProcessor pipeline as three separate implementations based on an integer pipeline. This includes Dequantization, IDCT, Subsampling, and Colorspace transforms.
- Scalar. (Old framework, odd devices, edge cases)
- Limited Intrinsics. (NET Core 2.1 with good migration path)
- Full Intrinsics (NET Core 3+, the future)
I would focus on 1 and 3 as a priority and refactor our current floating point implementation to fit via scaling at the beginning as you describe.
I would also suggest to keep the current floating point pipeline in the codebase as is, to avoid perf regressions for pre-3.0 users. I believe those platforms will be still relevant for many customers for a couple of other years.
The upgrade path from NET Core 2 to 3 is surprisingly simple for the most part. I recently ported 5 quite complex libraries in a matter of days with very little refactoring required so while 2.1 has LTS support until August 2021 I believe many customers will have moved on by then.
The benefit I see from cleanly sliced implementations are the following:
- You get to write each implementation without compromise. Scalar does not suffer from possible performance deficit floating point, and the full intrinsic pipeline can be optimized to maximum capacity.
- Cleaner, more easy to understand, and debug code. Easier to document also.
I appreciate that this is a lot of work but together I think we can do it. I'm also thinking V1 not RC as a milestone since all the APIs are internal.
Focusing on Jpeg decoding optimization for now I would advocate for being as radical as possible.
I propose redefining the entire JpegPostProcessor pipeline as three separate implementations based on an integer pipeline. This includes Dequantization, IDCT, Subsampling, and Colorspace transforms.
Now the bitter pill: the total amount of work for replacing the entire pipeline is huge. Think of at least 2-3 man-weeks of full time work, assuming that we know exactly what we are doing (I wouldn't dare to say so about myself). I make these estimations based on my own experiences and by lurking in
@dcommander's comments in libjpeg-turbo, especially in issues which are marked with "funding-needed"I'm also thinking V1 not RC as a milestone since all the APIs are internal.
Because of the amount of work, I don't think it makes sense to talk about a full scale rewrite within the V1 timeframe. We want to get the library released before 2021. Also: everything being internal doesn't mean that a significant rewrite is an option during a pre-release bugfix cycle. (Regressions, behavioral breaking changes.)
There are other problems about aiming full-integer pipeline everywhere:
- Scalar. (Old framework, odd devices, edge cases).
Expect a regression of an order of magnitude. This is basically the stuff we started with in 2016. And it's a big amount of extra work to do it properly.
- Limited Intrinsics. (NET Core 2.1 with good migration path)
This is not possible with integers because of missing intrinsics as fundamental as division.
Even if we started ImageSharp in 2019, I would say that FP pipelines are valuable and worth to implement:
- Performance does not suck on stuff != netcoreapp3.*
- Individual transformation code is straightforward and readable: it shows the calculation you are actually doing. No arbitrary bit magic. It can be used as a reference to understand what we need to calculate in other pipelines.
- Other libs also have junctions in their pipeline because of historical reasons and platform-specific optimizations. For me this is just a small part of the overall complexity.
Summary/TLDR
My opinion is exactly the opposite: We should be incremental and conservative when it's about refactoring. The low hanging fruits (II. b++ and III. b++) would bring visible and siginficant improvements. And by "low hanging" I mean: can be implemented in a couple of man-days, instead of weeks. There is no silver bullet for removing complexity in this stuff.While doing the optimizations, we can improve the understandability of the code by cleaning it up, adding comments, and introduce simplifications in the pipeline where the perf impacts are limited. Eg: at a certain point, we can consider dropping all the
Vector<T>code, sinceVector4is fast enough, and way easier to read. (=> Eliminate huge part of#if-s and other complex pipeline junctions.)EDIT
Typos, small additions.The low hanging fruits (II. b++ and III. b++) would bring visible and siginficant improvements. And by "low hanging" I mean: can be implemented in a couple of man-days, instead of weeks. There is no silver bullet for removing complexity in this stuff.
If you truly believe this then I'm with you all the way and we'll do it your way. You are, by far and away, the performance expert here.
Reacted by Anton Firszov, Sergio Pedri, Uwe Keim and samsosaClosing this as mostly obsolete.
The JPEG decoder has changed substantially since this was opened: the main decode path now stays on the float/SIMD pipeline, does YCbCr->RGB conversion in place, packs directly into the target pixel type (including Rgb24, so the old Rgba32 detour concern no longer applies to normal decode), and supports optimized decode-time downscaling via the JPEG decoder resize options.
Also, we do not want to move toward an int-based pipeline, so that part of the original proposal is no longer aligned with the decoder’s direction. There may still be smaller follow-up optimizations in specific scaled-copy/resize paths, but those are separate issues rather than this one.
Reacted by Anton Firszov
Introduction
Apart from the API simplification, the main intent of #907 was to enable new optimizations: it's possible to eliminate a bunch of unnecessary processing steps from the most common YCbCr Jpeg thumbnail making use-case. As it turned out in #1062, simply changing the pixel type to
Rgba24is not sufficient, we need to implement the processing pipeline optimizations enabled by the .NET Core 3.0 Hardware Intrinsic API, especially by the shuffle and permutation intrinsics which are allowowing fast conversion between different pixel type representations and component orders (eg.Rgba32<-->Rgb24), as well as fast conversion between Planar/SOA and Packed/AOS pixel representations. The latter is important because raw Jpeg data consists of 3 planes representing the YCbCr data, while an ImageSharpImageis always packed.This analyisis:
Rgb24slowdown in Add La16 and La32 IPixel formats. #1062Please let me know, if some pieces are still hard to follow. It's worth to check out all URL-s while reading.
TLDR
If you want to hear some good news before reading through the whole thing, jump to the Conclusion part 😄
Why is
Rgb24post processing slow in our current code?YCbCr->TPixelconversions, the generic caseJpegImagePostprocessoris processing the YCbCr data in two steps:Vector4RGBA buffers. The two operations are carried out together by the matchingJpegColorConverter. With the YCbCr colorspace which has only 3 components, this is already a sub-optimal, since the 4th alpha component (Vector4.W) is redundant.Vector4packing is done with non-vectorized code.Vector4buffer to pixel buffer, using the pixel specific implementation.Rgba32vsRgb24PostProcessIntoImage<Rgba32>PostProcessIntoImage<Rgb24>The difference is that
PixelOperations<Rgba32>.FromVector4()does not need to do any component shuffling, only expandingbytevalues tofloat-s, while inPixelOperations<Rgba32>.FromVector4(), we first convert the float buffers toRgba32buffers (fast), which is followed by anRgba32->Rgb24conversion using the sub-optimal default conversion implementation. This operation:JpegColorConverterwith a method to pack data intoVector3buffers, we could convertVector3data intoRgb24data exactly the same way we do theVector4->Rgba32conversion.Definition of Processing Pipelines
Personally, my memory is terrible and I always need to reverse engineer my own code when we want to understand what's happening and make decisions. Lack of comments and confusing terminology is also misleading. To get a good overview, it's really important to step back and abstract away implementation details, by thinking about our algorithms as PIPELINES composed of Data States and Transformations, where
This representation is only good for analyzing data flow for a specific configuration, eg. a well defined input image format + decoder configuration + output pixel type. To visualize the junctions, we need DAG-s 🤓.
Current floating point YCbCr Jpeg Color Processing & Resize pipelines, improvement opportunities
Presumtions:
netcoreapp2.1(enablesVector.Widen)Vector<T>-s are in fact AVX2 registers andVector<T>intrinsics are JIT-ed to AVX2 instructionsVector4operations are JIT-ed to SSE2 instructions(I.) Converting raw jpeg spectral data to YCbCr planes
CopyBlocksToColorBufferInt16jpeg components (3 xBuffer2D<Block8x8>, Y+Cb+Cr)Int16->Int32widening andInt32->floatconversion, both usingVector<T>, implemented inBlock8x8F.LoadFrom(Block8x8)floatjpeg components (3 xBuffer2D<Block8x8>, Y+Cb+Cr)Block8x8F.MultiplyInplace(DequantiazationTable)floatjpeg components (3 xBuffer2D<Block8x8>, Y+Cb+Cr)floatjpeg color channels (3 xBuffer2D<Block8x8>, Y+Cb+Cr)Vector<T>. Rounding is needed for better libjpeg compatibilityfloatjpeg color channels normalized to 0-255 (3 xBuffer2D<Block8x8>, Y+Cb+Cr)Block8x8.CopyTo())(super misleading name!)floatjpeg color channels normalized to 0-255 (3 xBuffer2D<float>, Y+Cb+Cr)(II. a) Converting the Y+Cb+Cr planes to an
Rgba32bufferRgba32buffer, done byConvertColorsIntofloatjpeg color channels normalized to 0-255 (3 xBuffer2D<float>, Y+Cb+Cr)Vector4bufferMemory<Vector4>Vector4buffer to anRgba32buffer. In theRgba32case case, the input buffer could be handled as homogenousfloatbuffer, where all individualfloatvalues should be converted tobyte-s. The conversion is implemented inBulkConvertNormalizedFloatToByteClampOverflows, utilizing AVX2 conversion and narrowing operations throughVector<T>Rgba32buffer(II. b) Converting the Y+Cb+Cr planes to an
Rgb24buffer, current sub-optimal pipelineRgb24buffer, done byConvertColorsIntofloatjpeg color channels normalized to 0-255 (3 xBuffer2D<float>, Y+Cb+Cr)Vector4bufferMemory<Vector4>Vector4buffer to anRgba32buffer, utilizingBulkConvertNormalizedFloatToByteClampOverflows, utilizing AVX2 conversion and narrow operations throughVector<T>Rgba32bufferPixelOperations<Rgb24>.FromRgba32()(sub-optimal, extra transformation!)Rgb24buffer(II. b++) Converting the Y+Cb+Cr planes to an
Rgba24buffer, IMPROVEMENT PROPOSALSee #1121
(III. a) Resize
Image<Rgba32>, current pipelineTODO
(III. b) Resize
Image<Rgb24>, current pipelineTODO.
Without any change, the current code shall run faster than for
Image<Rgba32>.(III. b++) Resize
Image<Rgb24>, IMPROVEMENT PROPOSALTODO
Integer-based SIMD pipelines
Although the Hardware Intrinsic API removes all theoretical boundaries to have 1:1 match with other high performance imaging libraries, for both Jpeg Decoder and Resize by utilizing AVX2 and SSE2 integer algorithms, there is a big practical challange: It's very hard to introduce these improvements in an iterative manner.
It's not possible to exchange the elements of the Jpeg pipeline at arbitrary points, because it would lead to insertion of extra
float<->Int16/32conversions. To overcome this, we should start introducing integer transformations and data states at the beginning and/or at the end of the pipeline. This could be done by replacing the transformations and the data states in subsequent PR-s while moving theInt16->floatconversion towards the bottom (when starting from the beginning), and thefloat->byteconversion towards the top (when starting from the end). EG:YCbCr24->Rgb24SIMD conversion firstConclusion
If we aim for low hanging fruits, I would start by implementing (II. b++) and (III. b++). After that, we can continue by introducing integer SIMD operations starting at the beginning or at the end of the Jpeg pipeline.
I would also suggest to keep the current floating point pipeline in the codebase as is, to avoid perf regressions for pre-3.0 users. I believe those platforms will be still relevant for many customers for a couple of other years.