Skip to content

Refactor JpegEncoder color conversion to enable SIMD optimization #810

Description

@antonfirsov

Would make the most sense to address this together with #808, AVX2 optimizing the YCbCr and grayscale path.

Activity

  1. JimBobSquarePants commented on Feb 23, 2019

    @JimBobSquarePants
    Member

    Recommend that we pause this until we get the new SIMD APIs in System.Runtime.Intrinsics.

  2. tannergooding commented on Mar 21, 2019

    @tannergooding
    Contributor

    I took a look at this briefly and I do think you will need HWIntrinsics to handle this properly. You might be able to get part of the way with Vector<T>, but it is lacking for things like register blending that would make it quite a bit more efficient.

    I think that the CalculateWeights function can also be vectorized, but it would require modifying the IResampler to expose a new method that took a Vector4.

    There might also be some small perf wins if the implementors of IResampler were made sealed (which helps with devirtualation) or if you made them structs (which might allow some generic specialization tricks).

  3. antonfirsov commented on Mar 22, 2019

    @antonfirsov
    MemberAuthor

    @tannergooding my understanding is that IResampler invocations are usually not on hot path (at least not in resize).

    I like to push towards extending such interfaces with bulk methods (eg resampler.SampleValues(sourceSpan, destSpan). (This way we can YAGNI it avoiding early large refactors)

    What I think we really need is: 100% simdified convolution (it's on our hottest resize path). For rhat we need shuffling intrinsics so we can at least temporarily switch from AOS to SOA layout.

  4. JimBobSquarePants commented on Mar 22, 2019

    @JimBobSquarePants
    Member

    IResampler is on a hot path in our affine and projective transforms. We can't use bulk operations there unfortunately either.

  5. antonfirsov commented on Mar 22, 2019

    @antonfirsov
    MemberAuthor

    We can't use bulk operations there unfortunately either.

    Why? Basically this is what we do in CalculateWeights()

    Or am I missing something?

  6. JimBobSquarePants commented on Mar 22, 2019

    @JimBobSquarePants
    Member

    The weights are currently calculated on the fly and per pixel. You have to calculate them based upon the transformed location not the input one which isn’t an integral vector. Always looking for a better way though

  7. antonfirsov commented on Mar 22, 2019

    @antonfirsov
    MemberAuthor

    Yeah, but you are still collecting them into a linear destination buffer. What we can do is to have method with similar semantics to CalculateWeights() right on IResampler. This way we can SIMD calculate sampler values in the IResampler implementation, for a given range of inputs. (Same way as we do in other bulk methods across the lib.)

    A method like SampleValues(Span<float> values, Span<float> result) or probably SampleValues(float startVal, float delta, int count, Span<float> result) can do the job. These API's could fit the usage in both ResizeKernelMap and TransformKernelMap.

  8. JimBobSquarePants commented on Mar 22, 2019

    @JimBobSquarePants
    Member

    I think is see where you are going here. I need coffee though, it's been a busy week!

  9. added this to the 1.0.0 milestone on Apr 24, 2020
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions