Skip to content

[$] Optimize NumPy SIMD algorithms for Power VSX #13393

Description

@edelsohn

NumPy contains SIMD vectorized code for x86 SSE and AVX. This issue is a feature request to implement native, equivalent enablement for Power VSX, achieving equivalent speedup appropriate for the SIMD vector width of VSX (128 bits).

EDIT (by @rgommers): link to bounty: https://www.bountysource.com/issues/73221262-optimize-numpy-simd-algorithms-for-power-vsx

The focus is PPC64LE Linux. If the optimization can be portable to AIX (big endian) that's great, but not a strict requirement. In other words, if AIX continues to use the scalar code for now, that's okay.

Activity

  1. charris commented on Apr 23, 2019

    @charris
    Member

    The main problem here is lack of expertise and hardware among the developers. Do you know anyone who could help?

  2. edelsohn commented on Apr 23, 2019

    @edelsohn
    Author

    This already has been discussed with Ralf Gommers and Julian Taylor, who suggested available developers. Of course, anyone is welcome to work on the issue and claim the financial bounty.

  3. seiko2plus commented on Apr 23, 2019

    @seiko2plus
    Member

    Can I work on it?

  4. charris commented on Apr 23, 2019

    @charris
    Member

    @seiko2plus Are you familiar with Power VSX and have the appropriate hardware to work on? This is one of those situations where we need an expert, but if you feel up to it, go ahead.

  5. seiko2plus commented on Apr 23, 2019

    @seiko2plus
    Member

    @charris, Yes, I'm familiar with Power/VSX, I have access to Power8&9 hardware and also I have a wide experience for other SIMD extensions too.

  6. charris commented on Apr 23, 2019

    @charris
    Member

    @seiko2plus Great! Go for it.

  7. seiko2plus commented on Apr 23, 2019

    @seiko2plus
    Member

    @charris , Thanks!

  8. rgommers commented on Apr 23, 2019

    @rgommers
    Member

    A bit of context for this:

    • This is the first time (as far as I know) that anyone has offered a bounty for implementing a feature in NumPy. It's a bit of an experiment for us as a project.
    • We have pre-discussed this on the steering committee mailing list. @juliantaylor has done a lot of work on this before, so if he wants to tackle this that would be good.
    • Keep in mind that it's not just about implementing - the feature has to be merged as well. For one, there's reviewer effort (usually in short supply; if part of the bounty would go to the reviewer or to the project that would probably help ....). Also, a requirement for merging should be that VSX support does not increase the maintainance burden substantially (this is possible).
    • Keep in mind also that there's open PRs on AVX support. Please check those, and build on them or discuss in case of possible conflicts. E.g. AVX for floats? #11113 is relevant, and ENH: Use AVX for float32 implementation of np.sin & np.cos #13368 is active right now.
    • We have powerpc CI on TravisCI, which is good. PR should ensure that VSX is exercised there. Also with GCC on AIX systems should work. EDIT: clarified by @edelsohn in the description that AIX isn't a hard requirement.

    I would recommend that if @seiko2plus or someone else tries to tackle this right now, they post a summary here of what they plan to do, or send an early WIP PR. It would be unfortunate if someone spends a lot of time and then in review we say that the changes should have been done differently.

    Second recommendation: please give @juliantaylor a bit of time (a couple of days) to respond before starting work @seiko2plus.

    EDIT: given that this is the first bounty for NumPy, I also posted a message to the numpy-discussion mailing list about this, https://mail.python.org/pipermail/numpy-discussion/2019-April/079371.html

  9. seiko2plus commented on Apr 23, 2019

    @seiko2plus
    Member

    @rgommers, Thank you for the clarification and sure I'm going to follow your recommendations.

  10. juliantaylor commented on Apr 24, 2019

    @juliantaylor
    Contributor

    I currently cannot do it, please consider other people.

  11. dmiller423 commented on Apr 25, 2019

    @dmiller423

    A solution to this should provide proper build changes to support the various architectures, right now it is focused on SSE/AVX only. I believe if it's is going to move forward it will require more than just cloning simd.inc and duplicating the code there for another architecture. Theoretically the vector built-in's for the compiler should have a default implementation, and specialization for x86*/ARM*/PPC64[LE] specific simd each a separate implementation that is built according to availability (decided by compiler/build ? not all builds are native) and/or at runtime via features testing? I started looking into this when I got an email about it 2 days ago, however now I see someone is working on it already.

  12. charris commented on Apr 25, 2019

    @charris
    Member

    @dmiller423 I don't think anybody is working on this yet, in fact, I think the scope/management has yet to be determined. So stay around.

  13. 20 remaining items

  14. edelsohn commented on Mar 27, 2020

    @edelsohn
    Author

    PPC64LE Linux has a set of SSE/SSSE3 emulation intrinsics headers. An IBM team built NumPy with the headers and saw a 15% performance improvement on benchmarks that stressed the SSE intrinsics code path, which matched the improvement for x86. This functional hack confirms the benefit of SIMD for Power.

    x86 derives more benefit from AVX intrinsics, for which emulation has not been written for Power. The SIMD infrastructure should address that.

  15. rgommers commented on Mar 29, 2020

    @rgommers
    Member

    Thanks @edelsohn, makes sense that there's a benefit but still good to see it confirmed.

    x86 derives more benefit from AVX intrinsics, for which emulation has not been written for Power. The SIMD infrastructure should address that.

    I think it does already, assuming you mean "AVX equivalent instructions as universal intrinsics".

  16. edelsohn commented on Mar 29, 2020

    @edelsohn
    Author

    By SIMD infrastructure, I mean the universal intrinsics. IBM observed that more of NumPy SIMD seems to utilize AVX/AVX2, for which there currently is no Power emulation. When NumPy is converted to universal intrinsics and the universal intrinsics are implemented for Power, then the AVX-equivalent benefit in NumPy will be exposed for Power as well.

  17. stweil commented on Dec 14, 2020

    @stweil

    Did anyone already try to use generic C/C++ code and optimize it using OpenMP? In tests with Tesseract such code was comparable to hand optimized code for AVX2, and the same source code worked for ARM NEON and PowerPC Altivec, too.

  18. rgommers commented on Dec 14, 2020

    @rgommers
    Member

    @stweil can you explain a bit more what you mean? SIMD instructions give faster single-threaded code, and when you say OpenMP I'm thinking about parallelism.

  19. mreineck commented on Dec 14, 2020

    @mreineck
    Contributor

    OpenMP gained support for vectorization in addition to parallelization some time ago. See for example
    https://www.openmp.org/spec-html/5.0/openmpsu42.html

  20. stweil commented on Dec 14, 2020

    @stweil

    Yes, OpenMP has different compiler pragmas for parallelism (not indended here) and for SIMD (that's what is desired here). C Code which uses these pragmas and which is compiled with the right compiler and compiler flags (optimization, cpu) automatically uses the SIMD machine instructions.

  21. mreineck commented on Dec 14, 2020

    @mreineck
    Contributor

    However I think (though I'm not sure) that OpenMP SIMD will only cover a small subset of what universal intrinsics can do. For example I don't see how things like shuffling, masked operations, horizontal addition etc. could be achieved with OpenMP pargmas.

  22. seiko2plus commented on Dec 14, 2020

    @seiko2plus
    Member

    @stweil, OpenMP is counting on the compiler level of auto-vectorization, and NumPy already provides similar functionality.

    /*
    * loop with contiguous specialization
    * op should be the code working on `tin in` and
    * storing the result in `tout *out`
    * combine with NPY_GCC_OPT_3 to allow autovectorization
    * should only be used where its worthwhile to avoid code bloat
    */
    #define BASE_UNARY_LOOP(tin, tout, op) \
    UNARY_LOOP { \
    const tin in = *(tin *)ip1; \
    tout *out = (tout *)op1; \
    op; \
    }
    #define UNARY_LOOP_FAST(tin, tout, op) \
    do { \
    /* condition allows compiler to optimize the generic macro */ \
    if (IS_UNARY_CONT(tin, tout)) { \
    if (args[0] == args[1]) { \
    BASE_UNARY_LOOP(tin, tout, op) \
    } \
    else { \
    BASE_UNARY_LOOP(tin, tout, op) \
    } \
    } \
    else { \
    BASE_UNARY_LOOP(tin, tout, op) \
    } \
    } \
    while (0)

  23. mattip commented on Jan 7, 2021

    @mattip
    Member

    I think we can close this issue and declare the bounty complete. The original bounty says:

    NumPy contains SIMD vectorized code for x86 SSE and AVX. This issue is a feature request to implement native, equivalent enablement for Power VSX, achieving equivalent speedup appropriate for the SIMD vector width of VSX (128 bits).

    We have implemented the infrastructure needed to enable the porting of loops via Universal SIMD, and even ported some of the loops, so the enablement is complete.

  24. edelsohn commented on Jan 8, 2021

    @edelsohn
    Author

    I completely agree. Sayed and the NumPy core developers have done an amazing job!

    Do you want me to close the issue? I have been deferring to the NumPy developers.

  25. mattip commented on Jan 8, 2021

    @mattip
    Member

    Thanks @edelsohn. Agreed that the solution exceeded expectations. I will close it then.

  26. rgommers commented on Jan 8, 2021

    @rgommers
    Member

    Awesome, thanks everyone for the excellent work!

  27. seiko2plus commented on Jan 9, 2021

    @seiko2plus
    Member

    Thank you for everything <3

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions