Repository navigation
[$] Optimize NumPy SIMD algorithms for Power VSX #13393
Description
Activity
The main problem here is lack of expertise and hardware among the developers. Do you know anyone who could help?
This already has been discussed with Ralf Gommers and Julian Taylor, who suggested available developers. Of course, anyone is welcome to work on the issue and claim the financial bounty.
Can I work on it?
@seiko2plus Are you familiar with Power VSX and have the appropriate hardware to work on? This is one of those situations where we need an expert, but if you feel up to it, go ahead.
Reacted by ThirteenMonths@charris, Yes, I'm familiar with Power/VSX, I have access to Power8&9 hardware and also I have a wide experience for other SIMD extensions too.
@seiko2plus Great! Go for it.
@charris , Thanks!
A bit of context for this:
- This is the first time (as far as I know) that anyone has offered a bounty for implementing a feature in NumPy. It's a bit of an experiment for us as a project.
- We have pre-discussed this on the steering committee mailing list. @juliantaylor has done a lot of work on this before, so if he wants to tackle this that would be good.
- Keep in mind that it's not just about implementing - the feature has to be merged as well. For one, there's reviewer effort (usually in short supply; if part of the bounty would go to the reviewer or to the project that would probably help ....). Also, a requirement for merging should be that VSX support does not increase the maintainance burden substantially (this is possible).
- Keep in mind also that there's open PRs on AVX support. Please check those, and build on them or discuss in case of possible conflicts. E.g. AVX for floats? #11113 is relevant, and ENH: Use AVX for float32 implementation of np.sin & np.cos #13368 is active right now.
- We have powerpc CI on TravisCI, which is good. PR should ensure that VSX is exercised there.
Also with GCC on AIX systems should work.EDIT: clarified by @edelsohn in the description that AIX isn't a hard requirement.
I would recommend that if @seiko2plus or someone else tries to tackle this right now, they post a summary here of what they plan to do, or send an early WIP PR. It would be unfortunate if someone spends a lot of time and then in review we say that the changes should have been done differently.
Second recommendation: please give @juliantaylor a bit of time (a couple of days) to respond before starting work @seiko2plus.
EDIT: given that this is the first bounty for NumPy, I also posted a message to the numpy-discussion mailing list about this, https://mail.python.org/pipermail/numpy-discussion/2019-April/079371.html
@rgommers, Thank you for the clarification and sure I'm going to follow your recommendations.
I currently cannot do it, please consider other people.
A solution to this should provide proper build changes to support the various architectures, right now it is focused on SSE/AVX only. I believe if it's is going to move forward it will require more than just cloning simd.inc and duplicating the code there for another architecture. Theoretically the vector built-in's for the compiler should have a default implementation, and specialization for x86*/ARM*/PPC64[LE] specific simd each a separate implementation that is built according to availability (decided by compiler/build ? not all builds are native) and/or at runtime via features testing? I started looking into this when I got an email about it 2 days ago, however now I see someone is working on it already.
@dmiller423 I don't think anybody is working on this yet, in fact, I think the scope/management has yet to be determined. So stay around.
20 remaining items
PPC64LE Linux has a set of SSE/SSSE3 emulation intrinsics headers. An IBM team built NumPy with the headers and saw a 15% performance improvement on benchmarks that stressed the SSE intrinsics code path, which matched the improvement for x86. This functional hack confirms the benefit of SIMD for Power.
x86 derives more benefit from AVX intrinsics, for which emulation has not been written for Power. The SIMD infrastructure should address that.
Thanks @edelsohn, makes sense that there's a benefit but still good to see it confirmed.
x86 derives more benefit from AVX intrinsics, for which emulation has not been written for Power. The SIMD infrastructure should address that.
I think it does already, assuming you mean "AVX equivalent instructions as universal intrinsics".
By SIMD infrastructure, I mean the universal intrinsics. IBM observed that more of NumPy SIMD seems to utilize AVX/AVX2, for which there currently is no Power emulation. When NumPy is converted to universal intrinsics and the universal intrinsics are implemented for Power, then the AVX-equivalent benefit in NumPy will be exposed for Power as well.
Reacted by Ralf GommersDid anyone already try to use generic C/C++ code and optimize it using OpenMP? In tests with Tesseract such code was comparable to hand optimized code for AVX2, and the same source code worked for ARM NEON and PowerPC Altivec, too.
@stweil can you explain a bit more what you mean? SIMD instructions give faster single-threaded code, and when you say OpenMP I'm thinking about parallelism.
OpenMP gained support for vectorization in addition to parallelization some time ago. See for example
https://www.openmp.org/spec-html/5.0/openmpsu42.htmlYes, OpenMP has different compiler pragmas for parallelism (not indended here) and for SIMD (that's what is desired here). C Code which uses these pragmas and which is compiled with the right compiler and compiler flags (optimization, cpu) automatically uses the SIMD machine instructions.
However I think (though I'm not sure) that OpenMP SIMD will only cover a small subset of what universal intrinsics can do. For example I don't see how things like shuffling, masked operations, horizontal addition etc. could be achieved with OpenMP pargmas.
@stweil,
OpenMPis counting on the compiler level of auto-vectorization, and NumPy already provides similar functionality.numpy/numpy/core/src/umath/fast_loop_macros.h
Lines 106 to 135 in de3fcf1
/* * loop with contiguous specialization * op should be the code working on `tin in` and * storing the result in `tout *out` * combine with NPY_GCC_OPT_3 to allow autovectorization * should only be used where its worthwhile to avoid code bloat */ #define BASE_UNARY_LOOP(tin, tout, op) \ UNARY_LOOP { \ const tin in = *(tin *)ip1; \ tout *out = (tout *)op1; \ op; \ } #define UNARY_LOOP_FAST(tin, tout, op) \ do { \ /* condition allows compiler to optimize the generic macro */ \ if (IS_UNARY_CONT(tin, tout)) { \ if (args[0] == args[1]) { \ BASE_UNARY_LOOP(tin, tout, op) \ } \ else { \ BASE_UNARY_LOOP(tin, tout, op) \ } \ } \ else { \ BASE_UNARY_LOOP(tin, tout, op) \ } \ } \ while (0) I think we can close this issue and declare the bounty complete. The original bounty says:
NumPy contains SIMD vectorized code for x86 SSE and AVX. This issue is a feature request to implement native, equivalent enablement for Power VSX, achieving equivalent speedup appropriate for the SIMD vector width of VSX (128 bits).
We have implemented the infrastructure needed to enable the porting of loops via Universal SIMD, and even ported some of the loops, so the enablement is complete.
I completely agree. Sayed and the NumPy core developers have done an amazing job!
Do you want me to close the issue? I have been deferring to the NumPy developers.
Thanks @edelsohn. Agreed that the solution exceeded expectations. I will close it then.
Awesome, thanks everyone for the excellent work!
Thank you for everything <3
NumPy contains SIMD vectorized code for x86 SSE and AVX. This issue is a feature request to implement native, equivalent enablement for Power VSX, achieving equivalent speedup appropriate for the SIMD vector width of VSX (128 bits).
EDIT (by @rgommers): link to bounty: https://www.bountysource.com/issues/73221262-optimize-numpy-simd-algorithms-for-power-vsx
The focus is PPC64LE Linux. If the optimization can be portable to AIX (big endian) that's great, but not a strict requirement. In other words, if AIX continues to use the scalar code for now, that's okay.