Repository navigation
How do we do move parts of the autodiff stack to the GPU? #1639
Description
Activity
Thanks for bringing this up.
I have been thinking about this a lot and I think how you drew it up is more or less how I envisioned it. I never got to writing the code out, but this example looks great. Adding var overloads and adding compiler support is the last phase to a fully OpenCL supported Stan. With Tadej's @t4c1 advances this seems like a great time to start.
Regarding when to move to the OpenCL device or not I think there are 2 stages:
-
the compiler figures out what is a good candidate to be moved. I think I have a good preliminary way that I can elaborate on, but there are a few more things we need to add before we come to this. But basically if you find a "fast" function, use the matrix_cl overload and then use matrix_cl overload for functions that build inputs and use outputs of this function if they exist and grow out from there. The "fast" functions are those that could potentially be much faster on the GPU. We probably dont want to move just A+B-C.
-
Add suport for a scheduler that would find if the OpenCL or C++ variant is faster. I worked on this a bit and also have a masters student trying some things out on this front. Basically:
use_opencl = select_device("section_name_or_just_ID") start_profile_section("section_name_or_just_ID", use_opencl) if(use_opencl){ // opencl overloads }else { // cpu code } end_profile_section("section_name_or_just_ID")
For (2.) we need to add profiling to the autodiff. I will a post design doc for profiling the log prob evaluation (which is basically the same as profiling the Stan models execution) over the weekend. Its not targeted specifically to the OpenCL parts because I think it could be useful outside of it. If it wont pass as a general solution we can think about hiding it behind STAN_OPENCL to not interfere with non-OpenCL code.
Sorry if I derailed the topic a bit, I have been thinking more on thoroughly on this stuff than the var overloads.
-
That will require LOTS of new code, but that might be unavoidable if we want to support this feature. One thing we should at least think about is whether we can in some way reuse existing rev code without duplicating it.
Another important question is how to make this use kernel generator in efficient way. I would like rev code to create kernel generator expressions instead of evaluating each operation as a separate kernel.
Reacted by Steve Bronder and Rok ČešnovarThese could be potentially useful concepts which work for across device ad tree like for mpi or threads. No?
^If we can figure out a good form for this then yeah possibly! It's going to matter how much we can/should specialize it for the OpenCL stuff. imo I like the idea of threads breaking up larger parts of the expression graph to compute in parallel (which if the gpu stuff is done right would compose both together nicely). Note the OpenCL stuff can also run on CPUs so having threads here may actually not be necessary
That will require LOTS of new code, but that might be unavoidable if we want to support this feature. One thing we should at least think about is whether we can in some way reuse existing rev code without duplicating it.
Another important question is how to make this use kernel generator in efficient way. I would like rev code to create kernel generator expressions instead of evaluating each operation as a separate kernel.
Agree with all points above. I'd like to re-use our current code but I'm having a hard time envisioning how that would work
The kernel generator portion in my brain makes sense to me, the chain method would just be building up the elementwise methods. Still mulling how literally to go about it tho
Reacted by Rok ČešnovarMunged around for an hour on this tonight. Maybe the first step is removing all the specializations for
vv_vari,vd_vari, etc. and having something liketemplate <typename... Types> class op_vari : public vari { protected: std::tuple<Types...> vi_; public: template <std::size_t Ind> auto& get() { return std::get<Ind>(vi_); } op_vari(double f, Types... args) : vari(f), vi_(std::make_tuple(args...)) {} };
Need to figure out some way to inherit nicely from a matrix_cl instead of a vari though
This can be closed, we are now at the "implement rev matrix_cl functions" stage. See #1854
Description
I was going to email @t4c1 and @rok-cesnovar seperately but I figured it would be better to discuss this in an open issue for clarity.
With the new compiler and a lot of @t4c1's work I think we can move whole chunks of the autodiff stack onto the GPU. So in the compiler when it sees a whole set of functions that can take in a
matrix_cl<var>I think we could get it to generate something likeBut I'm not totally sure what that would look like, maybe something like the below? We want a vari pointer to sit in the
matrix_cl<var>and the autodiff stack so when the reverse pass is called the correct chain method is also calledExample
I think that makes sense? Figuring out which variables are candidates to be moved over to the GPU on the compiler side would probably be a bigger issue. It would be something like how we would check for a
static_matrixlike bob mentionedCurrent Version:
v3.0.0