Skip to content

How do we do move parts of the autodiff stack to the GPU? #1639

Description

@SteveBronder

Description

I was going to email @t4c1 and @rok-cesnovar seperately but I figured it would be better to discuss this in an open issue for clarity.

With the new compiler and a lot of @t4c1's work I think we can move whole chunks of the autodiff stack onto the GPU. So in the compiler when it sees a whole set of functions that can take in a matrix_cl<var> I think we could get it to generate something like

// If a, b, and c are all large Eigen matrices of vars
auto a_cl = to_matrix_cl(a);
auto b_cl = to_matrix_cl(b);
auto c_cl = to_matrix_cl(c);
Eigen::Matrix<var, -1, -1> stuff = from_matrix_cl(add(a_cl, multiply(c_cl, d_cl)));

But I'm not totally sure what that would look like, maybe something like the below? We want a vari pointer to sit in the matrix_cl<var> and the autodiff stack so when the reverse pass is called the correct chain method is also called

Example

namespace internal {

// Adding two var matrices on gpu
class add_vv_vari_cl : public op_vv_vari_cl {
 public:
  add_vv_vari_cl(const matrix_cl<var>& avi, const matrix_cl<var>& bvi)
      : op_vv_vari_cl(avi.val() + bvi.val(), avi, bvi) {}
  void chain() {
      avi_.adj() += adj_;
      bvi_.adj() += adj_;
  }
};

matrix_cl<var> operator+(const matrix_cl<var>& a, const matrix_cl<var>& b) {
  // need a new vari constructor for matrix_cl<var>
  return matrix_cl<var>(new add_vv_vari_cl(a, b));
}

I think that makes sense? Figuring out which variables are candidates to be moved over to the GPU on the compiler side would probably be a bigger issue. It would be something like how we would check for a static_matrix like bob mentioned

Current Version:

v3.0.0

Activity

  1. rok-cesnovar commented on Jan 23, 2020

    @rok-cesnovar
    Member

    Thanks for bringing this up.

    I have been thinking about this a lot and I think how you drew it up is more or less how I envisioned it. I never got to writing the code out, but this example looks great. Adding var overloads and adding compiler support is the last phase to a fully OpenCL supported Stan. With Tadej's @t4c1 advances this seems like a great time to start.

    Regarding when to move to the OpenCL device or not I think there are 2 stages:

    1. the compiler figures out what is a good candidate to be moved. I think I have a good preliminary way that I can elaborate on, but there are a few more things we need to add before we come to this. But basically if you find a "fast" function, use the matrix_cl overload and then use matrix_cl overload for functions that build inputs and use outputs of this function if they exist and grow out from there. The "fast" functions are those that could potentially be much faster on the GPU. We probably dont want to move just A+B-C.

    2. Add suport for a scheduler that would find if the OpenCL or C++ variant is faster. I worked on this a bit and also have a masters student trying some things out on this front. Basically:

    use_opencl = select_device("section_name_or_just_ID") 
    start_profile_section("section_name_or_just_ID", use_opencl)
    if(use_opencl){
        // opencl overloads
    }else {
        // cpu code
    }
    end_profile_section("section_name_or_just_ID")

    For (2.) we need to add profiling to the autodiff. I will a post design doc for profiling the log prob evaluation (which is basically the same as profiling the Stan models execution) over the weekend. Its not targeted specifically to the OpenCL parts because I think it could be useful outside of it. If it wont pass as a general solution we can think about hiding it behind STAN_OPENCL to not interfere with non-OpenCL code.

    Sorry if I derailed the topic a bit, I have been thinking more on thoroughly on this stuff than the var overloads.

  2. t4c1 commented on Jan 23, 2020

    @t4c1
    Contributor

    That will require LOTS of new code, but that might be unavoidable if we want to support this feature. One thing we should at least think about is whether we can in some way reuse existing rev code without duplicating it.

    Another important question is how to make this use kernel generator in efficient way. I would like rev code to create kernel generator expressions instead of evaluating each operation as a separate kernel.

  3. wds15 commented on Jan 23, 2020

    @wds15
    Contributor

    These could be potentially useful concepts which work for across device ad tree like for mpi or threads. No?

  4. SteveBronder commented on Jan 23, 2020

    @SteveBronder
    CollaboratorAuthor

    ^If we can figure out a good form for this then yeah possibly! It's going to matter how much we can/should specialize it for the OpenCL stuff. imo I like the idea of threads breaking up larger parts of the expression graph to compute in parallel (which if the gpu stuff is done right would compose both together nicely). Note the OpenCL stuff can also run on CPUs so having threads here may actually not be necessary

    That will require LOTS of new code, but that might be unavoidable if we want to support this feature. One thing we should at least think about is whether we can in some way reuse existing rev code without duplicating it.

    Another important question is how to make this use kernel generator in efficient way. I would like rev code to create kernel generator expressions instead of evaluating each operation as a separate kernel.

    Agree with all points above. I'd like to re-use our current code but I'm having a hard time envisioning how that would work

    The kernel generator portion in my brain makes sense to me, the chain method would just be building up the elementwise methods. Still mulling how literally to go about it tho

  5. SteveBronder commented on Feb 9, 2020

    @SteveBronder
    CollaboratorAuthor

    Munged around for an hour on this tonight. Maybe the first step is removing all the specializations for vv_vari, vd_vari, etc. and having something like

    template <typename... Types>
    class op_vari : public vari {
     protected:
        std::tuple<Types...> vi_;
     public:
       template <std::size_t Ind>
       auto& get() {
         return std::get<Ind>(vi_);
       }
      op_vari(double f, Types... args)
          : vari(f), vi_(std::make_tuple(args...)) {}
    };

    Need to figure out some way to inherit nicely from a matrix_cl instead of a vari though

  6. rok-cesnovar commented on Dec 8, 2020

    @rok-cesnovar
    Member

    This can be closed, we are now at the "implement rev matrix_cl functions" stage. See #1854

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions