Repository navigation
Vectorize everything with broadcasting #202
Description
Activity
I'm reopening this issue because the pull request that was merged in only provides the infrastructure to test these vectorize functions. However, if you think that we should close this issue and open a new one, I can do that.
Either way is OK. You need to make the checklist of all
unary functions to which this applies. You can probably knock
them off in multiple issues.On Mar 15, 2016, at 4:52 PM, rayleigh [email protected] wrote:
I'm reopening this issue because the pull request that was merged in only provides the infrastructure to test these vectorize functions. However, if you think that we should close this issue and open a new one, I can do that.
—
You are receiving this because you were assigned.
Reply to this email directly or view it on GitHubOkay. Copying over your response to a question of where to put files:
- The function definitions so should go here:
stan/math/prim/mat/fun/<function-name>.hpp- For functions already partly vectorized, you need
to remove the original vectorized definitions (just get
rid of them and their associated tests). - The function signatures class is going to need a utility
method to declare these, presumably something like
``add_unary_vectorized("exp") and so on.
This should declare all the basic vectorizations and
perhaps up to 4 deep for arrays of all types.- You're going to need to write (or have me write)
a. a new intro section to the functions guide in the manual,
with some new syntax for this vectorization, and
b. update the doc for each of the added functions
Also, where should the test files go? Should they go in
/test/unit/math/mix/mat/fun/<function-name>_test.hpp?Thanks and yes, that's the right place for the test files, because
they depend on mix.On Mar 17, 2016, at 10:30 AM, rayleigh [email protected] wrote:
Okay. Copying over your response to a question of where to put files:
• The function definitions so should go here:
stan/math/prim/mat/fun/.hpp• For functions already partly vectorized, you need
to remove the original vectorized definitions (just get
rid of them and their associated tests).• The function signatures class is going to need a utility
method to declare these, presumably something like
``add_unary_vectorized("exp") and so on.This should declare all the basic vectorizations and
perhaps up to 4 deep for arrays of all types.• You're going to need to write (or have me write) a. a new intro section to the functions guide in the manual, with some new syntax for this vectorization, and b. update the doc for each of the added functions
Also, where should the test files go? Should they go in /test/unit/math/mix/mat/fun/_test.hpp?—
You are receiving this because you were assigned.
Reply to this email directly or view it on GitHubThanks. Looking through the 2.8.0 Stan manual, I noticed that
absis to be deprecated. Should I vectorize it or should I only vectorizefabs?Ack, I think we undeprecated abs and are going to
deprecate fabs (to make it more like math and less
like C++). So go ahead and do both and we can see where
the chips land.On Mar 21, 2016, at 11:45 AM, rayleigh [email protected] wrote:
Thanks. Looking through the 2.8.0 Stan manual, I noticed that abs is to be deprecated. Should I vectorize it or should I only vectorize fabs?
—
You are receiving this because you were assigned.
Reply to this email directly or view it on GitHubSounds good; I vectorized both. I've been trying to vectorize the function
cbrt. It's defined for 0, but its derivatives aren't (it returns NaN). From this, I realized that the testing framework assumed that if a function is defined for a value, its derivatives are as well. To handle this, does the vectorize function throw an error if an user enters 0? Or, do I write separate code to test this case?I think we should test derivatives.
We had this discussion in the past regarding defined values and undefined
derivatives. I think the functions should trap that error and fail early if
possible, even if it conflicts with other implementations of the function.
If we're using these functions primarily where they need derivatives, it
makes no sense to knowingly propagate NaNs.On Mon, Mar 21, 2016 at 7:00 PM, rayleigh [email protected] wrote:
Sounds good; I vectorized both. I've been trying to vectorize the function
cbrt. It's defined for 0, but its derivatives aren't (it returns NaN).
From this, I realized that the testing framework assumed that if a function
is defined for a value, its derivatives are as well. To handle this, does
the vectorize function throw an error if an user enters 0? Or, do I write
separate code to test this case?—
You are receiving this because you modified the open/close state.
Reply to this email directly or view it on GitHub
#202 (comment)Okay. Just to clarify, should the vectorize-all function or the individual scalar functions include the error checking? Or, both?
That was a general comment.
What test am I supposed to run? Can you provide the full path?
On Tue, Mar 22, 2016 at 11:34 AM, rayleigh [email protected] wrote:
Okay. Just to clarify, should the vectorize-all function or the individual
scalar functions include the error checking? Or, both?—
You are receiving this because you modified the open/close state.
Reply to this email directly or view it on GitHub
#202 (comment)14 remaining items
That conditional include is nasty. Maybe we should just use
boost::math::acosh independently of platform.Is there ever a reason to include both <math.h> and ?
And doesn't that <math.h> include have to go last?The error handling can be fixed by just adding it to the
code, even if we do use Boost's definition for computing
non-error inputs.On Apr 11, 2016, at 5:40 PM, rayleigh [email protected] wrote:
Thanks for pointing that out. I took another look and I think the includes might be the issue because for acosh.hpp, the includes are:
#include <math.h>
#include <stan/math/rev/core.hpp>
#include <boost/math/special_functions/fpclassify.hpp>
#include#ifdef _MSC_VER
#include <boost/math/special_functions/acosh.hpp>
using boost::math::acosh;
#endifSo, unless it's being compiled on a Visual C++ compiler, I don't think it'll use acosh.hpp from boost/math/special_functions, which is providing the error handling. Because Stan doesn't use C++11 and acosh.hpp is only available in C++11's cmath library, the vectorized version of acosh.hpp uses acosh.hpp from boost/math/special_functions as the base function. I think whether boost/math/special_functions/acosh.hpp is included explains the inconsistency in error handling that I'm seeing.
—
You are receiving this because you were assigned.
Reply to this email directly or view it on GitHub@rayleigh Is the current checklist above up to date? Is there a reason the dozen or so
(real): realsignature ones aren't implemented?Closing in favor of newer issue #347 given that it's partially done.
Reacted by Jesse Knight- modified the milestones: This milestone has been deleted, This milestone has been deleted
on Sep 7, 2016 Can we reopen this? The logical operators (at least) seem to remain un-vectorized
Binary infix operator == with functional interpretation logical_eq requires arguments of primitive type (int or real), found left type=int[ ], right arg type=int.Thanks,
@jessexknight Probably better to start a new issue specifically or logical operations, because it's not immediately obvious how they should behave. This issue was specifically for real-valued operations and used a specific vectorization assuming real valued functions with autodiff---here there are no derivatives because we get integer values out.
I can imagine two alternatives for the logical operators:
-
int logical_eq(reals x, reals y);which returns a single truth value conjoining the element wise result. -
int[] logical_eq(reals x, reals y);which returns a container of results element wise.
Both approaches can broadcast scalars to containers. If we go with (2), then we probably want functions
all()andany()like in NumPy that return the conjunction and disjunction of a container.This is related to the issue of comparing containers to elements to return booleans, the way that R allows. That'd give us something like this
int[] xs = { 1, 0, 2, 3, 0, -1}; int[] ys = (xs == 0); // after eval, ys = { 2, 5 }
-
Thanks, I've evidently added an issue...
Just want to add that I would strongly prefer (2) to allow more flexibility.
Not sure I follow the last issue you mentioned -- I'm not familiar with this behaviour in R, and personally wouldn't find this a priority vs the vectorization of the base logical functions.
Thanks for opening the issue. I edited it slightly to make it easier for our devs.
In R, you can do this:
> sex = c(0, 1, 1, 0, 0, 0, 1) > age = c(23, 29, 30, 12, 15, 18) > age[sex == 0] [1] 23 12 15 18 > age[sex == 1] [1] 29 30 NA > age = c(23, 29, 30, 22, 25, 18, 31) > sex = c(0, 1, 1 , 0, 0 , 0, 1) > sex == 0 [1] TRUE FALSE FALSE TRUE TRUE TRUE FALSE > age[sex == 0] [1] 23 22 25 18
As you can see, it's useful for picking subgroups out of parallel sequences. But it relies on really odd behavior where if you give R a list of boolean arguments, it'll include the ones that are
TRUEin the result. This doesn't make much sense, asTRUEevaluates to 1 andFALSEevaluates to 0, but you get very different results if you replace the booleans with integers here.> age[c(1, 0, 0, 1, 1, 1, 0)] [1] 23 23 23 23
This is a very R result in that it seems to just ignore the out of range
0inputs!Oh, I see. I think this logical indexing is available in Numpy and Matlab too. In fact, I think logical indexing can be faster in Numpy, besides the fact that
0is not out of bounds ;)I suppose this raises the idea of a logical data type in Stan, but I think this type of indexing might be one of the only use cases ...
logical indexing can be faster in Numpy
It's always going to be bound by having to evaluate the condition for every element of the container. You could potentially do it without constructing the intermediate
sex == 0array by evaluating it lazily with an expression template, which would be more efficient.I suppose this raises the idea of a logical data type in Stan
We've already committed to the C-style coding of 0 for false and everything else being true, so if we did go down the boolean route, we'd at least need the type to be promotable to
int.As a bit of background for this:
Reacted by Jesse Knight
Vectorize all functions for matching shapes and allow broadcasting of lower-dimensional types.
This will require
There is an example for the function
inv()instan/math/prim/mat/fun/inv.hppin the branchfeature/issue-202-vectorize-all.This issue involves the first step above for all of the Stan functions, so it may need to be broken down into stages, starting with unary functions.
For the stan-dev/stan issue: stan-dev/stan#1683
For each of thee functions, we need (1) functor, (2) general template definition, (3) doc in manual, (4) signatures [need function for this], (5) signature tests for all instantiations.
When I did the example for
inv(), I had to remove the templating forinv()itself and replace it withdoubleto remove the ambiguity. For functions that only have a template version (i.e., there is not also a version forfwdandrev), I believe there will need to be anenable_ifon the base definition that restricts application to the base class.We'll also need a testing framework. This should be easy to do following the way
apply_scalar_unaryis defined (acrossprim,rev, andfwd). It will require a simple templated testing functor for each function being tested.(int):int
absint_stepoperator!operator+operator-(int,int):int
maxminoperator*operator+operator-operator/operator<operator<=operator>operator>=operator==operator&&operator||(int,real):real
bessel_first_kindbessel_second_kindmodified_bessel_first_kindmodified_bessel_second_kind(real):int
int_stepis_infis_nanoperator!(real):real
absacosacoshasinasinhatanatanhcbrtceilcoscoshdigammaerferfcexpexp2expm1fabsfloorinvinv_clogloginv_logitinv_Phiinv_sqrtinv_squarelgammaloglog10log1mlog1m_explog1m_inv_logitlog1plog1p_explog2log_inv_logitlogitPhiPhi_approxroundsinsinhsqrtsquaresteptantanhtgammatrigammatrunc(real,real):int
operator<operator<=operator>operator>=operator==operator&&operator||(real,real):real
atan2binomial_coefficient_logfalling_factorialfdimfmaxfminfmodgamma_pgamma_qhypotlbetalog_diff_explog_falling_factoriallog_rising_factoriallog_sum_expmaxminmultiply_logoperator*operator+operator-operator/operator^owens_tpowrising_factorial(int, real):real
binary_log_losslmgamma(real,real,real):real
fmalog_mix(int,real,real):real
if_else