Test 9999.024 fails when running with 2+ threads.
After testing I can confirm that exact=T might result in rounding error on windows with 2+ threads, and that error can be actually larger than in case of exact=F. Which is strange, when looking at the core of that computation there is not much places where it can lose precision, maybe (double) long double cast? Iterations perfectly isolated, sharing only truehasna variable, which is always false for those tests as there are no NAs. Although the input is tricky having far outliers to stress rounding correction: sample(c(rnorm(1e3, 1e6, 5e5), 5e9, 5e-9)). If that would be the case then it should be also problem for single threaded process, unless windows openmp lib is causing problem.
const int threads = MIN(getDTthreads(), nx);
#pragma omp parallel num_threads(threads) shared(truehasna)
{
#pragma omp for schedule(static)
for (uint_fast64_t i=k-1; i<nx; i++) { // loop on every observation with complete window, partial already filled in single threaded section
if (narm && truehasna) continue; // if NAs detected no point to continue
long double w = 0.0;
for (int j=-k+1; j<=0; j++) { // sub-loop on window width
w += x[i+j]; // sum of window for particular observation
}
if (R_FINITE((double) w)) { // no need to calc roundoff correction if NAs detected as will re-call all below in truehasna==1
long double res = w / k; // keep results as long double for intermediate processing
long double err = 0.0; // roundoff corrector
for (int j=-k+1; j<=0; j++) { // nested loop on window width
err += x[i+j] - res; // measure difference of obs in sub-loop to calculated fun for obs
}
ans->ans[i] = (double) (res + (err / k)); // adjust calculated rollfun with roundoff correction
} else {
if (!narm) ans->ans[i] = (double) (w / k); // NAs should be propagated
truehasna = 1; // NAs detected for this window, set flag so rest of windows will not be re-run
}
}
} // end of parallel region
Test 9999.024 fails when running with 2+ threads.
After testing I can confirm that exact=T might result in rounding error on windows with 2+ threads, and that error can be actually larger than in case of exact=F. Which is strange, when looking at the core of that computation there is not much places where it can lose precision, maybe
(double) long doublecast? Iterations perfectly isolated, sharing onlytruehasnavariable, which is alwaysfalsefor those tests as there are no NAs. Although the input is tricky having far outliers to stress rounding correction:sample(c(rnorm(1e3, 1e6, 5e5), 5e9, 5e-9)). If that would be the case then it should be also problem for single threaded process, unless windows openmp lib is causing problem.