Repository navigation
Windows 10 is faster with -fno-openmp than setDTthreads(1) on many repeated calls #4527
Description
Activity
ColeMiller1 commented
on Jun 6, 2020 on Jun 6, 2020 · Hidden as outdatedAuthorshow commentMore actionsColeMiller1 commented
on Jun 6, 2020 on Jun 6, 2020 · Hidden as outdatedAuthorshow commentMore actionsColeMiller1 commented
on Jun 6, 2020 on Jun 6, 2020 · Hidden as outdatedAuthorshow commentMore actionsColeMiller1 commented
on Jun 7, 2020 on Jun 7, 2020 · Hidden as outdatedAuthorshow commentMore actionsColeMiller1 commented
on Jun 10, 2020 on Jun 10, 2020 · Hidden as outdatedAuthorshow commentMore actionsStill working on this - tried the throttle threads PR but it still did not work.
What is an example where OpenMP should shine? Using 1.12.8 and
x = sample(1e7L),order(x)is still 3x faster on 1 thread thanforder(x). Note, using multiple threads helps but it still does not catch up toorder(x).library(data.table) x = sample(1e7L) setDTthreads(1L) bench::mark(data.table:::forder(x), order(x), min_iterations = 10L) ### A tibble: 2 x 13 ## expression min median `itr/sec` mem_alloc `gc/sec` n_itr n_gc ## <bch:expr> <bch> <bch:> <dbl> <bch:byt> <dbl> <int> <dbl> ##1 data.table:::forder(x) 923ms 929ms 1.07 38.1MB 0.459 7 3 ##2 order(x) 354ms 355ms 2.79 38.1MB 1.19 7 3 setDTthreads(2L) ## expression min median `itr/sec` mem_alloc `gc/sec` n_itr n_gc ## <bch:expr> <bch> <bch:> <dbl> <bch:byt> <dbl> <int> <dbl> ##1 data.table:::forder(x) 614ms 657ms 1.50 38.1MB 0.376 8 2 ##2 order(x) 359ms 368ms 2.73 38.1MB 1.17 7 3
12 remaining items
Great info @ColeMiller1, many thanks for investigating here. Seems like the new throttle works on Linux but not as well on WIndows. So that's something that needs further work then, agree.
Can I clear up a few sentences on this one.
OpenMP support may cause issues with Windows performance, especially with many repeated calls.
Why the especially? This issue is just about repeated calls, and repeated calls on small data, right?
What is an example where OpenMP should shine?
Any of the examples in https://h2oai.github.io/db-benchmark/. They are single calls on large data that take more than a few seconds, and in some cases minutes, to run.
In generally,
[.data.tabledoesn't like being repeated (it has overhead). That's why we haveset()low-overhead alternative and why we use[[and$with data.table for ordinary operations. We'd like to one day to get to a state where the user doesn't need to know (for example by reworking[.data.table) but that hasn't happened yet.And in general data.table doesn't like doing things one row at a time, which is true in R is general, and Python too. Always work to do to make it easier for user but it's just not recommended practice in high level languages like R.
So yet we do want to work on one-row-at-a-time benchmarks like this one, but I wonder how high it should be on the list.
I agree that single row subsetting itself is not a large concern - the original example is related to #3735. But that's where I found another Windows user had slower subsets than Linux users for many repeated calls.
Why should we care about this? Operations by group can be affected, especially if when we subset by group (e.g.
dt[, .SD[1:2], grp]. For Windows users, there can be negative impacts even withsetDTthreads(1L).I hope my choice of words doesn't make us overlook that there are opportunities to make
data.tablethe best tool for both large data and small data regardless of platform. In working towards #852, I have a branch that takes down 14s to 5.3s formatrix(0L, nrow =1e5L, ncol= 10L)without OpenMP. With OpenMP, it goes from 24.5s to 13.6 s. So, 80us at a time can add up with optimizations.FWIW - this may be centered around where the OpenMP loop is. While it is extremely buggy, putting the
pragma omp parallel for ...in thesubsetDTinstead of thesubsetVectorRawchanges the time on this branch to 8.3s. The main difference being that whereas Linux has no overhead, Windows seems to have overhead anytime the process is forked and unforked. AssubsetDTis currently written, every column will get a new set of teams to go out and subset the vector.Old but related topic: http://forum.openmp.org/forum/viewtopic.php?f=3&t=1722
Also few hits on SO mentions visual studio, so looks to be windows specific.
https://stackoverflow.com/questions/16635814/reasons-for-omp-set-num-threads1-slower-than-no-openmp
https://stackoverflow.com/questions/16807766/may-compiler-optimizations-be-inhibited-by-multi-threadingJust hit a work issue where it was pretty natural to do rowwise
[:The basic idea is I need to collapse some columns into a metadata column, e.g.
NN = 1e5 DT = data.table(ID = 1:NN, V1 = rnorm(NN), V2=rnorm(NN))I tried:
DT[ , metadata := lapply(seq_len(.N), function(ii) .SD[ii]), .SDcols = c('V1', 'V2')]And it's indeed rather slow.
transposeessentially works in this case becauseV1andV2are the same type, but generally they could be different types (-> coercion) and include non-transpose-compatible types (e.g.list)@MichaelChirico thanks for the example. Using
bywould help speed this up some (maybe a dedicatednest()feature would be better):library(data.table) ##1.12.8 setDTthreads(1L) NN = 1e4 DT = data.table(ID = 1:NN, V1 = rnorm(NN), V2=rnorm(NN)) system.time(DT[ , metadata := lapply(seq_len(.N), function(ii) .SD[ii]), .SDcols = c('V1', 'V2')]) #> user system elapsed #> 1.86 0.12 1.98 system.time(DT[ , metadata2 := .(.(copy(.SD))), by = ID, .SDcols = c('V1', 'V2')]) #> user system elapsed #> 0.26 0.02 0.28 all.equal(DT$metadata, DT$metadata2) #> [1] TRUE
I had to use
copy()because otherwise the result has thelockedattribute.The
copy()part should be fixed in dev right?@MichaelChirico I just downloaded master - does not appear to be fixed. TBH, I'm not aware of an issue related to nesting
.SDby group:DT[, metadata3 := .(.(.SD)), by = ID, .SDcols = c("V1", "V2")] attributes(DT$metadata3[[1L]]) ##$row.names ##[1] 1 ##$class ##[1] "data.table" "data.frame" ##$.internal.selfref ##<pointer: 0x05772498> ##$names ##[1] "V1" "V2" ##$.data.table.locked ##[1] TRUE all.equal(DT$metadata2[[1L]], DT$metadata3[[1L]]) ##[1] "Datasets has different number of (non-excluded) attributes: target 2, current 3"
This can be observed on Linux as well:
this branch removes overhead: https://github.com/Rdatatable/data.table/compare/dogroups-overhead?expand=1
while this branch still has the overhead: https://github.com/Rdatatable/data.table/compare/dogroups-overhead2?expand=1@ColeMiller1 could confirm if #4558 resolves this issue?
@jangorecki sorry for the delay. Yes! (I even like that I do not have to add
setDTthreads(1L)at the top!)library(data.table) mat = matrix(0L, nrow =1e5L, ncol= 10L) DT = as.data.table(mat) DF = as.data.frame(mat) system.time(for (i in seq_len(nrow(DT))) {DT[i]}) #> user system elapsed #> 13.42 0.43 14.15 system.time(for (i in seq_len(nrow(DF))) {DF[i,]}) #> user system elapsed #> 12.66 0.03 12.97
The
system.timeis still a little higher - it would be great if we could add an argument tosubsetDTabout whether the indices need to be checked. From[.data.table, we already checked once so there's no need to convert zeros and negative IDs again.But, that's small overall and could be addressed by a follow-up issue. The subsetting performance was the only place I noticed the OpenMP issue. Please feel free to close this when the PR is merged.
Thanks for all of your help!
- linked a pull request that will close this issuetowards #4200, overhead in dogroups #4558
on Jun 21, 2020
OpenMP support may cause issues with Windows performance, especially with many repeated calls.
Note, I would expect using
setDTthreads(1L)would minimize any impacts to performance but that does not appear to be the case on Windows 10.