Skip to content

Windows 10 is faster with -fno-openmp than setDTthreads(1) on many repeated calls #4527

Description

@ColeMiller1

OpenMP support may cause issues with Windows performance, especially with many repeated calls.

library(data.table)
setDTthreads(1L)
allIterations <- data.table(v1 = runif(1e5), v2 = runif(1e5))
DoSomething <- function(row) someCalculation <- row[["v1"]] + 1

system.time(for (r in 1:nrow(allIterations)) DoSomething(allIterations[r, ]))
##1.12.8 
##   user  system elapsed 
##  17.09    4.48   22.14 

##master
##   user  system elapsed 
##  17.50    4.53   22.51 

##master w/ OpenMP lines deleted from subset.c
##   user  system elapsed 
##  12.37    0.01   12.85 

Note, I would expect using setDTthreads(1L) would minimize any impacts to performance but that does not appear to be the case on Windows 10.

Activity

  1. jangorecki commented on Jun 5, 2020

    @jangorecki
  2. ColeMiller1 commented on Jun 6, 2020

    @ColeMiller1
    Author
  3. jangorecki commented on Jun 6, 2020

    @jangorecki
  4. ColeMiller1 commented on Jun 6, 2020

    @ColeMiller1
    Author
  5. jangorecki commented on Jun 6, 2020

    @jangorecki
  6. jangorecki commented on Jun 6, 2020

    @jangorecki
  7. ColeMiller1 commented on Jun 6, 2020

    @ColeMiller1
    Author
  8. ColeMiller1 commented on Jun 7, 2020

    @ColeMiller1
    Author
  9. jangorecki commented on Jun 7, 2020

    @jangorecki
  10. ColeMiller1 commented on Jun 10, 2020

    @ColeMiller1
    Author
  11. jangorecki commented on Jun 10, 2020

    @jangorecki
  12. ColeMiller1 commented on Jun 13, 2020

    @ColeMiller1
    ContributorAuthor

    Still working on this - tried the throttle threads PR but it still did not work.

    What is an example where OpenMP should shine? Using 1.12.8 and x = sample(1e7L), order(x) is still 3x faster on 1 thread than forder(x). Note, using multiple threads helps but it still does not catch up to order(x).

    library(data.table)
    
    x = sample(1e7L)
    
    setDTthreads(1L)
    
     bench::mark(data.table:::forder(x),
     order(x),
     min_iterations = 10L)
    
    ### A tibble: 2 x 13
    ##  expression               min median `itr/sec` mem_alloc `gc/sec` n_itr  n_gc
     ## <bch:expr>             <bch> <bch:>     <dbl> <bch:byt>    <dbl> <int> <dbl>
    ##1 data.table:::forder(x) 923ms  929ms      1.07    38.1MB    0.459     7     3
    ##2 order(x)               354ms  355ms      2.79    38.1MB    1.19      7     3
    
    setDTthreads(2L)
    
    ##  expression               min median `itr/sec` mem_alloc `gc/sec` n_itr  n_gc
    ##  <bch:expr>             <bch> <bch:>     <dbl> <bch:byt>    <dbl> <int> <dbl>
    ##1 data.table:::forder(x) 614ms  657ms      1.50    38.1MB    0.376     8     2
    ##2 order(x)               359ms  368ms      2.73    38.1MB    1.17      7     3
  13. 12 remaining items

  14. mattdowle commented on Jun 15, 2020

    @mattdowle
    Member

    Great info @ColeMiller1, many thanks for investigating here. Seems like the new throttle works on Linux but not as well on WIndows. So that's something that needs further work then, agree.

    Can I clear up a few sentences on this one.

    OpenMP support may cause issues with Windows performance, especially with many repeated calls.

    Why the especially? This issue is just about repeated calls, and repeated calls on small data, right?

    What is an example where OpenMP should shine?

    Any of the examples in https://h2oai.github.io/db-benchmark/. They are single calls on large data that take more than a few seconds, and in some cases minutes, to run.

    In generally, [.data.table doesn't like being repeated (it has overhead). That's why we have set() low-overhead alternative and why we use [[ and $ with data.table for ordinary operations. We'd like to one day to get to a state where the user doesn't need to know (for example by reworking [.data.table) but that hasn't happened yet.

    And in general data.table doesn't like doing things one row at a time, which is true in R is general, and Python too. Always work to do to make it easier for user but it's just not recommended practice in high level languages like R.

    So yet we do want to work on one-row-at-a-time benchmarks like this one, but I wonder how high it should be on the list.

  15. ColeMiller1 commented on Jun 16, 2020

    @ColeMiller1
    ContributorAuthor

    I agree that single row subsetting itself is not a large concern - the original example is related to #3735. But that's where I found another Windows user had slower subsets than Linux users for many repeated calls.

    Why should we care about this? Operations by group can be affected, especially if when we subset by group (e.g. dt[, .SD[1:2], grp]. For Windows users, there can be negative impacts even with setDTthreads(1L).

    I hope my choice of words doesn't make us overlook that there are opportunities to make data.table the best tool for both large data and small data regardless of platform. In working towards #852, I have a branch that takes down 14s to 5.3s for matrix(0L, nrow =1e5L, ncol= 10L) without OpenMP. With OpenMP, it goes from 24.5s to 13.6 s. So, 80us at a time can add up with optimizations.

    FWIW - this may be centered around where the OpenMP loop is. While it is extremely buggy, putting the pragma omp parallel for ... in the subsetDT instead of the subsetVectorRaw changes the time on this branch to 8.3s. The main difference being that whereas Linux has no overhead, Windows seems to have overhead anytime the process is forked and unforked. As subsetDT is currently written, every column will get a new set of teams to go out and subset the vector.

  16. MichaelChirico commented on Jun 18, 2020

    @MichaelChirico
    Member

    Just hit a work issue where it was pretty natural to do rowwise [:

    The basic idea is I need to collapse some columns into a metadata column, e.g.

    NN = 1e5
    DT = data.table(ID = 1:NN, V1 = rnorm(NN), V2=rnorm(NN))
    

    I tried:

    DT[ , metadata := lapply(seq_len(.N), function(ii) .SD[ii]), .SDcols = c('V1', 'V2')]
    

    And it's indeed rather slow.

    transpose essentially works in this case because V1 and V2 are the same type, but generally they could be different types (-> coercion) and include non-transpose-compatible types (e.g. list)

  17. added this to the 1.12.11 milestone on Jun 18, 2020
  18. ColeMiller1 commented on Jun 18, 2020

    @ColeMiller1
    ContributorAuthor

    @MichaelChirico thanks for the example. Using by would help speed this up some (maybe a dedicated nest() feature would be better):

    library(data.table) ##1.12.8
    setDTthreads(1L)
    NN = 1e4
    DT = data.table(ID = 1:NN, V1 = rnorm(NN), V2=rnorm(NN))
    
    system.time(DT[ , metadata := lapply(seq_len(.N), function(ii) .SD[ii]), .SDcols = c('V1', 'V2')])
    #>    user  system elapsed 
    #>    1.86    0.12    1.98
    system.time(DT[ , metadata2 := .(.(copy(.SD))), by = ID, .SDcols = c('V1', 'V2')])
    #>    user  system elapsed 
    #>    0.26    0.02    0.28
    all.equal(DT$metadata, DT$metadata2)
    #> [1] TRUE

    I had to use copy() because otherwise the result has the locked attribute.

  19. MichaelChirico commented on Jun 18, 2020

    @MichaelChirico
    Member

    The copy() part should be fixed in dev right?

  20. ColeMiller1 commented on Jun 18, 2020

    @ColeMiller1
    ContributorAuthor

    @MichaelChirico I just downloaded master - does not appear to be fixed. TBH, I'm not aware of an issue related to nesting .SD by group:

    DT[, metadata3 := .(.(.SD)), by = ID, .SDcols = c("V1", "V2")]
    
    attributes(DT$metadata3[[1L]])
    ##$row.names
    ##[1] 1
    
    ##$class
    ##[1] "data.table" "data.frame"
    
    ##$.internal.selfref
    ##<pointer: 0x05772498>
    
    ##$names
    ##[1] "V1" "V2"
    
    ##$.data.table.locked
    ##[1] TRUE
    
    all.equal(DT$metadata2[[1L]], DT$metadata3[[1L]])
    ##[1] "Datasets has different number of (non-excluded) attributes: target 2, current 3"
  21. jangorecki commented on Jun 18, 2020

    @jangorecki
    Member

    This can be observed on Linux as well:
    this branch removes overhead: https://github.com/Rdatatable/data.table/compare/dogroups-overhead?expand=1
    while this branch still has the overhead: https://github.com/Rdatatable/data.table/compare/dogroups-overhead2?expand=1

  22. jangorecki commented on Jun 18, 2020

    @jangorecki
    Member

    @ColeMiller1 could confirm if #4558 resolves this issue?

  23. ColeMiller1 commented on Jun 19, 2020

    @ColeMiller1
    ContributorAuthor

    @jangorecki sorry for the delay. Yes! (I even like that I do not have to add setDTthreads(1L) at the top!)

    library(data.table)
    mat = matrix(0L, nrow =1e5L, ncol= 10L)
    DT = as.data.table(mat)
    DF = as.data.frame(mat)
    
    system.time(for (i in seq_len(nrow(DT))) {DT[i]})
    #>    user  system elapsed 
    #>   13.42    0.43   14.15
    system.time(for (i in seq_len(nrow(DF))) {DF[i,]})
    #>    user  system elapsed 
    #>   12.66    0.03   12.97

    The system.time is still a little higher - it would be great if we could add an argument to subsetDT about whether the indices need to be checked. From [.data.table, we already checked once so there's no need to convert zeros and negative IDs again.

    But, that's small overall and could be addressed by a follow-up issue. The subsetting performance was the only place I noticed the OpenMP issue. Please feel free to close this when the PR is merged.

    Thanks for all of your help!

  24. linked a pull request that will close this issuetowards #4200, overhead in dogroups #4558on Jun 21, 2020
  25. modified the milestones: 1.12.11, 1.12.9 on Jul 1, 2020
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions