Skip to content

groupby with dogroups (R expression) performance regression #4200

Description

@jangorecki

There is a performance regression (AFAIU) when doing by group computation where we run R's C eval by each group (q7 and q8 in db-benchmark).

    in_rows question_group                    question 20181206_da98fb2 20190913_35b0de3 20191115_92abb70 20191205_eba8704 20191209_6808d2c 20191212_e0140ea 20191229_d52b0d8 20200124_c005296
 1:     1e9          basic               sum v1 by id1           21.144           10.026            8.375            9.362            9.060            9.082            9.366            9.271
 2:     1e9          basic           sum v1 by id1:id2           38.914           11.746            9.243            9.327            9.331            9.978           10.813            9.220
 3:     1e9          basic       sum v1 mean v3 by id3           99.517           14.487           12.044           12.291           13.496           14.325           13.191           13.169
 4:     1e9          basic           mean v1:v3 by id4           26.593           17.357           15.135           15.157           15.278           16.761           16.724           16.754
 5:     1e9          basic            sum v1:v3 by id6          122.214           14.569           13.454           14.035           14.046           14.842           14.400           15.019
 6:     1e9       advanced  median v3 sd v3 by id4 id5               NA               NA          121.742          110.925          106.340          113.837          111.984          123.411
 7:     1e9       advanced      max v1 - min v2 by id3               NA               NA           98.680           91.596           87.005           93.749           91.294          299.863
 8:     1e9       advanced       largest two v3 by id6               NA               NA          234.926          215.574          213.824          216.152          211.295          411.241
 9:     1e9       advanced regression v1 v2 by id2 id4               NA           72.466           81.297           76.121           75.311           74.769           72.571           41.157
10:     1e9       advanced     sum v3 count by id1:id6               NA          180.257          196.403          187.816          177.257          187.702          190.282          187.403

worth to note that at the same time q9 x[, .(r2=cor(v1, v2)^2), by=.(id2, id4)] got nice speed up

Activity

  1. jangorecki commented on Jan 27, 2020

    @jangorecki
    MemberAuthor
    library(data.table)
    N = 1e6L
    set.seed(108)
    d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values
                   v1 = sample(5L, N, TRUE),
                   v2 = sample(5L, N, TRUE))
    system.time(d[, max(v1)-min(v2), by=id3])
    system.time(d[, max(v1)-min(v2), by=id3])
    
    ### using just 2 threads
    
    ## 1.12.9
    
    #   user  system elapsed
    #  3.630   0.452   4.122
    
    #   user  system elapsed
    #  3.681   0.354   4.098
    
    ## 1.12.8
    
    #   user  system elapsed
    #  2.056   0.020   2.102
    
    #   user  system elapsed
    #  1.989   0.024   2.042
  2. added this to the 1.12.9 milestone on Jan 27, 2020
  3. ColeMiller1 commented on Feb 16, 2020

    @ColeMiller1
  4. ColeMiller1 commented on Mar 10, 2020

    @ColeMiller1
  5. jangorecki commented on May 14, 2020

    @jangorecki
    MemberAuthor

    This regression is especially visible on G1_1e8_2e0_0_0 dataset. data.table Q7 and Q8 on that data case is now slower than other tools.

  6. modified the milestones: 1.12.11, 1.12.9 on May 26, 2020
  7. jangorecki commented on Jun 18, 2020

    @jangorecki
    MemberAuthor

    new throttle feature helped in that issue, but not yet resolves. Timings on 20/40 threads.

    library(data.table)
    N = 1e6L
    set.seed(108)
    d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values
                   v1 = sample(5L, N, TRUE),
                   v2 = sample(5L, N, TRUE))
    system.time(d[, max(v1)-min(v2), by=id3])
    system.time(d[, max(v1)-min(v2), by=id3])

    1.12.8

       user  system elapsed 
      1.463   0.031   0.964 
       user  system elapsed 
      1.396   0.036   0.953 
    
    

    1.12.9 before throttle 12586af

       user  system elapsed 
    276.056   1.332  13.914 
       user  system elapsed 
    274.565   1.319  13.832 
    

    1.12.9 after throttle

       user  system elapsed 
      2.390   0.444   2.342 
       user  system elapsed 
      2.433   0.418   2.338 
    

    PR #4558

       user  system elapsed 
      1.632   0.013   1.203 
       user  system elapsed 
      1.472   0.107   1.134 
    
  8. jangorecki commented on Jun 18, 2020

    @jangorecki
    MemberAuthor

    Regression was introduced by 4aadde8
    how to find it, beware using git reset --hard if you have uncommited changes!

    git reset --hard 12586afdc9fe68303cfa36b2182d75a9164bee32 ## any before throttle will do
    cat dogroups_4200.R
    #cc(FALSE)
    #N = 1e6L
    #set.seed(108)
    #d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values
    #               v1 = sample(5L, N, TRUE),
    #               v2 = sample(5L, N, TRUE))
    #nul = system.time(d[, max(v1)-min(v2), by=id3])
    #t = system.time(d[, max(v1)-min(v2), by=id3])
    #stopifnot(t[[3]] < 1.6)
    #cat("done\n")
    git bisect start master 1.12.8
    git bisect run Rscript dogroups_4200.R

    thanks to @MichaelChirico for teaching how to use bisect

  9. added a commit that references this issue on Jun 18, 2020
  10. jangorecki commented on Jun 18, 2020

    @jangorecki
    MemberAuthor

    Timings on G1_1e8_1e1_0_0, q7 max v1 - min v2 by id3:
    before regression: d52b0d8: 30.552s
    regression: c005296: 257.262s (up to 306.816s)
    PR #4558: 15f0598: 35.002s

  11. added a commit that references this issue on Jun 26, 2020
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions