Repository navigation
groupby with dogroups (R expression) performance regression #4200
Copy link
Copy link
Closed
Milestone
Description
Activity
library(data.table) N = 1e6L set.seed(108) d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values v1 = sample(5L, N, TRUE), v2 = sample(5L, N, TRUE)) system.time(d[, max(v1)-min(v2), by=id3]) system.time(d[, max(v1)-min(v2), by=id3]) ### using just 2 threads ## 1.12.9 # user system elapsed # 3.630 0.452 4.122 # user system elapsed # 3.681 0.354 4.098 ## 1.12.8 # user system elapsed # 2.056 0.020 2.102 # user system elapsed # 1.989 0.024 2.042
This regression is especially visible on
G1_1e8_2e0_0_0dataset. data.table Q7 and Q8 on that data case is now slower than other tools.new throttle feature helped in that issue, but not yet resolves. Timings on 20/40 threads.
library(data.table) N = 1e6L set.seed(108) d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values v1 = sample(5L, N, TRUE), v2 = sample(5L, N, TRUE)) system.time(d[, max(v1)-min(v2), by=id3]) system.time(d[, max(v1)-min(v2), by=id3])
1.12.8
user system elapsed 1.463 0.031 0.964 user system elapsed 1.396 0.036 0.9531.12.9 before throttle 12586af
user system elapsed 276.056 1.332 13.914 user system elapsed 274.565 1.319 13.8321.12.9 after throttle
user system elapsed 2.390 0.444 2.342 user system elapsed 2.433 0.418 2.338PR #4558
user system elapsed 1.632 0.013 1.203 user system elapsed 1.472 0.107 1.134Regression was introduced by 4aadde8
how to find it, beware usinggit reset --hardif you have uncommited changes!git reset --hard 12586afdc9fe68303cfa36b2182d75a9164bee32 ## any before throttle will do cat dogroups_4200.R #cc(FALSE) #N = 1e6L #set.seed(108) #d = data.table(id3 = sample(c(seq.int(N*0.9), sample(N*0.9, N*0.1, TRUE))), # 9e5 unq values # v1 = sample(5L, N, TRUE), # v2 = sample(5L, N, TRUE)) #nul = system.time(d[, max(v1)-min(v2), by=id3]) #t = system.time(d[, max(v1)-min(v2), by=id3]) #stopifnot(t[[3]] < 1.6) #cat("done\n") git bisect start master 1.12.8 git bisect run Rscript dogroups_4200.R
thanks to @MichaelChirico for teaching how to use bisect
- added a commit that references this issue
on Jun 18, 2020 - added a commit that references this issue
on Jun 26, 2020
There is a performance regression (AFAIU) when doing by group computation where we run R's C eval by each group (q7 and q8 in db-benchmark).
worth to note that at the same time q9
x[, .(r2=cor(v1, v2)^2), by=.(id2, id4)]got nice speed up