Repository navigation
implement guniqueN #1120
Description
Activity
Came looking for this. I run into this issue a lot - my recent case being unbearably slow. My case looks more like this
dt <- data.table( A=sample(100000, 1000000, replace=TRUE), B=sample(100000, 1000000, replace=TRUE), C=sample(1000000, 1000000, replace=TRUE) ) # slow system.time(result1 <- dt[, list(UniqueCs=uniqueN(C)), keyby=list(A, B)]) # user system elapsed # 12.132 0.038 12.178 # fast system.time(result2 <- dt[, list(1), keyby=list(A, B, C)][, list(UniqueCs=.N), keyby=list(A, B)]) # user system elapsed # 0.374 0.013 0.387I'd think
uniqueNshould take about as long as aggregating with its argument.Confirming timings of @ben519...
Ran on 1.9.6:
system.time(result <- dt[, list(UniqueCs=uniqueN(C)), keyby=list(A, B)]) # user system elapsed # 8.032 0.004 8.029 system.time(result <- dt[, list(UniqueCs=.N), keyby=list(A, B, C)]) # user system elapsed # 0.496 0.004 0.498Ran on 1.9.7:
system.time(result <- dt[, list(UniqueCs=uniqueN(C)), keyby=list(A, B)]) # user system elapsed # 11.764 0.488 9.706 system.time(result <- dt[, list(UniqueCs=.N), keyby=list(A, B, C)]) # user system elapsed # 0.100 0.008 0.109(I missed his edit, but the difference is marginal)
Update if improved:
Ideal case where
uniqueNis much faster than alternatives (list i.e. non-scalar input here)When used with with an
byargument and many groupsuniqueN()slowness is very bad:irisdt <- setDT(iris[sample(1:150, size = 10000, replace = TRUE), ]) irisdt[, Sepal.Width := Sepal.Width+ sample(0:50, size = 10000, replace = TRUE)] irisdt[, Sepal.Length:= Sepal.Width+ sample(0:5000, size = 10000, replace = TRUE)] microbenchmark::microbenchmark( irisdt[, uniqueN(Sepal.Width), Sepal.Length], irisdt[, length(unique(Sepal.Width)), Sepal.Length], times = 2 ) Unit: milliseconds expr min lq mean median uq max neval cld irisdt[, uniqueN(Sepal.Width), Sepal.Length] 3592.42280 3592.42280 3592.65173 3592.65173 3592.88065 3592.88065 2 b irisdt[, length(unique(Sepal.Width)), Sepal.Length] 73.84953 73.84953 79.74312 79.74312 85.63672 85.63672 2 aUsing
uniqueis the best way to goirisdt <- setDT(iris[sample(1:150, size = 10000, replace = TRUE), ]) irisdt[, Sepal.Width := Sepal.Width+ sample(0:50, size = 10000, replace = TRUE)] irisdt[, Sepal.Length:= Sepal.Width+ sample(0:5000, size = 10000, replace = TRUE)] microbenchmark::microbenchmark( irisdt[, uniqueN(Sepal.Width), Sepal.Length], irisdt[, length(unique(Sepal.Width)), Sepal.Length], unique(irisdt, by = c('Sepal.Length', 'Sepal.Width'))[ , .N, by = Sepal.Length], times = 100 ) # Unit: milliseconds # expr min lq # irisdt[, uniqueN(Sepal.Width), Sepal.Length] 235.857762 284.023470 # irisdt[, length(unique(Sepal.Width)), Sepal.Length] 56.797016 70.049278 # unique(irisdt, by = c("Sepal.Length", "Sepal.Width"))[, .N, by = Sepal.Length] 4.076486 4.652738 # mean median uq max neval # 370.566000 328.691682 392.016670 968.17539 100 # 73.354643 72.797490 74.845989 130.11590 100 # 5.489569 4.801915 5.080524 55.50387 100Or
setkey(irisdt, Sepal.Width, Sepal.Length) ; irisdt[, .N, by = .(Sepal.Width, Sepal.Length)][ , .N, by = Sepal.Length]which will be faster thanuniqueby about 30% and about X8 faster thanlength(unique())But this seem irrelevant to the fact that
uniqueNis about X70 slower thanlength(unique())Reacted by Sindri and chinsoon12Not sure about github etiquette... should I reply? Anyway, I just wanted to point out that uniqueN() performs particularly bad in this setting which is ok but one has come to expect anything data.table to outperform anything in almost any setting. So maybe there is an issue here? My actual application is kind of different but I'm doing fine using
uniqueN2 <- function(x) length(unique(x))which also does much better thandplyr::n_distinct().- you're absolutely right that there's a problem with uniqueN & thanks for the reproducible benchmark! I just wanted to suggest valid alternatives in the meantime…On Thu, Feb 14, 2019, 10:53 PM Sindri ***@***.*** wrote: Not sure about github etiquette... should I reply? Anyway, I just wanted to point out that uniqueN() performs particularly bad in this setting which is ok but one has come to expect anything data.table to outperform anything in almost any setting. So maybe there is an issue here? My actual application is kind of different but I'm doing fine using uniqueN2 <- function(x) length(unique(x)) which also does much better than dplyr::n_distinct(). — You are receiving this because you commented. Reply to this email directly, view it on GitHub <#1120 (comment)>, or mute the thread <https://github.com/notifications/unsubscribe-auth/AHQQdX1lqkCantKRBGeOW0w7rvE82YQbks5vNXhMgaJpZM4ECSul> .Reacted by Sindri and Jan Gorecki
Related #3395, #3438
Root of this problem is thatuniqueNis called for every group.uniqueNcallsforderwhich is multithreaded, thus for every group own group of omp threads has to be formed. This will be resolved when implementingguniqueNfunction.
Additionally what we could do is to force calls injwhich are notgfunto be single threaded (at least ours by locally setting DTthreads to 1) @mattdowle. That would "resolve" this and similar problems. Still might eventually result in slower performance if there are very few big groups.Reacted by Michael Chirico and Matt Dowle- changed the title
[-]uniqueN slower than length(unique())[/-][+]implement guniqueN[/+]on Mar 15, 2019 - addedGForceissues relating to optimized grouping calculations (GForce)issues relating to optimized grouping calculations (GForce)
on Mar 15, 2019 another case where setting threads to 1 would probably help is new fifelse function: 93cc9ab
- added a commit that references this issue
on Aug 3, 2019 - addedtop requestOne of our most-requested issuesOne of our most-requested issues
on Apr 14, 2024
Most recent data.table. Not always, but quite often...
Related SO: http://stackoverflow.com/a/29684533/2490497