I'm frequently presented with data.tables which have (a very small percentage of) duplicated keys which causes some trouble when I use them in j to merge.
In my application, it makes sense to just drop any of those observations because they can't reliably be distinguished and they're so infrequent;
unique seems perfectly suited to this end, except that it retains at least one of the observations in any duplicated group, while I prefer to cut them out completely because it's essential not to mess up the merge process, the results of which are crucial for the whole project.
There's a workaround which is quite verbose; let's use the sample in ?duplicated.data.table:
DT <- data.table(A = rep(1:3, each=4), B = rep(1:4, each=3), C = rep(1:2, 6), key = "A,B")
> unique(DT)
A B C
1: 1 1 1
2: 1 2 2
3: 2 2 1
4: 2 3 1
5: 3 3 1
6: 3 4 2
The troublesome observations are 1,3,4,6.
From what I can tell my only recourse at the moment is something rather elaborate like
> DT[.(DT[,.N,by=key(DT)][N==1,!"N",with=F])]
A B C
1: 1 2 2
2: 3 3 1
Something like unique(DT,only.unique=T) that achieves the same end seems like it would be easy to implement.
---EDIT 2015 May 27---
Arun's suggested workaround is much better than my approach, but comparison with unique suggests there's still considerable speed being lost:
> microbenchmark(times=1000L,
+ arun(),mike(),unique(DT))
Unit: microseconds
expr min lq mean median uq max neval cld
arun() 775.852 818.4715 950.2565 840.2605 865.327 47848.84 1000 b
mike() 2269.876 2346.1640 2953.5697 2413.6900 2478.700 50269.23 1000 c
unique(DT) 199.339 225.0725 289.1449 239.5555 253.971 46924.54 1000 a
Performance of the workaround also deteriorates substantially when .SD is large, while that of unique is barely affected:
DT[,paste0("V",1:100):=lapply(1:100,function(x)sample(.N))]
> microbenchmark(times=1000L,
+ arun(),unique(DT))
Unit: microseconds
expr min lq mean median uq max neval cld
arun() 3397.032 3517.9175 4686.5543 3631.659 3728.6725 56132.181 1000 b
unique(DT) 212.203 234.8935 256.9668 248.669 267.7265 623.812 1000 a
I'm frequently presented with
data.tables which have (a very small percentage of) duplicated keys which causes some trouble when I use them injto merge.In my application, it makes sense to just drop any of those observations because they can't reliably be distinguished and they're so infrequent;
uniqueseems perfectly suited to this end, except that it retains at least one of the observations in any duplicated group, while I prefer to cut them out completely because it's essential not to mess up the merge process, the results of which are crucial for the whole project.There's a workaround which is quite verbose; let's use the sample in
?duplicated.data.table:DT <- data.table(A = rep(1:3, each=4), B = rep(1:4, each=3), C = rep(1:2, 6), key = "A,B")The troublesome observations are 1,3,4,6.
From what I can tell my only recourse at the moment is something rather elaborate like
Something like
unique(DT,only.unique=T)that achieves the same end seems like it would be easy to implement.---EDIT 2015 May 27---
Arun's suggested workaround is much better than my approach, but comparison with
uniquesuggests there's still considerable speed being lost:Performance of the workaround also deteriorates substantially when
.SDis large, while that ofuniqueis barely affected: