Repository navigation
%notin% is not safe in the way it handles NA by default #5481
Description
Activity
I don't see the point in your example.
What would be your supposed output for
dt[x %notin% 1:3, y := z][]?
And what would be the supposed output for
dt[x %in% 1:3, y := z][]?My point is that
dt[x %notin% 1:3, y := z][]modifies y if x is missing and I think that it should not since we do not know if the unkown real values of missing values in x are contained in1:3or not.I expected the following
dt[x %notin% 1:3, y := z][] x y z <int> <int> <num> 1: 1 1 30 2: 2 2 10 3: 3 3 40 4: NA 4 80 5: 4 30 30 6: NA 6 80The output below is what I expected for
%in%; and in practice, the output ofdt[!x %in% 1:3]is not what users likely typically want (IMHO) if x contains missing values.dt[x %in% 1:3, y := z][] # note that this does not modify y if x is missing x y z <int> <int> <num> 1: 1 30 30 2: 2 10 10 3: 3 40 40 4: NA 4 80 5: 4 5 30 6: NA 6 80I don't really know how other users typically use %in% in negation but my tipical use is
DT[!x %in% c(val1, val2, ..., NA), ...](where val1, val2, ... are not missing) and notDT[!x %in% c(val1, val2, ...,valn), ...], so that rows where x is missing are not included in the output.@Kamgang-B I don't think this is an issue at all.
In your case ofdt[x %in% 1:3, y := z]with the note of "note that this does not modify y if x is missing". I think that it is pretty clear from the syntax thatNAis not in1:3, therefore, no modification. The note is superfluous. If I wanted the behavior that requireNAto be modified, I would have writedt[x %in% c(1:3, NA), y := z].
Similarly,dt[x %notin% c(1:3, NA), y := z]would bring out your preferred output. Writingdt[x %notin% 1:3, y := z]would mean it is true thatNAis not in1:3and modification is required.
The functional syntax for%notin%is`%notin%`(x, table)with backticks (similar to`+`(2, 3)is for2 + 3or`%notin%`(dt$x, 1:3)is fordt$x %notin% 1:3)I think it is a good question.
Should be addressed before releasing this feature to CRAN. By addressing I mean, well documenting, and eventually if there is agreement, changing behavior.
Personally my use cases were!x%in%values, but maybe I simply didn't have NAs inx. Both situations can be surprise to users but I think that simply negating IN may be a less frequently a surprise.Reacted by Kyle HaynesMaybe I'm missing the meaning... dt[x %notin% 1:3 implies that x is not in 1,2,3 and that would include NA because NA is not 1, 2, or 3
It would be great if R could handle a mixed type vector or list though I'm not sure that it can, I think NA being skipped in this case makes perfect sense. data.table is especially useful whereby you can assign a value in
jcontingent oniwhereby you can say
dt[is.na(b)==TRUE,y:=z]as a second step here.I would discourage chaining it or building "NA handling" into the awesome %notin% because you can repeat the above any number of times in very legible manner, without needing to learn the internal mechanics of the method and its NA handling or type safety.
I agree with @datocrats-org that the behavior makes sense because just as
NAis not%in%the vector1:3,NAshould be%notin%the vector1:3. To make things clear for users I'd be happy to add an example including NA in the left hand side as it is in the test cases@hdn012 @datocrats-org @mczek
I understand when you say thatNAis not in1:3. In this sense, you are right and%notin%is the logical complement of%in%. But my point is different. I am talking more bout the user intention/expectation when using%notin%. When doingx %notin% 1:3, does the user really expect the result to be TRUE when x is missing!? I feel (and maybe only me) like it's more natural to haveTRUEifxis not in1:3AND is not missing (even if it would no longer be consistent with!x %in% 1:3).The functional syntax for %notin% is
%notin%(x, table) with backticks (similar to+(2, 3) is for 2 + 3 or%notin%(dt$x, 1:3) is for dt$x %notin% 1:3)I think that I was not clear enough when talking about the functional form. I know that
`%notin%`(x, table)is the functional form ofdt$x %notin% 1:3. I am talking about something in the spirit of%like%andlike(the latter is more flexible because it provides more arguments (vector, pattern, ignore.case = FALSE, fixed = FALSE, perl = FALSE) than the former one (vector, pattern)).
So, the idea is that a new functionnotincould provide an additional argument that allows users to have better control ofNAs. Something like:notin = function(x, table, na=FALSE) { # na=FALSE/TRUE-->set NA to FALSE/TRUE if(na && !anyNA(table)) table = c(table, NA) x %notin% table # where %notin% is the current implement of %notin% }Both situations can be surprise to users
That's why I think that it would be nice to complement
%notin%with a more flexible version likenotin(x, table, na=FALSE)(similar to%like%andlike).Kindly consider exporting a new more flexible function
notinas a feature request (if the current implementation of%notin%is kept unchanged).Reacted by Hieu NguyenI like the idea
Reacted by Hieu NguyenIf there is nobody willing to submit such extended function
notinnow then we can proceed with documentation improvement, ideally also including examples
IMHO, the typical use of the function
%notin%is likely expected to beDT[lhs %notin% rhs), ...]where 1- rhs contains no missing value and 2- the user wants to return/modify rows where lhs contains only values in rhs.Also, I don't expect users to do something like
!lhs %notin%(since%in%is already convenient for this operation).For these reasons, I think that it is better to be on the safe side by allowing
DT[lhs %notin% rhs,...], to return/modify only rows whose values are in rhs. In doing so, the user will have to explicitly add NA to the rhs if he also wants to include rows with missing values.Consider the following example:
In doing this operation, I don't really think users expect the rows where x is NA to be modified.
So, even if %notin% is meant to provide a more memory-efficient version of
!lhs %in% rhs%(IIRW), I also think that it would better to handle missing values more safely.P.S.: I wonder if it's also possible to export a functional alternative of %notin%. something like
notin(x, table, nomatch=-1L).