Skip to content

Unintuitive output from a rolling join #1469

Description

@DavidArenburg

When performing a rolling join by different column names, it seems like the first column name is selected by default regardless where the values are coming from. This is very confusing to me. Is this by design? (sorry flying tonight in didn't have time to test on the GH version too)

library(data.table) # V 1.9.6+
set.seed(123)
DT <- data.table(X = 1:3)
DT2 <- data.table(Y = 3:5, Z = sample(3))
DT[DT2, on = c(X = "Y"), roll = -Inf]
#    X Z <~~~ The column name is X but values are from Y
#1: 3 1
#2: 4 2
#3: 5 3

Activity

  1. ben519 commented on Nov 15, 2016

    @ben519

    This. I couldn't agree more. I constantly work around this issue.

  2. jangorecki commented on Nov 15, 2016

    @jangorecki
    Member

    This issue is already covered in the linked one, it is also well explained there.
    Current recommendation below. It is unlikely to be changed any time soon, feel free to post your proposal into linked issue.

    DT[DT2, .(Y, Z), on = c(X = "Y"), roll = -Inf]
    # when column names are overlapping
    DT[DT2, .(Y=i.Y, Z=i.Z), on = c(X = "Y"), roll = -Inf]
    # now also 'x.' prefix should work the same as 'i.' before giving better control

    If I misinterpret this issue please re-open providing better example, which cannot be addressed by x. and i. prefixes when listing columns in j.

  3. added this to the milestone on Apr 10, 2018
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions