Skip to content

Compatibility with the future native pipe #4872

Description

@eliocamp

Right now, r-devel is implementing a native pipe that is incompatible with chaining in data.table. Using the proposed d => syntax I get an error using this R version

library(data.table)
data <- CJ(group1 = letters[1:5], group2 = letters[10:14])

data[, x := rnorm(.N)] |> 
  d => d[, mean(x), by = group1]
#> Error: function '[' not supported in RHS call of a pipe

If "[" won't be supported as the RHS of a pipe, then no syntax transformation can make data.table compatible with the |> pipe, if I understand correctly.

I see two ways of addressing this.

On the one side, we could talk with r-core to present this issue and see if they can change their implementation to make it work.

Another option would be to create a functional alias to data.table:::"[.data.table", lets say dt(). This, for example, seems to work.

dt <- data.table:::"[.data.table"

data[, x := rnorm(.N)] |> 
   dt(, mean(x), by = group1)
#>    group1          V1
#> 1:      a  0.58876302
#> 2:      b -0.24705765
#> 3:      c  0.07676786
#> 4:      d -0.33047608
#> 5:      e  0.54832829

(in fact, it also works with dt <- base::"[")

Personally, I don't mind that notation at all. It is true that if such a simple fix resolves the issue, then each user could do it in their own scripts, but I think it might be preferable if data.table provided a standardised alias so that other people's code stays readable. For the record, I don't like dt very much, it's just the first thing that came in mind.

Activity

  1. jangorecki commented on Jan 13, 2021

    @jangorecki
    Member

    This is base R limitation not related to data.table

    toupper(letters) |> x => x[1:5]
    #Error: function '[' not supported in RHS call of a pipe

    Is there a good reason for not supporting it? For me it almost classify as a bug.
    In case it is not going to be fixed in R-devel then this issue #641 should take care of it.

  2. tdeenes commented on Jan 13, 2021

    @tdeenes
    Member

    This is deliberately so. Special symbols are not allowed in the RHS of a pipe, see here.

    Nevertheless, only the name of the symbol is checked, so a workaround is to assign a new symbol to the given primitive function. E.g.:

    1:4 |> `+`(3)   # error 
    add <- `+`
    1:4 |> add(3) # OK
    
    extract <- `[`
    data.table(x = letters) |> extract(j = x)  # OK

    Of course the extract function above is not safe, e.g., you will get a warning if the input is a data.frame instead of a data.table:

    extract <- `[`
    data.frame(x = letters) |> extract(j = x) 
    # Warning message:
    # In `[.data.frame`(data.frame(x = letters), j = x) :
    #   named arguments other than 'drop' are discouraged
  3. eliocamp commented on Jan 13, 2021

    @eliocamp
    ContributorAuthor

    Yes, #641 should solve it. IMHO, dtquery() is too long a name, though. Especially with the goal of using the pipe in mind.

    mtcars |> 
       dtquery(mpg > 10) |> 
       dtquery(, disp2 := disp/wt) |> 
       dtquery(, .(hp = mean(hp)), by = cyl)

    To me, that code feel horrible to write and distracting to read, but it's purely a personal matter.

    Of course, since we are not talking about cache invalidation, the difficult thing here is naming the thing.

  4. jangorecki commented on Jan 13, 2021

    @jangorecki
    Member

    Agree, maybe sb from square bracket?

  5. eliocamp commented on Jan 13, 2021

    @eliocamp
    ContributorAuthor
    mtcars |> 
      sb(mpg > 10) |> 
      sb(, disp2 := disp/wt) |> 
      sb(, .(hp = mean(hp)), by = cyl)

    data.table symbols start with dot, perhaps something like that? It might be consistent.. It would also evoke a bit of the %>% .[] notation. .S (for "Square"?)

    mtcars |> 
      .S(mpg > 10) |> 
      .S(, disp2 := disp/wt) |> 
      .S(, .(hp = mean(hp)), by = cyl)

    Also, .t? because it kinda looks like .[

    mtcars |> 
      .t(mpg > 10) |> 
      .t(, disp2 := disp/wt) |> 
      .t(, .(hp = mean(hp)), by = cyl)
  6. grantmcdermott commented on Apr 29, 2021

    @grantmcdermott
    Contributor

    Adding my 2 cents:

    I really like the idea of a dot prefix, but would go for .d(). Fast to type and the most common letter assignment ("d") for a data frame-like object.

    Failing that, I would stick to well-known data.table naming conventions, which makes .DT() the next obvious candidate IMO.

  7. Kamgang-B commented on May 18, 2021

    @Kamgang-B
    Contributor

    I personally like the idea of .S ("Square") for square brackets but would prefer .s (lowercase) because it is easier to type. Because this is intended to be used with the pipe, I think that it should not only be as easier to type as possible but also as short as possible (possibly not more than 2 characters). This is just my opinion.

    Also, I noticed that the following expression does not print anything (even if the last function call is print());
    but, the output gets printed if the row with assignment ( .s(, b := -a) ) is removed.

    .s <- data.table:::`[.data.table`
    
    y <- list(a = 1:5)
    
    setDT(y) |>
        .s(, b := -a) |>
        print()
    

    I also tried the next chunk of code but it did not work either. I thought that it would have behaved like setDT(y)[, b := -a][] but it did not print anything.

    setDT(y) |>
        .s(, b := -a) |>
        .s()
    
  8. grantmcdermott commented on May 19, 2021

    @grantmcdermott
    Contributor

    Personally, I'm not so keen on either of .t() or .s()/.S().

    For .t(), I think it goes against the mnemonic that data.table uses for its other special operators. .SD, .GRP, etc. are all based on an abbreviation heuristic rather than the "shape" of the operator. I also think it's too close to the base transpose function, which is where my brain automatically goes when I see t... and invites unexpected behaviour in the case of a simple typo. (C.f. DT |> .t() vs DT |> t()).

    (dt() would be even worse ofc because that creates a namespace conflict with a base R function.)

    Similarly, .s()/.S() is too close to .SD and .SDcols in my view, without being at all related to the underlying functionality.

    Taking a step back, IMO the point of this symbol should be to encapsulate a data.table. I know I've already said this, but I genuinely think that .DT() is the obvious solution here. It's the widely used shorthand for constructed data.tables already, used throughout the docs, and should be intuitive enough for newcomers as well. (Although, I'll readily accept .d() or .D too ;-)

  9. philippechataignon commented on May 19, 2021

    @philippechataignon
    Contributor

    Personnally .dt() is my preference.

  10. added this to the 1.14.1 milestone on May 19, 2021
  11. hadley commented on May 19, 2021

    @hadley
    Contributor

    (FWIW the tidyverse team are planning to work with R core to figure out a placeholder syntax (or equivalent) for the base pipe that there's some hope that 4.2 might allow you do to dt|> .[i] (or similar))

  12. MichaelChirico commented on May 19, 2021

    @MichaelChirico
    Member

    that sounds great @hadley -- is there a bugzilla/r-devel thread we could use to track progress here to help gauge whether short-term investment in a wrapper is worthwhile?

  13. hadley commented on May 19, 2021

    @hadley
    Contributor

    Not that I know of, sorry.

  14. moodymudskipper commented on May 19, 2021

    @moodymudskipper

    If such a function is implemented it would be nice to have it work with non data.table data frames and return an output of the original class. I suppose we could make a case to always return a data.table but I like personally the idea of not forcing the user into data.table world if they only need one call (that'd be similar to https://github.com/moodymudskipper/withDT).

    I also wonder if it makes sense to modify by reference using the pipe, maybe it should always copy?

  15. MichaelChirico commented on May 20, 2021

    @MichaelChirico
    Member

    @moodymudskipper I don't think we're planning on maintaining any more complicated interface than a simple wrapper for convenience:

    <foo> = `[.data.table` = function(...) {...}
    
  16. eliocamp commented on May 20, 2021

    @eliocamp
    ContributorAuthor

    My sense is that iif there is some guarantee that 4.2.0 will support dt |> .[], then there's not much point in introducing a wrapper that could be easily be created by an user on a per-script basis. In the meantime, data.table's documentation could just suggest users to add the wrapper themselves.

  17. hadley commented on May 20, 2021

    @hadley
    Contributor

    I don't have 4.1.0 handy, but I'd think that this would work:

    mtcars |> 
       `[`(mpg > 10) |> 
       `[`(, disp2 := disp/wt) |> 
       `[`(, .(hp = mean(hp)), by = cyl)

    Given `[` is three characters, for a shortcut to make sense it would need to be either one or two characters. And of course that's easy to make yourself:

    X <- `[`
    mtcars |> 
      X(mpg > 10) |> 
      X(, disp2 := disp/wt) |> 
      X(, .(hp = mean(hp)), by = cyl)

    (I'm not proposing X as a name, just using it as an example.)

    There's no guarantee that 4.2.0 will have a placeholder syntax, but it's something that I hope we can persuade R core is useful and important. One of the main objections is that . is a very small symbol and easy to miss.

  18. deleted a comment from eliocamp on May 20, 2021
  19. eddelbuettel commented on May 21, 2021

    @eddelbuettel
    Contributor

    Under 4.1.0 the first example fails:

    > mtcars |> 
    +    `[`(mpg > 10) |> 
    +    `[`(, disp2 := disp/wt) |> 
    +    `[`(, .(hp = mean(hp)), by = cyl)
    Error: function '[' not supported in RHS call of a pipe
    > 

    The second one works if and only if mtcars was made into a data.table object:

    > X <- `[`
    > mtcars |> 
    +   X(mpg > 10) |> 
    +   X(, disp2 := disp/wt) |> 
    +   X(, .(hp = mean(hp)), by = cyl)
       cyl       hp
    1:   6 122.2857
    2:   4  82.6364
    3:   8 209.2143
    > 

    It is of course nothing but sugar over existing syntax:

    > mtcars[mpg > 10][, .(hp = mean(hp)), by = cyl]
       cyl       hp
    1:   6 122.2857
    2:   4  82.6364
    3:   8 209.2143
    > 
  20. modified the milestones: 1.14.9, 1.15.0 on Oct 29, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions