Skip to content

data.table coding style #1532

Description

@dselivanov

It would be great if we will have some coding style guidelines for data.table project (like contributing guide).
My 5 cents (for both C and R sources):

  1. avoid very long lines - use 80-100 characters.
  2. Place spaces around all infix operators. See for example this line.
  3. Always put a space after a comma, and never before (just like in regular English). See here.

Topic to discuss.

I'm not sure, that it is a good idea to use so long function like readfile in fread.c. Usually it is hard to debug/maintain such long functions.

Activity

  1. MichaelChirico commented on Feb 10, 2016

    @MichaelChirico
    Member

    adding a wiki about this and a link to it from the contribution guidelines is also quite important!

  2. dselivanov commented on Nov 12, 2017

    @dselivanov
    Author

    Partially addressed in #2420. Should be really part of Contributing policy.

  3. jangorecki commented on Apr 6, 2018

    @jangorecki
    Member

    aside from wiki/about we could provide (or recommend) to use specific style templates for RStudio, Emacs ESS, etc.

  4. MichaelChirico commented on May 5, 2019

    @MichaelChirico
    Member

    Started a small guide at Contributing Policy with the first couple points that came to mind...

    avoid very long lines - use 80-100 characters

    I've noticed @mattdowle has a much wider preference for lines... he must have laser sharp eyes 😄

  5. Atrebas commented on May 6, 2019

    @Atrebas

    I think this is also an important question from a user perspective. I do not have a CS background, but I try to get better at programming and to follow good coding practices.
    This is increasingly important, for example because:
    -production environments require better practices (e.g. with linting tools, SonarQube, ...)
    -when communicating code with others, consistency and clarity matter

    There is the tidyverse style guide that I think is useful. But I recently came across some questions with data.table.
    This may be a matter of taste, but I would be interested in knowing what other users do.
    There are also some advice distilled across the docs (mind integers, use on, ...), it could be useful to have them gathered in a single place.

    library(data.table)
    set.seed(1L)
    
    DT <- data.table(V1 = rep(c(1L,2L), 5),
                     V2 = round(rnorm(10), 2),
                     V3 = rep(c("A","B"), 5))
    
    DT
    
    # filter rows ------------------------
    
    # should we leave a comma in DT[i]?
    DT[3:4]
    
    DT[3:4,]
    DT[3:4, ] # is a white space 'mandatory' in this case?
    
    
    # select columns ------------------------
    
    # good
    DT[, list(V2)]
    DT[, .(V2)]   
    
    # bad
    DT[,list(V2)]
    DT[,.(V2)]   
    
    
    # new columns ------------------------
    
    # the following expression is concise
    DT[, .(sumv1 = sum(V1), sdv2 = sd(V2))]
    
    # but would this way be better?
    DT[, .(sumv1 = sum(V1),
           sdv2  = sd(V2))]
    
    # or even?
    DT[,
       .(sumv1 = sum(V1),
         sdv2  = sd(V2))
      ]
    
    # likewise with by, we can do a single line
    DT[1:5, .(sumV1 = sum(V1)), by = V3]
    
    # but it may be better to split with longer expressions?
    DT[V1 < 5 & V2 < 10,
       .(sumV1 = sum(V1),
         sdv2  = sd(V2)),
       by = V3]
    
    # as mentionned by @dzj_evalparse, we can reorder the arguments
    DT[by = V3,
       V1 < 5 & V2 < 10,
       .(sumV1 = sum(V1),
         sdv2  = sd(V2))
       ]
    
    # one more example where it would make sense to split/reorder
    # even if the code is < 80 characters
    DT[, lapply(.SD, mean), by = V3, .SDcols = c("V1", "V2")]
    
    DT[, by = V3,
         lapply(.SD, mean), 
         .SDcols = c("V1", "V2")
      ] # closing bracket here?
    
    
    # functional form recommended when adding several columns?
    
    DT[, c("V4","V5") := .("X", "Y")]
    
    # better to keep track of each column name and allows to comment
    DT[, ':='(V4 = "X",
              # V4 = "Z",
              V5 = "Y")]
    
    # chain expressions ------------------------
    # again with the brackets...
    
    DT[, by = V3,
         lapply(.SD, mean), 
         .SDcols = c("V1", "V2")
      ][
          order(-V2)
      ]
    
    # or?
    DT[, by = V3,
         lapply(.SD, mean), 
         .SDcols = c("V1", "V2")][
          order(-V2)]
    
    
    # better to use spaces around '=' for parameters
    shift(1:10, n = 1, fill = NA, type = "lag")
    
    
    # better to use 'on', even with keyed data
    # better to use 'explicit' NA, e.g. NA_character_
    # be aware of numeric/integer
    # ...
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions