Skip to content

Vignettes #944

Description

@arunsrinivasan

HTML vignette series:

Planned for v1.9.8


Future releases


Finished:


Minor:

  • Operations using integer64, and promoting it for large integers.

Notes (to update current vignettes based on feedbacks): Please let me know if I missed anything..

Introduction to data.table:

  • order in i.
  • Explain how to name columns in j while selecting/computing.
  • Emphasise that keyby is applied after obtaining the result on the computed result, not on the original data.table.
  • Mention new updates to .SDcols and cols in with=FALSE being able to select columns as colA:colB.

Reference semantics:

  • Also explain all other relevant set* functions here.. (setnames, setcolorder etc..)
  • Mainly set.
  • Explain that 1b) the := operator is just defining ways to use it - the example there doesn't work as it just shows two different ways of using it -- Following this comment.

Keys and fast binary search based subsets:

  • Add an example of subset using integer/double keys.
  • Difference in "nomatch" default in binary search based subsets.
  • replacing NAs with binary search based subsets possible?

FAQ (most appropriate here, I think).

  • Update FAQ with issue on external pointer being NULL when reading an R object from file, for example, using readRDS(). Update this SO post.
  • Explain with example, on over allocating the data.table using alloc.col(), and when to use it (when you need to create multiple columns), and why. Update this SO post.

Activity

  1. jangorecki commented on Nov 11, 2014

    @jangorecki
  2. arunsrinivasan commented on Nov 11, 2014

    @arunsrinivasan
    Author
  3. matthieugomez commented on Nov 14, 2014

    @matthieugomez
    Contributor

    I'm curious about what makes a cold by faster than say tapply. One part of the answer is gforce, but what about user written functions? I could not find anything about this. There's a nice post about panda : http://wesmckinney.com/blog/?p=489
    One could even compare it with sapply. For instance, suppose I start from a list of vectors. Is it ever worth it to append all the vectors in one column in a data.table and use by instead of sapply ?

  4. arunsrinivasan commented on Nov 14, 2014

    @arunsrinivasan
    Author
  5. added this to the v1.9.8 milestone on Nov 16, 2014
  6. gsee commented on Nov 19, 2014

    @gsee
  7. arunsrinivasan commented on Nov 24, 2014

    @arunsrinivasan
    Author
  8. markdanese commented on Nov 30, 2014

    @markdanese

    Being new to R and data.table (since March), I would say that there needs to be a basic outcome-oriented introduction as opposed to the current function-oriented one. In other words, it is one thing to read what each parameter in data.table does, but they often make little sense without having a use-case in mind. While there are examples of output, many people need to go the other direction. That is, they know what output they need, but they don't know what function/parameter/setting is most appropriate to use. It would be helpful to have a simple recipe approach to get them started.

    How to I create subsets of my data?
    How do I do an operation on subsets of my data to create a new or updated data set?
    How do I add a new column?
    How do I delete a column?
    How do I create a single variable?
    How do I create multiple variables?
    How do I do different operations on different subsets of my data? (.BY)
    How do I use data.table in a function and pass in data.table names and columns on which to operate?
    How do I do multiple sequential operations on the same data.table?
    Can I select a subset of data and do an operation on it at the same time?
    When do I need to be careful about creating/updating variables by reference?
    How do I select one observation per group (first, last)?
    How do I set a key and how is it different from setting an index?
    Under what conditions does my key get deleted when I do an operation on my data.table?
    Can I just use the regular "merge" syntax or do I need to use data.table syntax (Y[X])?
    How do I collapse a list of lists into one big data.table? What if the columns are in different order?

    There are probably a ton of other items all on SO that could be edited into a simple compilation of questions and answers.

  9. arunsrinivasan commented on Nov 30, 2014

    @arunsrinivasan
    Author
  10. vlulla commented on Dec 23, 2014

    @vlulla
  11. arunsrinivasan commented on Jan 17, 2015

    @arunsrinivasan
    Author
  12. 38 remaining items

  13. MichaelChirico commented on Aug 15, 2019

    @MichaelChirico
  14. zeomal commented on Apr 24, 2020

    @zeomal

    Wanted to chime in and ask if contributions to the vignette are accepted from non-code contributors (like me). I am particularly interested in contributing to the joins vignette as I had quite a bit of trouble with it initially and was guided to solutions from Arun's answers on Stackoverflow, and I'd like some guidance on how to do so, if allowed.

  15. Henrik-P commented on Apr 24, 2020

    @Henrik-P

    @arunsrinivasan I see that you have a point IDateTime vignette. Perhaps it could be included in the more general vignette suggested by @jangorecki: vignettes: timeseries - ordered observations?

    In addition, I am preparing a first draft on some of the topics suggested by jan. Perhaps parts of it may be relevant for a join vignette as well? I'm happy to share if anyone may find it useful.

  16. MichaelChirico commented on Apr 24, 2020

    @MichaelChirico
    Member

    @zeomal such a contribution would be highly valuable and much appreciated!

  17. zeomal commented on Apr 24, 2020

    @zeomal

    @MichaelChirico, thank you. @Henrik-P, will your brief on normal joins be comprehensive - i.e. will your focus be more on timeseries? If not, I can start work on it - I haven't used rolling joins yet, so no knowledge there. :)

  18. Henrik-P commented on Apr 24, 2020

    @Henrik-P

    @zeomal Hopefully I will be able to upload the first draft soon, so you can have a look at it. In my draft, I provide a simple example of a "normal" join on a single variable, time, where there are non-matching rows. I use nomatch = NA. (maaaybe also a quick example with nomatch = NULL)

    My idea was that this simple join could provide a context and a feeling for the problem, which I then treat more thoroughly in the following sections on rolling and non-equi joins et al.

    Thanks a lot for your willingness to contribute! .

  19. zeomal commented on Apr 24, 2020

    @zeomal
  20. jangorecki commented on Apr 24, 2020

    @jangorecki
  21. Henrik-P commented on Apr 25, 2020

    @Henrik-P

    @zeomal If you wish to check how brief my treatment on normal (equi) joins is, I just want to let you know that I posted a PR on a timeseries vignette.

  22. kjytay commented on Oct 11, 2021

    @kjytay
  23. MichaelChirico commented on Oct 11, 2021

    @MichaelChirico
  24. kjytay commented on Oct 11, 2021

    @kjytay
  25. MichaelChirico commented on Jul 11, 2025

    @MichaelChirico
    Member

    This issue is a bit too generic & dated; the main thing was the joins vignette which is done.

    Closing, please file specific issues now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions