Repository navigation
Vignettes #944
Description
Activity
arunsrinivasan commented
on Nov 11, 2014 on Nov 11, 2014 · Hidden as outdatedAuthorshow commentMore actionsI'm curious about what makes a cold by faster than say
tapply. One part of the answer is gforce, but what about user written functions? I could not find anything about this. There's a nice post about panda : http://wesmckinney.com/blog/?p=489
One could even compare it withsapply. For instance, suppose I start from a list of vectors. Is it ever worth it to append all the vectors in one column in a data.table and usebyinstead ofsapply?arunsrinivasan commented
on Nov 14, 2014 on Nov 14, 2014 · Hidden as outdatedAuthorshow commentMore actionsarunsrinivasan commented
on Nov 24, 2014 on Nov 24, 2014 · Hidden as outdatedAuthorshow commentMore actionsBeing new to R and data.table (since March), I would say that there needs to be a basic outcome-oriented introduction as opposed to the current function-oriented one. In other words, it is one thing to read what each parameter in data.table does, but they often make little sense without having a use-case in mind. While there are examples of output, many people need to go the other direction. That is, they know what output they need, but they don't know what function/parameter/setting is most appropriate to use. It would be helpful to have a simple recipe approach to get them started.
How to I create subsets of my data?
How do I do an operation on subsets of my data to create a new or updated data set?
How do I add a new column?
How do I delete a column?
How do I create a single variable?
How do I create multiple variables?
How do I do different operations on different subsets of my data? (.BY)
How do I use data.table in a function and pass in data.table names and columns on which to operate?
How do I do multiple sequential operations on the same data.table?
Can I select a subset of data and do an operation on it at the same time?
When do I need to be careful about creating/updating variables by reference?
How do I select one observation per group (first, last)?
How do I set a key and how is it different from setting an index?
Under what conditions does my key get deleted when I do an operation on my data.table?
Can I just use the regular "merge" syntax or do I need to use data.table syntax (Y[X])?
How do I collapse a list of lists into one big data.table? What if the columns are in different order?There are probably a ton of other items all on SO that could be edited into a simple compilation of questions and answers.
Reacted by Chitra M Saraswatiarunsrinivasan commented
on Nov 30, 2014 on Nov 30, 2014 · Hidden as outdatedAuthorshow commentMore actionsarunsrinivasan commented
on Jan 17, 2015 on Jan 17, 2015 · Hidden as outdatedAuthorshow commentMore actions38 remaining items
MichaelChirico commented
on Aug 15, 2019 on Aug 15, 2019 · Hidden as duplicateshow commentMore actionsWanted to chime in and ask if contributions to the vignette are accepted from non-code contributors (like me). I am particularly interested in contributing to the joins vignette as I had quite a bit of trouble with it initially and was guided to solutions from Arun's answers on Stackoverflow, and I'd like some guidance on how to do so, if allowed.
@arunsrinivasan I see that you have a point
IDateTime vignette. Perhaps it could be included in the more general vignette suggested by @jangorecki: vignettes: timeseries - ordered observations?In addition, I am preparing a first draft on some of the topics suggested by jan. Perhaps parts of it may be relevant for a join vignette as well? I'm happy to share if anyone may find it useful.
@zeomal such a contribution would be highly valuable and much appreciated!
@MichaelChirico, thank you. @Henrik-P, will your brief on normal joins be comprehensive - i.e. will your focus be more on timeseries? If not, I can start work on it - I haven't used rolling joins yet, so no knowledge there. :)
@zeomal Hopefully I will be able to upload the first draft soon, so you can have a look at it. In my draft, I provide a simple example of a "normal" join on a single variable, time, where there are non-matching rows. I use
nomatch = NA. (maaaybe also a quick example withnomatch = NULL)My idea was that this simple join could provide a context and a feeling for the problem, which I then treat more thoroughly in the following sections on rolling and non-equi joins et al.
Thanks a lot for your willingness to contribute! .
Reacted by zeomal@zeomal If you wish to check how brief my treatment on normal (equi) joins is, I just want to let you know that I posted a PR on a timeseries vignette.
Reacted by zeomalMichaelChirico commented
on Oct 11, 2021 on Oct 11, 2021 · Hidden as resolvedshow commentMore actionsThis issue is a bit too generic & dated; the main thing was the joins vignette which is done.
Closing, please file specific issues now.
Reacted by Jan Gorecki
HTML vignette series:
Planned for
v1.9.8i.colusage as filed in Docs: explain and document the i.col notation for joins #1038. d) Also cover about performance/advantages fromonperforming slower than doublesetkey#1232.[ ] Covercovered in programming on data.table #4304get()andmget(). E.g., http://stackoverflow.com/q/33785747/559784on=argument for joins #1623).Future releases
fread+rbindlist), ordering, ranking and set operationsdata.table()anddata.frame()somewhere - relevant issues: Creation of data.table using a list #968, data.table(x) != as.data.table(x) #877. Perhaps slightly more in detail in the FAQ.data.tableusage:fread+fwritevignette, include also Convenience features of fread wiki, also fread (and fwrite) vignette #2855Finished:
i, select / do injand aggregations usingby.iandbyin the same way as before)by=.EACHIuntil the vignette is done.Minor:
integer64, and promoting it for large integers.Notes (to update current vignettes based on feedbacks): Please let me know if I missed anything..
Introduction to data.table:
orderini.jwhile selecting/computing..SDcolsand cols inwith=FALSEbeing able to select columns ascolA:colB.Reference semantics:
set*functions here.. (setnames,setcolorderetc..)set.1b) the := operatoris just defining ways to use it - the example there doesn't work as it just shows two different ways of using it -- Following this comment.Keys and fast binary search based subsets:
FAQ (most appropriate here, I think).
readRDS(). Update this SO post.alloc.col(), and when to use it (when you need to create multiple columns), and why. Update this SO post.