Repository navigation
fread for directories #2582
Description
Activity
My thinking was that whenever
fread's input resolves to multiple files, then each of the files should be read in turn, and then returned as a list of DataTables. An attribute can be set on each of these DataTables to specify the name of its particular source. If one of the sources cannot be read, then it should be represented as an exception object in the list, while other sources should continue to parse (there could be an option to control whether to throw an error immediately, or perhaps to skip the bad files).However I don't think that automatically
rbinding all of the DataTables into a single result is a good idea. The reason being that in practice some of them may have irregularities (eg. different column names), or some of files are picked up that are not csv data at all, or perhaps a field needs to be added based on the name of the source file, etc. Adding options to support all of these intricacies would complicate the interface unnecessarily, will be time-consuming (if you get one of the options wrong you need to rescan all the files), and potentially fragile (new use cases may demand adding new options). Whereas if you just return a list of DataTables, then the user is free to do whatever she/he wants using familiar language constructs.Reacted by Frank and Michael Chirico- good feedback. actually it's a good point, as some users may want to use the forthcoming `cbindlist` as well, or to `Reduce(merge)` the items.
One use-case where I've found R wanting is when the directory contains a very large number of small files (i.e. 100,000 to 1,000,000 files of 1-10 kB). In such cases,
fread+rbindlistis not faster thanread.csv+rbindlistand both are orders of magnitude slower than using the command line:copy /b *.csv > out.csv. Difficulties arise when using the command line option when the columns are not in the same order (i.e.use.names = TRUEwould help) or when the column headers are present in each of the file (because concatenation results in a file with 100,000 headers interspersed throughout the file), but it's still much faster than the R alternatives I know.Interesting use case, did you try to investigate where the bottleneck is?
Is it in the constant overhead time that fread spends detecting format of each file? (Setting some of the parameters explicitly might reduce the time then)
Or is it in rbinding itself?
Or maybe there is significant overhead from R itself trying to read the directory?It looks something like this which is not as bad as I remember (Matt, can you stop improving the package? It's ruining my anecdotes.)
[System.IO.Directory]::GetFiles("address", "*.*").Count 463716
system.time(list.files(path = "address", pattern = "\\.csv$", full.names = TRUE)) # user system elapsed # 7.66 1.32 8.99 Files <- list.files(path = "address", pattern = "\\.csv$", full.names = TRUE) system.time(lapply(Files[1:100], fread, fill = TRUE)) # user system elapsed # 2.38 0.08 2.59 system.time(lapply(Files[1:1e2], fread, sep = ",", colClasses = "character", fill = TRUE)) # user system elapsed # 1.58 0.10 1.67 system.time(lapply(Files[1:1e4], fread, sep = ",", colClasses = "character", fill = TRUE)) # user system elapsed # 6.97 5.00 22.67
Reacted by Michael Chirico@st-pasha I assume it's in
freadoverhead, I seem to recall running a benchmark whereread.csvis faster for very small files (like <20 rows)also, relating to your first comment, returning the source attributes in the names would highlight the utility of #1948 as well for manipulating these objects in post.
Another use case I'm running into (.. not sure how different it is from the preceding):
I wrote a helper function to read a csv inside tar.gz like
fread("7z -so mycsv.tar.gz | 7z x -si -so -ttar")but now I have tar.gz containing multiple csvs (that should have identical column names and classes), and it seems I'll need to go another way (I guess: run the 7z call then lapply fread on the files it drops > confirm columns match > rbindlist).
Just to have some thoughts written down:
There's a pretty simple version of this where we just wrap
fread('/path/to/dir', ...)tolapply(list.files('/path/to/dir'), fread, ...)...there's also a substantially more involved version where directory-level
freadis all done in C, and the pre-amble stuff (ncol/nrow/type detection) is done first in a loop, then we allocate all the memory once & either (1) fill the table in parallel over files,nthread=1within files or (2) fill the table in serial over files, parallel within file.Definitely the first version should use the simple approach, but it almost surely won't be faster (for many use cases) than using the terminal to
catthe files totempfile()first & reading that. If we understand well when this latter approach is preferred, we might leveragefile.append(possibly excluding headers?)...IMO it is not good if
freadwould return a list rather than data.table or data.frame, unless we provide an extra argument. I mean that changing"dir"to"dir/file1.csv"should not change the class of returned object. Eventually when providing non-scalar filesc("dir/file1.csv","dir/file2.csv"), then it make sense to return a list of data.tables.
lapply(, fread)seems quite good to return list already.
If we want to fread a directory, maybe we could expect all files to be similar schema, and then extra argument on how tomerge/bindthose files could useful, so it can still return just a data.table.A simple
how = c('list', 'rbindlist', 'cbindlist', 'mergelist')(or similar) could be good once #4370 is doneReacted by Jan GoreckiIdle musing -- if
how='rbindlist', we should probably do something like: read schema from the first file, then supply that ascolClassesfor subsequent files for efficiency. As inspired hereUnless fill=T expected
Some file I/O APIs I've worked with have a simple idiom for reading full directories:
would be read as, e.g. in
spark,A basic idiom has developed for
freadto do this by adding bells and whistles to the following:It would be simple enough to wedge directory reading into the
freadAPI by changing:to (pseudocode around the
match.call()part)However, it might be nice to build in some flexibility to this, e.g. allowing the
list.filesto optionally be recursive, implementing some API for automatic source naming (if there are subdirectories, and the names of the subdirectories contain information, the manual version of this allows a bit more flexibility), specifyingidcolorfill, etc.So there's two questions here
fread? Just add...and post-process if theis.dirbranch is reached? Separate function call altogether?