Repository navigation
File-backed data.tables #1336
Description
Activity
Agreed. Probably for v2.0.0.. depending on how much time and motivation we've.
Reacted by Zach Deane-Mayer, Srikanth K S and Anantadinathzachmayer commented
on Sep 22, 2015 on Sep 22, 2015 · Hidden as duplicateAuthorshow commentMore actionszachmayer commented
on Jun 14, 2016 on Jun 14, 2016 · Hidden as outdatedAuthorshow commentMore actionsThe links in the original post of @zachmayer are not valid anymore. The GitHub repo of Graphlab/Dato/Turi can be found here. Because Graphlab/Dato/Turi has been acquired by Apple, this repo has been moved to here. It looks like it has evolved into a library for the development of machine learning models.
In case above two links stop working, I've created a fork in my own profile.
Reacted by Arun SrinivasanOne potential implementation strategy is via R's custom allocator mechanism. I constructed a file-backed
data.tablewith individual columns backed by mmap-d files based on the code here.See this gist, where I create the 2B row dataset (~75GB) from the benchmarks and run some aggregations on my laptop (16GB ram). There's many missing pieces that make this far from a user-friendly solution though. Among them: R's custom allocator is used for the entire array object, so there is an R implementation specific header prepended to the data; can't share even read-only between R sessions due to the former; can't hook data.table allocations for new objects (columns/indices) so they won't be memory-mapped; no support for real string columns; requires manual persistence of column attributes.
All those caveats aside, I've already found it to be quite useful when working with a large number of moderate sized datasets, where each is sequentially memory mapped, data.table is told they're already sorted (
attr(DT, 'order') = ...) and then performing a "roll" join to extract data with a given lookback, such that the only the data needed for the binary search and the subsequent values needs to be read from disk.Reacted by Jan Gorecki, DrOrrery, Daniel de Oliveira Arantes, Sean Ingerson and Aniwaynelapierre commented
on Feb 13, 2019 on Feb 13, 2019 · Hidden as duplicateshow commentMore actions- addedtop requestOne of our most-requested issuesOne of our most-requested issuesand removed
on Jun 7, 2020 @jonekeat disk.frame is possibly an alternative but I haven't tried it myself.
Reacted by Jone Keat Lim and GitHunter0Reacted by GitHunter0@jonekeat disk.frame is possibly an alternative but I haven't tried it myself.
disk.frameis the most promising R solution for this matter I've seen so far. It would be very interesting to seedata.tableanddisk.framecontributors working togetherReacted by Ruben Dries and AnantadinathAs a current-day workaround, what about the use of
arrow::open_datasetanddtplyror similar? The data is immutable so "saving" data would need to be an explicit step, but at least fast access to on-desk data should be feasible. (I recognize this does not fully address all likely use-cases for on-diskdata.tableoperations, mostly a technique for mitigating large-data operations.)This is currently out of scope https://github.com/Rdatatable/data.table/blob/master/GOVERNANCE.md#the-r-package and I don't think anyone has the time/interest/skill to implement, so I'm closing.
Reacted by r2evansI don't disagree, it's definitely big-scope. I offered my comment to illustrate alternative paths.
This is currently out of scope https://github.com/Rdatatable/data.table/blob/master/GOVERNANCE.md#the-r-package and I don't think anyone has the time/interest/skill to implement, so I'm closing.
to clarify I'd be glad to have scope expanded for this high-demand FR, but as noted current maintainer core has no time/ability to support this. outside contributions (and commitment to ownership) welcome.
Reacted by r2evans
SFrames are graphlab create's version of data.frames, and have some impressive performance benchmarks on single machines.
I'd really love to see something similar for data.table that could use disk rather than RAM to store the data.