Skip to content

File-backed data.tables #1336

Description

@zachmayer

SFrames are graphlab create's version of data.frames, and have some impressive performance benchmarks on single machines.

I'd really love to see something similar for data.table that could use disk rather than RAM to store the data.

Activity

  1. arunsrinivasan commented on Sep 22, 2015

    @arunsrinivasan
    Member

    Agreed. Probably for v2.0.0.. depending on how much time and motivation we've.

  2. zachmayer commented on Sep 22, 2015

    @zachmayer
    Author
  3. zachmayer commented on Jun 14, 2016

    @zachmayer
    Author
  4. mbacou commented on Jun 15, 2016

    @mbacou
  5. arunsrinivasan commented on Jul 3, 2016

    @arunsrinivasan
  6. clarkfitzg commented on Mar 27, 2017

    @clarkfitzg
  7. vors commented on Feb 11, 2018

    @vors
  8. jaapwalhout commented on Feb 12, 2018

    @jaapwalhout

    The links in the original post of @zachmayer are not valid anymore. The GitHub repo of Graphlab/Dato/Turi can be found here. Because Graphlab/Dato/Turi has been acquired by Apple, this repo has been moved to here. It looks like it has evolved into a library for the development of machine learning models.

    In case above two links stop working, I've created a fork in my own profile.

  9. aquasync commented on Jan 21, 2019

    @aquasync

    One potential implementation strategy is via R's custom allocator mechanism. I constructed a file-backed data.table with individual columns backed by mmap-d files based on the code here.

    See this gist, where I create the 2B row dataset (~75GB) from the benchmarks and run some aggregations on my laptop (16GB ram). There's many missing pieces that make this far from a user-friendly solution though. Among them: R's custom allocator is used for the entire array object, so there is an R implementation specific header prepended to the data; can't share even read-only between R sessions due to the former; can't hook data.table allocations for new objects (columns/indices) so they won't be memory-mapped; no support for real string columns; requires manual persistence of column attributes.

    All those caveats aside, I've already found it to be quite useful when working with a large number of moderate sized datasets, where each is sequentially memory mapped, data.table is told they're already sorted (attr(DT, 'order') = ...) and then performing a "roll" join to extract data with a given lookback, such that the only the data needed for the binary search and the subsequent values needs to be read from disk.

  10. DrOrrery commented on Jan 25, 2019

    @DrOrrery
  11. waynelapierre commented on Feb 13, 2019

    @waynelapierre
  12. added
    top requestOne of our most-requested issues
    and removed on Jun 7, 2020
  13. jonekeat commented on Dec 30, 2020

    @jonekeat

    Is something similar to what @aquasync proposed already implemented? I have tried to use mmap package to memory map each column in a list, then setDT, but it cannot work with data.table methods. I am looking for any alternatives before using databases/spark or rewrite into c/c++

  14. jangorecki commented on Dec 30, 2020

    @jangorecki
    Member

    @jonekeat disk.frame is possibly an alternative but I haven't tried it myself.

  15. GitHunter0 commented on Feb 22, 2021

    @GitHunter0

    @jonekeat disk.frame is possibly an alternative but I haven't tried it myself.

    disk.frame is the most promising R solution for this matter I've seen so far. It would be very interesting to see data.table and disk.frame contributors working together

  16. r2evans commented on Apr 9, 2024

    @r2evans
    Contributor

    As a current-day workaround, what about the use of arrow::open_dataset and dtplyr or similar? The data is immutable so "saving" data would need to be an explicit step, but at least fast access to on-desk data should be feasible. (I recognize this does not fully address all likely use-cases for on-disk data.table operations, mostly a technique for mitigating large-data operations.)

  17. tdhock commented on Apr 9, 2024

    @tdhock
    Member

    This is currently out of scope https://github.com/Rdatatable/data.table/blob/master/GOVERNANCE.md#the-r-package and I don't think anyone has the time/interest/skill to implement, so I'm closing.

  18. r2evans commented on Apr 9, 2024

    @r2evans
    Contributor

    I don't disagree, it's definitely big-scope. I offered my comment to illustrate alternative paths.

  19. MichaelChirico commented on Apr 9, 2024

    @MichaelChirico
    Member

    This is currently out of scope https://github.com/Rdatatable/data.table/blob/master/GOVERNANCE.md#the-r-package and I don't think anyone has the time/interest/skill to implement, so I'm closing.

    to clarify I'd be glad to have scope expanded for this high-demand FR, but as noted current maintainer core has no time/ability to support this. outside contributions (and commitment to ownership) welcome.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions