Skip to content

Support .gz file format for fread #717

Description

@renqian

I have several thousands of .gz files containing data in csv format - about 60GB in total in terms of .gz files. Decompressing them and load some pieces via fread turns out a huge pain in the first step. I'm wonder whether it is possible to improve the functionality of fread so that it can read compressed file formats just as read.table does?

Perhaps file connection issues are highly relevant, as mentioned in #341, #543, and #561.
Some other reference:

http://stackoverflow.com/questions/5764499/decompress-gz-file-using-r

http://blog.revolutionanalytics.com/2009/12/r-tip-save-time-and-space-by-compressing-data-files.html

Activity

  1. eantonya commented on Jul 7, 2014

    @eantonya
    Contributor

    You can just do fread('zcat file.gz'), or some loop variation, if you have many files.

  2. added and removed on Jul 8, 2014
  3. arunsrinivasan commented on Sep 28, 2014

    @arunsrinivasan
    Member

    :Bump: quite useful (and coming up frequently).

    Here's one SO post.

  4. gbonamy commented on Sep 28, 2014

    @gbonamy

    Yes this would be a very useful feature to have. Using a command line, is a temporary solution at best, since it relies on the underlying system to have the tools for decompression. For instance 'zcat' is not available on windows unless one installs cygwin etc.

    Since FRead is by far the best tool in R to read file, it would be a huge performance gain to read gzipped/bziped/... files directly.

  5. rsaporta commented on Sep 30, 2014

    @rsaporta
    Contributor

    I just saw @Arun's bump in my email, and literally a few hours ago I was ingesting 200+ such files. +1 for usefullness

  6. mattdowle commented on Oct 7, 2014

    @mattdowle
    Member

    Would have been useful here as well :
    http://www.magesblog.com/2014/10/visualising-seasonality-of-atlantic.html
    Will take a look.

  7. xiaodaigh commented on Nov 20, 2014

    @xiaodaigh

    I agree with @gbonamy readding directly from zip files would be a fantastic addition!!

  8. rmscriven commented on Mar 16, 2015

    @rmscriven

    Reading from a connection with unz() would also be quite useful. I have a function that downloads a zip file, and only reads one file then throws it away. So if I could use fread(unz(zipfile, file = file)) it would be a great addition.

  9. statquant commented on Apr 17, 2015

    @statquant

    I ++ about directly from gz files. I would personally use it every day.

  10. mspivakov commented on Apr 26, 2015

    @mspivakov

    +1 from me as well.

  11. 37 remaining items

  12. frenchja commented on May 12, 2017

    @frenchja

    One thing to note is that the zcat solution appears to only work if the file exists in the same directory that R is launched:

    Error in fread("zcat < data/directory/test.csv.gz") :
      File is empty: /var/folders/41/asdf_kj80000gn/T//RtmpwtAttt/fileebeb5e124cef
    
  13. map2085 commented on May 17, 2017

    @map2085

    i forgot about this issue and tried to fread a gz file, only to get a mysterious error causing me to waste time, again, searching for the solution.

    3 years later, still waiting for this elementary fix.

  14. frenchja commented on May 17, 2017

    @frenchja

    After further exploration, my error above only occurs when there are spaces in the directory name:

    fread("zcat < data/directory\ one/test.csv.gz"

    But not with underscores:

    fread("zcat < data/directory_two/test.csv.gz"

    And can be alleviated by escaping the backslash again:

    fread("zcat < data/directory\\ one/test.csv.gz"

    Hope this helps. Otherwise, the zcat solution works fine.

  15. jaapwalhout commented on Feb 27, 2018

    @jaapwalhout

    Another example on StackOverflow why this feature is needed:
    data.table fread error - gzip file - set temporary directory

  16. webbp commented on Mar 1, 2018

    @webbp

    How about:

    library(readr)
    DT = as.data.table(read_csv("myfile.gz"))
    
  17. mspivakov commented on Mar 1, 2018

    @mspivakov
  18. MichaelChirico commented on Mar 2, 2018

    @MichaelChirico
    Member
  19. malcook commented on Mar 18, 2018

    @malcook

    @frenchja - agree - though you might prefer to escape those spaces with R's shQuote

  20. byapparov commented on Apr 17, 2018

    @byapparov
  21. removed this from the milestone on May 10, 2018
  22. swvanderlaan commented on Jun 28, 2018

    @swvanderlaan

    @frenchja How would this work with past0? I have now the code below, but that throws an error:

    SOME_DIR = "/Users/swvanderlaan/some_dir"
    data <- fread('zcat < paste0(SOME_DIR,"/somedata.txt.gz")', 
                                                              header = TRUE, na.strings = "NA", 
                                                              verbose = TRUE, showProgress = TRUE)
    
    

    Ah got it, it should be this:

    data <- 
      fread(paste0("zcat < '", SOME_DIR,"/somedata.txt.gz","'"), 
                                                              header = TRUE, na.strings = "NA", 
                                                              verbose = TRUE, showProgress = TRUE)
    
  23. MichaelChirico commented on Jul 2, 2018

    @MichaelChirico
    Member

    @swvanderlaan I tend to use sprintf for cases like this; you should also use file.path and shQuote to be platform-robust:

    fread(sprintf('zcat %s', shQuote(file.path(SOME_DIR, 'somedata.txt.gz'))))
    
  24. added a commit that references this issue on Jul 4, 2018
  25. added this to the 1.11.8 milestone on Sep 29, 2018
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions