Repository navigation
Support .gz file format for fread #717
Description
Activity
You can just do
fread('zcat file.gz'), or some loop variation, if you have many files.Reacted by Gergely Daróczi, Alexander Grueneberg, Bartosz Czernecki, Webb Phillips, Dmytro Lituiev, Artem Klevtsov, Nathan Trujillo, Peter Hollows, Tim, António Mendes and 8 more:Bump: quite useful (and coming up frequently).
Yes this would be a very useful feature to have. Using a command line, is a temporary solution at best, since it relies on the underlying system to have the tools for decompression. For instance 'zcat' is not available on windows unless one installs cygwin etc.
Since FRead is by far the best tool in R to read file, it would be a huge performance gain to read gzipped/bziped/... files directly.
I just saw @Arun's bump in my email, and literally a few hours ago I was ingesting 200+ such files. +1 for usefullness
Would have been useful here as well :
http://www.magesblog.com/2014/10/visualising-seasonality-of-atlantic.html
Will take a look.I agree with @gbonamy readding directly from zip files would be a fantastic addition!!
Reading from a connection with unz() would also be quite useful. I have a function that downloads a zip file, and only reads one file then throws it away. So if I could use fread(unz(zipfile, file = file)) it would be a great addition.
Reacted by Rafael H M PereiraI ++ about directly from gz files. I would personally use it every day.
+1 from me as well.
37 remaining items
One thing to note is that the
zcatsolution appears to only work if the file exists in the same directory that R is launched:Error in fread("zcat < data/directory/test.csv.gz") : File is empty: /var/folders/41/asdf_kj80000gn/T//RtmpwtAttt/fileebeb5e124cefReacted by James Pirruccelloi forgot about this issue and tried to fread a gz file, only to get a mysterious error causing me to waste time, again, searching for the solution.
3 years later, still waiting for this elementary fix.
Reacted by Webb Phillips, Zhenguo Zhang, B. Ogan Mancarcı and Tokhir DadaevAfter further exploration, my error above only occurs when there are spaces in the directory name:
fread("zcat < data/directory\ one/test.csv.gz"But not with underscores:
fread("zcat < data/directory_two/test.csv.gz"And can be alleviated by escaping the backslash again:
fread("zcat < data/directory\\ one/test.csv.gz"Hope this helps. Otherwise, the
zcatsolution works fine.Reacted by SplitInf, Inventodiscoveo and Cat TriandafillouAnother example on StackOverflow why this feature is needed:
data.table fread error - gzip file - set temporary directoryReacted by map2085, Webb Phillips and Tokhir DadaevHow about:
library(readr) DT = as.data.table(read_csv("myfile.gz"))- This is considerably slower.…-------- Original message -------- From: Webb Phillips Date:2018/03/01 6:26 PM (GMT+00:00) To: "Rdatatable/data.table" Cc: Mikhail Spivakov , Comment Subject: Re: [Rdatatable/data.table] Support .gz file format for fread (#717) How about: dt = as.data.table(read_csv("myfile.gz"))``` — You are receiving this because you commented. Reply to this email directly, view it on GitHub<#717 (comment)>, or mute the thread<https://github.com/notifications/unsubscribe-auth/ABlQ65ZJPuwpxCfmJTuRZ1aBPh4Jejniks5taD06gaJpZM4CKNWu>. The Babraham Institute, Babraham Research Campus, Cambridge CB22 3AT Registered Charity No. 1053902. The information transmitted in this email is directed only to the addressee. If you received this in error, please contact the sender and delete this email from your system. The contents of this e-mail are the views of the sender and do not necessarily represent the views of the Babraham Institute. Full conditions at: www.babraham.ac.uk<http://www.babraham.ac.uk/terms>
- setDT will be faster than as.data.table. What tool does read_csv use for uncracking the .gz?…On Mar 2, 2018 2:32 AM, "mspivakov" ***@***.***> wrote: This is considerably slower. -------- Original message -------- From: Webb Phillips Date:2018/03/01 6:26 PM (GMT+00:00) To: "Rdatatable/data.table" Cc: Mikhail Spivakov , Comment Subject: Re: [Rdatatable/data.table] Support .gz file format for fread (#717) How about: dt = as.data.table(read_csv("myfile.gz"))``` — You are receiving this because you commented. Reply to this email directly, view it on GitHub<https://github.com/ Rdatatable/data.table#717#issuecomment-369684424>, or mute the thread<https://github.com/notifications/unsubscribe-auth/ ABlQ65ZJPuwpxCfmJTuRZ1aBPh4Jejniks5taD06gaJpZM4CKNWu>. The Babraham Institute, Babraham Research Campus, Cambridge CB22 3AT Registered Charity No. 1053902. The information transmitted in this email is directed only to the addressee. If you received this in error, please contact the sender and delete this email from your system. The contents of this e-mail are the views of the sender and do not necessarily represent the views of the Babraham Institute. Full conditions at: www.babraham.ac.uk<http://www. babraham.ac.uk/terms> — You are receiving this because you are subscribed to this thread. Reply to this email directly, view it on GitHub <#717 (comment)>, or mute the thread <https://github.com/notifications/unsubscribe-auth/AHQQdZa08PZo9LKMGbnzMnmUeaspKxgVks5taD6wgaJpZM4CKNWu> .
readris reading data from connection (?gzfile) in memory: https://github.com/tidyverse/readr/blob/6f0bb65296afa55709fd60cdc5d59a4c89623e36/src/connection.cppAnd it is parsed with
read_tokens_: https://github.com/tidyverse/readr/blob/6f0bb65296afa55709fd60cdc5d59a4c89623e36/src/read.cpp@frenchja How would this work with
past0? I have now the code below, but that throws an error:SOME_DIR = "/Users/swvanderlaan/some_dir" data <- fread('zcat < paste0(SOME_DIR,"/somedata.txt.gz")', header = TRUE, na.strings = "NA", verbose = TRUE, showProgress = TRUE)Ah got it, it should be this:
data <- fread(paste0("zcat < '", SOME_DIR,"/somedata.txt.gz","'"), header = TRUE, na.strings = "NA", verbose = TRUE, showProgress = TRUE)@swvanderlaan I tend to use
sprintffor cases like this; you should also usefile.pathandshQuoteto be platform-robust:fread(sprintf('zcat %s', shQuote(file.path(SOME_DIR, 'somedata.txt.gz'))))- added a commit that references this issue
on Jul 4, 2018
I have several thousands of
.gzfiles containing data incsvformat - about 60GB in total in terms of.gzfiles. Decompressing them and load some pieces viafreadturns out a huge pain in the first step. I'm wonder whether it is possible to improve the functionality offreadso that it can read compressed file formats just asread.tabledoes?Perhaps file connection issues are highly relevant, as mentioned in #341, #543, and #561.
Some other reference:
http://stackoverflow.com/questions/5764499/decompress-gz-file-using-r
http://blog.revolutionanalytics.com/2009/12/r-tip-save-time-and-space-by-compressing-data-files.html