Repository navigation
fwrite(): final items #1664
Description
Activity
Do people actually like having
quote=TRUEwhen writing to csv? I find it to be a big nuisance and would much prefer forfwriteto havequote=FALSEby default.Reacted by Arun Srinivasan, Dmitry Selivanov and Nathan MietkiewiczI find
quote = TRUEto be more robust -- you never know when you have aJAMES SMITH, JRin acharactercolumn and it can be a huge pain to get a .csv read when it has nuisance commas strewn about.Reacted by Otto Seiskari and Etan Schwartz@eantonya Agree. I prefer quote=FALSE too. The base R thinking I believe has numbers/ids with leading 0's stored as character format ... the default ensures they get read by Excel as character and the leading 0's not lost. But fwrite could detect that situation and quote just that situation by default. Where character columns contain letters and no embedded quotes, I really don't see why quotes are needed. Plus we save a bit on file size by saving the 2 extra quotes per field.
Reacted by Arun Srinivasan and smcinerney@MichaelChirico Agree with you too. fwrite can detect that and put the quotes in those situations. fwrite already does a first-pass through all strings to calculate maximum line length before allocating buffer sizes. It could test if there are any
seporquotein the string at that point. So I guess I'm suggestingquote='auto'by default.Reacted by Michael Chirico, Arun Srinivasan and eduard@mattdowle great, good point. Should only marginally affect speed then.
PS IIRC Excel converts "001" to 1 anyway :|
@MichaelChirico Now you mention it I do seem to remember Excel doing that. I haven't used Excel for many years now thankfully.
I would assume Excel behave in an inconsistent (os versions, office versions, os locales, office localces, 365s, etc.) way about that matter.
Reacted by Dmitry SelivanovAre you planning to include the
append = T? Please?! Anyway, congrats for the great job withdata.tablethat will become even greater withfwrite()!@rafapereirabr?
append = TRUEin fact works for me, are you suggesting that should be the default?@MichaelChirico , I didn't know it was already implemented ! I couldn't try it as I was planning to test it tonight. Just ignore my comment then. ps. I don't think this should be the default.
27 remaining items
excellent stuff Matt, thanks so much!!
Reacted by Matt Dowle- added a commit that references this issue
on Nov 11, 2016 Do we need to use "library(bit64)" with fwrite and fread when we have long numbers or not anymore?
Reacted by smcinerneyThank you very much for your work, this feature is really useful.
But last version looks not very stable.I caught 2 strange issues:
- showProgress should be set explicitly:
a <- c("1", "2", "3", "4", "5") d <- rep("2016-11-21", 5) c <- rep("a", 5) m <- rep("0.5", 5) data<-data.table(a, d, c, m) fwrite(data, "e:/tmp_buf/tmp.csv", sep="~", col.names=FALSE, append=FALSE, ..turbo = T, quote = F) #Error: isLOGICAL(showProgress) is not TRUE- eol delimiter does not work correctly
fwrite(data, "e:/tmp_buf/tmp.csv", sep="~", col.names=FALSE, append=FALSE, ..turbo = T, quote = F, showProgress = T) #result: #"1"~"2016-11-21"~"a"~"0.5""2"~"2016-11-21"~"a"~"0.5""3"~"2016-11-21"~"a"~"0.5""4"~"2016-11-21"~"a"~"0.5""5"~"2016-11-21"~"a"~"0.5" #No eol delimeters fwrite(data, "e:/tmp_buf/tmp.csv", sep="~", eol = "\r\n", col.names=FALSE, append=FALSE, ..turbo = T, quote = F, showProgress = T) #result: #"1"~"2016-11-21"~"a"~"0.5""2"~"2016-11-21"~"a"~"0.5""3"~"2016-11-21"~"a"~"0.5""4"~"2016-11-21"~"a"~"0.5""5"~"2016-11-21"~"a"~"0.5" #Still no eol delimeters- Also I would like to know, how can I write .csv files without scientific notation. For example,
93434234223523523.5converts to9.34342342235235E+016. But if I want to use this file for bulk insert I'll have problems. Can I set explicitly number of decimal places?
@stanislav-a It would be useful if you could provide your
sessionInfo()andread.dcf(system.file("DESCRIPTION", package="data.table"), "Commit"). And ideally re-run on latest version as there were lots of improvements made recently. I'm on linux and cannot reproduce problems you reported. Re scientific notation, it will round your number93434234223523523.5on writing, similarly towrite.csv. Exact floating point range is mentioned in manual?fwrite.
@skanskan how do you store your long numbers without bit64? if you store it as double, it will be processed as double and you don't need bit64.@jangorecki I reinstall package with last commit, now it works fine, thank you.
Reacted by Matt DowleI can confirm the first issue @stanislav-a has mentioned, I'm on the 1.9.8 release.
I use the following generated file, it's a simple csv file: https://gist.github.com/thvasilo/6edffdccda87f09572cbc4184662af47
surv_1k <- fread("surv_1k.csv") fwrite(surv_1k, "copy.csv") # Error: isLOGICAL(showProgress) is not TRUE
Session info:
> sessionInfo() R version 3.3.2 (2016-10-31) Platform: x86_64-pc-linux-gnu (64-bit) Running under: Ubuntu 16.04.1 LTS locale: [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C LC_TIME=en_US.UTF-8 [4] LC_COLLATE=en_US.UTF-8 LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8 [7] LC_PAPER=en_US.UTF-8 LC_NAME=C LC_ADDRESS=C [10] LC_TELEPHONE=C LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C attached base packages: [1] stats graphics grDevices utils datasets methods base other attached packages: [1] purrr_0.2.2 caret_6.0-73 ggplot2_2.2.0 lattice_0.20-34 data.table_1.9.8 loaded via a namespace (and not attached): [1] Rcpp_0.12.7 magrittr_1.5 splines_3.3.2 MASS_7.3-45 [5] munsell_0.4.3 colorspace_1.2-6 foreach_1.4.3 minqa_1.2.4 [9] stringr_1.1.0 car_2.1-4 plyr_1.8.4 tools_3.3.2 [13] parallel_3.3.2 nnet_7.3-12 pbkrtest_0.4-6 grid_3.3.2 [17] gtable_0.2.0 nlme_3.1-128 mgcv_1.8-16 quantreg_5.29 [21] MatrixModels_0.4-1 iterators_1.0.8 lme4_1.1-12 lazyeval_0.2.0 [25] assertthat_0.1 tibble_1.2 Matrix_1.2-7.1 nloptr_1.0.4 [29] reshape2_1.4.2 ModelMetrics_1.1.0 codetools_0.2-15 stringi_1.1.1 [33] scales_0.4.1 stats4_3.3.2 SparseM_1.74I haven't tried the latest master.
- please update, Matt just fixed this…On Dec 2, 2016 6:51 AM, "Theodore Vasiloudis" ***@***.***> wrote: I can confirm the first issue @stanislav-a <https://github.com/stanislav-a> has mentioned, I'm on the 1.9.8 release. I use the following generated file, it's a simple csv file: https://gist.github.com/thvasilo/6edffdccda87f09572cbc4184662af47 surv_1k <- fread("surv_1k.csv") fwrite(surv_1k, "copy.csv") # Error: isLOGICAL(showProgress) is not TRUE Session info: > sessionInfo() R version 3.3.2 (2016-10-31) Platform: x86_64-pc-linux-gnu (64-bit) Running under: Ubuntu 16.04.1 LTS locale: [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C LC_TIME=en_US.UTF-8 [4] LC_COLLATE=en_US.UTF-8 LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8 [7] LC_PAPER=en_US.UTF-8 LC_NAME=C LC_ADDRESS=C [10] LC_TELEPHONE=C LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C attached base packages: [1] stats graphics grDevices utils datasets methods base other attached packages: [1] purrr_0.2.2 caret_6.0-73 ggplot2_2.2.0 lattice_0.20-34 data.table_1.9.8 loaded via a namespace (and not attached): [1] Rcpp_0.12.7 magrittr_1.5 splines_3.3.2 MASS_7.3-45 [5] munsell_0.4.3 colorspace_1.2-6 foreach_1.4.3 minqa_1.2.4 [9] stringr_1.1.0 car_2.1-4 plyr_1.8.4 tools_3.3.2 [13] parallel_3.3.2 nnet_7.3-12 pbkrtest_0.4-6 grid_3.3.2 [17] gtable_0.2.0 nlme_3.1-128 mgcv_1.8-16 quantreg_5.29 [21] MatrixModels_0.4-1 iterators_1.0.8 lme4_1.1-12 lazyeval_0.2.0 [25] assertthat_0.1 tibble_1.2 Matrix_1.2-7.1 nloptr_1.0.4 [29] reshape2_1.4.2 ModelMetrics_1.1.0 codetools_0.2-15 stringi_1.1.1 [33] scales_0.4.1 stats4_3.3.2 SparseM_1.74 I haven't tried the latest master. — You are receiving this because you were mentioned. Reply to this email directly, view it on GitHub <#1664 (comment)>, or mute the thread <https://github.com/notifications/unsubscribe-auth/AHQQdTbq8P9plEY4BtUrdCKBRDlYvQP2ks5rEAY5gaJpZM4IL7OV> .
I just upgraded to 1.9.8 today and still have the same issue.
When I try and save a csv file using fwrite I still get "# Error: isLOGICAL(showProgress) is not TRUE"R version 3.3.2 (2016-10-31)
Platform: x86_64-w64-mingw32/x64 (64-bit)
Running under: Windows 7 x64 (build 7601) Service Pack 1locale:
[1] LC_COLLATE=English_United States.1252 LC_CTYPE=English_United States.1252 LC_MONETARY=English_United States.1252 LC_NUMERIC=C
[5] LC_TIME=English_United States.1252attached base packages:
[1] stats graphics grDevices utils datasets methods baseother attached packages:
[1] xtable_1.8-2 lubridate_1.6.0 ggrepel_0.6.3 data.table_1.10.0 cowplot_0.7.0 ggplot2_2.2.0 RPostgreSQL_0.4-1 DBI_0.5-1loaded via a namespace (and not attached):
[1] Rcpp_0.12.8 assertthat_0.1 grid_3.3.2 plyr_1.8.4 gtable_0.2.0 magrittr_1.5 scales_0.4.1 stringi_1.1.2 lazyeval_0.2.0 tools_3.3.2
[11] stringr_1.1.0 munsell_0.4.3 colorspace_1.3-0 knitr_1.15 tibble_1.2@MichaelChirico David is already on 1.10 according to session info.
@david-awam-jansen Please open new issue with you report. If possible include code to reproduce (at least on your machine), but please include only relevant part. Now I see in your session info you have many other unrelated packages loaded. Before reporting it is always good to ensure that issue is reproducible in clean session in R console. #1111 is same issue but onfread, you may try one solution from there:what solved the problem is the closing of all R sessions running on the computer before installing data.table
Reacted by Michael Chirico, Matt Dowle and George Walker
Include R_CheckUserInterrupt and test closes worker team ok1a4263fAdd "NEW:" item to startup bannersetthreadstosetDTthreadsso as not to affect other packages using OpenMP, add reference to its manual in fwrite man