Skip to content

fwrite(): final items #1664

Description

@mattdowle
  • malloc()s and write() need a thread safe way to error() if fail
  • Add progress % and ETA to appear after 2 seconds if more than 2 seconds are remaining
    Include R_CheckUserInterrupt and test closes worker team ok 1a4263f
  • Date, IDate and update this quesion
  • POSIXct (see fread/fwrite *base data types* directly for efficiency #1656)
  • ITime
  • integer64
    Add "NEW:" item to startup banner
  • Confirm fwrite() writes 10GB ok on Windows (it should do) to ensure 'big' file > 4GB ok. Thanks to Hugh Parsonage for testing as we don't have Windows other than via AppVeyor for test suite.
  • quote = 'auto'
  • sep2 (see fwrite to support secondary separator (sep2) for vectors in list columns #806)
  • match write.csv's scientific/decimal format exactly and add many tests. 6c1ed96
  • add dec='.' to R level and connect to already existing option at C level
  • add row.names for data.frames, default FALSE.
  • rename setthreads to setDTthreads so as not to affect other packages using OpenMP, add reference to its manual in fwrite man
  • refine ?fwrite
  • more tests qmethod= 'double' and 'escape'

Activity

  1. added this to the v1.9.8 milestone on Apr 20, 2016
  2. changed the title [-]fwrite final items[/-] [+]fwrite(): final items[/+] on Apr 20, 2016
  3. eantonya commented on Apr 20, 2016

    @eantonya
    Contributor

    Do people actually like having quote=TRUE when writing to csv? I find it to be a big nuisance and would much prefer for fwrite to have quote=FALSE by default.

  4. MichaelChirico commented on Apr 20, 2016

    @MichaelChirico
    Member

    I find quote = TRUE to be more robust -- you never know when you have a JAMES SMITH, JR in a character column and it can be a huge pain to get a .csv read when it has nuisance commas strewn about.

  5. mattdowle commented on Apr 20, 2016

    @mattdowle
    MemberAuthor

    @eantonya Agree. I prefer quote=FALSE too. The base R thinking I believe has numbers/ids with leading 0's stored as character format ... the default ensures they get read by Excel as character and the leading 0's not lost. But fwrite could detect that situation and quote just that situation by default. Where character columns contain letters and no embedded quotes, I really don't see why quotes are needed. Plus we save a bit on file size by saving the 2 extra quotes per field.

  6. mattdowle commented on Apr 20, 2016

    @mattdowle
    MemberAuthor

    @MichaelChirico Agree with you too. fwrite can detect that and put the quotes in those situations. fwrite already does a first-pass through all strings to calculate maximum line length before allocating buffer sizes. It could test if there are any sep or quote in the string at that point. So I guess I'm suggesting quote='auto' by default.

  7. MichaelChirico commented on Apr 20, 2016

    @MichaelChirico
    Member

    @mattdowle great, good point. Should only marginally affect speed then.

    PS IIRC Excel converts "001" to 1 anyway :|

  8. mattdowle commented on Apr 20, 2016

    @mattdowle
    MemberAuthor

    @MichaelChirico Now you mention it I do seem to remember Excel doing that. I haven't used Excel for many years now thankfully.

  9. jangorecki commented on Apr 20, 2016

    @jangorecki
    Member

    I would assume Excel behave in an inconsistent (os versions, office versions, os locales, office localces, 365s, etc.) way about that matter.

  10. rafapereirabr commented on Apr 25, 2016

    @rafapereirabr

    Are you planning to include the append = T ? Please?! Anyway, congrats for the great job with data.table that will become even greater with fwrite() !

  11. MichaelChirico commented on Apr 25, 2016

    @MichaelChirico
    Member

    @rafapereirabr? append = TRUE in fact works for me, are you suggesting that should be the default?

  12. rafapereirabr commented on Apr 25, 2016

    @rafapereirabr

    @MichaelChirico , I didn't know it was already implemented ! I couldn't try it as I was planning to test it tonight. Just ignore my comment then. ps. I don't think this should be the default.

  13. 27 remaining items

  14. added 2 commits that reference this issue on Nov 8, 2016
  15. MichaelChirico commented on Nov 11, 2016

    @MichaelChirico
    Member

    excellent stuff Matt, thanks so much!!

  16. juanpide commented on Nov 18, 2016

    @juanpide

    Do we need to use "library(bit64)" with fwrite and fread when we have long numbers or not anymore?

  17. stanislav-a commented on Nov 21, 2016

    @stanislav-a

    Thank you very much for your work, this feature is really useful.
    But last version looks not very stable.

    I caught 2 strange issues:

    • showProgress should be set explicitly:
    a <- c("1", "2", "3", "4", "5")
    d <- rep("2016-11-21", 5)
    c <- rep("a", 5)
    m <- rep("0.5", 5)
    
    data<-data.table(a, d, c, m)
    
    fwrite(data, "e:/tmp_buf/tmp.csv", sep="~",
           col.names=FALSE, append=FALSE, ..turbo = T, quote = F)
    
    
    #Error: isLOGICAL(showProgress) is not TRUE
    
    • eol delimiter does not work correctly
    fwrite(data, "e:/tmp_buf/tmp.csv", sep="~",
           col.names=FALSE, append=FALSE, ..turbo = T, quote = F, showProgress = T)
    
    #result:
    #"1"~"2016-11-21"~"a"~"0.5""2"~"2016-11-21"~"a"~"0.5""3"~"2016-11-21"~"a"~"0.5""4"~"2016-11-21"~"a"~"0.5""5"~"2016-11-21"~"a"~"0.5"
    #No eol delimeters
    
    fwrite(data, "e:/tmp_buf/tmp.csv", sep="~",
           eol = "\r\n",
           col.names=FALSE, append=FALSE, ..turbo = T, quote = F, showProgress = T)
    #result:
    #"1"~"2016-11-21"~"a"~"0.5""2"~"2016-11-21"~"a"~"0.5""3"~"2016-11-21"~"a"~"0.5""4"~"2016-11-21"~"a"~"0.5""5"~"2016-11-21"~"a"~"0.5"
    #Still no eol delimeters
    
    • Also I would like to know, how can I write .csv files without scientific notation. For example, 93434234223523523.5 converts to 9.34342342235235E+016. But if I want to use this file for bulk insert I'll have problems. Can I set explicitly number of decimal places?
  18. jangorecki commented on Nov 21, 2016

    @jangorecki
    Member

    @stanislav-a It would be useful if you could provide your sessionInfo() and read.dcf(system.file("DESCRIPTION", package="data.table"), "Commit"). And ideally re-run on latest version as there were lots of improvements made recently. I'm on linux and cannot reproduce problems you reported. Re scientific notation, it will round your number 93434234223523523.5 on writing, similarly to write.csv. Exact floating point range is mentioned in manual ?fwrite.
    @skanskan how do you store your long numbers without bit64? if you store it as double, it will be processed as double and you don't need bit64.

  19. stanislav-a commented on Nov 21, 2016

    @stanislav-a

    @jangorecki I reinstall package with last commit, now it works fine, thank you.

  20. thvasilo commented on Dec 2, 2016

    @thvasilo

    I can confirm the first issue @stanislav-a has mentioned, I'm on the 1.9.8 release.

    I use the following generated file, it's a simple csv file: https://gist.github.com/thvasilo/6edffdccda87f09572cbc4184662af47

    surv_1k <- fread("surv_1k.csv")
    
    fwrite(surv_1k, "copy.csv")
    
    # Error: isLOGICAL(showProgress) is not TRUE

    Session info:

    > sessionInfo()
    R version 3.3.2 (2016-10-31)
    Platform: x86_64-pc-linux-gnu (64-bit)
    Running under: Ubuntu 16.04.1 LTS
    
    locale:
     [1] LC_CTYPE=en_US.UTF-8       LC_NUMERIC=C               LC_TIME=en_US.UTF-8       
     [4] LC_COLLATE=en_US.UTF-8     LC_MONETARY=en_US.UTF-8    LC_MESSAGES=en_US.UTF-8   
     [7] LC_PAPER=en_US.UTF-8       LC_NAME=C                  LC_ADDRESS=C              
    [10] LC_TELEPHONE=C             LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C       
    
    attached base packages:
    [1] stats     graphics  grDevices utils     datasets  methods   base     
    
    other attached packages:
    [1] purrr_0.2.2      caret_6.0-73     ggplot2_2.2.0    lattice_0.20-34  data.table_1.9.8
    
    loaded via a namespace (and not attached):
     [1] Rcpp_0.12.7        magrittr_1.5       splines_3.3.2      MASS_7.3-45       
     [5] munsell_0.4.3      colorspace_1.2-6   foreach_1.4.3      minqa_1.2.4       
     [9] stringr_1.1.0      car_2.1-4          plyr_1.8.4         tools_3.3.2       
    [13] parallel_3.3.2     nnet_7.3-12        pbkrtest_0.4-6     grid_3.3.2        
    [17] gtable_0.2.0       nlme_3.1-128       mgcv_1.8-16        quantreg_5.29     
    [21] MatrixModels_0.4-1 iterators_1.0.8    lme4_1.1-12        lazyeval_0.2.0    
    [25] assertthat_0.1     tibble_1.2         Matrix_1.2-7.1     nloptr_1.0.4      
    [29] reshape2_1.4.2     ModelMetrics_1.1.0 codetools_0.2-15   stringi_1.1.1     
    [33] scales_0.4.1       stats4_3.3.2       SparseM_1.74     
    

    I haven't tried the latest master.

  21. MichaelChirico commented on Dec 2, 2016

    @MichaelChirico
    Member
  22. david-awam-jansen commented on Dec 5, 2016

    @david-awam-jansen

    I just upgraded to 1.9.8 today and still have the same issue.
    When I try and save a csv file using fwrite I still get "# Error: isLOGICAL(showProgress) is not TRUE"

    R version 3.3.2 (2016-10-31)
    Platform: x86_64-w64-mingw32/x64 (64-bit)
    Running under: Windows 7 x64 (build 7601) Service Pack 1

    locale:
    [1] LC_COLLATE=English_United States.1252 LC_CTYPE=English_United States.1252 LC_MONETARY=English_United States.1252 LC_NUMERIC=C
    [5] LC_TIME=English_United States.1252

    attached base packages:
    [1] stats graphics grDevices utils datasets methods base

    other attached packages:
    [1] xtable_1.8-2 lubridate_1.6.0 ggrepel_0.6.3 data.table_1.10.0 cowplot_0.7.0 ggplot2_2.2.0 RPostgreSQL_0.4-1 DBI_0.5-1

    loaded via a namespace (and not attached):
    [1] Rcpp_0.12.8 assertthat_0.1 grid_3.3.2 plyr_1.8.4 gtable_0.2.0 magrittr_1.5 scales_0.4.1 stringi_1.1.2 lazyeval_0.2.0 tools_3.3.2
    [11] stringr_1.1.0 munsell_0.4.3 colorspace_1.3-0 knitr_1.15 tibble_1.2

  23. jangorecki commented on Dec 5, 2016

    @jangorecki
    Member

    @MichaelChirico David is already on 1.10 according to session info.
    @david-awam-jansen Please open new issue with you report. If possible include code to reproduce (at least on your machine), but please include only relevant part. Now I see in your session info you have many other unrelated packages loaded. Before reporting it is always good to ensure that issue is reproducible in clean session in R console. #1111 is same issue but on fread, you may try one solution from there:

    what solved the problem is the closing of all R sessions running on the computer before installing data.table

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions