Skip to content

fread from 1.11: self healing regression for fill = T when unmatched quote occurs #2859

Description

@christianhomberg

For the attached file, the following code will not throw a warning in data.table 1.11 whereas we get an expected warning with 1.10.4. Background: the file contains an unmatched quote in line 20000:

dt_2859 = data.table::fread("dt_2859.csv", fill = T)
This results in warning with following sessionInfo():
R version 3.4.4 (2018-03-15)
Platform: x86_64-pc-linux-gnu (64-bit)
Running under: Ubuntu 18.04 LTS

Matrix products: default
BLAS: /usr/lib/x86_64-linux-gnu/blas/libblas.so.3.7.1
LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.7.1

locale:
 [1] LC_CTYPE=en_GB.UTF-8       LC_NUMERIC=C              
 [3] LC_TIME=en_GB.UTF-8        LC_COLLATE=en_GB.UTF-8    
 [5] LC_MONETARY=en_GB.UTF-8    LC_MESSAGES=en_GB.UTF-8   
 [7] LC_PAPER=en_GB.UTF-8       LC_NAME=C                 
 [9] LC_ADDRESS=C               LC_TELEPHONE=C            
[11] LC_MEASUREMENT=en_GB.UTF-8 LC_IDENTIFICATION=C       

attached base packages:
[1] stats     graphics  grDevices utils     datasets  methods   base     

loaded via a namespace (and not attached):
[1] compiler_3.4.4      magrittr_1.5        tools_3.4.4        
[4] yaml_2.1.16         data.table_1.10.4-3 rlang_0.2.0.9001   
[7] purrr_0.2.4    
whereas `fread` stays quiet with following sessionInfo():
R version 3.4.4 (2018-03-15)
Platform: x86_64-pc-linux-gnu (64-bit)
Running under: Ubuntu 18.04 LTS

Matrix products: default
BLAS: /usr/lib/x86_64-linux-gnu/blas/libblas.so.3.7.1
LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.7.1

locale:
 [1] LC_CTYPE=en_GB.UTF-8       LC_NUMERIC=C              
 [3] LC_TIME=en_GB.UTF-8        LC_COLLATE=en_GB.UTF-8    
 [5] LC_MONETARY=en_GB.UTF-8    LC_MESSAGES=en_GB.UTF-8   
 [7] LC_PAPER=en_GB.UTF-8       LC_NAME=C                 
 [9] LC_ADDRESS=C               LC_TELEPHONE=C            
[11] LC_MEASUREMENT=en_GB.UTF-8 LC_IDENTIFICATION=C       

attached base packages:
[1] stats     graphics  grDevices utils     datasets  methods   base     

other attached packages:
[1] markovchain_0.6.9.8-1 shiny_1.0.5           data.table_1.11.2    
[4] magrittr_1.5          ggplot2_2.2.1        

loaded via a namespace (and not attached):
 [1] Rcpp_0.12.16        compiler_3.4.4      pillar_1.2.2        later_0.7.2        
 [5] plyr_1.8.4          tools_3.4.4         digest_0.6.15       packrat_0.4.9-2    
 [9] jsonlite_1.5        evaluate_0.10.1     tibble_1.4.2        gtable_0.2.0       
[13] lattice_0.20-35     pkgconfig_2.0.1     rlang_0.2.0.9001    Matrix_1.2-14      
[17] igraph_1.2.1        parallel_3.4.4      yaml_2.1.19         expm_0.999-2       
[21] stringr_1.3.0       knitr_1.20          stats4_3.4.4        rprojroot_1.3-2    
[25] grid_3.4.4          flexdashboard_0.5.1 R6_2.2.2            rmarkdown_1.9      
[29] matlab_1.0.2        purrr_0.2.4         codetools_0.2-15    backports_1.1.2    
[33] scales_0.5.0        promises_1.0.1      htmltools_0.3.6     mime_0.5           
[37] colorspace_1.3-2    xtable_1.8-2        httpuv_1.4.2        stringi_1.2.2      
[41] RcppParallel_4.4.0  lazyeval_0.2.1      munsell_0.4.3

In version 1.11 only the first 20000 rows are read and at least for me rstudio crashes when trying to print the resulting dt_2859. With version 1.10.4 however we get a warning and know to set quote = "".

Activity

  1. st-pasha commented on May 10, 2018

    @st-pasha
    Contributor

    Thanks Christian, this is an interesting example.

    I notice that the file is read correctly under the default setting fill=FALSE. There is even a warning:

    Warning message:
    In fread("/Users/pasha/Downloads/dt_2859.csv", verbose = T) :
      Found and resolved improper quoting out-of-sample. First healed line 20001: <<20000,"20000,20000,20000,20000,20000,20000,20000,20000,20000>>. If the fields are not quoted (e.g. field separator does not appear within any field), try quote="" to avoid this warning.
    

    But if you do need fill=TRUE, the problems start to appear.

    1. In the verbose mode, a message is printed Column 2 ("V2") bumped from 'int32' to 'string' due to <<"20000,20000,20000,20000 ..., and that ... contains almost the entire file. Lucky it's just a megabyte, not gigabyte. So the problem is that the error message "Column %d (\"%.*s\") bumped from '%s' to '%s' due to <<%.*s>> on row %llu\n" does not sanitize the field through strlim().
    2. That entire field, starting from "20000, ... until the end of file becomes cell [20000, 1] of the resulting DT, and the rest of the elements on row 20000 become NAs (because of fill=TRUE). The problem here is that if there are no closing quote anywhere in the file, the field should not have parsed under QR=0; and under QR=3 it wouldn't have gobbled any newlines.
    3. Presumably, printing such a large string element crashes RStudio (this could actually be RStudio's problem). It also looks quite horribly in R console (truncate string fields that are too long when printing a DT?).

    Lastly, I'd like to point out that using fill=TRUE with a file which contains invalid quotes is quite dangerous. If there was another quote anywhere in the file (and frankly I'm surprised there is only one!), then everything between the two quotes would be converted into a single field (because that's what the CSV standard demands). You'd need to specify explicitly quoteRule=3 to parse such a file, and that's only after #2768 is implemented...

  2. added this to the 1.11.4 milestone on May 11, 2018
  3. christianhomberg commented on May 11, 2018

    @christianhomberg
    Author

    Regarding 3.: It's not only printing that crashes RStudio. In my original files that resulted in the issue I fread multiple files in a purrr::map iteration and that causes R to crash, I actually don't execute the script in RStudio by default. The resulting exception is:

    *** caught segfault ***
    address (nil), cause 'memory not mapped'
    
    Traceback:
    1: fread(.x, fill = T, skip = 1, integer64 = "character")
    2: cbind(fread(.x, fill = T, skip = 1, integer64 = "character"),     foo = "bar")
    3: .f(.x[[i]], ...)
    4: .Call(map_impl, environment(), ".x", ".f", "list")
    5: purrr::map(file_list$filename, function(.x) {    if (file.size(.x)== 0) {        cat(paste0("Skipping empty file ", .x), sep = "\n")        return(NULL)    }    cat(paste0("Loading file ", .x), sep = "\n")    cbind(fread(.x, fill = T, skip = 1, integer64 = "character"),         foo = "bar")})
    6: eval(lhs, parent, parent)
    7: eval(lhs, parent, parent)
    8: purrr::map(file_list$filename, function(.x) {    if (file.size(.x)== 0) {        cat(paste0("Skipping empty file ", .x), sep = "\n")        return(NULL)    }    cat(paste0("Loading file ", .x), sep = "\n")    cbind(fread(.x, fill = T, skip = 1, integer64 = "character"),         foo = "bar")}) %>% rbindlist(fill = T)
    An irrecoverable exception occurred. R is aborting now ...
    

    I agree that unmatched quotes are extremely dangerous and could crash any csv reader by default. In my actual files there are multiple quotes but there are a few thousand lines between them. Which is obviously terrible but I never knew because it applies to a small subset of files. Also, in the past I was on a data.table version which was not 1.10.4 or 1.11 which read the files just fine even without manually setting quote = "" (I guess it must have guessed) but I don't know anymore which version that could have been (probably 1.10.X).

    Eventually, I'd expect fread to automatically set quote = "" by scanning the file, just as humans would.

  4. modified the milestones: 1.11.4, 1.11.6 on May 24, 2018
  5. modified the milestones: 1.12.0, 1.11.6 on Jun 6, 2018
  6. brooksambrose commented on Aug 29, 2018

    @brooksambrose

    This appears to be a similar case, though it crashes whether fill is TRUE or FALSE. The file is naturally unquoted, but has " within fields. Setting quote='' avoids the crash.

    library(data.table) 
    system('wget -qO jstor.txt http://www.jstor.org/kbart/collections/all-archive-titles?contentType=journals'))
    jstor<-fread('jstor.txt',quote='') # works as expected
    jstor<-fread('jstor.txt',fill=F) # crashes R session
    jstor<-fread('jstor.txt',fill=T) # crashes R session
    > devtools::session_info()
    Session info ----------------------------------------------------------------
     setting  value                       
     version  R version 3.4.4 (2018-03-15)
     system   x86_64, linux-gnu           
     ui       RStudio (1.1.419)           
     language (EN)                        
     collate  en_US.UTF-8                 
     tz       Etc/UTC                     
     date     2018-08-29                  
    
    Packages --------------------------------------------------------------------
     package    * version date       source                                
     base       * 3.4.4   2018-04-21 local                                 
     compiler     3.4.4   2018-04-21 local                                 
     data.table * 1.11.5  2018-08-29 Github (Rdatatable/data.table@8d4ce7a)
     datasets   * 3.4.4   2018-04-21 local                                 
     devtools     1.13.4  2017-11-09 CRAN (R 3.4.4)                        
     digest       0.6.16  2018-08-22 cran (@0.6.16)                        
     graphics   * 3.4.4   2018-04-21 local                                 
     grDevices  * 3.4.4   2018-04-21 local                                 
     memoise      1.1.0   2017-04-21 CRAN (R 3.4.4)                        
     methods    * 3.4.4   2018-04-21 local                                 
     stats      * 3.4.4   2018-04-21 local                                 
     tools        3.4.4   2018-04-21 local                                 
     utils      * 3.4.4   2018-04-21 local                                 
     withr        2.1.2   2018-03-15 CRAN (R 3.4.4)                        
     yaml         2.2.0   2018-07-25 CRAN (R 3.4.4)                        
    
  7. jangorecki commented on Sep 4, 2018

    @jangorecki
    Member

    @dhanashreedeshpande indeed, jstor is not related to json. Please read fread manual for more details: https://rdatatable.gitlab.io/data.table/library/data.table/html/fread.html

  8. jangorecki commented on Sep 5, 2018

    @jangorecki
    Member

    @dhanashreedeshpande it means that documentation has been written, and maintained for the purposed of answering such a basic questions, and even more complex. Please do not continue "json" topic in this issue as it is unrelated to json, so it is polluting original topic. fread does not support reading json format, currently there are no plans to support it, there were no feature requests for it.

  9. MichaelChirico commented on Sep 5, 2018

    @MichaelChirico
    Member

    @dhanashreedeshpande please see tools specialized for JSON format, e.g. jsonlite.

  10. modified the milestones: 1.11.6, 1.12.0 on Sep 20, 2018
  11. modified the milestones: 1.12.0, 1.12.2 on Jan 11, 2019
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions