Skip to content

fread: fill=true doesn't work.  #4130

Description

@Befrancesco

Good morning,
I'd like to manage the csv of 3.5 gb, but I can't because in my case the fill=TRUE doesn't work.
My code is:

n_a<-fread("C:/xyz/file_names.csv",sep=";", fill = TRUE)

somebody can help me?

Thank you in advance.
.

Activity

  1. changed the title [-]fread fill=true doesn't work. [/-] [+]fread: fill=true doesn't work. [/+] on Dec 19, 2019
  2. MichaelChirico commented on Dec 19, 2019

    @MichaelChirico
    Member

    Hi @Befrancesco thanks for using data.table We'll need a bit more to go on if we're to have any chance at solving your issue.

    Can you please run fread again with verbose=TRUE and report the output?

    Ideally, you would be able to share with us your file. Barring that, you could try to reproduce the issue on a smaller subset of the table, e.g. using readLines and sample()ing lines to find some set of rows for which the error results. With a smaller subset, it would be easier to anonymize the data so you could share a facsimile version.

  3. Befrancesco commented on Dec 19, 2019

    @Befrancesco
    Author

    Thank you for your support :) :)
    Unfotunaly I can't share my file, because is confidential.

    omp_get_num_procs()==8
    R_DATATABLE_NUM_PROCS_PERCENT=="" (default 50)
    R_DATATABLE_NUM_THREADS==""
    omp_get_thread_limit()==2147483647
    omp_get_max_threads()==8
    OMP_THREAD_LIMIT==""
    OMP_NUM_THREADS==""
    data.table is using 4 threads. This is set on startup, and by setDTthreads(). See ?setDTthreads.
    RestoreAfterFork==true
    Input contains no \n. Taking this to be a filename to open
    [01] Check arguments
      Using 4 threads (omp_get_max_threads()=8, nth=4)
      NAstrings = [<<NA>>]
      None of the NAstrings look like numbers.
      show progress = 1
      0/1 column will be read as integer
    [02] Opening the file
      Opening file C:/xyz/file_names.csv
      File opened, size = 3.375GB (3624311815 bytes).
      Memory mapped ok
    [03] Detect and skip BOM
    [04] Arrange mmap to be \0 terminated
      \n has been found in the input and different lines can end with different line endings (e.g. mixed \n and \r\n in one file). This is common and ideal.
    [05] Skipping initial rows if needed
      Positioned on line 1 starting: <<xxx;yyy>>
    [06] Detect separator, quoting rule, and ncolumns
      Using supplied sep ';'
      sep=';'  with 29 fields using quote rule 0
      Detected 29 columns on line 1. This line is either column names or first data row. Line starts as: <<x;y>>
      Quote rule picked = 0
      fill=true and the most number of columns found is 29
    [07] Detect column types, good nrow estimate and whether first row is column names
      Number of sampling jump points = 100 because (3624311813 bytes from row 1 to eof) / (2 * 18462 jump0size) == 98155
      Type codes (jump 000)    : 55AA5557555555A2AAAAA55AA25AA  Quote rule 0
      Type codes (jump 001)    : 55AA5557555555A2AAAAA55AAA5AA  Quote rule 0
      Type codes (jump 036)    : 55AA5557555555A2AAAAA55AAAAAA  Quote rule 0
      A line with too-many fields (29/29) was found on line 22 of sample jump 43. Most likely this jump landed awkwardly so type bumps here will be skipped.
      Type codes (jump 100)    : 55AA5557555555A2AAAAA55AAAAAA  Quote rule 0
      'header' determined to be true due to column 1 containing a string on row 1 and a lower type (int32) in the rest of the 9967 sample rows
      =====
      Sampled 9967 rows (handled \n inside quoted fields) at 101 jump points
      Bytes from first data row on line 2 to the end of last row: 3624311480
      Line length: mean=197.19 sd=10.25 min=166 max=252
      Estimated number of rows: 3624311480 / 197.19 = 18379670
      Initial alloc = 20512103 rows (18379670 + 11%) using bytes/max(mean-2*sd,min) clamped between [1.1*estn, 2.0*estn]
      =====
    [08] Assign column names
    [09] Apply user overrides on column types
      After 0 type and 0 drop user overrides : 55AA5557555555A2AAAAA55AAAAAA
    [10] Allocate memory for the datatable
      Allocating 29 column slots (29 - 0 dropped) with 20512103 rows
    [11] Read the data
      jumps=[0..3456), chunk_size=1048701, total_size=3624311480
    |--------------------------------------------------|
    |=  Restarting team from jump 84. nSwept==0 quoteRule==1
      jumps=[84..3456), chunk_size=1048701, total_size=3624311480
      Restarting team from jump 84. nSwept==0 quoteRule==2
      jumps=[84..3456), chunk_size=1048701, total_size=3624311480
      Restarting team from jump 84. nSwept==0 quoteRule==3
      jumps=[84..3456), chunk_size=1048701, total_size=3624311480
    =================================================|
      jumps=[0..3456), chunk_size=1048701, total_size=3624311480
    |--------------------------------------------------|
    |==================================================|
    Read 458943 rows x 29 columns from 3.375GB (3624311815 bytes) file in 00:01.926 wall clock time
    [12] Finalizing the datatable
      Type counts:
            12 : int32     '5'
             1 : float64   '7'
            16 : string    'A'
    =============================
       0.001s (  0%) Memory map 3.375GB file
       0.008s (  0%) sep=';' ncol=29 and header detection
       0.000s (  0%) Column type detection using 9967 sample rows
       0.816s ( 42%) Allocation of 20512103 rows x 29 cols (3.362GB) of which 458943 (  2%) rows used
       1.101s ( 57%) Reading 3456 chunks (0 swept) of 1.000MB (each chunk 132 rows) using 4 threads
       +    0.132s (  7%) Parse to row-major thread buffers (grown 0 times)
       +    0.695s ( 36%) Transpose
       +    0.273s ( 14%) Waiting
       0.184s ( 10%) Rereading 2 columns due to out-of-sample type exceptions
       1.926s        Total
    Column 16 ("TITLE") bumped from 'bool8' to 'string' due to <<Frau>> on row 3856
    Column 22 ("LOCATIONTYPE") bumped from 'int32' to 'string' due to <<12.05.1977>> on row 437392
    Warning message:
    In fread("C:/x/xy/xyz/name_file.csv",  :
      Stopped early on line 458945. Expected 29 fields but found 30. Consider fill=TRUE and comment.char=. First discarded non-empty line: 
    
    
    
  4. MichaelChirico commented on Dec 19, 2019

    @MichaelChirico
    Member

    I don't see an error message there?

  5. Befrancesco commented on Dec 19, 2019

    @Befrancesco
    Author

    Thank you and sorry.
    I updated the error.

  6. wligtenberg commented on Dec 20, 2019

    @wligtenberg
    Contributor

    @Befrancesco The error message seem to tell you that using the ; as a separator, it found 30 columns on line 458945.
    Likely, there is a ; in one of your text fields, or that line is not correct.
    If separators appear in the text columns, then normally that field is quoted. (surrounded by ") like this: "text with; in it". However, using " for quoting is the default in data.table.
    You might want to inspect that line manually to see what is going on there.

  7. added this to the 1.16.0 milestone on Jan 5, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions