Repository navigation
fread: fill=true doesn't work. #4130
Description
Activity
- changed the title
[-]fread fill=true doesn't work. [/-][+]fread: fill=true doesn't work. [/+]on Dec 19, 2019 Hi @Befrancesco thanks for using
data.tableWe'll need a bit more to go on if we're to have any chance at solving your issue.Can you please run
freadagain withverbose=TRUEand report the output?Ideally, you would be able to share with us your file. Barring that, you could try to reproduce the issue on a smaller subset of the table, e.g. using
readLinesandsample()ing lines to find some set of rows for which the error results. With a smaller subset, it would be easier to anonymize the data so you could share a facsimile version.Thank you for your support :) :)
Unfotunaly I can't share my file, because is confidential.omp_get_num_procs()==8 R_DATATABLE_NUM_PROCS_PERCENT=="" (default 50) R_DATATABLE_NUM_THREADS=="" omp_get_thread_limit()==2147483647 omp_get_max_threads()==8 OMP_THREAD_LIMIT=="" OMP_NUM_THREADS=="" data.table is using 4 threads. This is set on startup, and by setDTthreads(). See ?setDTthreads. RestoreAfterFork==true Input contains no \n. Taking this to be a filename to open [01] Check arguments Using 4 threads (omp_get_max_threads()=8, nth=4) NAstrings = [<<NA>>] None of the NAstrings look like numbers. show progress = 1 0/1 column will be read as integer [02] Opening the file Opening file C:/xyz/file_names.csv File opened, size = 3.375GB (3624311815 bytes). Memory mapped ok [03] Detect and skip BOM [04] Arrange mmap to be \0 terminated \n has been found in the input and different lines can end with different line endings (e.g. mixed \n and \r\n in one file). This is common and ideal. [05] Skipping initial rows if needed Positioned on line 1 starting: <<xxx;yyy>> [06] Detect separator, quoting rule, and ncolumns Using supplied sep ';' sep=';' with 29 fields using quote rule 0 Detected 29 columns on line 1. This line is either column names or first data row. Line starts as: <<x;y>> Quote rule picked = 0 fill=true and the most number of columns found is 29 [07] Detect column types, good nrow estimate and whether first row is column names Number of sampling jump points = 100 because (3624311813 bytes from row 1 to eof) / (2 * 18462 jump0size) == 98155 Type codes (jump 000) : 55AA5557555555A2AAAAA55AA25AA Quote rule 0 Type codes (jump 001) : 55AA5557555555A2AAAAA55AAA5AA Quote rule 0 Type codes (jump 036) : 55AA5557555555A2AAAAA55AAAAAA Quote rule 0 A line with too-many fields (29/29) was found on line 22 of sample jump 43. Most likely this jump landed awkwardly so type bumps here will be skipped. Type codes (jump 100) : 55AA5557555555A2AAAAA55AAAAAA Quote rule 0 'header' determined to be true due to column 1 containing a string on row 1 and a lower type (int32) in the rest of the 9967 sample rows ===== Sampled 9967 rows (handled \n inside quoted fields) at 101 jump points Bytes from first data row on line 2 to the end of last row: 3624311480 Line length: mean=197.19 sd=10.25 min=166 max=252 Estimated number of rows: 3624311480 / 197.19 = 18379670 Initial alloc = 20512103 rows (18379670 + 11%) using bytes/max(mean-2*sd,min) clamped between [1.1*estn, 2.0*estn] ===== [08] Assign column names [09] Apply user overrides on column types After 0 type and 0 drop user overrides : 55AA5557555555A2AAAAA55AAAAAA [10] Allocate memory for the datatable Allocating 29 column slots (29 - 0 dropped) with 20512103 rows [11] Read the data jumps=[0..3456), chunk_size=1048701, total_size=3624311480 |--------------------------------------------------| |= Restarting team from jump 84. nSwept==0 quoteRule==1 jumps=[84..3456), chunk_size=1048701, total_size=3624311480 Restarting team from jump 84. nSwept==0 quoteRule==2 jumps=[84..3456), chunk_size=1048701, total_size=3624311480 Restarting team from jump 84. nSwept==0 quoteRule==3 jumps=[84..3456), chunk_size=1048701, total_size=3624311480 =================================================| jumps=[0..3456), chunk_size=1048701, total_size=3624311480 |--------------------------------------------------| |==================================================| Read 458943 rows x 29 columns from 3.375GB (3624311815 bytes) file in 00:01.926 wall clock time [12] Finalizing the datatable Type counts: 12 : int32 '5' 1 : float64 '7' 16 : string 'A' ============================= 0.001s ( 0%) Memory map 3.375GB file 0.008s ( 0%) sep=';' ncol=29 and header detection 0.000s ( 0%) Column type detection using 9967 sample rows 0.816s ( 42%) Allocation of 20512103 rows x 29 cols (3.362GB) of which 458943 ( 2%) rows used 1.101s ( 57%) Reading 3456 chunks (0 swept) of 1.000MB (each chunk 132 rows) using 4 threads + 0.132s ( 7%) Parse to row-major thread buffers (grown 0 times) + 0.695s ( 36%) Transpose + 0.273s ( 14%) Waiting 0.184s ( 10%) Rereading 2 columns due to out-of-sample type exceptions 1.926s Total Column 16 ("TITLE") bumped from 'bool8' to 'string' due to <<Frau>> on row 3856 Column 22 ("LOCATIONTYPE") bumped from 'int32' to 'string' due to <<12.05.1977>> on row 437392 Warning message: In fread("C:/x/xy/xyz/name_file.csv", : Stopped early on line 458945. Expected 29 fields but found 30. Consider fill=TRUE and comment.char=. First discarded non-empty line:I don't see an error message there?
Thank you and sorry.
I updated the error.@Befrancesco The error message seem to tell you that using the ; as a separator, it found 30 columns on line 458945.
Likely, there is a ; in one of your text fields, or that line is not correct.
If separators appear in the text columns, then normally that field is quoted. (surrounded by ") like this: "text with; in it". However, using " for quoting is the default in data.table.
You might want to inspect that line manually to see what is going on there.Reacted by Jan Gorecki
Good morning,
I'd like to manage the csv of 3.5 gb, but I can't because in my case the
fill=TRUEdoesn't work.My code is:
somebody can help me?
Thank you in advance.
.