Repository navigation
fread(integer64 = "double") not working for some data #2607
Description
Activity
DT = fread("A\n1.010203040506070809010203040506\n") # too precise for double, so read as character # TODO: add numerals=c("allow.loss", "warn.loss", "no.loss") from base::read.table typeof(DT$A)=="character" # TRUEI'm now using the latest release
1.11.0.Another data that causes exactly the same problem is as follows which further causes
rbindlistto fail:> p1 <- fread("~/data/fread-issue-sample.txt", integer64 = "double") > str(p1) Classes ‘data.table’ and 'data.frame': 8160 obs. of 2 variables: $ volume : int 13203 8201 8041 5391 7779 6079 5848 7249 6109 5406 ... $ turnover:integer64 53907549 33420919 32739957 21940344 31610610 24671207 23693960 29322253 ... - attr(*, ".internal.selfref")=<externalptr>Any idea on this @mattdowle?
When verbose mode is turned on, it shows the following:
... Column 2 ("turnover") bumped from 'int32' to 'int64' due to <<2402620023>> on row 2400Thus, this issue appears to be caused by #2749 : the "integer64" parameter is only applied during stage
[09] Apply user overrides on column types, but is not taken into account when an out-of-sample type bump occurs.Reacted by Matt Dowle and MattMaoYes
integer64=control is dealt with inuserOverride()currently. Looks like we'll need to passreadInt64Asdown tofread.cin order for it be used in out-of-sample type bump as well. It's not just a matter of disabling theint64parser unfortunately, because disabling it would only providereadInt64As="double"ability whereas the control allowsreadInt64As="character"too (skipping double in the type hierarchy).
If there might be similar requirements for other types, perhapsdisabled_parsersinfread.ccould be expanded. It is currentlyintholding 0/1 only. Instead it could hold the number of positions to skip.
So:readInt64As="integer64" => disabled_parsers[CT_INT64] == 0 readInt64As="double" => disabled_parsers[CT_INT64] == 1 readInt64As="character" => disabled_parsers[CT_INT64] == 4That 4 being due to needing to skip CT_FLOAT64, CT_FLOAT64_HEX and CT_FLOAT64_EXT to get to CT_STRING.
Or,disabled_parserscould hold which type to use instead. 0 would mean use that parser as it means now. Non zero value in positioniwould need to be>iotherwise a infinite loop would occur.readInt64As="integer64" => disabled_parsers[CT_INT64] == 0 readInt64As="double" => disabled_parsers[CT_INT64] == CT_FLOAT64 readInt64As="character" => disabled_parsers[CT_INT64] == CT_STRINGThis approach would save needing to maintain the skip values in
disabled_parsersas parsers are added and removed in future.I am experiencing the same issue in the release version (1.11.4): one of my columns is being bumped to integer64, despite
data.table::fread(..., integer64 = 'numeric').(If useful I could create a reproducible example, but it looks like you have that already.)
Reacted by Matt DowleI'm having the same problem with integer64 conversions. Did this problem get solved?
Reacted by Rahulis this related to int64 issues in
rbindlistor should it be addressed in another issue?Quick & Dirty workaround:
df[, names(.SD) := lapply(.SD, as.numeric), .SDcols = bit64::is.integer64]Reacted by Peter Reschenhofer, Nathan Mietkiewicz, XR97 and Elias L OhnebergIssue persists now in 1.17.0 and R version 4.4 with
integer64='double'or 'numeric'Confirming this is still present with a reprex:
set.seed(39439) DT = data.table(i = 1:1e5, x = as.integer64(sample(1e5))) DT[sample(1e5, 1), x := lim.integer64()[2]/2] fwrite(DT, tmp<-tempfile()) fread(tmp, integer64='double', verbose=TRUE) # <other output> [07] Detect column types, dec, good nrow estimate and whether first row is column names # <other output> # Type codes (jump 000) : 77 Quote rule 0 # Type codes (jump 100) : 77 Quote rule 0 # <other output> # [12] Finalizing the datatable # Type counts: # 1 : int32 '7' # 1 : string 'E' # <other output> # Column 2 <<x>> bumped from 'int32' to 'string' due to <<4611686018427387904>> on row 65130 # i x # <int> <char> # 1: 1 65323 # 2: 2 55334 # 3: 3 78673 # 4: 4 57248 # 5: 5 83957 # --- # 99996: 99996 47451 # 99997: 99997 54340 # 99998: 99998 25001 # 99999: 99999 56523 # 100000: 100000 52523
I'm testing the latest development version of data.table and find that
freaddoes not respectinteger64 = "double"for some of my data. There's no such problem in the release version.The test code is:
but the resulted
data.tablestill hasinteger64column:The data is attached below:
test_data.txt
The same happens on both macOS and Ubuntu as I tested.
Here's my session info: