Repository navigation
fread handling of row names with header=TRUE #5634
Description
Activity
I may be wrong here, but
freadis for regular "delimited" files. The text you've supplied falls under "fixed width file", and whilefreadis able to read it, it might just be because it got lucky.Reacted by Toby Dylan HockingIf you convert fixed width to tabular (single space separating each field, see below) it works, so I would suggest the input data should be fixed by the user before passing to fread. If that is ok can you please close @MLopez-Ibanez ?
> fread(text=gsub(" +", " ", txt)) V1 foo foo2 <int> <int> <lgcl> 1: 1 0 FALSE 2: 2 1 NA Warning message: In fread(text = gsub(" +", " ", txt)) : Detected 2 column names but the data has 3 columns (i.e. invalid file). Added 1 extra default column name for the first column which is guessed to be row names or an index. Use setnames() afterwards if this guess is not correct, or fix the file write command that created the file to create a valid file.
The docs ?fread do not mention support of fixed width files, maybe we could add a sentence that says "fixed width files are not supported by fread" or similar?
I see. Then the first line in the documentation is misleading (https://rdatatable.gitlab.io/data.table/reference/fread.html): "Similar to read.table but faster and more convenient."
It should say similar to
read.delim().- added a commit that references this issue
on May 4, 2023 Actually, on second thought, this may be a bug, at least with respect to the current documentation
header: Does the first data line contain column names? Defaults according to whether every non-empty field on the first data line is type character. If so, or TRUE is supplied, any empty column names are given a default name. ... strip.white: default is 'TRUE'. Strips leading and trailing whitespaces of unquoted fields. If 'FALSE', only header trailing spaces are removed.It seems like there is a special case when sep=" " and this should be added to the documentation. For example after reading the doc above for
strip.whiteargument, I expected the following to have some spaces in the data, but there are none:> str(fread(txt,strip.white=FALSE,colClasses="character")) Classes 'data.table' and 'data.frame': 2 obs. of 3 variables: $ V1 : chr "1" "2" $ foo : chr "0" "1" $ foo2: chr "false" NA - attr(*, ".internal.selfref")=<externalptr> Warning message: In fread(txt, strip.white = FALSE, colClasses = "character") : Detected 2 column names but the data has 3 columns (i.e. invalid file). Added 1 extra default column name for the first column which is guessed to be row names or an index. Use setnames() afterwards if this guess is not correct, or fix the file write command that created the file to create a valid file.Since you have columns with a lower amount of cells, I think you are supposed to use
fill=TRUElibrary(data.table) txt <- ' foo foo2 1 0 false 2 1 NA ' fread(text=txt, fill=TRUE) #> foo foo2 V3 #> 1: 1 0 FALSE #> 2: 2 1 NA
Ofc this doesn't give you your desired output. An option would be to implement different padding strategies for
filllikefront/back.Since
read.delimis just a wrapper aroundread.tablesetting different default arguments I do not think it is a good idea to change the doc in this direction. Especially when you consider that probably many users probably knowread.tablebut have never heard ofread.delim.Reacted by Jan Gorecki and Toby Dylan HockingSince
read.delimis just a wrapper aroundread.tablesetting different default arguments I do not think it is a good idea to change the doc in this direction. Especially when you consider that probably many users probably knowread.tablebut have never heard ofread.delim.My suggestion is based on comment #5634 (comment)
Users ofread.tablemay assume (like I did) thatfreadbehaves likeread.table, however, it does not. If the doc tells me thatfreadis similar to a function that I have never used, then I will not assume that I know how it works.Slightly off-topic but I'm not sure where to ask: Is data.table still developed or in low maintenance mode? There are 192 pull requests pending and no commits since February. I'm considering replacing a few large data.frames with data.tables in my own R package (https://github.com/MLopez-Ibanez/irace), but I am not sure yet if this is a good idea.
- added a commit that references this issue
on Jan 13, 2024 Marking as closed by #5635. If we think there's something further to do for this case, please open a new issue to keep the discussion focused.
Reacted by Toby Dylan Hocking- added a commit that references this issue
on Jan 14, 2024 - added a commit that references this issue
on Jul 14, 2025
read.table()handles the row names as expected.fread(header="auto")gives a warning and creates an extra column. Not great but can be fixed withsuppressWarnings()and removing the extra column.fread(header=TRUE)should do the same but instead gives:which is completely wrong.
Verbose output: