data set ezerai shared in #2431 exhibits a tie for detecting sep -- exactly the same number of columns from using sep='\t' and sep=',', but "human eye" can easily see sep='\t'.
One way we could nudge fread to getting it right would be to try using the detected column types to tip the balance -- in this case, sep=',' leads to parsing two char and one int column, vs one char and two numeric columns for sep='\t'. One char column is guaranteed because there's non-numeric/non-whitespace characters in the first column; returning two numeric columns is better than one-char/one-int.
Not sure if this can be regularized enough to make a Pareto improvement but I think it works for this file.
data set
ezeraishared in #2431 exhibits a tie for detectingsep-- exactly the same number of columns from usingsep='\t'andsep=',', but "human eye" can easily seesep='\t'.One way we could nudge
freadto getting it right would be to try using the detected column types to tip the balance -- in this case,sep=','leads to parsing twocharand oneintcolumn, vs onecharand twonumericcolumns forsep='\t'. Onecharcolumn is guaranteed because there's non-numeric/non-whitespace characters in the first column; returning twonumericcolumns is better than one-char/one-int.Not sure if this can be regularized enough to make a Pareto improvement but I think it works for this file.