Skip to content

Improve how fread's drop/select deal with duplicate column names #1899

Description

@MichaelChirico

Consider

fread("x1,x1,y\n2,3,4", drop = "x1")
#    x1 y
# 1:  3 4
fread("x1,x1,y\n2,3,4", select = "x1")
#    x1
# 1:  2

x1 is duplicated. If both x1 are chosen by select, both should be excluded by drop. At the very least, there should be a warning about multiple x1 detected.

Use case is I have a .csv with a lot of columns called filler which are pointless white space. It would be nice to just say drop = "filler" and exclude all of them.

Originally filed on SO.

Activity

  1. MichaelChirico commented on Nov 8, 2017

    @MichaelChirico
    MemberAuthor

    Poking around a bit; seems this comes down to the return behavior of chmatch:

    https://github.com/Rdatatable/data.table/blob/master/src/freadR.c#L275

    if (isString(dropSxp)) itemsInt = PROTECT(chmatch(dropSxp, colNamesSxp, NA_INTEGER, FALSE))
    

    This is the same as the behavior of base match:

    match('x1', c('x1', 'x1', 'y'))
    # [1] 1
    

    To skirt changing how chmatch works, I guess we have to (?) reverse the arguments and look for NA:

    match(c('x1', 'x1', 'y'), 'x1')
    # [1] 1 1 NA
    # ^ know to drop anything that's not NA.
    #  IIUC we can do this more easily setting `in = TRUE` in arguments to chmatch @ C level
    

    I know how I would do this in R (!names(x) %in% drop), but not sure the right way to go about it in C

  2. self-assigned this
    on Oct 17, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions