Skip to content

data.table(vector, matrix) can return duplicate column names #3193

Description

@MichaelChirico
data.table(1:5, matrix(6:15, nrow = 5L))
#    V1 V1 V2
# 1:  1  6 11
# 2:  2  7 12
# 3:  3  8 13
# 4:  4  9 14
# 5:  5 10 15

Not ideal to be creating data.tables with duplicate names from the get-go... those names should be V2_1 and V2_2 i guess.

Fixes with check.names but I'm not sure creating duplicates like this by default is the right behavior; my only hesitation is I'm sure there are other more complicated cases that are still creating duplicates even if we fix this.

The fix to this specific issue can be very simple, however, just add 7 characters to this line:

https://github.com/Rdatatable/data.table/blob/master/R/data.table.R#L98

if (any(tt)) namesi[tt] = paste0("V", i, '_', which(tt))

Activity

  1. Atrebas commented on Feb 3, 2019

    @Atrebas

    Sowewhat related, I have noticed several cases when colnames are duplicated.
    These may be some specific cases, when working with toy datasets and colnames V1, V2, ...
    Some examples below for information.

    library(data.table)
    
    ## case 1: setDT duplicates V1 and V2
    lst = list(V1=1, V2=1, 1, 1)
    setDT(lst)[] #check.names=TRUE solves this
    
    
    DT = data.table(V1 = c(1L,2L),
                    V2 = 1:9,
                    V3 = round(rnorm(3),2),
                    V4 = LETTERS[1:3])
    
    ## case 2: computation in j duplicates V1
    DT[, sum(V3), by = "V1"] 
    
    ## case 3 (linked to case 2): rollup duplicates the V1 colname and bugs
    rollup(DT, j = sum(V3), by = c("V4", "V2"))   # ok
    rollup(DT, j = sum(V3), by = c("V4", "V1"))   # error
    
    # Error in groupingsets.data.table(x, by = by, sets = sets, .SDcols = .SDcols,  : 
    # There exists duplicated column names in the results, 
    # ensure the column passed/evaluated in `j` and those in `by` are not overlapping.
    
    # So specifying a different name works:
    rollup(DT, j = .(SumV3 = sum(V3)), by = c("V4", "V1"))   # ok
    
  2. jangorecki commented on Feb 4, 2019

    @jangorecki
    Member

    @Atrebas thanks for investigation. Groupings sets functions just wraps around ordinary grouping, thus are affected.

  3. jangorecki commented on Apr 5, 2020

    @jangorecki
    Member

    Related to #4124

  4. self-assigned this
    on Apr 7, 2020
  5. jangorecki commented on Apr 7, 2020

    @jangorecki
    Member

    @Atrebas

    • case 1: setDT is designed for speed, so the less the checks the better, any checks on quality of data/metadata should be on user side
    • case 2: has a dedicated issue duplicate names on simple operation #3644
    • case 3: inherits from case 2, as explained before.
  6. jangorecki commented on Apr 7, 2020

    @jangorecki
    Member

    This issue is caused by a fact that data.table has check.names=FALSE by default, while data.frame has check.names=TRUE. I filled #4362

  7. MichaelChirico commented on Jul 11, 2025

    @MichaelChirico
    MemberAuthor

    I think this is fixed by #4363, though the breaking change means it's an "eventually" thing.

  8. removed their assignment
    on Jul 12, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions