Skip to content

setDT should accept named vectors #1244

Description

@MichaelChirico

Just came across a situation where I want to use a named vector in a way most easily handled by a join, but unfortunately:

join_vector <- c("id1"=1, "id2"=2, "id3"=3)
setDT(join_vector, keep.rownames=T)

Is an error: Argument 'x' to 'setDT' should be a 'list', 'data.frame' or 'data.table'

However, first casting join_vector as a data.frame works:

setDT(data.frame(join_vector), keep.rownames=T)

(though this returns an invisible object, since data.frame(join_vector) isn't assigned to anything).

Is there any reason not to allow setDT to re-class vectors like the above?

Activity

  1. franknarf1 commented on Jul 24, 2015

    @franknarf1
    Contributor

    setDT modifies by reference, as do all the set* functions.

    To turn something that is not a list into a data.table, you cannot simply modify its attributes by reference, e.g., your is.list(join_vector) # FALSE, so it's not clear what desired result you want. To create a data.table, use a different function. I don't want to have to wrap my head around "The set* functions all modify by reference, EXCEPT..."

  2. MichaelChirico commented on Jul 24, 2015

    @MichaelChirico
    MemberAuthor

    I do expect set* to continue to update by reference for vectors; why is this impossible? Curious. Could be that I don't have a solid enough understanding of how by-reference updates work.

  3. franknarf1 commented on Jul 24, 2015

    @franknarf1
    Contributor

    If you type setDT in the console, you can see how it works. All of the substantial steps use other set* functions or alloc.col (which also modifies by reference). I don't think you could add another block in here to modify a named vector by reference, but you could try. Alternately, you could just make a habit of working with lists... because why wouldn't you? (Just a rhetorical question.)

  4. eantonya commented on Jul 24, 2015

    @eantonya
    Contributor

    I was actually looking at an SO question today that's very similar to this - http://stackoverflow.com/q/30344192/817778. So you can put a vector inside a data.table without copying any data, which I think qualifies as "in-place" conversion.

  5. franknarf1 commented on Jul 24, 2015

    @franknarf1
    Contributor

    @eantonya Very cool! However, that operation treats the vector like a column vector, not like a full data.table (creating one column per element of the vector).

  6. eantonya commented on Jul 24, 2015

    @eantonya
    Contributor

    @franknarf1 that's what both data.table(your_vec) and as.data.table(your_vec) do, so why would setDT do smth else?

  7. franknarf1 commented on Jul 24, 2015

    @franknarf1
    Contributor

    @eantonya Oh, I thought your first comment was suggesting that you supported the proposal at hand (which is what I described: vector => column per element) because it's possible to modify vectors by reference. Anyway, that is a cool trick.

  8. MichaelChirico commented on Jul 25, 2015

    @MichaelChirico
    MemberAuthor

    Neat. If I understand correctly, then, why not just update setDT's native behavior to (when encountering a vector) wrap it with list and get keep.rownames to work by assigning the same names? Expected result of setDT(join_vector, keep.rownames=T) being something like:

        rn join_vector
    1: id1           1
    2: id2           2
    3: id3           3
    
  9. eantonya commented on Jul 26, 2015

    @eantonya
    Contributor

    @franknarf1 I don't think a row-vector is what OP is looking for - they're looking for a regular column vector, just with names added.

    @MichaelChirico yep, I think it's doable, and seems reasonable to me.

  10. franknarf1 commented on Jul 26, 2015

    @franknarf1
    Contributor

    Oh I see, my mistake. Sorry for the noise.

  11. MichaelChirico commented on Sep 18, 2015

    @MichaelChirico
    MemberAuthor

    @eantonya I'm not sure I understand your update to your previously referenced SO post. Does this mean we won't be able to do this so easily after all?

    I took a stab at this today and the current version appears to make a copy (as evidenced by checking address before and after use of setDT). See my fork.

  12. jangorecki commented on Sep 18, 2015

    @jangorecki
    Member

    @MichaelChirico regarding your 41d4a77 commit.
    I would avoid recursive call when it is not necessary and go with something like this:

    ...
    } else if (is.list(x) || is.atomic(x)) {
        if(is.atomic(x)) x = list(x)
        ....

    but this is just my (not so much skilled) opinion.

    Still I'm not sure if it is correct to make setDT on atomic vector, might be better to put that optimization in as.data.table method for atomic types?

  13. MichaelChirico commented on Sep 18, 2015

    @MichaelChirico
    MemberAuthor

    @jangorecki yes, I abandoned the recursive call when I realized it'd be much simpler to just imitate the list approach given the simplification that the length as a list is one.

  14. MichaelChirico commented on Sep 19, 2015

    @MichaelChirico
    MemberAuthor

    @jangorecki I was thinking of that, but the keep.rownames option seems to be a natural (read: concise and parallel to usage for other types) fit for the case when we want to convert a named vector to a two-column data.table. Of course if there's no way to convert an atomic type by reference, perhaps we should just focus our energies on adding keep.rownames to as.data.table and/or data.table().

  15. added this to the v1.9.8 milestone on Sep 22, 2015
  16. modified the milestones: , v1.9.8 on Nov 17, 2015
  17. removed this from the milestone on May 10, 2018
  18. jangorecki commented on Mar 16, 2019

    @jangorecki
    Member

    I would close that one. Vector is not compatible data type to be converted to data.table by reference, it is better leave atomic vector support or as.data.table.

  19. MichaelChirico commented on Apr 13, 2019

    @MichaelChirico
    MemberAuthor

    Sure, makes sense, the main thing i was going for was a method to go from named vector to two-column data.table. By-reference part would just be a bonus & I was holding out hope for what Eddi mentioned as a possibility. For a single vector, the efficiency aspect should be really marginal.

    Though, I happened to notice that apparently copy (and hence as.data.table) was really slow on e.g. 1B-length vector:

    system.time(x <- character(1e9))
       ユーザ   システム       経過  
         1.677      2.035      3.714 
    system.time(copy(x))
       ユーザ   システム       経過  
         7.940     18.811     28.952 
    
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions