Skip to content

Implement mutate follwoing group_by #331

Description

@nassuphis
require(disk.frame)
setup_disk.frame()
options(future.globals.maxSize = Inf)

test_frame<-
  tidyr::expand(
    data=tibble(),
    "Ticker":=LETTERS[1:5],
    "date":=Sys.Date()-(1:100)
  ) %>%
  mutate(minute=map(1:n(),~sample(20:120,1)+1:sample(20:120,1))) %>%
  unchop(minute) %>%
  mutate(px=runif(n())) %>%
  as.disk.frame() %>%  
  hard_group_by(Ticker,date,outdir="df_test")


res1<-
  test_frame %>%
  group_by(Ticker,date) %>%
  mutate(max_m=max(minute)) %>%
  summarize(
      max_of_max=max(max_m),
      max_minute=max(minute)
  ) %>%
  collect()
  
res2<-
  disk.frame("df_test") %>%
    group_by(Ticker,date) %>%
    mutate(max_m=max(minute)) %>%
    summarize(
      max_of_max=max(max_m),
      max_minute=max(minute)
    ) %>%
    collect()

  all(res1$max_of_max==res1$max_minute)
  all(res2$max_of_max==res2$max_minute)

Activity

  1. nassuphis commented on Apr 3, 2021

    @nassuphis
    Author
    >  all(res1$max_of_max==res1$max_minute)
    [1] TRUE
    >   all(res2$max_of_max==res2$max_minute)
    [1] FALSE
    > 
    
  2. nassuphis commented on Apr 3, 2021

    @nassuphis
    Author

    it seems the res2 max calculation does not respect the grouping specified by group_by

    > res2
    # A tibble: 500 x 4
    # Groups:   Ticker [5]
       Ticker date       max_of_max max_minute
       <chr>  <date>          <int>      <int>
     1 A      2020-12-24        238        171
     2 A      2020-12-25        238         49
     3 A      2020-12-26        238        133
     4 A      2020-12-27        238        162
     5 A      2020-12-28        238         99
     6 A      2020-12-29        238        125
     7 A      2020-12-30        238        181
     8 A      2020-12-31        238         96
     9 A      2021-01-01        238        115
    10 A      2021-01-02        238        156
    # … with 490 more rows
    
  3. nassuphis commented on Apr 3, 2021

    @nassuphis
    Author

    max_of_max is always larger than max_minute, and equal to the max across all dates
    despite the group_by(Ticker,date)

    this only happens if the disk.frame is on a file, if the same calculation is off a pipeline the result is ok.

  4. xiaodaigh commented on Apr 4, 2021

    @xiaodaigh
    Collaborator

    group_by(Ticker,date) %>%
    mutate(max_m=max(minute))

    group_by and mutate cannot be combined this way in a disk.frame. To do ANY summary you must use summarize. I think mutate will compute the max within each chunk so it's not correct.

  5. nassuphis commented on Apr 5, 2021

    @nassuphis
    Author

    doesnt hard_group_by(Ticker,date,outdir="df_test") ensure that chunks are exactly what is needed to compute
    the group-wise max?
    if chunks exactly match the Ticker,date grouping, then max(minute) is both the max of the group and max of chunk?
    what does hard_group_by do? does it not create a chunking from the grouping?

  6. xiaodaigh commented on Apr 6, 2021

    @xiaodaigh
    Collaborator

    It doesn't but disk.frame doesn't know how to handle mutate after group_by. I think in the next minor version I will throw an error is group_by is not followed by summarization and in the next major version I will handle mutate.

    The issue is that mutate is basically group_by and then join back to the original in one step which can be handled in disk.frame

  7. changed the title [-]disk.frame saved to disk does not match disk.frame that was saved[/-] [+]Implement `mutate` follwoing `group_by`[/+] on Apr 6, 2021
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions