Repository navigation
Implement mutate follwoing group_by #331
Description
Activity
> all(res1$max_of_max==res1$max_minute) [1] TRUE > all(res2$max_of_max==res2$max_minute) [1] FALSE >it seems the res2 max calculation does not respect the grouping specified by group_by
> res2 # A tibble: 500 x 4 # Groups: Ticker [5] Ticker date max_of_max max_minute <chr> <date> <int> <int> 1 A 2020-12-24 238 171 2 A 2020-12-25 238 49 3 A 2020-12-26 238 133 4 A 2020-12-27 238 162 5 A 2020-12-28 238 99 6 A 2020-12-29 238 125 7 A 2020-12-30 238 181 8 A 2020-12-31 238 96 9 A 2021-01-01 238 115 10 A 2021-01-02 238 156 # … with 490 more rowsmax_of_max is always larger than max_minute, and equal to the max across all dates
despite the group_by(Ticker,date)this only happens if the disk.frame is on a file, if the same calculation is off a pipeline the result is ok.
group_by(Ticker,date) %>%
mutate(max_m=max(minute))group_byandmutatecannot be combined this way in adisk.frame. To do ANY summary you must usesummarize. I thinkmutatewill compute the max within each chunk so it's not correct.doesnt
hard_group_by(Ticker,date,outdir="df_test")ensure that chunks are exactly what is needed to compute
the group-wise max?
if chunks exactly match the Ticker,date grouping, then max(minute) is both the max of the group and max of chunk?
what doeshard_group_bydo? does it not create a chunking from the grouping?It doesn't but
disk.framedoesn't know how to handlemutateaftergroup_by. I think in the next minor version I will throw an error isgroup_byis not followed bysummarizationand in the next major version I will handlemutate.The issue is that
mutateis basicallygroup_byand thenjoinback to the original in one step which can be handled indisk.frame- changed the title
[-]disk.frame saved to disk does not match disk.frame that was saved[/-][+]Implement `mutate` follwoing `group_by`[/+]on Apr 6, 2021