Repository navigation
Comparative underperformance (vs quantmod & xts) for large-N time-series exercise. #4870
Description
Activity
Thank you for reporting. For completeness I am pasting related code below.
library(data.table) N = 1e6 spot = 100 r = 0.01 sigma = 0.02 dt = data.table( V1 = seq.Date(as.Date('1970-01-01'), by = 1, length.out = N), V2 = spot * exp(cumsum((r - 0.5 * sigma**2) * 1/N + (sigma * (sqrt(1/N)) * rnorm(N, mean = 0, sd = 1)))) ) library(zoo) system.time(dt[, .(EOM = last(V2)), .(Month = as.yearmon(V1))][, .(Month, Return = EOM/shift(EOM, fill = dt[, first(V2)]) - 1)]) # user system elapsed # 2.196 0.000 2.080 system.time(as.yearmon(dt$V1)) # user system elapsed # 2.160 0.001 2.159
From the timings above you can see that it is not really data.table operations takes so much time but
zoo::as.yearmon. I don't think we provide high performance alternative to this zoo function. There is a pending PR #4868 but it is not yet the implementation that will address this gap because it uses POSIXlt rather than operating on numerics.Reacted by Jacob and Xianying TanConsidering benchmarks an integer solutions for
Dateslooks favorable (basic idea in R below). However, I do not get why @mattdowle gsub work around (from here) is so slow (maybe a Windows thing?).convert_date = function(x, type = "year"){ x = as.integer(x) days = x - 11017 years400_days = 146097 years100_days = 36524 years4_days = 1461 years400 = days %/% years400_days days = days %% years400_days years100 = days %/% years100_days days = days %% years100_days years4 = days %/% years4_days days = days %% years4_days years = days %/% 365 days = days %% 365 year = 2000 + years + 4*years4 + 100*years100 + 400*years400 year = year + (days > 305) if (type == "year"){ return(year) } else if (type == "month"){ days_mon = data.table( x = c(30, 60, 91, 121, 152, 183, 213, 244, 274, 305, 336, 365) , mon = c(3:12, 1:2) ) dt = data.table(days = days) month = days_mon[dt, on=.(x = days), roll=-Inf][["mon"]] return(month) } } N = 1e6 spot = 100 r = 0.01 sigma = 0.02 dt = data.table( V1 = seq.Date(as.Date('1970-01-01'), by = 1, length.out = N), V2 = spot * exp(cumsum((r - 0.5 * sigma**2) * 1/N + (sigma * (sqrt(1/N)) * rnorm(N, mean = 0, sd = 1)))) ) library(zoo) system.time(dt[, .(EOM = last(V2)), .(Month = as.yearmon(V1))]) # user system elapsed # 1.86 0.05 1.90 system.time(dt[, .(EOM = last(V2)), .(Month = convert_date(V1,"month"), Year = convert_date(V1, "year"))]) # user system elapsed # 0.42 0.03 0.44 system.time(dt[, .(as.integer(gsub('-', '', V1)))][, .(as.integer(V1/100))]) # user system elapsed # 8.20 0.11 8.31 sessionInfo() # R version 4.0.2 (2020-06-22) # Platform: x86_64-w64-mingw32/x64 (64-bit) # Running under: Windows 10 x64 (build 17763) # Matrix products: default # attached base packages: # [1] stats graphics grDevices utils datasets methods base # other attached packages: # [1] zoo_1.8-8 data.table_1.13.1 # loaded via a namespace (and not attached): # [1] compiler_4.0.2 grid_4.0.2 lattice_0.20-41
My guess is gsub is slow because it needs to populate global string cache with new strings that are going to be used only for
as.integer.Closing this, since it is fixed since #5300
I built a time-series benchmark of a simple period return calculation (daily prices to monthly returns) across common libraries: https://github.com/griipen/TSBenchmark/
Curiously, data.table's competitiveness does not appear to hold at scale in this particular case; it underperforms dramatically vis-a-vis xts and quantmod for N>1e5. As I understand following a chat with the authors, it may be related to data.table's internal date handling.
From my issue at griipen/TSBenchmark#1:
Extended results, supporting figures and full code are found in my master branch. Below are the functions for xts, quantmod and data.table for a quick reference: