Given a data set with POSIXct times, I find that expanding the time series using the seq() function to be successful using 'by' operations in data.table.
However, the same execution when using the interval() function, from the lubridate package, does not give the desired output...
Dataset:
| ID |
TimeStamp |
| 1 |
2010-09-11 21:55:38 |
| 1 |
2010-09-12 15:48:13 |
| 1 |
2010-09-12 19:36:31 |
| 1 |
2010-09-13 07:36:43 |
| 1 |
2010-09-13 07:44:16 |
| 1 |
2010-09-13 15:46:44 |
| 2 |
2010-10-15 02:11:21 |
| 2 |
2010-10-15 18:00:00 |
| 2 |
2010-10-16 02:51:37 |
| 2 |
2010-10-16 08:43:15 |
| 2 |
2010-10-16 08:53:19 |
| 2 |
2010-10-17 02:55:27 |
| 3 |
2012-03-24 23:17:51 |
| 3 |
2012-03-25 14:26:12 |
| 3 |
2012-03-25 14:31:48 |
| 3 |
2012-03-26 02:52:20 |
| 3 |
2012-03-26 11:52:25 |
| 3 |
2012-03-26 11:58:06 |
'
Unsuccessful interval calculations:
data[, .(Ranges = range(TimeStamp)), by = ID][, .(Interval = interval(Ranges[1],Ranges[2])) ,by = ID]
The intervals are all calculated based on the first input IDs min range, seen here, with a end time based on the first input IDs max range date, but variable time:
| ID |
Interval |
| 1 |
2010-09-11 21:55:38 EDT--2010-09-13 15:46:44 EDT |
| 2 |
2010-09-11 21:55:38 EDT--2010-09-13 22:39:44 EDT |
| 3 |
2010-09-11 21:55:38 EDT--2010-09-13 10:35:53 EDT |
'
However, other functions show that the data.table can capture/process using this methodolgy successfully, as seen when expanding the sequence of ranges with the seq() function.
Successful expansion:
data[, .(Ranges = range(TimeStamp)), by = ID][, .(TimeStamp = seq(Ranges[1], Ranges[2], "days")), by=ID]
| ID |
TimeStamp |
| 1 |
2010-09-11 21:55:38 |
| 1 |
2010-09-12 21:55:38 |
| 2 |
2010-10-15 02:11:21 |
| 2 |
2010-10-16 02:11:21 |
| 2 |
2010-10-17 02:11:21 |
| 3 |
2012-03-24 23:17:51 |
| 3 |
2012-03-25 23:17:51 |
Given a data set with POSIXct times, I find that expanding the time series using the seq() function to be successful using 'by' operations in data.table.
However, the same execution when using the interval() function, from the lubridate package, does not give the desired output...
Dataset:
'
Unsuccessful interval calculations:
data[, .(Ranges = range(TimeStamp)), by = ID][, .(Interval = interval(Ranges[1],Ranges[2])) ,by = ID]The intervals are all calculated based on the first input IDs min range, seen here, with a end time based on the first input IDs max range date, but variable time:
'
However, other functions show that the data.table can capture/process using this methodolgy successfully, as seen when expanding the sequence of ranges with the seq() function.
Successful expansion:
data[, .(Ranges = range(TimeStamp)), by = ID][, .(TimeStamp = seq(Ranges[1], Ranges[2], "days")), by=ID]