Skip to content

sum(value, na.rm = FALSE) returns unexpected 0s for integer64 field within NAs in group #7571

Description

@rweberc

Problem: It's not clear to me why sum(value, na.rm = FALSE) does not always return NA when there is an NA value in an integer64 column within a given id group: dt_short[, result_problem := sum(value, na.rm = FALSE), by = id]. In the examples below, the groups each contain one value, but behavior doesn't seem look to be limited to that scenario.

Resolved by: Casting the value in question to an integer, rather than the initial integer64 value.

# [Minimal reproducible example]

library(data.table)
library(bit64)

# Expected that sum(., na.rm = FALSE) would return "NA" for id's 1-3 in the example below.

dt_short <- data.table(
  id = c(1:8, 9, 9),
  value = c(rep(NA_integer64_, 3), 4:10)
)

# Shouldn't the first 4 values of result_problem be 3 NA's followed by a 4?
dt_short[, result_problem := sum(value, na.rm = FALSE), by = id]
dt_short 
#>        id value result_problem
#>     <num> <i64>          <i64>
#>  1:     1  <NA>           <NA>
#>  2:     2  <NA>              0
#>  3:     3  <NA>              0
#>  4:     4     4              0
#>  5:     5     5              5
#>  6:     6     6              6
#>  7:     7     7              7
#>  8:     8     8              8
#>  9:     9     9             19
#> 10:     9    10             19

# works as expected when cast as an integer (instead of an integer64)
dt_short[, result_asinteger := sum(as.integer(value), na.rm = FALSE), by = id]
dt_short

#>        id value result_problem result_asinteger
#>     <num> <i64>          <i64>            <int>
#>  1:     1  <NA>           <NA>               NA
#>  2:     2  <NA>              0               NA
#>  3:     3  <NA>              0               NA
#>  4:     4     4              0                4
#>  5:     5     5              5                5
#>  6:     6     6              6                6
#>  7:     7     7              7                7
#>  8:     8     8              8                8
#>  9:     9     9             19               19
#> 10:     9    10             19               19

# Note: For some reason, the issue does not seem to affect the first id.
# And there seems to be some offset effect since the sum for id 4 ends up being "0" instead of "4".  

# Other scenarios
# see the issue here when just returning the by and new column
dt_short[, .(result_problem = sum(value, na.rm = FALSE)), by = id]
#>       id result_problem
#>    <num>          <i64>
#> 1:     1           <NA>
#> 2:     2              0
#> 3:     3              0
#> 4:     4              0
#> 5:     5              5
#> 6:     6              6
#> 7:     7              7
#> 8:     8              8
#> 9:     9             19

# Don't see the issue when casting the value that gets summed as an integer first
dt_short[, .(result_problem = sum(as.integer(value), na.rm = FALSE)), by = id]
#>       id result_problem
#>    <num>          <int>
#> 1:     1             NA
#> 2:     2             NA
#> 3:     3             NA
#> 4:     4              4
#> 5:     5              5
#> 6:     6              6
#> 7:     7              7
#> 8:     8              8
#> 9:     9             19

# Don't see the issue here, when including another column in the returned values
dt_short[, .(value, result_problem = sum(value, na.rm = FALSE)), by = id]
#>        id value result_problem
#>     <num> <i64>          <i64>
#>  1:     1  <NA>           <NA>
#>  2:     2  <NA>           <NA>
#>  3:     3  <NA>           <NA>
#>  4:     4     4              4
#>  5:     5     5              5
#>  6:     6     6              6
#>  7:     7     7              7
#>  8:     8     8              8
#>  9:     9     9             19
#> 10:     9    10             19

# Output of sessionInfo()
#> R version 4.3.1 (2023-06-16)
#> Platform: x86_64-pc-linux-gnu (64-bit)
#> Running under: Ubuntu 20.04.6 LTS
#>
#> Matrix products: default
#> BLAS: /usr/lib/x86_64-linux-gnu/openblas-pthread/libblas.so.3
#> LAPACK: /usr/lib/x86_64-linux-gnu/openblas-pthread/liblapack.so.3; LAPACK version 3.9.0
#>
#> locale:
#> [1] LC_CTYPE=en_US.UTF-8 LC_NUMERIC=C
#> [3] LC_TIME=en_US.UTF-8 LC_COLLATE=en_US.UTF-8
#> [5] LC_MONETARY=en_US.UTF-8 LC_MESSAGES=en_US.UTF-8
#> [7] LC_PAPER=en_US.UTF-8 LC_NAME=C
#> [9] LC_ADDRESS=C LC_TELEPHONE=C
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C
#>
#> time zone: America/New_York
#> tzcode source: system (glibc)
#>
#> attached base packages:
#> [1] stats graphics grDevices utils datasets methods base
#>
#> other attached packages:
#> [1] bit64_4.6.0-1 bit_4.6.0 data.table_1.17.8
#>
#> loaded via a namespace (and not attached):
#> [1] digest_0.6.37 fastmap_1.2.0 xfun_0.52 glue_1.8.0
#> [5] knitr_1.50 htmltools_0.5.8.1 rmarkdown_2.29 lifecycle_1.0.4
#> [9] cli_3.6.5 reprex_2.1.1 withr_3.0.2 compiler_4.3.1
#> [13] rstudioapi_0.17.1 tools_4.3.1 evaluate_1.0.4 yaml_2.3.10
#> [17] rlang_1.1.6 fs_1.6.6


<sup>Created on 2025-12-20 with [reprex v2.1.1](https://reprex.tidyverse.org)</sup>

Activity

  1. added
    GForceissues relating to optimized grouping calculations (GForce)
    on Jan 7, 2026
  2. MichaelChirico commented on Jan 7, 2026

    @MichaelChirico
    Member

    Shouldn't the first 4 values of result_problem be 3 NA's followed by a 4?

    Yes; confirming this is a {data.table} bug, not {bit64}:

    do.call(c, lapply(split(dt_short, ~id), \(x) sum(x$value, na.rm=FALSE)))
    # integer64
    #    1    2    3    4    5    6    7    8    9 
    # <NA> <NA> <NA>    4    5    6    7    8   19

    And also that this is a GForce bug:

    options(datatable.optimize=0)
    dt_short[, result_problem := sum(value, na.rm = FALSE), by = id][]
    #        id value result_problem
    #     <num> <i64>          <i64>
    #  1:     1  <NA>           <NA>
    #  2:     2  <NA>           <NA>
    #  3:     3  <NA>           <NA>
    #  4:     4     4              4
    #  5:     5     5              5
    #  6:     6     6              6
    #  7:     7     7              7
    #  8:     8     8              8
    #  9:     9     9             19
    # 10:     9    10             19
  3. MichaelChirico commented on Jan 7, 2026

    @MichaelChirico
    Member

    The behavior is quite strange:

    dt_short[1:4, sum(value), by=id]
    #       id    V1
    #    <num> <i64>
    # 1:     1  <NA>
    # 2:     2  <NA>
    # 3:     3  <NA>
    # 4:     4     0
    
    dt_short[1:5, sum(value), by=id]
    #       id    V1
    #    <num> <i64>
    # 1:     1  <NA>
    # 2:     2     0
    # 3:     3  <NA>
    # 4:     4     0
    # 5:     5     5

    That makes it looks like a parallelization issue, but setDTthreads(1) doesn't fix it either.

  4. manmita commented on Jan 7, 2026

    @manmita
    Contributor

    Hello,
    Can I pickup this one?

    The problem is occuring at optimize >= 2, but not at 0 and 1. Still checking the issue,

    Thanks

  5. MichaelChirico commented on Jan 7, 2026

    @MichaelChirico
    Member

    Yes, please have a look. It's a pretty technical part of the code in gsumm.c. I spent about 30 minutes trying to spot something obvious and didn't see anything.

  6. manmita commented on Jan 7, 2026

    @manmita
    Contributor

    Hello @MichaelChirico ,

    in the following code - line no: 495 of gsum.c file, removing the break fixes the issue
    Will raise a PR after writing the tests.

    Thanks

    if (!narm) {
              #pragma omp parallel for num_threads(getDTthreads(highSize, false))
              for (int h=0; h<highSize; h++) {
                int64_t *restrict _ans = ansp + (h<<bitshift);
                for (int b=0; b<nBatch; b++) {
                  const int pos = counts[ b*highSize + h ];
                  const int howMany = ((h==highSize-1) ? (b==nBatch-1?lastBatchSize:batchSize) : counts[ b*highSize + h + 1 ]) - pos;
                  const int64_t *my_gx = gx + b*batchSize + pos;
                  const uint16_t *my_low = low + b*batchSize + pos;
                  for (int i=0; i<howMany; i++) {
                    const int64_t elem = my_gx[i];
                    if (elem!=INT64_MIN) {
                      _ans[my_low[i]] += elem;
                    } else {
                      _ans[my_low[i]] = INT64_MIN;
                      break;
                    }
                  }
                }
              }
    
    
    > dt_short <- data.table(
      id = c(1:8, 9, 9),
      value = c(rep(NA_integer64_, 3), 4:10)
    )
    > dt_short
           id value
        <num> <i64>
     1:     1  <NA>
     2:     2  <NA>
     3:     3  <NA>
     4:     4     4
     5:     5     5
     6:     6     6
     7:     7     7
     8:     8     8
     9:     9     9
    10:     9    10
    > t_short[, result_problem := sum(value, na.rm = FALSE), by = id]
    Error: object 't_short' not found
    > dt_short[, result_problem := sum(value, na.rm = FALSE), by = id]
    > dt_short
           id value result_problem
        <num> <i64>          <i64>
     1:     1  <NA>           <NA>
     2:     2  <NA>           <NA>
     3:     3  <NA>           <NA>
     4:     4     4              4
     5:     5     5              5
     6:     6     6              6
     7:     7     7              7
     8:     8     8              8
     9:     9     9             19
    10:     9    10             19
    > dt_short[, .(result_problem = sum(value, na.rm = FALSE)), by = id]
          id result_problem
       <num>          <i64>
    1:     1           <NA>
    2:     2           <NA>
    3:     3           <NA>
    4:     4              4
    5:     5              5
    6:     6              6
    7:     7              7
    8:     8              8
    9:     9             19
    > dt_short
           id value result_problem
        <num> <i64>          <i64>
     1:     1  <NA>           <NA>
     2:     2  <NA>           <NA>
     3:     3  <NA>           <NA>
     4:     4     4              4
     5:     5     5              5
     6:     6     6              6
     7:     7     7              7
     8:     8     8              8
     9:     9     9             19
    10:     9    10             19
    > dt_short[, .(result_problem = sum(as.integer(value), na.rm = FALSE)), by = id]
          id result_problem
       <num>          <int>
    1:     1             NA
    2:     2             NA
    3:     3             NA
    4:     4              4
    5:     5              5
    6:     6              6
    7:     7              7
    8:     8              8
    9:     9             19
    > dt_short[1:5, sum(value), by=id]
          id    V1
       <num> <i64>
    1:     1  <NA>
    2:     2  <NA>
    3:     3  <NA>
    4:     4     4
    5:     5     5
    
    
  7. manmita commented on Jan 7, 2026

    @manmita
    Contributor

    Added a check if elem is not NA and ans is NA to make sure ans is not edited if NA
    Removed the break statement if the element is NA, as that was causing bug on consecutive NA values in the same partition but different group.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    GForceissues relating to optimized grouping calculations (GForce)bit64bug

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions