Skip to content

revdep bbotk fails to create a column by group after unserializing a data.table #7498

Description

@aitap

From their tests:

library(bbotk)
library(data.table)
  search_space = domain = ps(
    x1 = p_dbl(-5, 10),
    x2 = p_dbl(0, 15)
  )

  fun = function(xdt, instances) {
    data.table(y = branin(xdt[["x1"]], xdt[["x2"]], noise = as.numeric(instances)))
  }

  objective = ObjectiveRFunDt$new(fun = fun, domain = domain)

  instance = OptimInstanceBatchSingleCrit$new(
    objective = objective,
    search_space = search_space,
    terminator = trm("evals", n_evals = 96))


  optimizer = opt("irace", instances = rnorm(10, mean = 0, sd = 0.1))

optimizer$optimize(instance)
Error in `[.data.table`(log, , `:=`("step", rleid("instance")), by = "iteration") :
  Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. (omitted)
Enter a frame number, or 0 to exit 

 1: optimizer$optimize(instance)
 2: .__OptimizerBatch__optimize(self = self, private = private, super = super, inst = inst)
 3: optimize_batch_default(inst, self)
 4: tryCatch({
    get_private(optimizer)$.optimize(instance)
}, terminated_error = function(cond) {
})
 5: tryCatchList(expr, classes, parentenv, handlers)
 6: tryCatchOne(expr, names, parentenv, handlers[[1]])
 7: doTryCatch(return(expr), name, parentenv, handler)
 8: get_private(optimizer)$.optimize(instance)
 9: .__OptimizerBatchIrace__.optimize(self = self, private = private, super = super, inst = inst)
10: log[, `:=`("step", rleid("instance")), by = "iteration"]
11: `[.data.table`(log, , `:=`("step", rleid("instance")), by = "iteration")

Browse[1]> str(log)
Classes ‘data.table’ and 'data.frame':  96 obs. of  5 variables:
 $ iteration    : int  1 1 1 1 1 1 1 1 1 1 ...
 $ instance     : int  1 1 1 1 1 2 2 2 2 2 ...
 $ configuration: int  1 2 3 4 5 1 2 3 4 5 ...
 $ cost         : num  3.52 71.51 61.13 94.09 22.35 ...
 $ time         : num  NA NA NA NA NA NA NA NA NA NA ...
 - attr(*, ".internal.selfref")=<pointer: (nil)> 
Browse[1]> .Internal(inspect(log, 0))
@55995985eba8 19 VECSXP g0c4 [OBJ,REF(4),gp=0x20,ATT] (len=5, tl=0, gr)
ATTRIB:
  @55994fef2410 02 LISTSXP g0c0 [REF(1)] 
Browse[1]> .Internal(inspect(attr(log, '.internal.selfref'), 1))
@55994fef2528 22 EXTPTRSXP g0c0 [REF(65535)] <(nil)>
PROTECTED:
  @55994fef2560 22 EXTPTRSXP g0c0 [REF(2)] <(nil)>
TAG:
  @55995985f2a8 16 STRSXP g0c4 [REF(1),gp=0x20] (len=5, tl=0, gr)

The code is trying to add a new column by group to a freshly unserialized data.table. This used to work (replacing the table with an over-allocated shallow copy in the caller's environment) but now doesn't (with a different error message):

as.data.table(mtcars) |> serialize(NULL) |> unserialize() -> x
x[, foo := mean(mpg), by = "cyl"]
# Error in `[.data.table`(x, , `:=`(foo, mean(mpg)), by = "cyl") : 
#   This data.table has either been loaded from disk (e.g. using readRDS()/load()) or constructed manually (e.g. using structure()). Please run setDT() or setalloccol() on it first (to pre-allocate space for new columns) before assigning by reference to it.

This is similar to #7488, but in a different code path.

Activity

  1. added theissue type on Dec 22, 2025
  2. aitap commented on Dec 22, 2025

    @aitap
    MemberAuthor

    This is due to #6789. A check for ok < 1 || truelength(x) < ncol(x)+length(newnames) used to follow the ok == 0 check in the common case. Now it's been moved inside the if (missingby || bynull || (!byjoin && !length(byval))) branch, so it doesn't happen for group-by operations.

    Too bad dogroups has to assign the result by itself, so there's no post-factum check on the R side after jval is known.

  3. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    A similar error for {circhelp}:

    circhelp::Pascucci_et_al_2019_data[, prev_ori := shift(orientation), by = observer]
    # Error in `[.data.table`(circhelp::Pascucci_et_al_2019_data, , `:=`(prev_ori,  : 
    #   This data.table has either been loaded from disk (e.g. using readRDS()/load()) or constructed manually (e.g. using structure()). Please run setDT() or setalloccol() on it first (to pre-allocate space for new columns) before assigning by reference to it.
    
  4. ben-schwen commented on Dec 22, 2025

    @ben-schwen
    Member

    Maybe revert #6789 since it is obviously causing the most problems?

  5. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    Another similar error in {DramaAnalysis}:

    DramaAnalysis::utteranceStatistics(DramaAnalysis::rksp.0)
    # Error in `[.data.table`(text, , `:=`(dl, .N), .(corpus, drama)) : 
    #   Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. Please report to the data.table issues tracker.
    
  6. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    Again similar for {hicp}

    library(data.table)
    download.file(
      "https://github.com/cran/hicp/raw/refs/heads/master/vignettes/data/hicp_itemweights.RData",
      tmp<-tempfile())
    load(tmp)
    
    item.weights[, "t1" := hicp::tree(id=coicop, w=values, flag=TRUE, settings=list(w.tol=0.1)), by=c("geo","time")]
    # Error in `[.data.table`(item.weights, , `:=`("t1", hicp::tree(id = coicop,  : 
    #   Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. Please report to the data.table issues tracker.
  7. ben-schwen commented on Dec 22, 2025

    @ben-schwen
  8. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
  9. TysonStanley commented on Dec 22, 2025

    @TysonStanley
    Member

    Looks like {webtrackR} is also having the same issue:

    Running ‘testthat.R’
    ERROR
        Running the tests in ‘tests/testthat.R’ failed.
        Last 13 lines of output:
           1. └─webtrackR::add_duration(wt) at test-summarize.R:62:5
           2.   ├─...[]
           3.   └─data.table:::`[.data.table`(...)
          ── Error ('test-summarize.R:108:5'): sum_durations testdt_specific ─────────────
          Error in ``[.data.table`(wt, , `:=`(duration, as.numeric(data.table::shift(timestamp, n = 1, type = "lead", fill = NA) - timestamp)), by = "panelist_id")`: Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. Please report to the data.table issues tracker.
          Backtrace:
              ▆
           1. └─webtrackR::add_duration(wt) at test-summarize.R:108:5
           2.   ├─...[]
           3.   └─data.table:::`[.data.table`(...)
          
          [ FAIL 12 | WARN 22 | SKIP 7 | PASS 107 ]
          Error:
          ! Test failures.
          Execution halted
    
  10. aitap commented on Dec 22, 2025

    @aitap
    MemberAuthor

    We still need to detect column deletion before attempting to resize the table. Some of it could be detected statically (before evaluating jsub), but not all (think foo[, bar := if (runif(1) < .5) baz else NULL]). We're stuck between the Scilla of some revdeps attempting to delete columns on non-selfrefok tables (which used to work but now doesn't) and the Charybdis of this check only being fully possible after jval is known. #7502 is a kludge that passes our R CMD check and fixes bbotk, but that's not much.

  11. TysonStanley commented on Dec 22, 2025

    @TysonStanley
    Member

    And {ume}, {SeaVal}, {metR}

  12. TysonStanley commented on Dec 22, 2025

    @TysonStanley
    Member

    And {blocking}

  13. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    Ditto for {slim}:

    slim::dialysis[, fully_observed := (length(month) == 5), by = id]
    # Error in `[.data.table`(slim::dialysis, , `:=`(fully_observed, (length(month) ==  : 
    #   Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. Please report to the data.table issues tracker.
  14. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    A slightly different error message for {tidyrules}, but seems at first glance like the same issue:

    tidyrules::tidy(partykit::ctree(bill_length_mm ~ ., data = palmerpenguins::penguins))
    # Error in `[.data.table`(~.df, , `:=`(average = weighted.mean(response,  : 
    #   Internal error in assign: input dt has not been allocated enough column slots. l=3, tl=3, adding 1. Please report to the data.table issues tracker.
  15. MichaelChirico commented on Dec 22, 2025

    @MichaelChirico
    Member

    Ditto {tidytable}:

    tidytable::add_count(tidytable::tidytable(g = c(1, 2, 2, 2), val = c(1, 1, 2, 3)), g)
    # Error in `[.data.table`(~.df, , `:=`(n = .N), by = "g", keyby = FALSE) : 
    #   Internal error in dogroups: Trying to add new column by reference but tl is full; setalloccol should have run first at R level before getting to this point. Please report to the data.table issues tracker.
  16. aitap commented on Dec 24, 2025

    @aitap
    MemberAuthor
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    revdepReverse dependencies

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions