Skip to content

frollapply could optionally use mirai+mori rather than base R parallel's fork #7891

Description

@jangorecki

frollapply now uses base R's parallel package fork for parallel computation of FUN. Adding mirai and mori packages together it could provide an alternative backend for parallel processing in frollapply.
Cost of it means 2 extra suggested dependency (3 pkgs in total, including nanonext, dep of mirai).
Benefits are that parallel processing will work on Windows, and other OSes where fork is not available.
Moreover, posit workbenchmark is not handling fork well, so that could be an alternative for posit workbench users.

Fast rolling user-defined function (\emph{UDF}) to calculate on a sliding window. Experimental. Please read, at least, \emph{caveats} section below. For "time-aware" (irregularly spaced time series) rolling function see \code{\link{frolladapt}}.

note that frollapply is still experimental which gives more freedom to re-work

Activity

  1. jangorecki commented on Sep 22, 2026

    @jangorecki
    MemberAuthor

    @MichaelChirico @aitap @ben-schwen @tdhock just wanted to check with you if we all agree that we are open for adding mirai and mori to suggested dependencies?

  2. MichaelChirico commented on Sep 22, 2026

    @MichaelChirico
    Member

    Keeping it Suggests SGTM. Is the proposal to fall back to {parallel} if {mirai} unavailable, or to disable parallelism in that case?

    The former means continuing to maintain two code paths, while the latter means possibly less functionality if {mirai} is difficult to install in some environments (I'm not familiar enough to say if that's moot).

    Has R declared similar functionality as {mirai} offers as out-of-scope / won't be supported in {parallel}?

  3. jangorecki commented on Sep 23, 2026

    @jangorecki
    MemberAuthor

    Is the proposal to fall back to {parallel} if {mirai} unavailable, or to disable parallelism in that case?

    I would say that on linux default should be still parallel. mirai has to spawn fresh R session for each worker so overhead of that will be still there.

    Has R declared similar functionality as {mirai} offers as out-of-scope / won't be supported in {parallel}?

    I think it is supported, but we are not using high level wrappers that could let us use mirai through parallel, we are using fork-related functions exported from parallel.

  4. aitap commented on Sep 23, 2026

    @aitap
    Member

    It should be fine to add both packages to Suggests.

    mori is as minimal as it gets, I'm actually not sure why it Suggests: mirai; there seem to be no references to it inside the package. We could reimplement our own shared-memory ALTREP classes, but it's not fun, and mori already does the job. (Anything non-ALTREP'pable is still serialized. mori requires R ≥ 4.3, although character / integer / logical / numeric ALTREP classes appeared in 3.5.0 and raw / complex appeared in 3.6.0. Support for lists does require 4.3.0 or above.)

    (Edit: shared memory is actually not zero-copy, but copy-once: to share a vector, we have to copy it into a new shared memory region. Not even Linux is crazy enough to let a program mmap() someone else's address space through /proc/PID/mem for zero-copy sharing.)

    mirai requires nanonext (libnng), which has problems with TLS and HTTP in some configurations, but should work fine for our purposes (as a replacement for parallel::mcparallel on Windows).

    Has R declared similar functionality as {mirai} offers as out-of-scope / won't be supported in {parallel}?

    Support for mirai clusters in parallel has to do with snow-style (distributed) clusters which data.table doesn't use.

    Many people have wanted something like parallel::mcparallel on Windows, but the exact semantics are very hard to replicate (you'd have to be Cygwin). Running something in a background R process, on the other hand, is relatively easy (start with tools::R but use system2(wait = FALSE) and some other means of discovering whether the worker process is done). mirai's added value is in doing everything much faster: libnng instead of temporary files for IPC and a custom serialization format for common object kinds instead of base serialize(). In theory, we could use just mori and hand-launched child processes.

    Users whose R is not forkable (RStudio, Ark, ...) will probably prefer the non-mcparallel approach even on non-Windows.

  5. tdhock commented on Sep 23, 2026

    @tdhock
    Member

    Suggests sounds ok to me (although I don’t use frollapply).
    For parallel interfaces I tend to do something like this

    frollapply = function(…, LAPPLY=lapply){
    …
    }

    then the user can provide a parallel function similar to lapply (future_lapply etc), or just keep regular lapply (useful for debugging).
    See for example proj_compute_all() in https://github.com/tdhock/mlr3resampling/blob/main/R/proj.R

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions