Skip to content

valgrind issue in 1.18.0 #7546

Description

@MichaelChirico

From CRAN:

https://www.stats.ox.ac.uk/pub/bdr/memtests/valgrind/data.table/00check.log

  Test 6001.732 ran without errors but failed check that x equals y:
  > x = frollsd(y, 3)[4L] 
  First 1 of 1 (type 'double'): 
  [1] 1.825012e-08
  > y = 0 
  First 1 of 1 (type 'double'): 
  [1] 0
  Mean relative difference: 1
  Error in test.data.table(script = "froll.Rraw") : 
    1 error out of 1316. Search tests/froll.Rraw.bz2 for test number 6001.732. Duration: 00:03:47 elapsed (00:03:20 cpu).
  Calls: test.data.table -> stopf -> raise_condition -> signal
  Execution halted
  ==3229069== 
  ==3229069== HEAP SUMMARY:
  ==3229069==     in use at exit: 92,363,323 bytes in 16,467 blocks
  ==3229069==   total heap usage: 174,670 allocs, 158,203 frees, 332,134,037 bytes allocated
  ==3229069== 
  ==3229069== 432 bytes in 1 blocks are possibly lost in loss record 358 of 2,244
  ==3229069==    at 0x484B133: calloc (/builddir/build/BUILD/valgrind-3.24.0/coregrind/m_replacemalloc/vg_replace_malloc.c:1675)
  ==3229069==    by 0x4011F63: UnknownInlinedFun (/usr/src/debug/glibc-2.39-38.fc40.x86_64/elf/../include/rtld-malloc.h:44)
  ==3229069==    by 0x4011F63: allocate_dtv (/usr/src/debug/glibc-2.39-38.fc40.x86_64/elf/../elf/dl-tls.c:395)
  ==3229069==    by 0x4012A61: _dl_allocate_tls (/usr/src/debug/glibc-2.39-38.fc40.x86_64/elf/../elf/dl-tls.c:673)
  ==3229069==    by 0x557CC03: allocate_stack (/usr/src/debug/glibc-2.39-38.fc40.x86_64/nptl/allocatestack.c:431)
  ==3229069==    by 0x557CC03: pthread_create@@GLIBC_2.34 (/usr/src/debug/glibc-2.39-38.fc40.x86_64/nptl/pthread_create.c:660)
  ==3229069==    by 0x54B0076: gomp_team_start (/usr/src/debug/gcc-14.2.1-3.fc40.x86_64/obj-x86_64-redhat-linux/x86_64-redhat-linux/libgomp/../../../libgomp/team.c:859)
  ==3229069==    by 0x54A60A0: GOMP_parallel (/usr/src/debug/gcc-14.2.1-3.fc40.x86_64/obj-x86_64-redhat-linux/x86_64-redhat-linux/libgomp/../../../libgomp/parallel.c:176)
  ==3229069==    by 0x173DE573: frollfunR (packages/tests-vg/data.table/src/frollR.c:208)
  ==3229069==    by 0x4A711D: R_doDotCall (svn/R-devel/src/main/dotcode.c:790)
  ==3229069==    by 0x4E1283: bcEval_loop (svn/R-devel/src/main/eval.c:8682)
  ==3229069==    by 0x4F13D7: bcEval (svn/R-devel/src/main/eval.c:7515)
  ==3229069==    by 0x4F13D7: bcEval (svn/R-devel/src/main/eval.c:7500)
  ==3229069==    by 0x4F170A: Rf_eval (svn/R-devel/src/main/eval.c:1167)
  ==3229069==    by 0x4F348D: R_execClosure (svn/R-devel/src/main/eval.c:2389)
  ==3229069== 
  ==3229069== LEAK SUMMARY:
  ==3229069==    definitely lost: 0 bytes in 0 blocks
  ==3229069==    indirectly lost: 0 bytes in 0 blocks
  ==3229069==      possibly lost: 432 bytes in 1 blocks
  ==3229069==    still reachable: 92,360,875 bytes in 16,445 blocks
  ==3229069==         suppressed: 0 bytes in 0 blocks
  ==3229069== Reachable blocks (those to which a pointer was found) are not shown.
  ==3229069== To see them, rerun with: --leak-check=full --show-leak-kinds=all
  ==3229069== 
  ==3229069== For lists of detected and suppressed errors, rerun with: -s
  ==3229069== ERROR SUMMARY: 375255 errors from 4 contexts (suppressed: 0 from 0)

Possibly, it just means we need to include another test to be skipped:

  **** Skipping 7 NaN/NA algo='exact' tests because .Machine$longdouble.digits==53 (!=64); e.g. under valgrind

Activity

  1. added this to the 1.18.2 milestone on Dec 29, 2025
  2. self-assigned this
    on Dec 29, 2025
  3. aitap commented on Dec 29, 2025

    @aitap
    Member
  4. MichaelChirico commented on Dec 30, 2025

    @MichaelChirico
    MemberAuthor

    R doesn't struggle with computing sd(c(1e8, 1e8, 1e8)) to be 0 under Valgrind.

    Gemini helped me with the intuition here:

    • We shouldn't expect the same result as sd(x) because, for efficiency, we use the previous window's results and update. So the c(1e8+eps, 1e8, 1e8) window's tiny variance "taints" the calculation for the c(1e8, 1e8, 1e8) window.
    • And then it's just a matter of tolerances/numerical error. On full-precision platforms, there's probably still a tiny non-0 value in the fast result, but small enough all.equal() passes. But when switching to 53 bytes, the tiny difference becomes too large & bubbles out.
  5. jangorecki commented on Jan 3, 2026

    @jangorecki
    Member

    There are two separate issues.

    One is that under Valgrind, data.table::frollsd(c(1e8+2.980232e-8, 1e8,
    1e8, 1e8), 3) returns c(NA, NA, 1.82501207499443e-08,
    1.82501207499443e-08) instead of c(NA, NA, 1.72058537479884e-08, 0).
    This should be solved by #7548, although base R doesn't struggle with
    computing sd(c(1e8, 1e8, 1e8)) to be 0 under Valgrind.

    The other is that when frollmedian(1:10, 3) tries to compare

    if (n[A]!=tail && m[A] == n[A]) {

    at froll.c:1710, n[0] is uninitialised:

    (gdb) p A
    $26 = 0
    (gdb) p m
    $27 = (int *) 0x8875700
    (gdb) monitor xb 0x8875700 4
    00 00 00 00 <-- all valid
    0x8875700: 0x02 0x00 0x00 0x00
    (gdb) p n
    $28 = (int *) 0x88c6140
    (gdb) monitor xb 0x88c6140 4
    ff ff ff ff <-- whole word invalid
    0x88C6140: 0x00 0x00 0x00 0x00

    Possibly because the assignment at froll.c:1622

    if (even)
      n[j] = o[j*k+h+1];
    

    was skipped due to 'even' being false.

    The second one is still not resolved afaik

  6. added 2 commits that reference this issue on Jan 12, 2026
    5987179
    f11e0ff
  7. added a commit that references this issue on Jan 14, 2026
    21bd6b0
  8. added a commit that references this issue on Jan 15, 2026
    d3d0019
  9. added a commit that references this issue on Jan 27, 2026
    afe1218
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions