Skip to content

SUGGESTION: Run the OpenCL tests under Oclgrind in CI to catch data races and invalid memory accesses in kernels #3430

Description

@jachymb

Description

Motivation. CI runs the OpenCL tests on one GPU. A kernel bug that this device happens to tolerate stays invisible. Example: #3429, where the cholesky_decompose kernel returns a wrong factor on an AMD GPU because its barriers fence local memory only.

Oclgrind is an OpenCL device simulator. It needs no GPU and checks what real devices do not:

  • data races between work items, including a barrier with the wrong fence flag (--data-races),
  • reads and writes outside a buffer (always on),
  • use of uninitialized values (--uninitialized).

It runs our test binaries unchanged:

oclgrind --data-races --log races.log ./test/unit/math/opencl/cholesky_decompose_test

On develop this reports read-write races in cholesky_decompose, already for the 3×3 case, and nothing once the barriers are fixed. A first local pass also reports races in tridiagonalization_householder.

Suggested steps.

  1. Add oclgrind to the CI image.
  2. Add a CPU-only stage that builds test/unit/math/opencl with STAN_OPENCL=true and runs each test binary through oclgrind --data-races --log <file>. runTests.py already prefixes MPI tests with mpirun; the same mechanism can prefix oclgrind.
  3. Fail the stage when a log is not empty.
  4. Start it as non-blocking and make it required once the existing reports are fixed.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions