Skip to content

OpenCL: single-group launches exceeding per-kernel work-group limit fail silently on PowerVR #246

Description

@Redkin-Aleksey

Describe the bug

spla sizes single-group work-groups from CL_DEVICE_MAX_WORK_GROUP_SIZE
instead of CL_KERNEL_WORK_GROUP_SIZE. On Imagination PowerVR
(PowerVR B-Series BXE-2-32, EMBEDDED_PROFILE, kernels built with
-cl-std=CL1.2) the oversized launch is refused with
CL_OUT_OF_RESOURCES, the status is never checked, and the failure is
silent. E.g. belgium_osm (Belgian road network, about 1.4M vertices — any graph
reaching 481..512 intermediate products triggers it): SSSP stops early, the
feedback vector comes back empty while the frontier is non-empty, distances
stay at FLT_MAX (99.7% of vertices mismatched, 1,436,362 VERIFY lines
total):

CHECK cpu
CHECK acc
 VERIFY: 40 expected 465 actual 3.40282e+38
 VERIFY: 41 expected 464 actual 3.40282e+38
 VERIFY: 42 expected 416 actual 3.40282e+38
 ... (1,436,362 lines total)

(Full crash log is 69MB; head and tail with exact omitted-line counts
attached as base_belgium_crash_excerpt.log.)

To Reproduce

  1. Build sssp for RISC-V (spla upstream main @ b81d1e29, Syntacore
    toolchain) and run it on the device above.
  2. Execute ./sssp --mtxpath=belgium_osm.mtx --niters=1.
  3. Get an early stop with a flood of VERIFY: ... actual 3.40282e+38 lines.
  4. Observe the mechanism isolated from spla with the standalone reproducer
    (single work-group of align(n, 32) threads running reduce_by_key_small):
    Gist reproducer
    (full output attached as test_reduce_small_limits.log): with n
    in [481, 512] the launch returns CL_OUT_OF_RESOURCES, with n <= 480
    it succeeds.

Expected behavior

Full convergence matching the reference: the in-tree verification
(CHECK acc, which compares the OpenCL backend against the CPU backend
and the naive host reference) reports no mismatches. Both reference
implementations live outside src/opencl/* touched by this defect, so the
check is independent of the changed code.

Environment

  • OS: Bianbu 3.0.1 (Ubuntu-derived SpacemiT distro), Linux 6.6.63, riscv64
    (Milk-V Jupiter)
  • GPU: PowerVR B-Series BXE-2-32, Imagination Technologies
  • GPU driver version: 24.2@6603887
  • OpenCL 3.0 runtime (EMBEDDED_PROFILE), Device OpenCL C 1.2, kernels
    built for OpenCL 1.2

Build Configuration

  • Compiler: RISC-V GCC 14.2.0 (Syntacore toolchain)
  • SDK: Khronos headers bundled in-tree (deps/opencl-headers), no vendor SDK
  • CMake 4.4.3, Ninja 1.13.2
  • OpenCL 1.2 (kernel language version)

Additional context

spla computes local = align(n, 32), so n <= 480 maps to groups of at
most 480 threads and n >= 481 maps to exactly 512 — no gap, no overlap.
The per-kernel limit on this device is 480 for reduce_by_key_small (each
kernel reports its own value; the bitonic kernels in
cl_sort_by_key_bitonic (src/opencl/cl_sort_by_key.hpp) share the same
single-group pattern and are gated the same way as a preventive measure —
no failure was observed there). The driver is within its rights to refuse
the oversized launch; with the status unchecked, outputs stay zeroed and a
stale counter is read back. Per the spec, CL_KERNEL_WORK_GROUP_SIZE is
"the maximum work-group size that can be used to execute the kernel",
computed from "the resource requirements of the kernel (register usage
etc.)".

Suggested fix: gate single-group launches by CL_KERNEL_WORK_GROUP_SIZE
queried from the built kernel (cl_reduce_by_key, bitonic sort; helper
kernel_group_limit declared in src/opencl/cl_program_builder.hpp,
defined in src/opencl/cl_program_builder.cpp), falling back to the existing
small-group paths (offsets+scan+scalar, radix sort). No vendor checks, no
magic numbers.

Verified with the fix: belgium_osm converges in 1459 iterations
(identical to the unmodified CPU backend count), roadNet-CA converges
fully; no mismatches in any run (full logs attached as final_test_*.log).

test_reduce_small_limits.log

base_belgium_crash_excerpt.log

final_test_roadNet-CA.log

final_test_belgium_osm.log

Activity

  1. changed the title [-][bug] OpenCL single-group launches exceeding per-kernel work-group limit fail silently on PowerVR[/-] [+]OpenCL: single-group launches exceeding per-kernel work-group limit fail silently on PowerVR[/+] on Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions