Describe the bug
spla sizes single-group work-groups from CL_DEVICE_MAX_WORK_GROUP_SIZE
instead of CL_KERNEL_WORK_GROUP_SIZE. On Imagination PowerVR
(PowerVR B-Series BXE-2-32, EMBEDDED_PROFILE, kernels built with
-cl-std=CL1.2) the oversized launch is refused with
CL_OUT_OF_RESOURCES, the status is never checked, and the failure is
silent. E.g. belgium_osm (Belgian road network, about 1.4M vertices — any graph
reaching 481..512 intermediate products triggers it): SSSP stops early, the
feedback vector comes back empty while the frontier is non-empty, distances
stay at FLT_MAX (99.7% of vertices mismatched, 1,436,362 VERIFY lines
total):
CHECK cpu
CHECK acc
VERIFY: 40 expected 465 actual 3.40282e+38
VERIFY: 41 expected 464 actual 3.40282e+38
VERIFY: 42 expected 416 actual 3.40282e+38
... (1,436,362 lines total)
(Full crash log is 69MB; head and tail with exact omitted-line counts
attached as base_belgium_crash_excerpt.log.)
To Reproduce
- Build
sssp for RISC-V (spla upstream main @ b81d1e29, Syntacore
toolchain) and run it on the device above.
- Execute
./sssp --mtxpath=belgium_osm.mtx --niters=1.
- Get an early stop with a flood of
VERIFY: ... actual 3.40282e+38 lines.
- Observe the mechanism isolated from spla with the standalone reproducer
(single work-group of align(n, 32) threads running reduce_by_key_small):
Gist reproducer
(full output attached as test_reduce_small_limits.log): with n
in [481, 512] the launch returns CL_OUT_OF_RESOURCES, with n <= 480
it succeeds.
Expected behavior
Full convergence matching the reference: the in-tree verification
(CHECK acc, which compares the OpenCL backend against the CPU backend
and the naive host reference) reports no mismatches. Both reference
implementations live outside src/opencl/* touched by this defect, so the
check is independent of the changed code.
Environment
- OS: Bianbu 3.0.1 (Ubuntu-derived SpacemiT distro), Linux 6.6.63, riscv64
(Milk-V Jupiter)
- GPU:
PowerVR B-Series BXE-2-32, Imagination Technologies
- GPU driver version:
24.2@6603887
- OpenCL 3.0 runtime (
EMBEDDED_PROFILE), Device OpenCL C 1.2, kernels
built for OpenCL 1.2
Build Configuration
- Compiler: RISC-V GCC 14.2.0 (Syntacore toolchain)
- SDK: Khronos headers bundled in-tree (
deps/opencl-headers), no vendor SDK
- CMake 4.4.3, Ninja 1.13.2
- OpenCL 1.2 (kernel language version)
Additional context
spla computes local = align(n, 32), so n <= 480 maps to groups of at
most 480 threads and n >= 481 maps to exactly 512 — no gap, no overlap.
The per-kernel limit on this device is 480 for reduce_by_key_small (each
kernel reports its own value; the bitonic kernels in
cl_sort_by_key_bitonic (src/opencl/cl_sort_by_key.hpp) share the same
single-group pattern and are gated the same way as a preventive measure —
no failure was observed there). The driver is within its rights to refuse
the oversized launch; with the status unchecked, outputs stay zeroed and a
stale counter is read back. Per the spec, CL_KERNEL_WORK_GROUP_SIZE is
"the maximum work-group size that can be used to execute the kernel",
computed from "the resource requirements of the kernel (register usage
etc.)".
Suggested fix: gate single-group launches by CL_KERNEL_WORK_GROUP_SIZE
queried from the built kernel (cl_reduce_by_key, bitonic sort; helper
kernel_group_limit declared in src/opencl/cl_program_builder.hpp,
defined in src/opencl/cl_program_builder.cpp), falling back to the existing
small-group paths (offsets+scan+scalar, radix sort). No vendor checks, no
magic numbers.
Verified with the fix: belgium_osm converges in 1459 iterations
(identical to the unmodified CPU backend count), roadNet-CA converges
fully; no mismatches in any run (full logs attached as final_test_*.log).
test_reduce_small_limits.log
base_belgium_crash_excerpt.log
final_test_roadNet-CA.log
final_test_belgium_osm.log
Describe the bug
spla sizes single-group work-groups from
CL_DEVICE_MAX_WORK_GROUP_SIZEinstead of
CL_KERNEL_WORK_GROUP_SIZE. On Imagination PowerVR(
PowerVR B-Series BXE-2-32,EMBEDDED_PROFILE, kernels built with-cl-std=CL1.2) the oversized launch is refused withCL_OUT_OF_RESOURCES, the status is never checked, and the failure issilent. E.g.
belgium_osm(Belgian road network, about 1.4M vertices — any graphreaching 481..512 intermediate products triggers it): SSSP stops early, the
feedback vector comes back empty while the frontier is non-empty, distances
stay at
FLT_MAX(99.7% of vertices mismatched, 1,436,362VERIFYlinestotal):
(Full crash log is 69MB; head and tail with exact omitted-line counts
attached as
base_belgium_crash_excerpt.log.)To Reproduce
ssspfor RISC-V (spla upstreammain@b81d1e29, Syntacoretoolchain) and run it on the device above.
./sssp --mtxpath=belgium_osm.mtx --niters=1.VERIFY: ... actual 3.40282e+38lines.(single work-group of
align(n, 32)threads runningreduce_by_key_small):Gist reproducer
(full output attached as
test_reduce_small_limits.log): withnin
[481, 512]the launch returnsCL_OUT_OF_RESOURCES, withn <= 480it succeeds.
Expected behavior
Full convergence matching the reference: the in-tree verification
(
CHECK acc, which compares the OpenCL backend against the CPU backendand the naive host reference) reports no mismatches. Both reference
implementations live outside
src/opencl/*touched by this defect, so thecheck is independent of the changed code.
Environment
(Milk-V Jupiter)
PowerVR B-Series BXE-2-32, Imagination Technologies24.2@6603887EMBEDDED_PROFILE), Device OpenCL C 1.2, kernelsbuilt for OpenCL 1.2
Build Configuration
deps/opencl-headers), no vendor SDKAdditional context
spla computes
local = align(n, 32), son <= 480maps to groups of atmost 480 threads and
n >= 481maps to exactly 512 — no gap, no overlap.The per-kernel limit on this device is 480 for
reduce_by_key_small(eachkernel reports its own value; the bitonic kernels in
cl_sort_by_key_bitonic(src/opencl/cl_sort_by_key.hpp) share the samesingle-group pattern and are gated the same way as a preventive measure —
no failure was observed there). The driver is within its rights to refuse
the oversized launch; with the status unchecked, outputs stay zeroed and a
stale counter is read back. Per the spec,
CL_KERNEL_WORK_GROUP_SIZEis"the maximum work-group size that can be used to execute the kernel",
computed from "the resource requirements of the kernel (register usage
etc.)".
Suggested fix: gate single-group launches by
CL_KERNEL_WORK_GROUP_SIZEqueried from the built kernel (
cl_reduce_by_key, bitonic sort; helperkernel_group_limitdeclared insrc/opencl/cl_program_builder.hpp,defined in
src/opencl/cl_program_builder.cpp), falling back to the existingsmall-group paths (
offsets+scan+scalar, radix sort). No vendor checks, nomagic numbers.
Verified with the fix:
belgium_osmconverges in 1459 iterations(identical to the unmodified CPU backend count),
roadNet-CAconvergesfully; no mismatches in any run (full logs attached as
final_test_*.log).test_reduce_small_limits.log
base_belgium_crash_excerpt.log
final_test_roadNet-CA.log
final_test_belgium_osm.log