Skip to content

Pytorch CUDA P100 GPU Incompatibility #1546

Description

@JH766

🐛 Bug

The model fails to train on GPU due to a CUDA capability mismatch. The installed version of PyTorch requires CUDA capability sm_70 or higher, but the available GPU (Tesla P100-PCIE-16GB) only supports sm_60. Here is traceback:

/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning: 
    Found GPU0 Tesla P100-PCIE-16GB which is of cuda capability 6.0.
    Minimum and Maximum cuda capability supported by this version of PyTorch is
    (7.0) - (12.0)
    
  queued_call()
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning: 
    Please install PyTorch with a following CUDA
    configurations:  12.6 following instructions at
    https://pytorch.org/get-started/locally/
    
  queued_call()
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:435: UserWarning: 
Tesla P100-PCIE-16GB with CUDA capability sm_60 is not compatible with the current PyTorch installation.
The current PyTorch install supports CUDA capabilities sm_70 sm_75 sm_80 sm_86 sm_90 sm_100 sm_120.
If you want to use the Tesla P100-PCIE-16GB GPU with PyTorch, please check the instructions at https://pytorch.org/get-started/locally/

  queued_call()

To Reproduce

Use P100 GPU on Kaggle, move Pytorch neural network model to GPU and start training loop.

Expected behavior

Successful forward training pass without errors.

Additional context

This was the error given:

AcceleratorError: CUDA error: no kernel image is available for execution on the device
Search for `cudaErrorNoKernelImageForDevice' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.

Code worked fine for me a few weeks ago, but I'm guessing a change in the Docker environment broke something.

Activity

  1. HxH-h commented on Apr 1, 2026

    @HxH-h

    You can execute the following command to revert to a previous version

    pip uninstall -y torch torchvision torchaudio
    pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118
    
  2. Mr-Neutr0n commented on Jun 14, 2026

    @Mr-Neutr0n

    Also confirming this issue on our end. We spent several hours debugging this across multiple Kaggle notebook versions.

    What we found:

    • The issue affects BOTH Tesla P100 (sm_60) and Tesla T4 (sm_75) GPUs on Kaggle
    • System PyTorch 2.10.0+cu128 is missing CUDA kernels for these architectures
    • Simple tensor allocation works (torch.zeros(1).cuda()) but any actual operation (.ne(), relu(), model forward pass) crashes with:
      AcceleratorError: CUDA error: no kernel image is available for execution on the device

    Workaround we found (hacky but functional):
    Install PyTorch 2.10.0+cu126 to a custom directory and load it before the system one:

    import sys, subprocess, os
    torch_dir = "/tmp/torch_working"
    os.makedirs(torch_dir, exist_ok=True)
    subprocess.check_call([
        sys.executable, "-m", "pip", "install", "--target", torch_dir, "--no-cache-dir",
        "torch==2.10.0", "--index-url", "https://download.pytorch.org/whl/cu126"
    ])
    sys.path.insert(0, torch_dir)
    import torch  # This loads the working PyTorch, not the broken system one

    Caveat: This workaround has issues with --target installation for packages with compiled C extensions (e.g., torchaudio fails with OSError: undefined symbol). We had to skip torchaudio and torchvision entirely.

    Real fix needed: Kaggle needs to update the Docker image to either:

    1. PyTorch 2.10+cu126 (supports sm_60 and sm_75), or
    2. PyTorch 2.11+ (which reportedly fixes the missing kernel issue)

    This is a significant regression — Kaggle GPU is effectively unusable for PyTorch training out of the box right now. The suggested workaround of pip install torch torchvision torchaudio --index-url https://download.pytorch.org/whl/cu118 from the thread above is not ideal (cu118 is quite old at this point).

  3. Mr-Neutr0n commented on Jun 14, 2026

    @Mr-Neutr0n

    Update: We found a better workaround than --target install. The issue is that the broken PyTorch (2.10+cu128) is loaded into memory by the notebook runner before our code executes.

    Better fix: Uninstall the broken system PyTorch before importing it, install 2.10+cu126 to the system path, then import fresh. This avoids the C extension reload problem entirely.

    import subprocess, sys
    subprocess.run([sys.executable, "-m", "pip", "uninstall", "-y", "torch", "torchvision", "torchaudio"])
    subprocess.run([sys.executable, "-m", "pip", "install", "--no-cache-dir", "torch==2.10.0", "torchvision", "torchaudio", "--index-url", "https://download.pytorch.org/whl/cu126"])
    # NOW import torch for the first time
    import torch

    Verified on P100 (sm_60) and works. torch.zeros(1).cuda(), tensor.ne(), relu(), all pass. The --target install approach we tried earlier had issues with shared library linking for compiled extensions like torchaudio.

    Still hoping Kaggle can update the Docker image to save everyone this workaround.

  4. Mr-Neutr0n commented on Jun 14, 2026

    @Mr-Neutr0n

    Update: The issue is P100-specific, not universal.

    After testing both GPU types on Kaggle with the default PyTorch 2.10.0+cu128:

    • Tesla T4 (sm_75): Everything works natively. No fix needed. torch.matmul, nn.Linear, .sum(), .backward() all pass.
    • Tesla P100 (sm_60): Fails on .sum() and .backward() with torch.AcceleratorError: CUDA error: no kernel image is available for execution on the device. Matmul forward passes because it routes through cuBLAS (NVIDIA's library), but PyTorch's custom kernels for reduction ops are missing sm_60 in the cu128 build.

    PyTorch 2.10.0+cu128 literally dropped sm_60 support. The warning says: "The current PyTorch install supports CUDA capabilities sm_70 sm_75 sm_80 sm_86 sm_90 sm_100 sm_120". Minimum is now sm_70 (V100).

    Workaround for P100 users:

    import subprocess, sys
    subprocess.run([sys.executable, '-m', 'pip', 'uninstall', '-y', 'torch', 'torchvision', 'torchaudio'])
    subprocess.run([sys.executable, '-m', 'pip', 'install', '--no-cache-dir', 'torch==2.10.0', 'torchvision', 'torchaudio', '--index-url', 'https://download.pytorch.org/whl/cu126'])

    PyTorch 2.10.0+cu126 still includes sm_60 kernels and fixes the issue. But T4 users should skip this -- the default cu128 is actually faster and has newer CUDA features.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugbug & failures with existing packageshelp wanted

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions