Repository navigation
Cuda 3.0? #25
Description
Activity
Officially, Cuda compute capability 3.5 and 5.2 are supported. You can try to enable other compute capability by modifying the build script:
Reacted by PAUL MCQUESTENThanks! Will try it and report here.
This is not officially supported yet. But if you want to enable Cuda 3.0 locally, here are the additional places to change:
https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/common_runtime/gpu/gpu_device.cc#L610
https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/common_runtime/gpu/gpu_device.cc#L629
Where the smaller GPU device is ignored.The official support will eventually come in a different form, where we make sure the fix works on all different computational environment.
Reacted by PAUL MCQUESTENI made the changes to the lines above, and was able to compile and run the basic example on the Getting Started page: http://tensorflow.org/get_started/os_setup.md#try_your_first_tensorflow_program - it did not complain about gpu, but it didn't report using the gpu either.
How can I help with next steps?
Reacted by PAUL MCQUESTENinfojunkie@, could you post your step and upload the log?
If you were following this example:
bazel build -c opt --config=cuda //tensorflow/cc:tutorials_example_trainer
bazel-bin/tensorflow/cc/tutorials_example_trainer --use_gpuIf you see the following line, the GPU logic device is being created:
Creating TensorFlow device (/gpu:0) -> (device: ..., name: ..., pci bus id: ...)
If you want to be absolutely sure GPU was used, set CUDA_PROFILE=1 and enable Cuda profiler. If the Cuda profiler logs were generated, it was a sure sign GPU was used.
http://docs.nvidia.com/cuda/profiler-users-guide/#command-line-profiler-control
Reacted by PAUL MCQUESTENI got the following log:
I tensorflow/core/common_runtime/local_device.cc:25] Local device intra op parallelism threads: 8 I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:888] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero I tensorflow/core/common_runtime/gpu/gpu_init.cc:88] Found device 0 with properties: name: GeForce GT 750M major: 3 minor: 0 memoryClockRate (GHz) 0.967 pciBusID 0000:02:00.0 Total memory: 2.00GiB Free memory: 896.49MiB I tensorflow/core/common_runtime/gpu/gpu_init.cc:112] DMA: 0 I tensorflow/core/common_runtime/gpu/gpu_init.cc:122] 0: Y I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_region_allocator.cc:47] Setting region size to 730324992 I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/gpu/gpu_device.cc:643] Creating TensorFlow device (/gpu:0) -> (device: 0, name: GeForce GT 750M, pci bus id: 0000:02:00.0) I tensorflow/core/common_runtime/local_session.cc:45] Local session inter op parallelism threads: 8I guess it means the GPU was found and used. I can try the CUDA profiler if you think it's useful.
Please prioritize this issue. It is blocking gpu usage on both OSX and AWS's K520 and for many people this is the only environments available.
Thanks!Reacted by Haoyuan GeFor reference, here's my very primitive patch to work with Cuda 3.0: https://gist.github.com/infojunkie/cb6d1a4e8bf674c6e38e
Reacted by PAUL MCQUESTEN@infojunkie I applied your fix, but I got lots of nan's in the computation output:
$ bazel-bin/tensorflow/cc/tutorials_example_trainer --use_gpu 000006/000003 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000004/000003 lambda = 2.000027 x = [79795.101562 -39896.468750] y = [159592.375000 -79795.101562] 000005/000006 lambda = 2.000054 x = [39896.468750 -19947.152344] y = [79795.101562 -39896.468750] 000001/000007 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000002/000003 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000009/000008 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000004/000004 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000001/000005 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000006/000007 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000003/000006 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000] 000006/000006 lambda = -nan x = [0.000000 0.000000] y = [0.000000 0.000000]@markusdr, this is very strange. Could you post the completely steps you build the binary?
Could what GPU and OS are you running with? Are you using Cuda 7.0 and Cudnn 6.5 V2?
Just +1 to fix this problem on AWS as soon as possible. We don't have any other GPU cards for our research.
Hi, not sure if this is a separate issue but I'm trying to build with a CUDA 3.0 GPU (Geforce 660 Ti) and am getting many errors with --config=cuda. See the attached file below. It seems unrelated to the recommended changes above. I've noticed that it tries to compile a temporary compute_52.cpp1.ii file which would be the wrong version for my GPU.
I'm on Ubuntu 15.10. I modified the host_config.h in the Cuda includes to remove the version check on gcc. I'm using Cuda 7.0 and cuDNN 6.5 v2 as recommended, although I have newer versions installed as well.
Yes, I was using Cuda 7.0 and Cudnn 6.5 on an EC2 g2.2xlarge instance with this AIM:
cuda_7 - ami-12fd8178
ubuntu 14.04, gcc 4.8, cuda 7.0, atlas, and opencv.
To build, I followed the instructions on tensorflow.org.94 remaining items
- added a commit that references this issue
on Dec 6, 2019 - added a commit that references this issue
on Feb 1, 2021 - added 5 commits that reference this issue
on Apr 9, 2025 - added 5 commits that reference this issue
on Jul 28, 2025 - added a commit that references this issue
on Nov 26, 2025

Are there plans to support Cuda compute capability 3.0?