Repository navigation
High cpu usage by kernel while using inv and solve from linalg #8120
Description
Activity
There seems to be a significant difference between running the code under ipython and standalone (<10% processor load in the latter case). Can you reproduce the issue with explicit concurrency (say using threads)?
Can you explain a little be more what you want me to do ? What do you mean by "explicit concurrency" ?
That was just a guess at what to research (how is ipython spawning 8 processes). How about something like the following, does it consume all of your resources? (Run as
python <filename>; seems to give load proportional to the number of processes, significantly lower under Python 3).from multiprocessing import Process import time import numpy as np def f(i, A, b): print('Worker {}'.format(i)) while True: np.linalg.solve(A, b) time.sleep(.01) if __name__ == '__main__': np.random.seed(0) A = np.asarray(np.random.random((56, 56))) b = np.eye(A.shape[0]) for i in range(3): Process(target=f, args=(i, A, b)).start()
@vsamy: So for context, what you're observing here isn't numpy itself, it's the BLAS library that numpy's calling to do the actual linear algebra operations. This doesn't mean that we can't help, but it makes debugging a little more complicated. For example, the difference you see between numpy 1.8 and numpy 1.11 is almost certainly not any change in numpy, but that your two installs are being configured differently.
I'm a little unclear on why you think there's a problem -- if you right a tight
forloop doing some linear algebra, then you should expect this to use CPU :-). Some BLAS libraries use multiple threads internally, so you should expect linear algebra to use all your CPUs. This is generally a good thing...?Where did your numpy come from? (anaconda? pip install? linux package manager? something else?) Can you paste the output of
np.show_config()?It is a good thing to use multiple threads but getting all the cpu ressources seem a little bit extreme.
This chunk of code should be running in parallel to another bigger code which must run fast enough.Numpy has been installed through a classic apt-get. Here is the output:
lapack_info: libraries = ['lapack', 'lapack'] library_dirs = ['/usr/lib'] language = f77 lapack_opt_info: libraries = ['lapack', 'lapack', 'blas', 'blas'] library_dirs = ['/usr/lib'] language = c define_macros = [('NO_ATLAS_INFO', 1), ('HAVE_CBLAS', None)] openblas_lapack_info: NOT AVAILABLE blas_info: libraries = ['blas', 'blas'] library_dirs = ['/usr/lib'] define_macros = [('HAVE_CBLAS', None)] language = c atlas_3_10_blas_threads_info: NOT AVAILABLE atlas_threads_info: NOT AVAILABLE atlas_3_10_threads_info: NOT AVAILABLE atlas_blas_info: NOT AVAILABLE atlas_3_10_blas_info: NOT AVAILABLE atlas_blas_threads_info: NOT AVAILABLE openblas_info: NOT AVAILABLE blas_mkl_info: NOT AVAILABLE blas_opt_info: libraries = ['blas', 'blas'] library_dirs = ['/usr/lib'] language = c define_macros = [('NO_ATLAS_INFO', 1), ('HAVE_CBLAS', None)] atlas_info: NOT AVAILABLE atlas_3_10_info: NOT AVAILABLE lapack_mkl_info: NOT AVAILABLE mkl_info: NOT AVAILABLEWhat does
ls -l /etc/alternatives/libblas.sosay?ls -l /etc/alternatives/libblas.sooutput/etc/alternatives/libblas.so -> /usr/lib/libblas/libblas.so
ls -l /usr/lib/libblas/libblas.sooutput/usr/lib/libblas/libblas.so -> libblas.so.3.6.0@njsmith So i ask a colleague to test it on his machine and it works perfectly fine. He has the same ubuntu/kernel/blas/numpy/python version than me. So maybe it is something wrong with my machine or my system configuration ?
what is the output of
ls -l /etc/alternatives/libblas.so.3andcat /proc/cpuinfols -l /etc/alternatives/libblas.so.3
output:/etc/alternatives/libblas.so.3 -> /usr/lib/openblas-base/libblas.so.3
if i follow the link i end up tolibblas.so.3.6.0cat /proc/cpuinfo:processor : 0 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2552.265 cache size : 8192 KB physical id : 0 siblings : 8 core id : 0 cpu cores : 4 apicid : 0 initial apicid : 0 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 1 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2800.109 cache size : 8192 KB physical id : 0 siblings : 8 core id : 1 cpu cores : 4 apicid : 2 initial apicid : 2 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 2 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2800.000 cache size : 8192 KB physical id : 0 siblings : 8 core id : 2 cpu cores : 4 apicid : 4 initial apicid : 4 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 3 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 3387.234 cache size : 8192 KB physical id : 0 siblings : 8 core id : 3 cpu cores : 4 apicid : 6 initial apicid : 6 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 4 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2463.781 cache size : 8192 KB physical id : 0 siblings : 8 core id : 0 cpu cores : 4 apicid : 1 initial apicid : 1 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 5 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2800.546 cache size : 8192 KB physical id : 0 siblings : 8 core id : 1 cpu cores : 4 apicid : 3 initial apicid : 3 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 6 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2526.781 cache size : 8192 KB physical id : 0 siblings : 8 core id : 2 cpu cores : 4 apicid : 5 initial apicid : 5 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management: processor : 7 vendor_id : GenuineIntel cpu family : 6 model : 60 model name : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz stepping : 3 microcode : 0x17 cpu MHz : 2474.500 cache size : 8192 KB physical id : 0 siblings : 8 core id : 3 cpu cores : 4 apicid : 7 initial apicid : 7 fpu : yes fpu_exception : yes cpuid level : 13 wp : yes flags : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts bugs : bogomips : 5586.71 clflush size : 64 cache_alignment : 64 address sizes : 39 bits physical, 48 bits virtual power management:you are using openblas, it is expected that that will use all cpu cores and is super wasteful on small matrices.
You should be able to control it with the OPENBLAS_NUM_THREADS environment variable.Reacted by Victor Kristof, heli, Rayan and Joshua J. CogliatiTo expand a bit:
- the difference between you and your colleague is that your colleague doesn't have openblas installed
- openblas is generally much faster than the alternatives for large matrices; it's faster on a single thread, and then on top of that it can use multiple threads (but you can disable the multiple threads if you want)
- for small matrices it has pointless overhead, sigh, and forcing it to use 1 thread instead will reduce this overhead.
- this may or may not matter, depending on whether your actual code really does involve running a tight loop of nothing but small matrix linear algebra operations. Small matrix operations are fast to start with, so if they go 2x slower while your large matrix operations go 10x faster, then this might be a win overall. Or not. It really depends on your machine and exactly what your code is doing.
- ...there have occasionally been bugs where openblas on small-matrix operations has been more like 100x too slow than like 2x too slow. If you encounter that then you should let us know so we can take it up with the openblas folks. (Specifically: if you can find a reproducible example where setting OPENBLAS_NUM_THREADS=1 makes your code much much much faster then we can probably get them to fix that.)
Hope that helps. Don't think there's anything else here for numpy to do right now, so closing, but feel free to re-open if I'm wrong.
Reacted by Victor Kristof, Nguyễn Văn Viết, baixiaohuang, Tianrui Wang (王天锐) and IdoMaHi @juliantaylor @njsmith, I observed same phenomenon here using scipy.sparse.linalg.spsolve, single thread is much much faster than taking all resources (mainly kernel threads). Do you know what these kernel threads do (communicating between process?) in multithreading cases? Thank you so much for the help.
I'm not really familiar with the internals of
spsolve. You should file an issue on scipy; they should be able to help you.Reacted by baoruxiaoWhen the computation load is small, numpy will use only one thread continuously.

Hi guys,
So i run into a big trouble while using linalg.solver or linalg.inv. All my 8 cpus are running at 100% where most part is due to the kernel.
First the program:
Here is what shows htop:

I am running under ubuntu 16.04.1.
Kernel version: 4.4.0-38-generic
Python version: 2.7.12
Numpy version: 1.11.0
Testing it on ubuntu 14.04 with older version of python (2.7.6) and numpy (1.8) seems to work fine.
Any help would be appreciated :)