Skip to content

High cpu usage by kernel while using inv and solve from linalg #8120

Description

@vsamy

Hi guys,

So i run into a big trouble while using linalg.solver or linalg.inv. All my 8 cpus are running at 100% where most part is due to the kernel.

First the program:

import numpy as np
import time

np.random.seed(0)
n = 56
A = np.asarray(np.random(n,n))
b = np.eye(A.shape[0])
while True:
               c = np.linalg.solve(A, b)
               time.sleep(1e-2)

Here is what shows htop:
htop

I am running under ubuntu 16.04.1.
Kernel version: 4.4.0-38-generic
Python version: 2.7.12
Numpy version: 1.11.0

Testing it on ubuntu 14.04 with older version of python (2.7.6) and numpy (1.8) seems to work fine.

Any help would be appreciated :)

Activity

  1. wrwrwr commented on Oct 6, 2016

    @wrwrwr
    Contributor

    There seems to be a significant difference between running the code under ipython and standalone (<10% processor load in the latter case). Can you reproduce the issue with explicit concurrency (say using threads)?

  2. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    Can you explain a little be more what you want me to do ? What do you mean by "explicit concurrency" ?

  3. wrwrwr commented on Oct 6, 2016

    @wrwrwr
    Contributor

    That was just a guess at what to research (how is ipython spawning 8 processes). How about something like the following, does it consume all of your resources? (Run as python <filename>; seems to give load proportional to the number of processes, significantly lower under Python 3).

    from multiprocessing import Process
    import time
    
    import numpy as np
    
    
    def f(i, A, b):
        print('Worker {}'.format(i))
        while True:
            np.linalg.solve(A, b)
            time.sleep(.01)
    
    
    if __name__ == '__main__':
        np.random.seed(0)
        A = np.asarray(np.random.random((56, 56)))
        b = np.eye(A.shape[0])
        for i in range(3):
            Process(target=f, args=(i, A, b)).start()
  4. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    i have the same results with this code

    htop2

    When i stop the program the number of threads display by htop passes from ~432 to ~411.

    During program 25 threads are running. (only 1 when no python program)

  5. njsmith commented on Oct 6, 2016

    @njsmith
    Member

    @vsamy: So for context, what you're observing here isn't numpy itself, it's the BLAS library that numpy's calling to do the actual linear algebra operations. This doesn't mean that we can't help, but it makes debugging a little more complicated. For example, the difference you see between numpy 1.8 and numpy 1.11 is almost certainly not any change in numpy, but that your two installs are being configured differently.

    I'm a little unclear on why you think there's a problem -- if you right a tight for loop doing some linear algebra, then you should expect this to use CPU :-). Some BLAS libraries use multiple threads internally, so you should expect linear algebra to use all your CPUs. This is generally a good thing...?

    Where did your numpy come from? (anaconda? pip install? linux package manager? something else?) Can you paste the output of np.show_config()?

  6. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    It is a good thing to use multiple threads but getting all the cpu ressources seem a little bit extreme.
    This chunk of code should be running in parallel to another bigger code which must run fast enough.

    Numpy has been installed through a classic apt-get. Here is the output:

    lapack_info:
        libraries = ['lapack', 'lapack']
        library_dirs = ['/usr/lib']
        language = f77
    lapack_opt_info:
        libraries = ['lapack', 'lapack', 'blas', 'blas']
        library_dirs = ['/usr/lib']
        language = c
        define_macros = [('NO_ATLAS_INFO', 1), ('HAVE_CBLAS', None)]
    openblas_lapack_info:
      NOT AVAILABLE
    blas_info:
        libraries = ['blas', 'blas']
        library_dirs = ['/usr/lib']
        define_macros = [('HAVE_CBLAS', None)]
        language = c
    atlas_3_10_blas_threads_info:
      NOT AVAILABLE
    atlas_threads_info:
      NOT AVAILABLE
    atlas_3_10_threads_info:
      NOT AVAILABLE
    atlas_blas_info:
      NOT AVAILABLE
    atlas_3_10_blas_info:
      NOT AVAILABLE
    atlas_blas_threads_info:
      NOT AVAILABLE
    openblas_info:
      NOT AVAILABLE
    blas_mkl_info:
      NOT AVAILABLE
    blas_opt_info:
        libraries = ['blas', 'blas']
        library_dirs = ['/usr/lib']
        language = c
        define_macros = [('NO_ATLAS_INFO', 1), ('HAVE_CBLAS', None)]
    atlas_info:
      NOT AVAILABLE
    atlas_3_10_info:
      NOT AVAILABLE
    lapack_mkl_info:
      NOT AVAILABLE
    mkl_info:
      NOT AVAILABLE
    
  7. njsmith commented on Oct 6, 2016

    @njsmith
    Member

    What does ls -l /etc/alternatives/libblas.so say?

  8. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    ls -l /etc/alternatives/libblas.so output /etc/alternatives/libblas.so -> /usr/lib/libblas/libblas.so
    ls -l /usr/lib/libblas/libblas.so output /usr/lib/libblas/libblas.so -> libblas.so.3.6.0

  9. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    @njsmith So i ask a colleague to test it on his machine and it works perfectly fine. He has the same ubuntu/kernel/blas/numpy/python version than me. So maybe it is something wrong with my machine or my system configuration ?

  10. juliantaylor commented on Oct 6, 2016

    @juliantaylor
    Contributor

    what is the output of ls -l /etc/alternatives/libblas.so.3 and cat /proc/cpuinfo

  11. vsamy commented on Oct 6, 2016

    @vsamy
    Author

    ls -l /etc/alternatives/libblas.so.3
    output: /etc/alternatives/libblas.so.3 -> /usr/lib/openblas-base/libblas.so.3
    if i follow the link i end up to libblas.so.3.6.0

    cat /proc/cpuinfo:

    processor   : 0
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2552.265
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 0
    cpu cores   : 4
    apicid      : 0
    initial apicid  : 0
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 1
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2800.109
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 1
    cpu cores   : 4
    apicid      : 2
    initial apicid  : 2
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 2
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2800.000
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 2
    cpu cores   : 4
    apicid      : 4
    initial apicid  : 4
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 3
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 3387.234
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 3
    cpu cores   : 4
    apicid      : 6
    initial apicid  : 6
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 4
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2463.781
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 0
    cpu cores   : 4
    apicid      : 1
    initial apicid  : 1
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 5
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2800.546
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 1
    cpu cores   : 4
    apicid      : 3
    initial apicid  : 3
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 6
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2526.781
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 2
    cpu cores   : 4
    apicid      : 5
    initial apicid  : 5
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
    processor   : 7
    vendor_id   : GenuineIntel
    cpu family  : 6
    model       : 60
    model name  : Intel(R) Core(TM) i7-4900MQ CPU @ 2.80GHz
    stepping    : 3
    microcode   : 0x17
    cpu MHz     : 2474.500
    cache size  : 8192 KB
    physical id : 0
    siblings    : 8
    core id     : 3
    cpu cores   : 4
    apicid      : 7
    initial apicid  : 7
    fpu     : yes
    fpu_exception   : yes
    cpuid level : 13
    wp      : yes
    flags       : fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush dts acpi mmx fxsr sse sse2 ss ht tm pbe syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon pebs bts rep_good nopl xtopology nonstop_tsc aperfmperf eagerfpu pni pclmulqdq dtes64 monitor ds_cpl vmx smx est tm2 ssse3 sdbg fma cx16 xtpr pdcm pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand lahf_lm abm epb tpr_shadow vnmi flexpriority ept vpid fsgsbase tsc_adjust bmi1 hle avx2 smep bmi2 erms invpcid rtm xsaveopt dtherm ida arat pln pts
    bugs        :
    bogomips    : 5586.71
    clflush size    : 64
    cache_alignment : 64
    address sizes   : 39 bits physical, 48 bits virtual
    power management:
    
  12. juliantaylor commented on Oct 6, 2016

    @juliantaylor
    Contributor

    you are using openblas, it is expected that that will use all cpu cores and is super wasteful on small matrices.
    You should be able to control it with the OPENBLAS_NUM_THREADS environment variable.

  13. njsmith commented on Oct 7, 2016

    @njsmith
    Member

    To expand a bit:

    • the difference between you and your colleague is that your colleague doesn't have openblas installed
    • openblas is generally much faster than the alternatives for large matrices; it's faster on a single thread, and then on top of that it can use multiple threads (but you can disable the multiple threads if you want)
    • for small matrices it has pointless overhead, sigh, and forcing it to use 1 thread instead will reduce this overhead.
    • this may or may not matter, depending on whether your actual code really does involve running a tight loop of nothing but small matrix linear algebra operations. Small matrix operations are fast to start with, so if they go 2x slower while your large matrix operations go 10x faster, then this might be a win overall. Or not. It really depends on your machine and exactly what your code is doing.
    • ...there have occasionally been bugs where openblas on small-matrix operations has been more like 100x too slow than like 2x too slow. If you encounter that then you should let us know so we can take it up with the openblas folks. (Specifically: if you can find a reproducible example where setting OPENBLAS_NUM_THREADS=1 makes your code much much much faster then we can probably get them to fix that.)

    Hope that helps. Don't think there's anything else here for numpy to do right now, so closing, but feel free to re-open if I'm wrong.

  14. baoruxiao commented on Dec 3, 2018

    @baoruxiao

    Hi @juliantaylor @njsmith, I observed same phenomenon here using scipy.sparse.linalg.spsolve, single thread is much much faster than taking all resources (mainly kernel threads). Do you know what these kernel threads do (communicating between process?) in multithreading cases? Thank you so much for the help.

  15. njsmith commented on Dec 3, 2018

    @njsmith
    Member

    I'm not really familiar with the internals of spsolve. You should file an issue on scipy; they should be able to help you.

  16. liuaibin commented on Aug 22, 2023

    @liuaibin

    When the computation load is small, numpy will use only one thread continuously.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions