Skip to content

Default int type is platform dependent #9464

Description

@eric-wieser

np.array([1]).dtype is platform-dependant, presumably because it defaults to np.int_

  1. Is this by design?
  2. If not, can we force it to int64?

Activity

  1. njsmith commented on Jul 26, 2017

    @njsmith
    Member

    It is by design – the idea is that numpy's default int type matches the range of python 2's int, which in turn matches the platform C compiler's long.

    Whether this is a good design is another question, especially since python 3 has eliminated this. There have been intermittent discussion about changing it before that you can probably dig up – especially the confusing and error prone way the default is 32 bits on win64.

    I suppose one way to move that discussion forward would be to test whether any major packages break if you do make that change.

  2. pv commented on Jul 26, 2017

    @pv
    Member

    One thing that may break is if someone is using dtype=int and assumes this is somehow related to C long type...

  3. juliantaylor commented on Sep 1, 2017

    @juliantaylor
    Contributor

    Changing the default int type on windows 64 to 64 bit would imo be an important enough change to warrant breaking software.
    The current behavior just causes too many bugs.

    That the default int type on 32 bit is 32 bit int is probably not so bad, as it does at least cover the full addressable range and changing it could have performance impact.

  4. shoyer commented on Mar 20, 2018

    @shoyer
    Member

    We should seriously consider changing this.

    In my experience, if a Python library of moderate complexity that uses NumPy does not run Windows specific tests, it probably broken for this reason.

  5. thomasmansencal commented on Sep 15, 2018

    @thomasmansencal
    Contributor

    @shoyer : We ran into this exact problem on Windows with @MichaelMauderer on colour-science/colour#431.

    I was assuming incorrectly that np.int_ was platform independent.

  6. eric-wieser commented on Sep 15, 2018

    @eric-wieser
    MemberAuthor

    Perhaps we should drop this default at the same time as python 2, since the sole reason for defaulting to np.int_ was that it matched the size of builtins.int, which in python 3 is not even true.

  7. lawrence858 commented on Oct 10, 2018

    @lawrence858

    Ideally numpy should behave the same way across platforms. A colleague of mine uses Windows and recently had to spend some time trying to figure out why a program was yielding different results on his machine than on my Mac. IMO performance considerations pale in comparison to getting correct and consistent results.

  8. gojomo commented on Jul 7, 2020

    @gojomo

    Is there any runtime workaround a user could execute, before their other code, to force numpy-on-Windows default types to the same widths as elsewhere? (Perhaps, a data-driven, tamperable mapping of Python types to numpy types?)

    As a fresh example of some of the resulting craziness, specifically asking for a array of a type compatible with type(2**32) results in an array that can't store 2**32:

    2020-07-07T06:53:20.9528159Z     def testTiny(self):
    2020-07-07T06:53:20.9528423Z         a = np.empty(1, dtype=type(2**32))
    2020-07-07T06:53:20.9529046Z >       a[0] = 2**32
    2020-07-07T06:53:20.9529318Z E       OverflowError: Python int too large to convert to C long
    
  9. adeak commented on Jul 7, 2020

    @adeak
    Contributor

    @gojomo I'm not sure that's a right approach anyway. On python 3 type(2**32) is guaranteed to be int, so that's just a more complicated way of saying dtype=int. If you're using a literal like that anyway you could of course use explicit dtype=np.int64.

    To make it more dynamic, does dtype=np.array(2**32).dtype work? (Odds are there are even more idiomatic ways to do this.)
    EDIT: np.empty_like(2**32, shape=...) is probably it, assuming that works.

  10. seberg commented on Jul 7, 2020

    @seberg
    Member

    No, I had a PR to add one, maybe I can open that again now that we decided to start the deprecation on some of the aliases: #16535

    So either use dtype=np.intp which gives you 32bit on 32bit systems and 64bit on 64bit systems, or use dtype=np.int64 to begin with. That PR made dtype=np.intp the default, which is the simpler change, because intp is fairly common in NumPy already.

  11. adeak commented on Jul 7, 2020

    @adeak
    Contributor

    I was thinking that if NEP 31 ever happens, it would also make this kind of replacing defaults easily opt-in.

  12. gojomo commented on Jul 7, 2020

    @gojomo

    @adeak My snippet's not a literal example; my actual issue is that I've got a list of many ints, which eventually reach 2**32, but a numpy array typed based on the first int breaks on Windows when it reaches 2**32, but works everywhere else.

    (I was hoping the snippet highlighted some of the on-the-face absurdity of the Python-to-numpy interaction: shouldn't a reported type for a specific number specifically-communicate a corresponding type wide enough to store it? But I suppose Python is an equal contributor to the problem, as 2**65 & 2**129 have the same problem of reporting as simple int. So it's more a brain-teaser than a guide to better behavior.)

    I'd answer the "is numpy's choice a good design?" question in @njsmith's 2017 comment as: "Reasonable way back when, but not anymore, with Python3, & the primacy of 64bit systems, and Microsoft's own phasing-out of WIndows 10's support for 32-bit systems."

    Traffic on this issue since looks like it has referenced many places this has caused problems for people, but not yet any extant examples of code that'd break with a changed default. (There's probably some, somewhere.)

    If the plunge of changing the default in one swoop is too risky, a call that opts-in to some minimum-width default (or user-chosen default) for all subsequent mappings of Python's int might help. (And then at some later date with warning to Windows users, change the default, but give laggards an option to change it back for a while.)

  13. 11 remaining items

  14. fmaussion commented on Feb 15, 2023

    @fmaussion

    Thanks @seberg this sounds reasonable - I'll see if I can find the bandwith to get things started but I'll certainly need help.

  15. seberg commented on Feb 15, 2023

    @seberg
    Member

    Don't hesitate to get in touch with me. There are too many things for the core team to push, so someone helping championing such change makes it much likely to happen!

  16. njh219 commented on Feb 15, 2023

    @njh219
  17. xor2k commented on Feb 24, 2023

    @xor2k
    Contributor

    Hi everybody! I just experienced this problem in my project npy-append-array, compare

    xor2k/npy-append-array#6

    The problem was that while with MacOs and Linux, int64 is the default, for Windows it is int32, even if it is a 64 bit operating system. My solution was to basically replace all Numpy functions with their corresponding Python functions, like numpy.multiply.reduce with prod and numpy.ceil with ceil. I could also have specified dtype=np.int64 or dtype=np.int64 but that would be quite explicit and hopefully not necessary anymore in the future.

    If this is API breaking or so, maybe it would be something for Numpy 2.0, wouldn't it?

  18. xor2k commented on Feb 24, 2023

    @xor2k
    Contributor

    @fmaussion writing a brief NEP would be good. At this point, we should probably consider including such change in a major release (but it might still be good to summarize things in a NEP!), since I hope that isn't too far off.

    I also suspect that the sane choice is probably (unfortunately) to switch to intp as default (i.e. 64bit on 64bit windows and not attempting any change e.g. on 32bit linux). But NEP would be the place to summarize that.

    I added a switch for "use NumPy 2 behavior" very recently, so once there is some general consensus to push for this, there is also a path to start implementing it as planned for the major release.

    I have no bandwidth either, but if you need someone to write that NEP, I can give it a shot.

  19. joaoe commented on Apr 17, 2023

    @joaoe
    Contributor

    hi.
    This issue causes serious compatibility problems between Windows 64 and Linux/Mac. E.g., another one unionai-oss/pandera#726

  20. albertopasqualetto commented on Aug 18, 2023

    @albertopasqualetto

    I think that this issue should be written in the numpy.array documentation page, because the phrase "[...] NumPy will try to use a default dtype that can represent the values (by applying promotion rules when necessary.)" is misleading.

  21. rgommers commented on Jun 24, 2024

    @rgommers
    Member

    This has been addressed in NumPy 2.0 by changing the default integer type to intp, which has the practical consequence of changing the default on 64-bit Windows from int32 to int64. See:

    This solves the key issue of Windows behaving differently than other platforms; after the change all 64-bit platforms return int64 for np.array([1]).dtype, and all 32-bit platforms return int32.

    This issue has been addressed to the extent possible, and I don't think we'll be making another change to it after this, so I'll close this issue. Thanks all!

  22. added this to the 2.0.0 release milestone on Jun 24, 2024
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions