Skip to content

astype(int) silently returns wrong answer under Windows #17640

Description

@Rik-de-Kort

Casting an array of float64 to int using astype with as argument int will yield the value -2147483648 under windows.

Reproducing code example:

x = 2384351503.0
np.testing.assert_array_equal(np.array([x]).astype(int), np.array([int(x)]))

I expect this test to go through, but it fails with the following message:

Traceback (most recent call last):
  File "C:\Users\koRR\AppData\Local\conda\conda\envs\LRE\lib\site-packages\IPython\core\interactiveshell.py", line 3417, in run_code
    exec(code_obj, self.user_global_ns, self.user_ns)
  File "<ipython-input-10-725c5d2f7abb>", line 1, in <module>
    np.testing.assert_array_equal(np.array([x]).astype(int), np.array([int(x)]))
  File "C:\Users\koRR\AppData\Local\conda\conda\envs\LRE\lib\site-packages\numpy\testing\_private\utils.py", line 931, in assert_array_equal
    verbose=verbose, header='Arrays are not equal')
  File "C:\Users\koRR\AppData\Local\conda\conda\envs\LRE\lib\site-packages\numpy\testing\_private\utils.py", line 840, in assert_array_compare
    raise AssertionError(msg)
AssertionError: 
Arrays are not equal
Mismatched elements: 1 / 1 (100%)
Max absolute difference: 4531835151
Max relative difference: 1.90065733

Using dtypes np.int64 works fine, already np.int32 returns the wrong answer (although that's 0 and the above result occurs only with the builtin int or np.int).

NumPy/Python version information:

Numpy version: 1.19.1
System version: '3.7.6 (default, Jan 8 2020, 20:23:39) [MSC v.1916 64 bit (AMD64)]'
Bug is taking place under Windows 10. Cannot reproduce on my Arch system with numpy 1.19.2 and python 3.8.5 built using GCC.

Activity

  1. eric-wieser commented on Oct 26, 2020

    @eric-wieser
    Member

    Your example is not the same as your error message - the test uses np.array([2384351503], dtype=int), while the error message shows you are using np.array([238435103]). At least on 1.17, the former raises an OverflowError, so I can't even run your example code.

  2. eric-wieser commented on Oct 26, 2020

    @eric-wieser
    Member

    Bug is taking place under Windows 10. Cannot reproduce on my Arch system with numpy 1.19.2 and python 3.8.5 built using GCC.

    int gives you a C long - and long is 32-bit on windows but 64-bit everywhere else.

  3. Rik-de-Kort commented on Oct 27, 2020

    @Rik-de-Kort
    Author

    Sorry about that, I was typing the message from outside the VM where I got it. Very sloppy on my part. I updated the OP with a proper example.

    and long is 32-bit on windows but 64-bit everywhere else.

    Thanks for the info. int has unbounded precision in Python so this really took me by surprise. Also it's weird that it didn't raise an overflow error.

  4. charris commented on Oct 27, 2020

    @charris
    Member

    so this really took me by surprise.

    Python 2 had two types of integer: int (C long) and long (unbounded precision). Python 3 only kept the second.

  5. finoptimal-dev commented on Jun 1, 2021

    @finoptimal-dev

    Is there a workaround for this?

  6. BvB93 commented on Jun 1, 2021

    @BvB93
    Member

    Is there a workaround for this?

    Since the issue is caused by an overflow of np.int32 (i.e. the default int-type on windows) you can use np.int64 instead.

  7. finoptimal-dev commented on Jun 1, 2021

    @finoptimal-dev

    Thanks, BvB93. Is there a way to map it globally? I have a codebase that works on linux with astype(int) all over it; now I have a Windows developer. Before I consider changing it everywhere it exists, I'm wondering if there's a solution that spares needing to take that route.

  8. BvB93 commented on Jun 2, 2021

    @BvB93
    Member

    Thanks, BvB93. Is there a way to map it globally?

    I think you might be out of luck here. As far as I'm aware there is no way of manually setting the default integer type to something different from what is specified by the platform in question.

  9. charris commented on Jun 3, 2021

    @charris
    Member

    There has been some discussion about changing the default type. I believe the use of c_long is inherited from early python.

  10. seberg commented on Jun 3, 2021

    @seberg
    Member

    I had a PR once to make it np.intp at least, so that 64bit windows would at least use 64bits (making the rule "64bit on 64bit systems").

    That was mainly for discussion, but I am more and more seeing a NumPy 2.0 in any case, and this might be important enough to fold in if it happens. (Even if I would prefer to not do too many of such changes at once and rather make this type of "breaking a bit" releases every few years so that the number of such changes are small each time.)

  11. mattip commented on Jun 3, 2021

    @mattip
    Member

    Is this done correctly in the NEP 47 array API?

  12. seberg commented on Jun 3, 2021

    @seberg
    Member

    @asmeurer can you check this? From what I remember of the pass I did just today, it probably isn't.

  13. asmeurer commented on Jun 3, 2021

    @asmeurer
    Member

    I think this issue is relevant data-apis/array-api#151

  14. seberg commented on Jun 3, 2021

    @seberg
    Member

    @asmeurer the point is that the standard probably wants a well defined "default integer"? But NumPy's default integer is currently not well defined:

    • It differs for windows 64bit compared to linux 64 bit (because long is defined differently) and is also 32bit on 32bit platforms.
    • We automatically "spill" into long long or unsigned long long as noted here. EDIT np.array(2**63)

    If we add a new, clean namespace we should try to make good use of it and fix both of these. (unless we are sure we fix both of them in NumPy proper, which I would like to try but is more difficult.)

  15. asmeurer commented on Jun 3, 2021

    @asmeurer
    Member

    I actually mean this change specifically https://github.com/data-apis/array-api/pull/167/files#diff-7e75cfe3133de16126433bc962f9fe14f216bd682e3870efa965bab40d08322fR9. Based on what is there, I guess we should make the "default" integer dtype int64 on 64-bit Windows. Note that in the array API itself the default integer dtype is only relevant in a couple of places.

  16. seberg commented on Jun 3, 2021

    @seberg
    Member

    @asmeurer good that it is spelled out nicely!

    But your PR does not conform to this for asarra([1, 2, 3, 4]). We could probably make it conform from within NumPy at least with an abstract DType right now. Or you would have to write a light-weight asarray yourself, I guess.

  17. asmeurer commented on Jun 3, 2021

    @asmeurer
    Member

    You mean specifically on Windows? Or does asarray also use some value based casting in some cases?

    What is an abstract DType? If there is some trick that would avoid having to reimplement asarray, that would be ideal.

  18. seberg commented on Jun 3, 2021

    @seberg
    Member

    We automatically "spill" into long long or unsigned long long as noted here. EDIT np.array(2**63)

    This part is not windows specific.

    What is an abstract DType

    A DType class as per NEP 42. That could customize the dtype discovery during array coercion as per the NEP. But, it we may have to fix casting to/for abstract DTypes first. (Shouldn't be super hard, it should have an additional "common DType" based path, I think – that is probably already mentioned in NEP 42.)

  19. Rik-de-Kort commented on Jun 4, 2021

    @Rik-de-Kort
    Author

    Thanks, BvB93. Is there a way to map it globally? I have a codebase that works on linux with astype(int) all over it; now I have a Windows developer. Before I consider changing it everywhere it exists, I'm wondering if there's a solution that spares needing to take that route.

    In my codebase this is exactly what I have done. Explicit is better than implicit after all. No problems changing over, no problems after. Would recommend!

  20. seberg commented on Nov 3, 2023

    @seberg
    Member

    Closing, we are trying to switch to 64bit on 64bi tplatforms for NumPy 2.0. See the dev release notes for example. Which will change the default on windows specifically.
    (And if we may have to undo that, there is probably not much we can do about this.)

  21. seberg commented on Nov 3, 2023

    @seberg
    Member

    Note that for now, NumPy will still happily return uints and object for out of bound ints, this also doesn't affect 32bit platforms, just because it would be a bigger, more difficult to deal with change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions