Repository navigation
np.dtype(int) should be np.longlong on python 3 #12332
Description
Activity
builtins.longis mapped tonp.longlong(Clong long), sincebuiltins.longis infinite precision and that's the largest available integer typeThat's not quite the explanation, I don't think.
np.int128is usually available (though it may be true thatnp.longlongis more universally available).However, I think the actual explanation is a bit more pragmatic. I think that the real explanation is that we used
builtins.longforsize_t-sized integers (like.shapetuple entries) on Win64 sincebuiltins.intcouldn't hold some of the 64-bit values. In order to do some kind of round-tripping, we needed to mapbuiltins.longto something at leastsize_t-sized on Win64 platforms.I think for the Python 3 case, we chose
np.long_because that's the expectation that's been embedded in everyone's C code for so long. If I could turn back time, I'd probably suggest that we makessize_tthe default integer type. Barring time travel, Python 3 Win64 is going to experience a problem no matter which one we pick. Either it can't round tripnp.int64->builtins.int->np.int64without explicit typing, or some C-implemented code that's depending on getting Clongs is going to break. I'm wondering if the latter is not more likely. On the other hand, it's more noisy while the former can break silently.np.int128 is usually available
I have never seen this. In the case where it's available, are you sure that
np.int128 is np.longlongis not true?I think for the Python 3 case, we chose
np.long_Do you mean
np.int_? There is nonp.long_.or some C-implemented code that's depending on getting C longs is going to break
You could make exactly the same argument about
c_func(np.array(..., dtype=long))breaking after passing through2to3, if it expectslong longsnp.int128 is usually available
I have never seen this. In the case where it's available, are you sure that np.int128 is np.longlong is not true?
Hmm, okay, not sure what I was thinking. Conflating in my mind my experience of writing C code for
int128_ts and the general availability ofnp.float128. Never mind.Nonetheless, I think the need to round-trip
np.intp-sized integers was the motivating factor. If we just cared about the one-way Python->numpy conversion, I'm sure we would have left it restricted at a Clongfor consistency. If you're going to cut off a countably infinite number of possible values, it doesn't really matter where that cutoff is. :-)I think for the Python 3 case, we chose np.long_
Do you mean np.int_? There is no np.long_.
I meant
np.clong; I forgot how we de-conflicted that name.or some C-implemented code that's depending on getting C longs is going to break
You could make exactly the same argument about c_func(np.array(..., dtype=long)) breaking after passing through 2to3, if it expects long longs
Sure, but there are many more
c_funcs written for Clongs because that's the default integer type than Clong longs. If such ac_funcdid exist and depended on the user specifying annp.longlongarray on the Python side, I would expect that the Python-side user would be expected to usedtype=np.longlongin any case.The arguments are all true, but do we have any idea about bugs for C extensions? Cython code that just uses
longsuddenly breaking everywhere, etc. Those are probably crash bugs, but still.Frankly, it feels too ambitious, maybe that is partially because I do not feel the pain of windows normally. How about instead:
- Write a short NEP (the problem is easy, so mostly to suggest the next steps).
- 1.16 won't change, but we could add a compile time or run time flag to switch the behaviour to allow testing both.
- Change it in a future version, maybe 1.17 (or 2.0) or a bit later depending on opinion.
If we have a reasonable way for warnings that might be good and could change the approach.
My palantír is somewhat dim these days (and this is a hard thing to search the archives for), but I don't recall anyone running into actual problems with the status quo, just hypotheticals.
I mean, I'd love it if we could just standardize on
np.int64for all platforms and get away with it…Cython code that just uses long suddenly breaking everywhere
If this is the case, then this code will break already if passed a python 2.7
longI forgot how we de-conflicted that name.
As
np.int_, confusingly. There is nonp.clongeither.@eric-wieser If I write a cython or worse a C-extension that is naively typed as
longfor all input arrays and forgets to check (or is not specialized for anything else), and then feed it the default created arrays it will go from working to breaking. All reasonable code should not be doing that, but there is a lot of unreasonable code out there – whether for this particular case or not, I have no idea. It also might double someones memory usage silently, etc.The thing is, I have currently no clue at all how much code could break. Probably all larger packages are fine, since they target systems with different defaults anyway, but all the small scripts out there are a different matter.
I am fine with trying to change it. Heck I am in favor, but trying to do it in on short notice now seems ambitious. At least I would like to have some idea that this is indeed very unlikely, and I simply do not have it.
So, maybe we can get some confidence that nothing bad will happen, but I doubt that is easy or even possible and, thus, I would prefer looking for ways to do such a change slower. Heck, we can even do a FU pre-release if it helps switching the behaviour in the rc just to see what happens, I just don't think we should rush into switching it in a release.
@rkern I think I do remember sklearn or so complaining about it, though it is probable that most/many of such things were things where
np.intpshould have been used.So the summary is, we care more about preserving the behavior of
np.dtype(type(1))than we do about preservingnp.dtype(long)? I suppose that's fair.Maybe I have been reading this a bit wrong. I thought what you suggested effectively changes the default integer type to int64, and I like that but it seems to me should take it slower. Is it something quite different you are suggesting?
I agree with @eric-wieser here. We should avoid the use of C
longwhenever possible in any case because its precision varies between platform, andlong longseems a better match for Python infinite precision integers in any case. The possible downside is upcasting, but I think that is already handled for the Python 2 case.I thought what you suggested effectively changes the default [numpy] integer type to int64
I am suggesting this, but only for python 3, since that's consistent with the python 3 behavior of the default integer type now being arbitrary precision, rather than matching the C
long.Ideally, we would have made this transition when we first started supporting python 3. Obviously it's too late for that, but if we're going to fix this ever, the final transition from 2 to 3 seems like a convenient point
Good, because I fully agree to everything being 64 bit ints. I am just a bit wary of pushing the switch for the (final) 1.16 release and wonder if we can't make some progress (e.g. build with switch) before to get a bit of a better idea about it.
10 remaining items
An alternative would be to change the default to
intpinstead ofint64Hmm, be nice to settle this, but pushing off to 1.18 just because it doesn't look like a blocker for 1.17.
I still somewhat feel we may want to try it with a major version. Although for libraries it doesn't matter, it would only matter for scripts running on windows machines and even then probably only if they eat a lot of memory or so...
Theintpsolution is likely the much less tricky one (the other seems like it requires touching all functions to make sure they prefer 64bit ints). That still gives different precision, but at least it solves the "large arrays break on windows 64" problem, and is easier to reason about.We could start adding some such things as environment variable switches if it doesn't get too ugly? Just to see where it goes.
Although for libraries it doesn't matter
I can count at least 4 times I've run into bugs for libaries on Windows because...
dtype=intdid the wrong thing on Windows. 🙁@hameerabbasi what I meant is: For libraries we could just switch over to higher always higher precisions without anybody noticing (and if anything, fixing bugs). I think it is only end users who such a change can possibly break existing (albeit brittle) code.
Welp, going to push this off again as time to the 1.18 release is getting short.
- added60 - Major releaseIssues that need or may be better addressed in a major releaseIssues that need or may be better addressed in a major releaseand removed
on Dec 2, 2020 I don't think this is relevant anymore
This has come up before, but I'm not sure we have a canonical issue.
Historically in python 2:
builtins.intis mapped tonp.int_(Clong), sincebuiltins.intis stored as a Clongbuiltins.longis mapped tonp.longlong(Clong long), sincebuiltins.longis infinite precision and that's the largest available integer typeNote the above output only reflects the state of windows - other platforms have
sizeof(long) == sizeof(long long), so the distinction isn't important.In python 3,
builtins.inthas been removed, andbuiltins.longhas been renamed to__builtins__.int. Yet:This means that python code translated from 2 to 3 by replacing
longwithintwill start behaving differently:Since this affects users transitioning from 2 to 3, I think it's important that we get it fixed in 1.16, which will be the last version that transitioning users can test both version of python against.
The current implementation, introduced in aa7be88 by @pv, is:
I'd propose it should have been:
Ie, treating
PyLong_Typein python 3 just as we always did in python 2.