Repository navigation
Discussion: uint / int_ / float_ / complex_ in NEP 52 #24743
Description
Activity
- added62 - Python APIChanges or additions to the Python API. Mailing list should usually be notified.Changes or additions to the Python API. Mailing list should usually be notified.
on Sep 18, 2023 Thanks for opening this issue Jake.
But I think the consistency of having default-width dtype specifiers available for each dtype kind outweighs the cost of having overlapping names for those default types.
Can you clarify what you mean by "default-width dtype specifiers"? Are these the ones you mention above?
I'll note that gh-24651 is still in progress. The table there has none of the four aliases. I think two are removed and the other two can probably also be removed.
Can you clarify what you mean by "default-width dtype specifiers"? Are these the ones you mention above?
Yes, my understanding has always been that
uint,int_,float_, andcomplex_are how you determine the (possibly platform-dependent) default width for each type class.I think two are removed and the other two can probably also be removed.
If this is the case, then I'd be happy with the result!
Reacted by Sebastian BergYes, my understanding has always been that
uint,int_,float_, andcomplex_are how you determine the (possibly platform-dependent) default width for each type class.I'd say that there isn't really such a concept in NumPy, at least explicitly. It became relevant in JAX/PyTorch/etc. because there is more variation. For NumPy, I don't believe there has ever been a case of platform-dependent floating-point defaults. For integers, I'd say that
intpanduintpare the most obvious way, as those are the canonical dtypes for integers used for indexing (which matches the default bit width for the OS + Python interpreter).Removing
int_actually has at least the slight advantage that we would also like to change its meaning on windows (my open PR). So removing it and propagatingnp.longornp.dtype("long")to fetch the info (where you need the old one) seems good; while the "new" one would then beintpindeed.On the other hand, I would also be happy to just hide
int_away or slowly deprecated it if it seems used a fair bit by downstream.PS: Any hint towards difficulties of
intpandPy_ssize_t/ssize_tmismatching on relevant platforms would be useful (and problematic).I'd say that there isn't really such a concept in NumPy, at least explicitly.
My memory may be fuzzy, but I thought that ca. 2006 or so,
np.float_wasfloat32orfloat64depending on system architecture. Anyway, if there's no explicit concept of "default scalar type", it seems removing these four is the most consistent option here.No, it wasn't. It might have been the case, very briefly, that
numpy.float_ornumpy.floatcorresponded to a Cfloata lanp.longand friends, etc., but it was never platform-dependent, nor would it have been a marker of a default floating-point width.The intent of
float_was to correspond withnp.dtype(float), aka "the type that Python'sfloatis represented by", and the same forcomplex_.The same was also true of
int_in python 2, but now that meaning is irrelevant.The intent of
float_was to correspond withnp.dtype(float), aka "the type that Python'sfloatis represented by", and the same forcomplex_.The answer here has always been, and probably will always be,
float64andcomplex128, right? No need for a separate alias for that as far as I can see, because it's extremely unlikely to ever change.Reacted by Eric WieserI noted this in my first comment, actually:
I understand the argument about platform-dependence of integer dtypes and non-platform dependence of inexact dtypes.
I'm not totally clear on the evolving state of platform dependence of integer dtypes, but I maintain that numpy should either define all four of these aliases, or define none of them, and anything in between is confusing.
ping @mtsokol, there seems to be consensus about removing all of the aliases, would you like to take on following up in the code and docs?
Reacted by Mateusz SokółWhat's the proposed spelling to get a C long if
np.int_is removed? Perhaps I'm out of the loop and we finally have it atnp.long.Last I checked numpy had five signed integer types corresponding to C's 5 types; like C++, even though two happen to coincide in bit-width they are still treated as different types in things likE PEP3118.
If we keep only
int8,int16,int32,int64, then we have a secret 5th type that is not part of thenp.namespace.What's the proposed spelling to get a C long if
np.int_is removed? Perhaps I'm out of the loop and we finally have it atnp.longThis PR's description has an up-to-date table that maps Python canonical names to C types: #24651 (comment). It should be
np.intp.That table doesn't seem to reflect platform-dependence;
intpisint_ptr_twhich certainly doesn't coincidence with Clongon Windows.Oh huh, that's a good point that
np.longis missing and thatintp_tisn't quite the same thing. I think including the stdint pointer int types in the table makes sense along with splitting off long into its own row.How about instead of removing
int_we rename itlong? Seems a more sensible name - it's a the size of a C long on all platforms right?The current
int_name is because in python 2,intwas represented by a Clong, which is the same reason thatnp.float_actually refers to a Cdouble.There was a long period where
np.longwas not available as a name, because long agonp.long = builtins.long, which eventually becamenp.long = builtins.intin python 3.I think reclaiming
np.longfor the the C name is probably in scope for 2.0 if we want it to be.However, there is the argument that the C type names are a niche use case (and worse, confusing when comparing python
intto Cint); so maybe onlynp.intXXshould get the obvious names, and the C type names should be harder to find as something likenp.c_long, which matchesctypes.c_long.The current "C-like" names we have for integer types in the python API are
byte,short,intc,int_, andlonglong. So I guesslongcis also an option since it follows the precedent ofintc. Not that eitherintcorlongcare particularly intuitive names...I personally think
longis fine, we already haveshortandlonglongso exposing one more C type name just makes the naming in the API a little more consistent, we would haveinttoo if that didn't step on a Python type name.Does it matter that it steps on a python type name?
I personally don't think it matters if it steps on a Python 2 type name.
Oh I see, for
intc. I think havingnp.intavailable and not correspond to a python int would introduce new confusion, especially since it's int32 which is very different from python's extended precision ints. There's nolongtype in Python anymore so there's no confusion with a python type if we havenp.long.I think I have a preference for
longas well, since it's the most obvious name and matches most other C-like names (e.g.,longlong).longcseems a bit ugly,clongcould be confused withclongdouble& co where the c is for "complex", andc_longwould be yet another naming convention.FWIW, I had also added
longto "replace"int_in the PR to change the default integer (I don't think I had removedint_, except in Cython where it cannot exist asint_t).longhas been awkward since Python 2 is gone (or longer), so I am fine with reusing it. I think I would lean to not havenp.intat all.Reacted by Ralf GommersI've got one question: Do we want to keep
np.dtype("int")andnp.dtype("uint")?
If yes, should it correspond tonp.intcornp.long?I think it can still correspond to
np.longmixing C and Python names, similarly asnp.dtype("float")andnp.dtype("double"), that also map tonp.float64.Please note that
np.dtype("intc")andnp.dtype("long")will be both available.It would correspond to the default integer (currently
long). I don't much like it, because its confusing with the C types, but I half suspect there is such a huge amount ofdtype="int"code in practice that I am not willing to touch it.Looking at @mtsokol changes, there seems to be very few uses of
int_in actual documentation which is good news, because in actual use replacing it withlongis correct, but not nice if we change the default tointp. (At which point we don't want anyone to uselongunless they are interacting with C explicitly!)I am not sure to I like the idea of replacing
np.int_([1, 2, 3])withnp.intp([1, 2, 3])either. So TBH, I would also be just fine living with the fact that we have the underscore for the default int, but not the default float/complex (because those are always clear). Is it really that bad?It may be OK either way because asking for "default integer" is really a bit silly in most places. (So the few times someone needs it, it is maybe fine if they have to use
long/intp.)Reacted by Ralf GommersWe ended up landing on keeping
int_anduint_, at least for a few more releases. We still need a way for users and library authors to write "the platform-dependent default numpy integer type". The plan for NumPy 2.0 is forint_to correspond tointpon all platforms, so we could removeint_and tell people to replace it withintp, but we ended up deciding that is not a big enough clarity improvement to justify the code churn. The argument that this introduces an inconsistency in the API due to the lack offloat_wasn't persuasive to others on the call when I brought that up.I'm going to close this now, reflecting the decision that was made at the last community meeting.
Reacted by Sebastian Berg- added 2 commits that reference this issue
on May 20, 2024
Pulling-out a discussion of #24376 (comment), which is buried in a long thread.
I have been working on updating JAX for compatibility with NEP 52, and the inconsistency of treatment of default-width dtypes (
uint,int_,float_,complex_) is something I'm finding confusing.For consistency, I think either all four should be present in NumPy 2.0, or all four should be removed. The current state, where
uintandint_are present, butfloat_andcomplex_are not, is inconsistent and confusing.I understand the argument about platform-dependence of integer dtypes and non-platform dependence of inexact dtypes. I understand the goal of eliminating unnecessary aliases to types. But I think the consistency of having default-width dtype specifiers available for each dtype kind outweighs the cost of having overlapping names for those default types.