Skip to content

Discussion: uint / int_ / float_ / complex_ in NEP 52 #24743

Description

@jakevdp

Pulling-out a discussion of #24376 (comment), which is buried in a long thread.

I have been working on updating JAX for compatibility with NEP 52, and the inconsistency of treatment of default-width dtypes (uint, int_, float_, complex_) is something I'm finding confusing.

For consistency, I think either all four should be present in NumPy 2.0, or all four should be removed. The current state, where uint and int_ are present, but float_ and complex_ are not, is inconsistent and confusing.

I understand the argument about platform-dependence of integer dtypes and non-platform dependence of inexact dtypes. I understand the goal of eliminating unnecessary aliases to types. But I think the consistency of having default-width dtype specifiers available for each dtype kind outweighs the cost of having overlapping names for those default types.

Activity

  1. added
    62 - Python APIChanges or additions to the Python API. Mailing list should usually be notified.
    on Sep 18, 2023
  2. rgommers commented on Sep 18, 2023

    @rgommers
    Member

    Thanks for opening this issue Jake.

    But I think the consistency of having default-width dtype specifiers available for each dtype kind outweighs the cost of having overlapping names for those default types.

    Can you clarify what you mean by "default-width dtype specifiers"? Are these the ones you mention above?

    I'll note that gh-24651 is still in progress. The table there has none of the four aliases. I think two are removed and the other two can probably also be removed.

  3. jakevdp commented on Sep 18, 2023

    @jakevdp
    ContributorAuthor

    Can you clarify what you mean by "default-width dtype specifiers"? Are these the ones you mention above?

    Yes, my understanding has always been that uint, int_, float_, and complex_ are how you determine the (possibly platform-dependent) default width for each type class.

    I think two are removed and the other two can probably also be removed.

    If this is the case, then I'd be happy with the result!

  4. rgommers commented on Sep 19, 2023

    @rgommers
    Member

    Yes, my understanding has always been that uint, int_, float_, and complex_ are how you determine the (possibly platform-dependent) default width for each type class.

    I'd say that there isn't really such a concept in NumPy, at least explicitly. It became relevant in JAX/PyTorch/etc. because there is more variation. For NumPy, I don't believe there has ever been a case of platform-dependent floating-point defaults. For integers, I'd say that intp and uintp are the most obvious way, as those are the canonical dtypes for integers used for indexing (which matches the default bit width for the OS + Python interpreter).

  5. seberg commented on Sep 19, 2023

    @seberg
    Member

    Removing int_ actually has at least the slight advantage that we would also like to change its meaning on windows (my open PR). So removing it and propagating np.long or np.dtype("long") to fetch the info (where you need the old one) seems good; while the "new" one would then be intp indeed.

    On the other hand, I would also be happy to just hide int_ away or slowly deprecated it if it seems used a fair bit by downstream.

    PS: Any hint towards difficulties of intp and Py_ssize_t/ssize_t mismatching on relevant platforms would be useful (and problematic).

  6. jakevdp commented on Sep 19, 2023

    @jakevdp
    ContributorAuthor

    I'd say that there isn't really such a concept in NumPy, at least explicitly.

    My memory may be fuzzy, but I thought that ca. 2006 or so, np.float_ was float32 or float64 depending on system architecture. Anyway, if there's no explicit concept of "default scalar type", it seems removing these four is the most consistent option here.

  7. rkern commented on Sep 19, 2023

    @rkern
    Member

    No, it wasn't. It might have been the case, very briefly, that numpy.float_ or numpy.float corresponded to a C float a la np.long and friends, etc., but it was never platform-dependent, nor would it have been a marker of a default floating-point width.

  8. eric-wieser commented on Sep 19, 2023

    @eric-wieser
    Member

    The intent of float_ was to correspond with np.dtype(float), aka "the type that Python's float is represented by", and the same for complex_.

    The same was also true of int_ in python 2, but now that meaning is irrelevant.

  9. rgommers commented on Sep 20, 2023

    @rgommers
    Member

    The intent of float_ was to correspond with np.dtype(float), aka "the type that Python's float is represented by", and the same for complex_.

    The answer here has always been, and probably will always be, float64 and complex128, right? No need for a separate alias for that as far as I can see, because it's extremely unlikely to ever change.

  10. jakevdp commented on Sep 20, 2023

    @jakevdp
    ContributorAuthor

    I noted this in my first comment, actually:

    I understand the argument about platform-dependence of integer dtypes and non-platform dependence of inexact dtypes.

    I'm not totally clear on the evolving state of platform dependence of integer dtypes, but I maintain that numpy should either define all four of these aliases, or define none of them, and anything in between is confusing.

  11. ngoldbaum commented on Sep 22, 2023

    @ngoldbaum
    Member

    ping @mtsokol, there seems to be consensus about removing all of the aliases, would you like to take on following up in the code and docs?

  12. eric-wieser commented on Sep 22, 2023

    @eric-wieser
    Member

    What's the proposed spelling to get a C long if np.int_ is removed? Perhaps I'm out of the loop and we finally have it at np.long.

    Last I checked numpy had five signed integer types corresponding to C's 5 types; like C++, even though two happen to coincide in bit-width they are still treated as different types in things likE PEP3118.

    If we keep only int8, int16, int32, int64, then we have a secret 5th type that is not part of the np. namespace.

  13. mtsokol commented on Sep 22, 2023

    @mtsokol
    Member

    What's the proposed spelling to get a C long if np.int_ is removed? Perhaps I'm out of the loop and we finally have it at np.long

    This PR's description has an up-to-date table that maps Python canonical names to C types: #24651 (comment). It should be np.intp.

  14. eric-wieser commented on Sep 22, 2023

    @eric-wieser
    Member

    That table doesn't seem to reflect platform-dependence; intp is int_ptr_t which certainly doesn't coincidence with C long on Windows.

  15. ngoldbaum commented on Sep 22, 2023

    @ngoldbaum
    Member

    Oh huh, that's a good point that np.long is missing and that intp_t isn't quite the same thing. I think including the stdint pointer int types in the table makes sense along with splitting off long into its own row.

    How about instead of removing int_ we rename it long? Seems a more sensible name - it's a the size of a C long on all platforms right?

  16. eric-wieser commented on Sep 22, 2023

    @eric-wieser
    Member

    The current int_ name is because in python 2, int was represented by a C long, which is the same reason that np.float_ actually refers to a C double.

    There was a long period where np.long was not available as a name, because long ago np.long = builtins.long, which eventually became np.long = builtins.int in python 3.

    I think reclaiming np.long for the the C name is probably in scope for 2.0 if we want it to be.

    However, there is the argument that the C type names are a niche use case (and worse, confusing when comparing python int to C int); so maybe only np.intXX should get the obvious names, and the C type names should be harder to find as something like np.c_long, which matches ctypes.c_long.

  17. ngoldbaum commented on Sep 22, 2023

    @ngoldbaum
    Member

    The current "C-like" names we have for integer types in the python API are byte, short, intc, int_, and longlong. So I guess longc is also an option since it follows the precedent of intc. Not that either intc or longc are particularly intuitive names...

    I personally think long is fine, we already have short and longlong so exposing one more C type name just makes the naming in the API a little more consistent, we would have int too if that didn't step on a Python type name.

  18. eric-wieser commented on Sep 22, 2023

    @eric-wieser
    Member

    Does it matter that it steps on a python type name?

  19. ngoldbaum commented on Sep 23, 2023

    @ngoldbaum
    Member

    I personally don't think it matters if it steps on a Python 2 type name.

  20. ngoldbaum commented on Sep 23, 2023

    @ngoldbaum
    Member

    Oh I see, for intc. I think having np.int available and not correspond to a python int would introduce new confusion, especially since it's int32 which is very different from python's extended precision ints. There's no long type in Python anymore so there's no confusion with a python type if we have np.long.

  21. rgommers commented on Sep 23, 2023

    @rgommers
    Member

    I think I have a preference for long as well, since it's the most obvious name and matches most other C-like names (e.g., longlong). longc seems a bit ugly, clong could be confused with clongdouble & co where the c is for "complex", and c_long would be yet another naming convention.

  22. seberg commented on Sep 23, 2023

    @seberg
    Member

    FWIW, I had also added long to "replace" int_ in the PR to change the default integer (I don't think I had removed int_, except in Cython where it cannot exist as int_t). long has been awkward since Python 2 is gone (or longer), so I am fine with reusing it. I think I would lean to not have np.int at all.

  23. mtsokol commented on Sep 25, 2023

    @mtsokol
    Member

    I've got one question: Do we want to keep np.dtype("int") and np.dtype("uint")?
    If yes, should it correspond to np.intc or np.long?

    I think it can still correspond to np.long mixing C and Python names, similarly as np.dtype("float") and np.dtype("double"), that also map to np.float64.

    Please note that np.dtype("intc") and np.dtype("long") will be both available.

  24. seberg commented on Sep 25, 2023

    @seberg
    Member

    It would correspond to the default integer (currently long). I don't much like it, because its confusing with the C types, but I half suspect there is such a huge amount of dtype="int" code in practice that I am not willing to touch it.

  25. seberg commented on Sep 25, 2023

    @seberg
    Member

    Looking at @mtsokol changes, there seems to be very few uses of int_ in actual documentation which is good news, because in actual use replacing it with long is correct, but not nice if we change the default to intp. (At which point we don't want anyone to use long unless they are interacting with C explicitly!)

    I am not sure to I like the idea of replacing np.int_([1, 2, 3]) with np.intp([1, 2, 3]) either. So TBH, I would also be just fine living with the fact that we have the underscore for the default int, but not the default float/complex (because those are always clear). Is it really that bad?

    It may be OK either way because asking for "default integer" is really a bit silly in most places. (So the few times someone needs it, it is maybe fine if they have to use long/intp.)

  26. ngoldbaum commented on Oct 17, 2023

    @ngoldbaum
    Member

    We ended up landing on keeping int_ and uint_, at least for a few more releases. We still need a way for users and library authors to write "the platform-dependent default numpy integer type". The plan for NumPy 2.0 is for int_ to correspond to intp on all platforms, so we could remove int_ and tell people to replace it with intp, but we ended up deciding that is not a big enough clarity improvement to justify the code churn. The argument that this introduces an inconsistency in the API due to the lack of float_ wasn't persuasive to others on the call when I brought that up.

    I'm going to close this now, reflecting the decision that was made at the last community meeting.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    62 - Python APIChanges or additions to the Python API. Mailing list should usually be notified.Numpy 2.0 API Changes

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions