Repository navigation
int.to_bytes() and int.from_bytes(): raise ValueError when bytes count is zero #71810
Description
Activity
As you can see, these conversions are not consistent. What is use case to allow that?
===================
In [22]: (-1).to_bytes(0, 'big', signed=True)
Out[22]: b''
In [23]: (0).to_bytes(0, 'big', signed=True)
Out[23]: b''
As you can see, two different values serialized to same empty bytes sequence.
===================
In [28]: int.from_bytes(b'', 'big', signed=True)
Out[28]: 0
In [29]: int.from_bytes(b'', 'big', signed=False)
Out[29]: 0
Anyway, -1 can not be deserialized.
===================
So, as I think, it must ValueError when bytes count is zero. This is like division by zero.
This is actually a problem in Objects/longobject.c, in the _PyLong_AsByteArray function. It should have given an overflow error, because -1 cannot be encoded in 0 bytes.
Isn't it possible to just add a small line of code that checks if length is less than or equal to 0, and if it is, call the necessary c functions to have python raise a valueerror...? Sorry if this is giving a solution without actually submitting the patch - but this is all very new to me. I have never contributed to anything yet (fourth year CS-student), and I am as fresh to the process as can be. I registered here just now. This seems like an issue I could handle.. I just need to take the time to learn about the process of how things are done around here.
Here's a patch that fixes this.
I think, we should deny:
- Passing
0toto_bytes(even if integer is equal to zero) - Passing empty string to
from_bytes
I do not see any use cases for this to work. It was never guaranteed to work earlier. Everyone pass constant to to_bytes that is > 0. And passing empty buffer to from_bytes just indicate error in logic (i.e. something was not read correctly from file/socket e.t.c).
struct.pack/unpack does not support zero-byte types.
I don't use this feature enough to have a clear opinion, however it's specifically accounted for in the code and has a test for it. It might be a good idea to bring this up on Python-ideas. It's very likely to break some code, but I wonder if said code wasn't already broken to begin with :)
I agree that the signed conversion cases should be an error.
However the unsigned case would break working code that I have written for bijective numeration. See _bytes_to_int() and _int_to_bytes() in bpo-20132, inc-codecs.diff, for example. Since non-zero unsigned conversions work by converting
N bytes <-> 0 <= value < 2^N
For N = 0, there is only one possible value, 0.
I agree with Martin. The ambiguous signed conversion cases should be an error, the unambiguous unsigned conversion case should be supported (especially if there are tests for this).
The ambiguous signed conversion cases should be an error, the unambiguous unsigned conversion case should be supported
+1. A signed representation *requires* 1 bit for the sign (regardless of whether the number being represented is negative or nonnegative), so it should be an error to encode into zero bytes. But there's nothing wrong with trying to encode 0 in zero bytes when using an unsigned representation.
So, am I to understand that the only corner case we should fix is that
>>> (-1).to_bytes(0, 'big', signed=True)
should raise an overflow error (currently it returns b'') ?
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs