Skip to content

int.to_bytes() and int.from_bytes(): raise ValueError when bytes count is zero #71810

Description

@socketpair
BPO 27623
Nosy @mdickinson, @socketpair, @vadmium, @serhiy-storchaka, @Vgr255, @lucasem, @Phaqui
Files
  • int_to_bytes_overflow_1.patch
  • int_to_bytes_overflow_cornercase.patch: Throws OverflowError on (-1).to_bytes(0, 'big', signed=True)
  • int_to_bytes_overflow_cornercase2.patch: (-x).to_bytes(0, 'big', signed=True), x=0 or -1 now throws OverflowError; also updated the tests. They run successfully.
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2016-07-26.06:26:20.726>
    labels = ['interpreter-core', 'type-bug']
    title = 'int.to_bytes() and int.from_bytes(): raise ValueError when bytes count is zero'
    updated_at = <Date 2016-07-30.01:21:55.777>
    user = 'https://github.com/socketpair'

    bugs.python.org fields:

    activity = <Date 2016-07-30.01:21:55.777>
    actor = 'martin.panter'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['Interpreter Core']
    creation = <Date 2016-07-26.06:26:20.726>
    creator = 'socketpair'
    dependencies = []
    files = ['43913', '43925', '43942']
    hgrepos = []
    issue_num = 27623
    keywords = ['patch']
    message_count = 14.0
    messages = ['271329', '271330', '271476', '271479', '271484', '271485', '271487', '271499', '271504', '271506', '271564', '271565', '271657', '271660']
    nosy_count = 7.0
    nosy_names = ['mark.dickinson', 'socketpair', 'martin.panter', 'serhiy.storchaka', 'abarry', 'lucasem', 'Phaqui']
    pr_nums = []
    priority = 'normal'
    resolution = None
    stage = 'patch review'
    status = 'open'
    superseder = None
    type = 'behavior'
    url = 'https://bugs.python.org/issue27623'
    versions = ['Python 3.6']

    Linked PRs

    Activity

    socketpair commented on Jul 26, 2016

    socketpairmannequin
    MannequinAuthor

    As you can see, these conversions are not consistent. What is use case to allow that?

    ===================

    In [22]: (-1).to_bytes(0, 'big', signed=True)
    Out[22]: b''

    In [23]: (0).to_bytes(0, 'big', signed=True)
    Out[23]: b''

    As you can see, two different values serialized to same empty bytes sequence.

    ===================

    In [28]: int.from_bytes(b'', 'big', signed=True)
    Out[28]: 0

    In [29]: int.from_bytes(b'', 'big', signed=False)
    Out[29]: 0

    Anyway, -1 can not be deserialized.

    ===================

    added
    stdlibStandard Library Python modules in the Lib/ directory
    type-bugAn unexpected behavior, bug, or error
    on Jul 26, 2016

    socketpair commented on Jul 26, 2016

    socketpairmannequin
    MannequinAuthor

    So, as I think, it must ValueError when bytes count is zero. This is like division by zero.

    lucasem commented on Jul 28, 2016

    lucasemmannequin
    Mannequin

    This is actually a problem in Objects/longobject.c, in the _PyLong_AsByteArray function. It should have given an overflow error, because -1 cannot be encoded in 0 bytes.

    phaqui commented on Jul 28, 2016

    phaquimannequin
    Mannequin

    Isn't it possible to just add a small line of code that checks if length is less than or equal to 0, and if it is, call the necessary c functions to have python raise a valueerror...? Sorry if this is giving a solution without actually submitting the patch - but this is all very new to me. I have never contributed to anything yet (fourth year CS-student), and I am as fresh to the process as can be. I registered here just now. This seems like an issue I could handle.. I just need to take the time to learn about the process of how things are done around here.

    Vgr255 commented on Jul 28, 2016

    Vgr255mannequin
    Mannequin

    Here's a patch that fixes this.

    added
    interpreter-core(Objects, Python, Grammar, and Parser dirs)
    and removed
    stdlibStandard Library Python modules in the Lib/ directory
    on Jul 28, 2016

    socketpair commented on Jul 28, 2016

    socketpairmannequin
    MannequinAuthor

    I think, we should deny:

    • Passing 0 to to_bytes (even if integer is equal to zero)
    • Passing empty string to from_bytes

    I do not see any use cases for this to work. It was never guaranteed to work earlier. Everyone pass constant to to_bytes that is > 0. And passing empty buffer to from_bytes just indicate error in logic (i.e. something was not read correctly from file/socket e.t.c).

    struct.pack/unpack does not support zero-byte types.

    Vgr255 commented on Jul 28, 2016

    Vgr255mannequin
    Mannequin

    I don't use this feature enough to have a clear opinion, however it's specifically accounted for in the code and has a test for it. It might be a good idea to bring this up on Python-ideas. It's very likely to break some code, but I wonder if said code wasn't already broken to begin with :)

    vadmium commented on Jul 28, 2016

    @vadmium
    Member

    I agree that the signed conversion cases should be an error.

    However the unsigned case would break working code that I have written for bijective numeration. See _bytes_to_int() and _int_to_bytes() in bpo-20132, inc-codecs.diff, for example. Since non-zero unsigned conversions work by converting

    N bytes <-> 0 <= value < 2^N

    For N = 0, there is only one possible value, 0.

    serhiy-storchaka commented on Jul 28, 2016

    @serhiy-storchaka
    Member

    I agree with Martin. The ambiguous signed conversion cases should be an error, the unambiguous unsigned conversion case should be supported (especially if there are tests for this).

    mdickinson commented on Jul 28, 2016

    @mdickinson
    Member

    The ambiguous signed conversion cases should be an error, the unambiguous unsigned conversion case should be supported

    +1. A signed representation *requires* 1 bit for the sign (regardless of whether the number being represented is negative or nonnegative), so it should be an error to encode into zero bytes. But there's nothing wrong with trying to encode 0 in zero bytes when using an unsigned representation.

    phaqui commented on Jul 28, 2016

    phaquimannequin
    Mannequin
    So, am I to understand that the only corner case we should fix is that
    >>> (-1).to_bytes(0, 'big', signed=True)
    should raise an overflow error (currently it returns  b'') ?

    8 remaining items

    added 2 commits that reference this issue on Sep 11, 2025
    added a commit that references this issue on Sep 11, 2025
    added a commit that references this issue on Sep 12, 2025
    added 2 commits that reference this issue on Sep 13, 2025
    added 2 commits that reference this issue on Sep 14, 2025
    added a commit that references this issue on Sep 14, 2025
    added a commit that references this issue on Oct 7, 2025
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      interpreter-core(Objects, Python, Grammar, and Parser dirs)type-bugAn unexpected behavior, bug, or error

      Projects

      No projects

        Milestone

        No milestone

        Relationships

        None yet

        Development

        No branches or pull requests

        Issue actions