Skip to content

bug: tobytes() byteorder with multibyte samples, bigendian #385

Description

@manisandro

Short story and end result: numpy tests are failing on bigendian arches, as reported here: https://bugzilla.redhat.com/show_bug.cgi?id=1019656

Long story: The problem is that the test image 12bit.cropped.tif is a tiff image with data stored in little-endian format and 16bit samples. Now, Image.tobytes() (which is called when doing numpy.array(img)) returns the raw data, without byteswapping. This means that i.e.

img = Image.open('12bit.cropped.tif')
print img.tobytes()[0:2]

returns '\xe0\x01' on both big- and little-endian platforms, which corresponds (correctly) to 480 on little-endian, but on big-endian it corresponds to 57345.

FIxing the issue itself should not be too much trouble, the main issue is where to fix it.
Some possibilities (without very in-depth study of the code):

  • The test is wrong. tobytes() returns the raw data, and if that appens to be in little-endian format, then it is up to the user to handle this.
  • TiffDecode should byteswap if the endianness of the image data differs from the platform endinanness.

Any other ideas?

Activity

  1. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    (recognizing my code there...)

    I think the decoder should be putting the bytes into native order, i.e. what we get out of the tiff decoder should always be a I;16N rawmode, no matter if it's a II or MM tiff. I'd bet the libtiff decoder already does this.

    On the other hand, I wouldn't bet against the test being wrong, as I didn't test it against a big endian machine. Time to fire up the old mini I guess.

    I'll dig some on this. (oh, it would be nice if there was documentation of what the in memory bits and assumptions were supposed to be. I'm so sure I've been over this before)

  2. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    Are there any places to get a bigendian virtual machine to test against?

  3. manisandro commented on Oct 18, 2013

    @manisandro
    MemberAuthor

    I've been granted access to a ppc64 cluster by a Fedora contributor. If you post your public ssh key and give a username, I can see whether they can give you access also.

  4. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    Thanks, wiredfool works as a username.

    https://gist.github.com/wiredfool/c441e27403d09239a0e9

    A little googling led to qemu + a ppc distro + mucking about as well, so there may be a plan b.

  5. manisandro commented on Oct 18, 2013

    @manisandro
    MemberAuthor

    Request posted https://bugzilla.redhat.com/show_bug.cgi?id=1019656#c4

    I tried with qemu a while back, little success (if memory serves me right, there was some major component not fully working for ppc back then, and possibly still).

  6. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    Thanks for the account request.

    My standard estimation on the qemu approach is that it's a day of chasing dependencies. (based on no one giving detailed directions for it working out of the box).

  7. manisandro commented on Oct 18, 2013

    @manisandro
    MemberAuthor

    Dependency-wise it's pretty ok on Fedora, just # yum install qemu and then grab a ppc iso (i.e. from http://mirrors.fedoraproject.org/publiclist/Fedora/19/ppc/) and then

    qemu-system-ppc64 -cdrom Fedora-19-ppc64-netinst.iso
    

    Last time I only saw a yellow screen, but now something actually seems to happen (albeit very slow).

  8. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    I'm on ubuntu for a host os, and after tracking down an openbios and mucking with the network configuration, I've got a black screen and 100% processor usage. for 15 minutes. It's either glacial or broken. Hopefully your guy will come through.

  9. gustavold commented on Oct 18, 2013

    @gustavold

    wiredfool, please try ssh chinook.fedora.osuosl.org and let me know if you have any problem logging in.

  10. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    that would be @gustavold , it's asking for a pw. sorry gustavoid.

  11. gustavold commented on Oct 18, 2013

    @gustavold

    Oh, sorry. I think it was a file permissions issue. Try again, please.

  12. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    Still not working.

    ssh [email protected] 
    [email protected]'s password: 
    Permission denied, please try again.
    [email protected]'s password: 
    Permission denied, please try again.
    [email protected]'s password: 
    Permission denied (publickey,gssapi-keyex,gssapi-with-mic,password).
    
  13. gustavold commented on Oct 18, 2013

    @gustavold

    My bad. It is working now. Sorry about that.

  14. wiredfool commented on Oct 18, 2013

    @wiredfool
    Member

    \o/

  15. wiredfool commented on Oct 19, 2013

    @wiredfool
    Member

    Ok, I've got Pillow built with enough libraries to replicate this on chinook.

    My understanding at this point is:

    • to bytes should return the in memory representation.
    • 16 bit integer tiffs don't do any endian shuffling. (with the native decoder, I haven't checked the libtiff one yet)
    • there aren't any conversions from I;16B to I;16L, or vice versa.
    • Loading a MSB version of this image results in swapped bytes on both platforms. \x01\xe0
    • The mode for an LSB image is I;16, for MSB is I;16B.
    • img.getpixel((0,0)) returns 480 in all 4 combinations of machine/image endianness.
    • It appears that the internal image format is image native, which could be either endian.
  16. wiredfool commented on Oct 21, 2013

    @wiredfool
    Member

    More understanding:

    • The numpy tests were right. The type strings for numpy were wrong. 366f9a5
    • Libtiff decoding does bit twiddling, so tiffs are stored in native order. I believe this is a bug.
      • There is now a new failing test for libtiff interpreting little endian tiffs on big endian machines. 1945fec
      • Native tiff decoding does not do bit twiddling, so tiffs are stored in image byte order. This seems correct.

    I'm keeping this in my bigendian branch: https://github.com/wiredfool/Pillow/tree/bigendian

  17. added a commit that references this issue on Sep 24, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Big-endianBig-endian processors

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions