Repository navigation
bug: tobytes() byteorder with multibyte samples, bigendian #385
Description
Activity
(recognizing my code there...)
I think the decoder should be putting the bytes into native order, i.e. what we get out of the tiff decoder should always be a I;16N rawmode, no matter if it's a II or MM tiff. I'd bet the libtiff decoder already does this.
On the other hand, I wouldn't bet against the test being wrong, as I didn't test it against a big endian machine. Time to fire up the old mini I guess.
I'll dig some on this. (oh, it would be nice if there was documentation of what the in memory bits and assumptions were supposed to be. I'm so sure I've been over this before)
Are there any places to get a bigendian virtual machine to test against?
I've been granted access to a ppc64 cluster by a Fedora contributor. If you post your public ssh key and give a username, I can see whether they can give you access also.
Thanks, wiredfool works as a username.
https://gist.github.com/wiredfool/c441e27403d09239a0e9
A little googling led to qemu + a ppc distro + mucking about as well, so there may be a plan b.
Request posted https://bugzilla.redhat.com/show_bug.cgi?id=1019656#c4
I tried with qemu a while back, little success (if memory serves me right, there was some major component not fully working for ppc back then, and possibly still).
Thanks for the account request.
My standard estimation on the qemu approach is that it's a day of chasing dependencies. (based on no one giving detailed directions for it working out of the box).
Dependency-wise it's pretty ok on Fedora, just
# yum install qemuand then grab a ppc iso (i.e. from http://mirrors.fedoraproject.org/publiclist/Fedora/19/ppc/) and thenqemu-system-ppc64 -cdrom Fedora-19-ppc64-netinst.isoLast time I only saw a yellow screen, but now something actually seems to happen (albeit very slow).
I'm on ubuntu for a host os, and after tracking down an openbios and mucking with the network configuration, I've got a black screen and 100% processor usage. for 15 minutes. It's either glacial or broken. Hopefully your guy will come through.
wiredfool, please try ssh chinook.fedora.osuosl.org and let me know if you have any problem logging in.
that would be @gustavold , it's asking for a pw. sorry gustavoid.
Oh, sorry. I think it was a file permissions issue. Try again, please.
Still not working.
ssh [email protected] [email protected]'s password: Permission denied, please try again. [email protected]'s password: Permission denied, please try again. [email protected]'s password: Permission denied (publickey,gssapi-keyex,gssapi-with-mic,password).My bad. It is working now. Sorry about that.
\o/
Ok, I've got Pillow built with enough libraries to replicate this on chinook.
My understanding at this point is:
- to bytes should return the in memory representation.
- 16 bit integer tiffs don't do any endian shuffling. (with the native decoder, I haven't checked the libtiff one yet)
- there aren't any conversions from I;16B to I;16L, or vice versa.
- Loading a MSB version of this image results in swapped bytes on both platforms.
\x01\xe0 - The mode for an LSB image is I;16, for MSB is I;16B.
- img.getpixel((0,0)) returns 480 in all 4 combinations of machine/image endianness.
- It appears that the internal image format is image native, which could be either endian.
More understanding:
- The numpy tests were right. The type strings for numpy were wrong. 366f9a5
- Libtiff decoding does bit twiddling, so tiffs are stored in native order. I believe this is a bug.
- There is now a new failing test for libtiff interpreting little endian tiffs on big endian machines. 1945fec
- Native tiff decoding does not do bit twiddling, so tiffs are stored in image byte order. This seems correct.
I'm keeping this in my bigendian branch: https://github.com/wiredfool/Pillow/tree/bigendian
Short story and end result: numpy tests are failing on bigendian arches, as reported here: https://bugzilla.redhat.com/show_bug.cgi?id=1019656
Long story: The problem is that the test image 12bit.cropped.tif is a tiff image with data stored in little-endian format and 16bit samples. Now, Image.tobytes() (which is called when doing numpy.array(img)) returns the raw data, without byteswapping. This means that i.e.
returns
'\xe0\x01'on both big- and little-endian platforms, which corresponds (correctly) to 480 on little-endian, but on big-endian it corresponds to 57345.FIxing the issue itself should not be too much trouble, the main issue is where to fix it.
Some possibilities (without very in-depth study of the code):
Any other ideas?