Repository navigation
Image.open() fails on a JPEG2000 image #9455
Description
Activity
Thanks, that was fast! However, the JPEG2000 image has a CMYK palette. So 1) while PNG does not support CMYK, wouldn't it be more accurate to convert just the palette to RGB, i.e. to have an RGB-paletted PNG image on save()? 2) Pillow itself does not support CMYK-paletted images (not sure why). So, there's some CMYK->RGB conversion happening even before the save(). Again, is the conversion on the palette level, i.e. are we going to get mode 'P' after Image.open()? Also, the JPEG2000 image may have an ICC profile embedded (not sure this image has one, though; just considering). Will the ICC CMYK -> RGB conversion take into account the ICC profile? And if there's no ICC profile, what implicit ICC profile does Pillow use in such cases? Asking this since, apparently, this question is not discussed in the docs.
Are you able to attach a copy of what the image should look like? Also, are we able to include the original image in our test suite and distribute it under the Pillow license?
Here are 2 paletted, JPEG2000 images I created using an original image which is available under CC license:
https://openverse.org/image/b267ce10-c211-4be4-8537-76908afcbba4?q=cat&p=4
The images are:
cat-rgb-palette.jp2 — uses an RGB palette;
cat-cmyk-palette.jp2 — uses a CMYK palette.To get the most robust reference representation of these two images I embedded them inside a PDF file — the JPX streams in the PDF files are identical to the .jp2 files. Also, the reference PDF of the CMYK palette variant comes in two version: the one using an uncalibrated CMYK colorspace (/DeviceCMYK in PDF) and one with a calibrated (ICC profile-based one). Now, this profile is added to the PDF, but it's not in the .jp2 file / JPX image stream in PDF. It is added simply because without it (i.e. when a PDF uses an uncalibrated CMYK color space) the viewer app is free to choose whatever ICC profile it wishes as a "default" one. So to avoid the dependence of the viewed results on the viewer app I embedded in cat-cmyk-palette-icc.pdf the default CMYK profile that Acrobat uses to show PDFs with uncalibrated CMYK. Therefore: a) the two CMYK reference PDFs should appear identical if both viewed in Acrobat; b) cat-cmyk-palette-icc.pdf should, in principle, look the same in Acrobat and in any other viewer.
To clarify: the CMYK and RGB examples strongly differ in color (the colors are warm/cold for CMYK/RGB, resp.) because I used different palettes for the two — where "different" means in the absolute (physical) color sense — i.e., if one correctly remaps the CMYK-palette from the (calibrated) CMYK color space to RGB (=sRGB) one would still not get anything close to the RGB-palette used in my example. This is done on purpose to avoid confusion.
Now, I wasn't yet able to produce JPEG 2000 along the same lines, but with an ICC profile emedded in .jp2 as well (i.e., a situation, where the embedded palette maps from indices to an ICC-based color space). But let's deal with the palette issue first.
P.S. If I could remind: it would be great if Pillow exposed the index data, i.e. the 1-component pixel data before the palette-based mapping is applied — I am hinting here on making Pillow support CMYK-palettes generically, so that Image.open('file.jp2') would create a 'P' mode. With glymur this is doable with a trick (might be useful to whomever might write a patch for this issue):
def get(item, id:str): return next(b for b in item.box if b.box_id == id) jp2 = glymur.Jp2k('file.jp2') # call print(jp2k) to inspect the boxes pclr = None if jp2h := get(jp2, 'jp2h'): pclr = get(jp2h, 'pclr') if pclr: # palette is present jp2c = get(jp2, 'jp2c') with open(T('temp.jp2'), "rb") as f: f.seek(jp2c.offset + 8) # skip jp2c header codestream_bytes = f.read(jp2c.length - 8) open(T('temp2.jp2'), 'wb').write(codestream_bytes) j2k = glymur.Jp2k(T('temp2.jp2')) data = j2k[:] else: # no palette data = jp2[:] # This parses all JPEG2000 image slices for a full resolution image # get the icc_profile try: # this should be rewritten based on box names, not indices, like with pclr above icc_profile = jp2.box[3].box[1].icc_profile except: icc_profile = None
So, we dump the jp2c box as a separate .jp2 file, and when we reopen it with glymur the latter thinks it's a grayscale image, and that's how we get the index data. I suspect that Pillow uses OpenJPEG under the hood (just like glymur), and so something like this trick could avoid the necessity to dig into the OpenJPEG's source code.
These are the same JPEG2000 images converted to RGB/CMYK TIFF resp., i.e. non-paletted (=palette-mapped). Better to use these for reference than the PDFs in my previous post.




Running:
on image.jp2.zip fails:
My set-up:
MacOS 10.14.6, Python 3.10.18
pillow-12.1.1-cp310-cp310-macosx_10_10_x86_64.whl