Repository navigation
Large JPEG with DAC compression gives OSError: broken data stream when reading image file #9162
Description
Activity
From https://gitlab.gnome.org/GNOME/glycin/-/issues/179#note_2525412
Okay, so DAC stands for define-arithmetic-coding-conditioning according to T.81 (the original JPEG standard.) It's an alternative way to DHT (define-Huffman-tables) for the second compression step in JPEGs. I don't see it mentioned in later JPEG standards.
libjpeg-turbo has been the benchmark for JPEG decoders so far, and it doesn't support this compression method. That's why so many apps can't handle it. We have to ask jpeg-zune if they are interested in supporting this. But I wouldn't be surprised if they are not interested in implementing an additional decoding algorithm.
Could an unsupported compression format cause
OSError: broken data stream when reading image filein Pillow? Supporting the format would be nice, but even if it doesn't make sense to add support for it, that seems like a misleading error message.If adding the format doesn't make sense, would it make sense to add a new knob like
PIL.ImageFile.LOAD_TRUNCATED_IMAGESthat's specific to unsupported formats? All I really want to do is get the image dimensions and some EXIF metadata after runningPIL.ImageOps.exif_transpose(), so I don't actually need the image data itself.If you aren't concerned about the pixel data, then
LOAD_TRUNCATED_IMAGESwill actually work perfectly fine.from PIL import Image, ImageFile, ImageOps ImageFile.LOAD_TRUNCATED_IMAGES = True im = Image.open("P1080392-raw-P1080403-raw.jpg") im.load() print(im.size) transposed_im = ImageOps.exif_transpose(im) print(transposed_im.getexif())
gives
(13704, 4563) {256: 13704, 257: 4563, 258: 8, 296: 1, 530: (1, 1), 531: 1, 282: 1.0, 283: 1.0}If you aren't concerned about the pixel data, then
LOAD_TRUNCATED_IMAGESwill actually work perfectly fine.It works for this image, but couldn't it hide unrelated issues in other images that do affect metadata too? That's probably an ok enough trade-off for what I'm doing, but it would be nicer to only ignore things that can't affect metadata.
Re supporting DAC (if that is the cause of this issue), the other bug I filed in glycin had a comment about libjpeg-turbo/libjpeg-turbo#120. Any chance that Pillow is feeding data to libjpeg-turbo in chunks, and just giving it the entire file all at once would fix this issue?
If you aren't concerned about the pixel data, then
LOAD_TRUNCATED_IMAGESwill actually work perfectly fine.It works for this image, but couldn't it hide unrelated issues in other images that do affect metadata too? That's probably an ok enough trade-off for what I'm doing, but it would be nicer to only ignore things that can't affect metadata.
The only case I can think of where
LOAD_TRUNCATED_IMAGESwould affect metadata would be for PNG images, where the setting skips checksum validation - but in that case, adding this setting would give the user more metadata.This is an image that isn't loading properly. For most formats (PNG being a potential exception), metadata occurs before the image data, so the metadata and unsupported image data seem unrelated to me.
Reacted by David MandelbergRe supporting DAC (if that is the cause of this issue), the other bug I filed in glycin had a comment about libjpeg-turbo/libjpeg-turbo#120. Any chance that Pillow is feeding data to libjpeg-turbo in chunks, and just giving it the entire file all at once would fix this issue?
You are correct, passing the entire file at once would fix this.
from PIL import Image, ImageFile ImageFile.MAXBLOCK = 49339501 im = Image.open("P1080392-raw-P1080403-raw.jpg") im.load()
That isn't something that to change permanently in Pillow itself, because we want to have more reasonable limits on how resources are used when opening random and potentially untrusted images, but if you would like to adjust the setting in your own code because you know you are dealing with large images, you can.
I don't see
MAXBLOCKat https://pillow.readthedocs.io/en/stable/reference/ImageFile.html. Is that meant to be a stable-ish part of the API for code to set? From a glance at the code, it looks like it doesn't support setting toNoneto disable the limit?Makes sense about the default limits. Since libjpeg-turbo doesn't seem likely to add support for partial DAC-coded images, would it make sense to just change the error message in Pillow to explain that
MAXBLOCKcan be increased if needed?- changed the title
[-]JPEG, maybe with DAC header, gives `OSError: broken data stream when reading image file`[/-][+]Large JPEG with DAC compression gives `OSError: broken data stream when reading image file`[/+]on Aug 19, 2025 I've created #9163 to document
MAXBLOCK. See https://pillow--9163.org.readthedocs.build/en/9163/reference/ImageFile.html#PIL.ImageFile.MAXBLOCK for what the documentation would look like.I'm not completely on board with the idea of having a security measure, but telling people to workaround it at the first sign of trouble. As I think you've gathered, this particular data format is uncommon, but we do get truncated image reports semi-frequently. I'm reluctant to tell people to weaken their security if it might not actually the cause of the problem.
Thanks for the documentation!
I'm guessing it's not easy to detect that the error comes from a format that doesn't support partial data? I was thinking of only changing the error message in that case.
A bit out of scope of this bug, but I wonder if it would make sense to have some sort of security profiles. Something like a server profile and a client profile, where only the server profile protects against resource exhaustion1. Probably not worth it for just this, but maybe worth it if disabling MAX_IMAGE_PIXELS is common in client apps?
Any chance that there's enough of a pattern to the truncated image reports, that it would make sense to add something to the documentation about what can cause the error, with this as one of the less-prominent options?
Footnotes
-
That's the only security reason for MAXBLOCK, right? I don't really care about resource exhaustion attacks via files I'm choosing to open locally since I can always kill the program and not open that file again, but I do care about arbitrary code execution. ↩
-
I'm guessing it's not easy to detect that the error comes from a format that doesn't support partial data? I was thinking of only changing the error message in that case.
We could detect if the DAC marker is present, and if a "broken data stream" error is present, and raise a different message accordingly.
However, I don't think there's any way to say conceptually if the error is because of lack of suspension support, or because an image is truncated. Surely a suitably truncated image would trigger exactly the same response from libjpeg-turbo. Pillow doesn't necessarily know if an image has more data or not.
A bit out of scope of this bug, but I wonder if it would make sense to have some sort of security profiles. Something like a server profile and a client profile, where only the server profile protects against resource exhaustion1. Probably not worth it for just this, but maybe worth it if disabling MAX_IMAGE_PIXELS is common in client apps?
It sounds like you'd like to add a new mechanism just to avoid changing two separate settings? If you're open to the possibility of exhausting your local resources, that's fine, but Pillow would also be used a library that end users run on their machines. If someone glanced at Pillow documentation and chose 'client profile', thinking they weren't going to run this on a server, I wouldn't want them to unintentionally make their users vulnerable. If there's a common Python config tool that implements that sort of thing that you think is relevant, feel free to create a new issue suggesting it.
Any chance that there's enough of a pattern to the truncated image reports, that it would make sense to add something to the documentation about what can cause the error, with this as one of the less-prominent options?
I've looked through our "broken data stream when reading image file" issues, and they generally end with either acceptance that an image is truncated, or a pull request on our part that fixes support for the image.
If you would like specific documentation around this, my suggestion would be at https://pillow.readthedocs.io/en/stable/handbook/image-file-formats.html#jpeg, since there is already mention of truncated JPEG images there.
However, I don't think there's any way to say conceptually if the error is because of lack of suspension support, or because an image is truncated.
That's too bad.
It sounds like you'd like to add a new mechanism just to avoid changing two separate settings? If you're open to the possibility of exhausting your local resources, that's fine, but Pillow would also be used a library that end users run on their machines. If someone glanced at Pillow documentation and chose 'client profile', thinking they weren't going to run this on a server, I wouldn't want them to unintentionally make their users vulnerable. If there's a common Python config tool that implements that sort of thing that you think is relevant, feel free to create a new issue suggesting it.
TL;DR: I was thinking about it wrong before. I still think something like this might make sense, but not enough to push for it.
Yeah, I think client/server was the wrong way to think about it. Probably something more like whether the primary goal of the program is to load the image or not. E.g., a server probably has a more important goal of not letting users crash the server, and a web browser also has a more important goal of not letting a malicious website crash the browser. But for a dedicated image viewer, it's more important to show the requested image if at all possible than to avoid crashing, since crashing in that case has a pretty similar effect to not showing the requested image; either way, the user couldn't view the image, and a crashing image viewer shouldn't take anything else down with it.
For the program I'm writing now, setting 2 things is obviously easier than making any changes to Pillow itself. The only reason I was thinking about this is that as a potential user of programs that use Pillow, I would only want those programs to protect against resource exhaustion if the purpose of the program is better served by refusing to load an image.
I can see how it could get complicated though, e.g., if an image browser crashes, the user might forget what folder they were in. So it's probably better to only disable the protections on simple programs that do exactly one thing that depends entirely on loading one requested image.
I've looked through our "broken data stream when reading image file" issues, and they generally end with either acceptance that an image is truncated, or a pull request on our part that fixes support for the image.
If you would like specific documentation around this, my suggestion would be at https://pillow.readthedocs.io/en/stable/handbook/image-file-formats.html#jpeg, since there is already mention of truncated JPEG images there.
Gotcha. I kinda doubt I'd think to look at the JPEG docs specifically if I only saw an error about truncation. I did search the existing bugs though. So maybe it's better to just close this, but ask anybody else who comes across it in the future to leave a comment? Then if a few more users come across this it might make sense to figure out a better way to communicate what to do with DAC, and if not, then it's not worth the effort.
(My image is from Hugin, which I think is the most popular open source panorama creator. I don't know if it still produces DAC images, or if it depends on the settings though. And panoramas do tend to be large by their very nature. So I'd expect there are at least some other images like this floating around, but I don't have a sense of just how uncommon they are.)
Sure, if people come across this and let us know that they also found a problem here, we can revisit it.
What did you do?
I tried to load an image.
What did you expect to happen?
The image to load. gThumb and ImageMagick can load it, but Pillow and Loupe can't. Loupe says
Parsing of the following header `DAC` is not supported, which makes me think that it might be an unsupported format issue rather than a corrupt file?What actually happened?
What are your OS, Python and Pillow versions?
The image is 48MB, which seems to be too large to upload to GitHub, but I was able to upload it to another bug report about the same image.