Repository navigation
Use CFFI for the pixel access object #248
Description
Activity
Cool 👍
Is this a 2.2.0 or some other milestone target?
@wiredfool Will this happen for 2.3.0?
Maybe. I think I know what has to happen.
Last call for 2.3.0!
Ok, made some progress on this one. I'm able to read and write one pixel of an rgb image through cffi without segfaulting or going through the _imaging c API for the get pixel/put pixel part.
- added a commit that references this issue
on Jan 5, 2014 Still some undones, but it nominally works.
Time to merge?
No. This demonstrates that it's possible. It's not complete, one of the tests is failing, documentation is iffy at best, and it hasn't proved that's it's any faster than the old version either. I don't expect that it'll be slower, it just hasn't been tested.
The crux on this was that the ffi.buffer call returns new memory that it owns, and duplicating the memory is a non-starter for the pixel access methods. The bit that I was missing is that you can use
ffi.cast('foo *', a_pointer)and use the already allocated memory at a_pointer* as a foo.That gets us everything we need, so long as we can get the pointer out of the c-api side, and that can be done with a 'n' format in Py_BuildValue.This patch can be cut down pretty dramatically, as we probably don't need most of the parts of the ImageMemoryInstance object, we're just using xsize, ysize, and one of the image8/image32 pointers. (and possibly the image pointer). We can avoid creating an ImagingMemoryInstance object at all, and just use the one pointer that we need.
Benchmarks:
pypy2.2(vpypy22)erics:~/Pillow$ python Tests/bench_cffi_access.py --installed Size: 128x128 PyAccess - get: 0.0087 s PyAccess - set: 0.0227 s C-api - get: 0.1372 s C-api - set: 0.2737 spy2.7:
(vpy27)erics:~/Pillow$ python Tests/bench_cffi_access.py --installed Size: 128x128 PyAccess - get: 0.0285 s PyAccess - set: 0.0475 s C-api - get: 0.0027 s C-api - set: 0.0056 s okSo, a win for pypy, not so much for c-python.
Those times are way too small for PyPy - please run it in a loop so the total is at least a second
Ok, longer, running 5k iterations or 10 seconds:
pypy2.2(vpypy22)erics:~/Pillow$ python Tests/bench_cffi_access.py --installed Size: 128x128 PyAccess - get: 0.7001 s 0.000140 per iteration PyAccess - set: 1.3735 s 0.000275 per iteration C-api - get: breaking at 82 iterations, 0.122830 per iteration C-api - set: breaking at 37 iterations, 0.273431 per iterationpy2.7
(vpy27)erics:~/Pillow$ python Tests/bench_cffi_access.py --installed Size: 128x128 PyAccess - get: breaking at 454 iterations, 0.022029 per iteration PyAccess - set: breaking at 217 iterations, 0.046280 per iteration C-api - get: breaking at 3600 iterations, 0.002779 per iteration C-api - set: breaking at 1757 iterations, 0.005694 per iterationSo, pypy is 2 OOM faster when running longer? I wonder if the JIT is optimizing something important out of the loop, since the entire benchmark reduces to a bunch of noops.
And while I've got benchmarks on the brain, I'll note that test_image_point.py is really slow on pypy (2m6.9s) and not on cpython (.6 seconds).
Hoisted from #67, post-close, with additional brain dumping.
@fijal said:
The crux that I see is that (IIRC) the data isn't necessarily in a continuous block. If it were, the raw data pointer would be easy.
The backbrain is thinking that the solution might be related to checking the strides from #224. If the data's not continuous, we can fall back to slow access.
If a cffi approach would work, we'd need to require cffi and libffi-dev, but it's really only necessary on pypy, and anyone doing this on pypy would likely be ok an additional dependency.
Docs are here: http://cffi.readthedocs.org/en/latest/index.html
It looks like this is what we'd need:
So, Assuming that I can get a pointer from something like PyLong_FromVoidPtr, this should give higher performance access to the pixel_access_object.