Repository navigation
Conversation
|
Thank you for your contribution to Astropy! 🌌 This checklist is meant to remind the package maintainers who will review this pull request of some common things to look for.
|
|
👋 Thank you for your draft pull request! Do you know that you can use |
|
Yes using dask for tile decompression makes perfect sense, and more generally having better support to load arrays with dask. |
|
I don't have any preference either way, but one idea I had if we don't want to be married to a particular library (such as dask), and wanted the api to support other array types in the future (maybe cupy, xarray, etc.) is to use numpy's solution to this type of problem by introducing a |
|
Lovely, and indeed magically easy! But agree with @byrdie that ideally one would have something more generic, e.g., pass in the class of object one wants to use or a callable or so (maybe with strings for commonly used ones). May need to check whether there is anything relevant in the array specification work (https://data-apis.org/array-api/2021.12/index.html)... @nstarman has thought much more about this, so pinging him. |
|
I don't think the array-api specification is super relevant here, but I agree with both @byrdie and @mhvk that it would be great to be array-library generic. Looking at the PR thus far, the switch happens here (following code snippet). If if self._use_dask:
import dask.array as da
data = da.from_array(self.section, chunks=_tile_shape(self._header))I would suggest def open(..., format: str | Callable[...]="numpy"):
....And the magic sauce could look something like this: _METHOD_REGISTRY: dict[str, Callable[...]] # eventually public scope a method to register.
if callable(format):
method = format
elif format in _METHOD_REGISTRY:
method = _METHOD_REGISTRY[format]
else:
raise Exception("something useful")
data = method(self, ...)This PR would only refactor the "numpy" pre-registered callable and add the "dask" pre-registered callable. |
|
@astrofrog this needs rebasing when you get a chance. |
…dask array (only on CompImageHDU for now)
|
I just wanted to raise another idea for how to approach this - one of the issues at the moment with my proposed approach here is that it assumes that a single data 'format' can apply to all HDU types when in fact this is not necessarily the case, or not always unambiguous. For instance if a user wants dask for a binary table, they could ask for example want a dask dataframe but they might also instead want an astropy table with dask columns. So in principle the available 'formats' might depend on the HDU type. I wonder if it might therefore make more sense to not change the type of the In [3]: hdu.view_data_as('dask-array')
Out[3]: dask.array<array, shape=(10000, 10000), dtype=int32, chunksize=(2000, 2000), chunktype=numpy.ndarray>Then for binary tables, one could have e.g. The advantages are:
What do people think about this idea? As a side note, for now, I think I would avoid implementing any kind of public registry for the converters, because for them to be efficient and not simply call |
|
I think that's much better indeed! In that case, One thing: Regardless of those implementation details, this is I think much better indeed than passing something in during |
|
@mhvk - thanks for the suggestions! I have now pushed a new commit that changes the API. A few notes/questions (for you and others):
I have only implemented this on the compressed HDU for now but once we have converged on the above, I can copy the method over to other HDU types for now, just implementing 'numpy', and we can add dask and other formats in subsequent PRs. I'll also work on adding some docs here. |
|
All your points make sense. For |
mhvk
left a comment
There was a problem hiding this comment.
This looks good!
Now that I see it, though, I'm less sure about the name, numpy's astype is really for changing to a different dtype... Could it be as simple as get_data or to_array? That would work especially if combined with the property (with a possible set_data or from_array counterparts?).
| from .conftest import FitsTestCase | ||
| from .test_table import comparerecords | ||
|
|
||
| try: |
There was a problem hiding this comment.
Bit off-topic, but should dask become an optional dependency? It seems at the same level as pandas.
| "pytest": ("https://docs.pytest.org/en/stable/", None), | ||
| "ipython": ("https://ipython.readthedocs.io/en/stable/", None), | ||
| "pandas": ("https://pandas.pydata.org/pandas-docs/stable/", None), | ||
| "dask": ("https://docs.dask.org/en/stable/", None), |
| ) | ||
| ) | ||
|
|
||
| def data_astype(self, data_type): |
There was a problem hiding this comment.
My sense would be to use this function for the property, i.e., here do data_type='numpy' and then have data = property(data_astype).
I think you can still do @data.setter after, though it may be an idea to already know make this a 2-way process, i.e., have a set_data() (class) method. But probably fine to postpone that (though for my baseband package, I have found that writing converters in two directions really helps to get rid of bugs!).
|
Hi humans 👋 - this pull request hasn't had any new commits for approximately 4 months. I plan to close this in 30 days if the pull request doesn't have any new commits by then. In lieu of a stalled pull request, please consider closing this and open an issue instead if a reminder is needed to revisit in the future. Maintainers may also choose to add keep-open label to keep this PR open but it is discouraged unless absolutely necessary. If this PR still needs to be reviewed, as an author, you can rebase it to reset the clock. If you believe I commented on this pull request incorrectly, please report this here. |
|
I'm going to close this pull request as per my previous message. If you think what is being added/fixed here is still important, please remember to open an issue to keep track of it. Thanks! If this is the first time I am commenting on this issue, or if you believe I closed this issue incorrectly, please report this here. |
This is work towards part 2 of the proposal for funding Improve FITS compressed image performance and Dask integration in io.fits
This should be rebased once #14430 is merged
This is an experimental PR to show how we could easily (at least for some FITS HDU classes) have an option to get dask arrays - for now I'm focusing on
CompImageHDUbecause there are clear gains to be made since tiles can be decompressed in parallel but we can extend this to other HDU classes if we agree on the approach/API. With this PR, we can control whether.datais a Numpy array or a dask array. To demonstrate this, we can first generate a test dataset:We can now read this in using the default options:
With this PR, we can specify
use_dask=Truein theopenoptions:And we can then choose to use e.g. the multi-threaded scheduler to compute the array, which is faster (the wall time is the relevant time here):
A factor of 2x or more faster compared to the non-dask version! The reason this works is because in #14430 we made it so the GIL is released inside the C extensions.
In any case, I think it would be useful to make it easy for users to get dask arrays out of HDUs, but I'm open as to how we achieve this. We can either use
use_daskas done here, or we could instead have a propertydask_dataor similar that exists onCompImageHDU. We could add this to all HDU classes even though for some of them there might not be a benefit yet (but we can always improve the efficiency of the implementation on different HDU classes over time).Having this would address (I think) the last point in #3895 which was to implement multi-threaded decompression of tiles - rather than implement a thread pool ourselves, I think it would be much cleaner to simply use dask to achieve this.
@saimn - do you have any thoughts on this? The changes in this PR are actually trivial (see last commit) - the rest is commits that will go away once other PRs are merged.
Obviously if we go ahead I would add docs etc to demonstrate this (and tests etc).