Repository navigation
Unicode support for astropy units #9347
Description
Activity
"Somewhere" is.... sympy?
No, I really mean "somewhere". For me it was a webpage with copy / paste.
Ah, okay. 😅
@maxnoe , given it is Hacktoberfest, you want to attempt an exploratory PR? 😉
Yeah sure, can you direct me to an entry point?
Maybe here? Thanks!
Line 2160 in 4deafe9
si_prefixes = [ And the contribution guide, in case you need it: https://docs.astropy.org/en/latest/development/workflow/development_workflow.html
p.s. I am not sure where subscript is done, maybe @mhvk knows.
@maxnoe - this would be lovely!
There are two places where we could deal with unicode input. One is where @pllim pointed to - which is for the prefixes of individual units. I think this would get you the micro's - essentially just add it to the
shortlist for micro. Note that this also means that one will be able to dou.µg!!! (nothing in python stopping us from that). Though it is a bit annoying that there are multiple code points for micro - some string translation may be needed.For anything related to composition, thus including superscripts, the work is done by a string parser in
units.formats.generic.Right now, that does not deal with unicode at all. However, unicode is dealt with in some detail for parsing of angles from strings in
coordinates.angle_parser, so that might help to get an idea of how to do this.Ok, the micro part was quite easy, see #9348
Reacted by P. L. LimThe superscripts should be directly replaceable by their non-superscript counterparts.
Would you think this is a good approach?
unicode_translation_table = { ord('\N{SUPERSCRIPT MINUS}'): '-', ord('\N{SUPERSCRIPT ONE}'): '1', } s = s.translate(unicode_translation_table)
(see discussion in PR)
For adding support for stuff like
u.µΩ, I found a neat solution using module level__getattr__.What would you think about this (
astropy/units/__init__.py):def __getattr__(name): # python normalizes identifiers to ''NFKC", this applies the same # normalization if someone uses `getattr(u, name)` instead of `u.name` normalized = unicodedata.normalize('NFKC', name) try: return Unit(name) except ValueError: raise AttributeError(f"module '{__name__}' has no attribute '{name}'") from None
I wondered about that as well, and like your solution - but
module.__getattr__was introduced in python 3.7 and for astropy 4.0 we still support python 3.6...Okay, but it does not harm in 3.6, just a 3 line function doing nothing.
However, as I found out that python normalizes all identifiers to NFKC, the problem is much smaller than I thought.
So just adding three characters for Angstrom, Ohm and Micro would be all, the degree sign is not a valid identifier, neither are the super scripts.
My first reaction is not super-excited about non-ASCII attributes on units. I don't feel really strongly and can be convinced, but the units attribute list is already quite long, and I don't think that most users can easily type those when coding.
This seems different from being able to parse them as a string in
Unit(), which I think is wonderful.I concur, this is why I thought it would be a good compromise to use the getattr version, which would enable it via falling back to string parsing.
#9348 has implemented this
Dealing with unit strings received from somewhere, it would be nice if astropy supported common unicode symbols.
E.g.
Probably there is more...