Repository navigation
RFC: Built-in module extending and removing weak links / umodules #9018
Description
Activity
- addedenhancementFeature requests, new feature implementationsFeature requests, new feature implementations
on Aug 5, 2022 Some other misc thoughts that don't quite fit into the above:
- It would be good to move towards only providing a subset of CPython functionality in our built-in modules, and to that end we should clarify exactly what the rules are. Entirely new methods/classes -- not ok. Different functionality -- not ok. Adding extra kwargs to existing methods -- maybe ok?
- Should we find new homes for existing non-CPython extensions (e.g. time.ticks* --> ticks.ticks*?, os.mount --> storage.mount?, asyncio.sleep_ms --> uhh... uasyncio.sleep_ms?) and how long should we keep aliases to the existing ones.
- Is it important to be able to replace as well as extend functions in existing modules. i.e. builtin foo.bar might not support some arg or kwarg, so we can replace it at runtime with one that does.
MicroPython is not CPython. I do not like the idea of changing names for long established classes, methods, functions. That would by a big pain for people actually using micropython for their projects and might want to update the firmware e.g. for bug fixes.
Reacted by David Lechner, Amirreza Hamzavi and Jos VerlindeMicroPython is not CPython. I do not like the idea of changing names for long established classes, methods, functions. That would by a big pain for people actually using micropython for their projects and might want to update the firmware e.g. for bug fixes.
@robert-hh Just to clarify, I'm not proposing changing the names of any classes, methods, or functions. Only modules.
In terms of backwards compatibility it would be very easy to continue to provide the existing weak links / aliases for ufoo->foo (except in the reverse direction to the current implementation, such that the module would now be called
foo). Additionally, of course if we moved e.g.os.mounttostorage.mountthen we would continue to alias it inos.The main breaking change I'm proposing would be that builtins always take precedence and are never searched on the filesystem.
Reacted by Andrew LeechThanks for the really clear, well thought-through write-up @jimmo!
+1 to remove the ufoo/weak link model. I've found it to be one of the most confusing topics for beginner/intermediate MicroPython users - and one of the more difficult to explain.
Adding extra kwargs to existing methods -- maybe ok?
I'd prefer not - but I can imagine some compelling cases that could be convincing.
Should we find new homes for existing non-CPython extensions...how long should we keep aliases...
Yes to finding new homes. Not long would be my preference - make painful changes quickly. More important to me: Making breaking changes like this - and their workarounds - very clear in release notes.
and the other historical reasons for the "u" prefix aren't compelling either.
A very compelling use case for me is that the "u" prefix allows us to use existing Python intellesene tools in IDEs to get correct code completion and type hints for the MicroPython version of a module. Existing tools don't have a way of saying "I want to replace the Python standard library with my own
.pyifiles". So being able toimport ufoosidesteps this and allows one to very easily benefit from coding tools without any extra work (other than creating theufoo.pyifiles for the MicroPython API).In terms of backwards compatibility it would be very easy to continue to provide the existing weak links / aliases for ufoo->foo
As long as this is possible, then the proposed changes aren't really breaking then are they (other than the aliases may be disabled by default instead of enabled)?
- It's worked so far for CircuitPython, and it's entirely possible that many MicroPython users have never needed/wanted it either (I know this isn't quite true).
Here is my perspective and what I use to guide CircuitPython.
- Having the builtin modules without the u prefix is almost always what people use because they have some background from CPython. If they don't, then teaching the non-u version is still more portable knowledge.
- CircuitPython code should run in CPython without changes. CPython code running in CircuitPython is not a goal because CPython written code tends to treat RAM as effectively infinite.
So, for 1, we've dropped the utime prefixes and for 2 we've dropped or moved all non-CPython APIs to other modules like
storageormultiterminal. The CPython-compatible versions (builtins and libraries) shouldn't have anything added that will cause an error in CPython (extra kwargs would for some functions I believe.). We want imports to represent what people are using and cause errors early on start up rather than some unknown time later at first use. Moving extra functionality to other modules also makes it easy to provide that functionality in CPython with a library (see Blinka.)Regarding allowing extending builtins, do you really need it? Why can't the calling code be changed to use a different module name or simply
import foo_ext as foo?foo_extcan stillimport footo use the builtin version.I also have run into u* libraries that may have intended to match their CPython equivalent but then varied. I think the u* naming in libraries is a bit of a scapegoat to be similar instead of a strict subset. I find that variation confusing because it doesn't match one's expectations. Leaving the u* naming behind will lead to more different names for different modules and I think that'd be good.
I'm 100% in favour of this. Extending builtins is a surprisingly common practice in even cpython (in my experience), and even more so in micropython. As @jimmo mentioned this aids in keeping the built-in modules lean while allowing users to extend as needed for their application.
The
uprefix is confusing for a lot of new users I encounter though.It's true that the both the current mechanism and the proposed change comes with a ram cost in the duplication of the module table with
from ufoo import *, though this could be replaced with amod.getattrfunction pattern similar to uasyncio to import on demand.Similarly, the new system could be implemented in C with the built-in getattr/setattr functions; make
mod.setattrcreate a__dict__if needed and store the added features there, with getattr checking that. If__dict__exists then it should get checked before the rom table else you wouldn't be able to override built-in functions.Why can't the calling code be changed to use a different module name or simply import foo_ext as foo?
I've run into a number of cases where porting cpython code to micropython where the
ufoobuilt-in is missing some needed features. Currently I can create afoopython module that provides these missing parts and not need to modify the module being ported. So I'm strongly against needing a_extstyle renamed module to provide cpython compat. Often using this approach I've been able to use third party cpython code without any changes which is far better than needing to fork/patch/maintain just to supportimportchanges.Regarding allowing extending builtins, do you really need it?
Definitely yes, main reason being that most builtins are not 100% CPython-comaptible: the choice is then to either write a new module and provide functions, or to just do what is most convenient and user-friendly in all possible ways (code is CPython compatible, can just use standard CPython docs, no need to learn about another module and/or function name, ...)
to provide different extensions to the os module, and the user can choose which ones to install, and then "activate" them by importing them once at the top of main.py (or even boot.py)
For the unix-lik ports this would require some new mechanism (preferrably CPython-compatible, site.py perhaps) to have something run before everything else. At least I don't think we have that now. Possible (perf not measured) disadvantage is that in practice for most projects I have this would mean importing like 10 modules, always, even when none needed.
Option 1 is a no-go as far as I'm concerned, 3 only fixes the
uprefix itself so not sure if it's worth it but for the rest it's interesting because it doesn't change a lot.2 and 4 on the other hand are technically nice but have the issue mentioned above, and are not super convenient to write in general, and it also hurts discoverability somewhat (e.g. now if you want to know what's in
os, you open os.py and/or moduos.c, not os_ext.py or so). Not sure which of 2, 4 and Andrew's approach with getattr would be 'best'. Would be interesting to see actual code as was shown for 4, just to be able to compare.Thanks for the comments!
So being able to import ufoo sidesteps this and allows one to very easily benefit from coding tools without any extra work (other than creating the ufoo.pyi files for the MicroPython API).
@dlech OK interesting. This is something I hadn't considered, nor do I have any experience with.
My initial reaction is along the lines "surely there's got to be a better way" and to be honest we're already trying to make "use foo not ufoo" the default guidance, and so complicating that with "but ufoo fixes autocompletion" is unfortunate (and also means that you have to choose between supporting extensions-from-Python or autocompletion).
How bad is it that the autocomplete just uses Python's full completion? Is this a limitation of all IDE that they can't have project-aware auto-completion sources?
As long as this is possible, then the proposed changes aren't really breaking then are they (other than the aliases may be disabled by default instead of enabled)?
Yes, the only breaking change would be that we would stop the existing extension mechanism from working (i.e. builtins would now always take precedence).
Yes. Others will correct me if I'm wrong but I think our philosophy is much more: "MicroPython code can run in CPython with necessary shims, e.g. micropython.const, etc, and CPython fragments and occasionally whole modules should run unmodified.".
(@tannewt) Regarding allowing extending builtins, do you really need it? Why can't the calling code be changed to use a different module name or simply import foo_ext as foo?
(@andrewleech) I've run into a number of cases where porting cpython code to micropython where the ufoo built-in is missing some needed features.
(@tannewt) CircuitPython code should run in CPython without changes. CPython code running in CircuitPython is not a goal because CPython written code tends to treat RAM as effectively infinite.
@tannewt @andrewleech Yes, I guess this is exactly the crux of this conversation. This whole goal of extending built-in modules is only really necessary to make an existing file work completely unmodified.
In the "import foo_ext as foo" case, foo_ext still needs to do "from foo import *" (duplicate dict cost) or the module
__getattr__trick (extra code cost). It also doesn't work for more than one extension.It's true that the both the current mechanism and the proposed change comes with a ram cost in the duplication of the module table with from ufoo import *, though this could be replaced with amod.getattr function pattern similar to uasyncio to import on demand.
This is a good idea, thanks @andrewleech . I think the whole "getattr + builtinimport hook" could be combined into a utility module too, so os_ext.py could look like:
import builtin_extension # This gives us a __getattr__ forwarding to the existing os # module (either built-in or a previous extesnion), and hooks # "import os" for further usage to return this module instead. builtin_extension.apply("os", __name__) def foo(): pass
Here is an implementation of builtin_extension.py to support this -- https://gist.github.com/jimmo/b918b73abc2f5c107b102d2ddb7a7976.
Similarly, the new system could be implemented in C with the built-in getattr/setattr functions; make mod.setattr create a dict if needed and store the added features there, with getattr checking that. If dict exists then it should get checked before the rom table else you wouldn't be able to override built-in functions.
Yes, exactly. This is along the lines of how the existing
sysandbuiltinoverrides work. Some consideration to how to do this with the absolute minimum RAM overhead (i.e. you need a pointer for that dict, and you don't want to do that for ~30-40 modules). There are ways to solve this though.Definitely yes, main reason being that most builtins are not 100% CPython-comaptible: the choice is then to either write a new module and provide functions, or to just do what is most convenient and user-friendly in all possible ways (code is CPython compatible, can just use standard CPython docs, no need to learn about another module and/or function name, ...)
@stinos Does Scott's point above with "import foo_ext as foo" work for you? Or, like Andrew, do you want the CPython-compatible file to run exactly as-is without any modifications to import.
For the unix-lik ports this would require some new mechanism (preferrably CPython-compatible, site.py perhaps) to have something run before everything else. At least I don't think we have that now.
@stinos I'm not quite sure I follow.
I had imagined that the top of main.py (or for Unix/Windows, whatever your entry point is) would do this.
Possible (perf not measured) disadvantage is that in practice for most projects I have this would mean importing like 10 modules, always, even when none needed.
You could also imagine that some sort of automatically maintained script that imported all your installed extension packages could be useful. But conceptually I see this more as "I'm explicitly enabling the extra functionality I need for my app".
it also hurts discoverability somewhat (e.g. now if you want to know what's in os, you open os.py and/or moduos.c, not os_ext.py or so).
Do any of the approaches solve this? If we're going to make it possible to extend a module, it needs to be implemented across multiple files.
Not sure which of 2, 4 and Andrew's approach with getattr would be 'best'. Would be interesting to see actual code as was shown for 4, just to be able to compare.
Which one do you want to see code for? The implementation of mutable builtin-modules for 2?
Or, like Andrew, do you want the CPython-compatible file to run exactly as-is without any modifications to import.
Yes I'd really prefer this to stay as it is now, such that we can write code which is CPython compatible with the least friction, i.e. just
import osoros.pathetc (just to name the two probably used most examples of files which we have extensions to builtin functionality for).edit just to give an idea, there are 43 matches for
import osin our main codebase; so in principle that's just a find/replace to change that intoimport os_ext as osbut that's just a step back, and a lot of frictionI had imagined that the top of main.py (or for Unix/Windows, whatever your entry point is) would do this.
When prototyping, writing (unit)tests etc I just want to create a new file and start writing 'normal' code, then run it (1). I wouldn't want to have to manually add
import foo_extto every single file which needs builtinfoo+ extension. Meaning instead I'd want whatever command (1) is to do that automatically. Which eventually boils down to at interpreter startup importing a custom module automatically which in turn has all theimport xxx_extstatements. There are a couple of ways to do that, but perhaps makes sense to provide something for the unix port in the mainline.Do any of the approaches solve this? If we're going to make it possible to extend a module, it needs to be implemented across multiple files.
No, but it's a minor issue. Also because as long as the extension is written in Python most text editor's 'go to symbol' etc will find it anyway.
Which one do you want to see code for? The implementation of mutable builtin-modules for 2?
Yes, to see how it compares in practice with the approach for 4
How bad is it that the autocomplete just uses Python's full completion?
You end up using modules/classes/methods/args that don't exist in MicroPython and your code fails at run time. Then you have to go find the relevant MicroPython documentation to figure out what is going on. This slows down development and is confusing for inexperienced programmers.
Is this a limitation of all IDE that they can't have project-aware auto-completion sources?
Generally, the way these tools work is that you say I want to use X Python runtime. This is specified by the absolute path to python executable and will use that find all site packages, etc this way. Usually this a virtual environment where you have all of the dependencies for your project installed. But since MicroPython is not fully CPython compatible, you can't just plug in the path to MicroPython here or create a MicroPython virtual environment.
I had imagined that the top of main.py (or for Unix/Windows, whatever your entry point is) would do this.
When prototyping, writing (unit)tests etc I just want to create a new file and start writing 'normal' code, then run it (1). I wouldn't want to have to manually add import foo_ext to every single file which needs builtin foo + extension.
I concur, currently if you have all the libraries "upip installed" or equivalent then just running files have all the libs on the path ready to import. The
_extscheme would not work this way so needs a pre-process step to get these on the path.
Cpython has the concept of.pthfiles that can be in a virtualenv that get run at the start of every python startup that allow python code to setup the env. On boards, the boot.py can also to this. On Unix etc I think it works be good to find/run a boot.py (or similar) automatically at startup if it exists on path.You end up using modules/classes/methods/args that don't exist in MicroPython and your code fails at run time. Then you have to go find the relevant MicroPython documentation to figure out what is going on. This slows down development and is confusing for inexperienced programmers.
Is there a particular ide / autocomplete package you're using currently? Autocomplete and code navigation is something I use daily myself but have just gotten used to the sub-par behaviour of cpython stubs. This certainly isn't ideal though.
A proper micropython autocomplete setup is something I'm keen to get working. I know there are some stubbing tools it there that should do most of the work for us, perhaps arranging the outputs of that into something an ide could recognise as a python env would be possible? Maybe that's something that could be built on top of the new manifest system to include copies/links to libraries installed as well.
Is there a particular ide / autocomplete package you're using currently?
We're using Pylance(based on Pyright) in VSCode and also Jedi.
A proper micropython autocomplete setup is something I'm keen to get working.
Microbit forked Pyright to make it work in the browser and with their MicroPython API. You can see it in action at https://python.microbit.org/v/beta. Making this work in general for MicroPython with some sort of manifest system as you have describe would be ideal IMHO.
2 remaining items
Thanks @laurensvalk
... PEP399...
I'm not arguing to do this exactly, but it seems like there is some precedent that could provide inspiration here.Yes, this is exactly how
_uasyncioprovides an "accelerated" Task etc. Although our implementation is a bit different because we just don't includetask.pyat all when_uasynciois available.It's also relevant to #8968 where one of the options being considered is to freeze ssl.py and provide
_sslas the low-level implementation. (But in this case it's conceptually a bit different because there is no pure-Python version).It's a bit difficult to apply this as a precedent to this issue though because you still can't override
osin CPython. And the way accelerator modules are used are definitely not at all "micro" in its implementation (i.e. the whole Python implementation gets loaded only to be replaced in its entirety with the accelerator module).So I guess my point is that having u-modules in the implementation makes sense and has some CPython precedent.
As mentioned above, end-users should be able to import foo and the documentation just needs to have foo. It is certainly an (ongoing) effort to make it work like users expect, but I don't see this is an argument against having the u-modules in the implementation.My goal here is that I don't think our umodule/weak link system is actually the best way to provide the functionality that users want/need.
- Many (most?) users don't need to extend builtins. So we should make that fast and efficient. (e.g. avoid the ~2-5ms import cost of searching the filesystem for a module that isn't there for each built-in module).
- Some users need a small number of extra functions that aren't provided by our built-ins. As @tannewt said, in many cases this doesn't even need to be made available as a built-in extension, i.e. they need the method, how it gets imported isn't such a big deal.
- Some users want to use a CPython library verbatim. In that case having some mechanism to extend a built-in with missing functionality is required to support this use case.
- We want to provide the "missing functions" in micropython-lib (and third party packages too), but at the same time it's wasteful to write large packages that provide 10 extra methods when the user just wanted one thing, rather it would be more efficient to write fine-grained extension packages. We also might want to combine extensions from multiple sources. So we need a way to compose extensions (this isn't currently possible with ufoo/weak links).
- Some "built-in" things that require extending aren't built-ins (e.g. frozen modules like uasyncio). It would be really good to have a single standard way to extend any library with whatever is needed for a given app.
- Using
from ufoo import *is wasteful as an implementation of a built-in. (We could recommend the__getattr__pattern as an alternative though)
but library builders can still use them.
I'd be interested to know more about your specific use case with pybricks. Do you provide pre-built firmware with frozen in extensions? Or do your users add them to the filesystem? Are they writing their own or using micropython-lib ones (or other sources?). Would any of the options outlined in this issue work for you & your users?
Thanks for your response. That led me to a subtlety that may be worth clarifying --- Reading back through this thread, it seems that not all comments about extending builtins are about the same thing.
If I understand the posts in favor of extending builtins correctly, most would like to extend uos, not necessarily extend os.
But both variants are implicitly or explicitly discussed in this thread, and I'm not sure the replies here always respond to the same thing.
I'll follow up with my personal opinion in the next post.
So with that in mind, I fully agree with the motivation in @tannewt's post, but with a subtle clarification of the implication --- moving away from the
uosname could be fine, so long as it doesn't becomeos. I believe it would be better to reserveosfor something that actually behaves likeos, however it gets implemented.I don't think our umodule/weak link system is actually the best way to provide the functionality that users want/need.
Agreed. My comments so far are mainly about keeping the internal implementation modules (e.g.
uos) available under a distinct name, so thatosremains reserved something closer to the real deal.Some notes on the performance / memory
I understand the concerns on RAM and build size, but how important is import time?
I'd be interested to know more about your specific use case with pybricks. Do you provide pre-built firmware with frozen in extensions? Or do your users add them to the filesystem? Are they writing their own or using micropython-lib ones (or other sources?). Would any of the options outlined in this issue work for you & your users?
Most of this doesn't affect Pybricks directly, at least not right now. I wrote this from a generic point of view, and with the backdrop of having written several coding books for kids.
Several posts have mentioned that some solutions are easy to explain. I believe it's even better when it's easy to understand. So if two things work differently, even subtely, then in my view they shouldn't have the same name.
I understand that this doensn't resolve the question at hand, but I'd have to read up a bit more on existing efforts before commenting on any of the proposed solutions.
Several posts have mentioned that some solutions are easy to explain. I believe it's even better when it's easy to understand.
I think the best/easiest case is when there is nothing there at all, when there's nothing to explain. And removing u-naming is a way to simplify things such that there's no longer any difference to CPython and hence nothing to document/teach.
Of course that's not the whole story because we still need a way to override built-in modules, so there will be something extra and something to document/teach. But for most use cases / most users / 80% of the time / to first order, things should just behave like CPython without having to worry about differences.
import osshould just work. On top of that, for the other 20%, to extendos, you need to learn something.At the moment it feels like it's the other way around, you need to first learn about u-naming before you can do anything.
Reacted by Laurens ValkSorry if I've missed this in the discussion, but did you consider the case where a frozen module adds bits/extends a built-in module ? For example if
foomodule in micropython-lib withfrom ufoo import *, is frozen via manifest.py, this as far as I know currently works, but how would it work after this change ? Can options 3/4 detect frozen vs built-in and give precedence to frozen modules ?- Our native apis are all documented as python stubs now. Folks have been adding type annotations now. See shared-bindings in the repo. (On my phone now)…On Sat, Aug 6, 2022, at 4:47 PM, David Lechner wrote: > Is there a particular ide / autocomplete package you're using currently? > We're using Pylance(based on Pyright) in VSCode and also Jedi. > A proper micropython autocomplete setup is something I'm keen to get working. > Microbit forked Pyright to make it work in the browser and with their MicroPython API. You can see it in action at https://python.microbit.org/v/beta. Making this work in general for MicroPython with some sort of manifest system as you have describe would be ideal IMHO. — Reply to this email directly, view it on GitHub <#9018 (comment)>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/AAAM3KMSC3IYPZRBNENJPJTVX32QRANCNFSM55V4ZESA>. You are receiving this because you were mentioned.Message ID: ***@***.***>
I maintain similar stubs for MicroPython. Most of them are generated automatically from the repo, and try to consider variations across, ports and boards.
Pyright has been adopted to be able to use these as there were some issues with overriding some of the stdlib modules.
Once settled on a model it should be straightforward to adjust the stubs accordingly and publish them, install the stubs in a folder or venv , and the tools should pick them up from there.See #11456 which opens another option for extending built-ins from Python.
#11456 brought me here: I had no idea what "u"modules were for and have been using them interchangeably and passing on that code smell to whoever uses/learns-from our examples. (Granted, my existence in mostly the make-everything-a-C-module space of MicroPython does not lend itself well to finding and understanding these quirks, which is why I'm lurking the GitHub and trying to expand my knowledge.)
It's pretty easy to see the extent of this across our MicroPython examples: https://github.com/search?q=repo%3Apimoroni%2Fpimoroni-pico+%22import+u%22&type=code
Granted not all of these are a problem. But it does make
uasyncioandurequestsandumqttrather confusing. Are those a cosmetic prefix or an intentional disambiguation from similarly-named CPython modules?In our case, I guess I should raise an issue/PR to switch from
uostoos,ujsontojsonand so on.#11456 brought me here: I had no idea what "u"modules were for and have been using them interchangeably and passing on that code smell to whoever uses/learns-from our examples.
It's pretty easy to see the extent of this across our MicroPython examples: https://github.com/search?q=repo%3Apimoroni%2Fpimoroni-pico+%22import+u%22&type=code
I think Damien's comment #9018 (comment) is about exactly this point -- we should just make it simple and obvious.
Granted not all of these are a problem. But it does make
uasyncioandurequestsandumqttrather confusing. Are those a cosmetic prefix or an intentional disambiguation from similarly-named CPython modules?There's an additional piece of historical context (and a third function of the u-prefix) there which was that it disambiguated the package name on PyPI. This is no longer an issue.
There is an open issue to rename
requests. See micropython/micropython-lib#540
The main decision there is the best way to provide backwards compatibility, but I think we have decent options and we should do that at the same time as the built-in module rename discussed here.We should do the same thing for
asyncio. (This one is super confusing because it's a built-in in behavior, i.e. it's frozen, but it's not actually a built-in so doesn't get the "weak links" treatment).In our case, I guess I should raise an issue/PR to switch from
uostoos,ujsontojsonand so on.Yep, see the official guidance here:
https://docs.micropython.org/en/latest/library/index.html#extending-built-in-libraries-from-pythonSpecifically: "Other than when you specifically want to force the use of the built-in module, we recommend always using
import modulerather thanimport umodule."See #9069 for an implementation of this.
In the end I went with a fifth option which doesn't change any behavior but still allows removes the u-prefix from the module objects.
- For now
import umodulecontinues to work, but the preferred mechanism to force a built-in import is to temporarily clearsys.pathas described and implemented in py/builtinimport: Allow built-in modules to be packages (v3) #11456. There is scope to reduce the number of allocations (see py/builtinimport: Allow built-in modules to be packages (v3) #11456 (comment)).
This means that there should be no user-visible change other the output of
help("modules")and what you see if you print a module object (i.e.<module 'os'>rather than<module 'uos'>).Later we can consider removing the handling for
umodule(or at least making it anmpconfigoption).- For now
- added a commit that references this issue
on Oct 22, 2023
We use the "u" prefix for built-in modules for several reasons (see #7499 (comment) for the full story). The primary reason though is to allow
foo.pyto exist on the filesystem and extend the built-inufoo. This feature is called "weak links" because the namefooused to be a "weak link" toufoo(i.e. the link is "broken" by having the filefoo.pyon the filesystem). Today the feature, when enabled, is automatic for any built-in namedufoo.There's a few drawbacks:
import foois slower than necessary because it must search the filesystem to (usually) not findfoo.pyinsys.pathonly to eventually findufooin the builtins table.builtin.__import__)import foo; print(foo)you getufoo. Alsohelp('modules')lists everything asufoo.import asyncio, and also it should be possible to extend asyncio with optional features.Our goal here is usually just to provide "optional" implementations of functionality that aren't general-purpose enough to include in standard firmware, or possibly are just better suited to implementation in Python. (i.e. they could still be good candidates for being frozen?). But in general we're "filling in" missing functionality from the CPython equivalent module, and therefore there's no precedent for this in CPython. Note that CircuitPython does not currently enable the "weak link" feature and doesn't provide a way to extend built-in modules from Python.
I would argue that based on the reasons above, the "ufoo" mechanism doesn't support this use case particularly well, and the other historical reasons for the "u" prefix aren't compelling either. So let's say that in a future release (2.0?) that we remove all traces of the "u" prefix for built-in modules (i.e. rename all built-ins, remove the weak links feature, remove the last remaining traces from the documentation). This is obviously a hugely breaking change, would need to be done gradually.
So that means we need to solve what to do about extending built-in modules. There's four options that I can think of:
Don't allow this. It's worked so far for CircuitPython, and it's entirely possible that many MicroPython users have never needed/wanted it either (I know this isn't quite true).
Allow a way to "runtime patch" built-in modules. So instead of replacing "foo", you could import a module (e.g.
import foo_ext) that would install/replace its own additional methods/classes into some built-in modulefoo. This is currently impossible because the modules' globals tables are in ROM. We have a partial mechanism for this that is selectively enabled for thesysmodule (to allow e.g.sys.ps1), and also thebuiltinsmodule is special cased to allow this for all elements.It seems feasible to implement this for all built-in modules without too much RAM or runtime performance cost. It composes well -- for example we can provide a bunch of different
os_ext_foo,os_ext_bar,os_ext_bazto provide different extensions to theosmodule, and the user can choose which ones to install, and then "activate" them by importing them once at the top of main.py (or even boot.py). It's perhaps a bit awkward that you need this activation step though, but from that point onimport oswill have all the extensions ready.It also addresses the asyncio case. We can rename the frozen uasyncio, and then provide optional extensions which install themselves at runtime (except for asyncio it's simple because it's frozen not built-in so the modules are already mutable).
In more detail, the idea would be to have lazy-constructed
MP_STATE_VM(module_overrides), (either a single layer ofmodule+qstr -> objor two-layermodule->(qstr->obj)) that is used preferentially for the module's globals dict. As a performance optimisation (especially for tab completion which does a lot of module lookups) we need a single bit somewhere to indicate that a module has had at least one override.from ufoo import *approach. One simple example would befrom builtins.foo import *orfrom micropython.foo import *(this is a fairly straightforward mod toobjmodule.candbuiltinimport.cto detect the "builtins." prefix and resolve the module viamp_module_get_builtin). Another alternative is to make builtins available inmicropython.builtinsor something (this would also provide a mechanism to programatically query built-ins which is occasionally requested). Another way could be to empty sys.path before importing (although this doesn't currently work, an empty sys.path is equivalent to['']).This doesn't solve the composability, or the fact that we still need to search the filesystem, or how to extend asyncio.
It also has a RAM cost for extending a builtin-module because the "from builtin.foo import *" needs to copy the entire globals dict of the built-in into the Python module that's replacing it.
builtin.__import__. As mentioned above, it is actually technically possible in CPython to hook import, even for builtins, and then return an extended module. For example, in "os_ext_foo.py".then
This composes well, avoids the filesystem search for builtin import, works for asyncio,
but has a fairly high RAM cost for duplicating the globals table (potentially once for each extension)(edit: solution below), as well as the cost for the hook code.CC a few people who have been involved in this topic in the past: @mattytrentini @andrewleech @tannewt @jepler @dhalbert @dlech @stinos