Skip to content

webassembly: standard variant fails 5 of its own port tests, any gc.collect() aborts #19670

Description

@andrewleech

Summary

I'm building LVGL into the webassembly port as a user C module, so micropython can render an LVGL framebuffer straight into a canvas in a web UI, for live interactive cross-platform GUI demos. While working on this I've run into a problem with garbage collection: with a widget tree on the heap the runtime hangs at ~100% CPU as soon as anything triggers a collection.

Tracking that down, it isn't LVGL specific and it isn't mine, gc_collect() on the standard variant suspends the wasm stack in a ccall that isn't expecting it. The port's own test suite already shows it, 5 of the 43 tests under tests/ports/webassembly/ fail on VARIANT=standard today and all five call gc.collect(). CI only builds VARIANT=pyscript, which takes a different GC path, so nothing catches it.

Port, board and/or hardware

ports/webassembly, VARIANT=standard (emscripten 3.1.73, node 20.18.0, linux)

MicroPython version

master at 791ba6e, unmodified.

Reproduction

make -C ports/webassembly VARIANT=standard
cd tests
MICROPY_MICROPYTHON_MJS=../ports/webassembly/build-standard/micropython.mjs \
    ./run-tests.py -t webassembly -d ports/webassembly

One test on its own is enough if you'd rather skip the suite:

node tests/ports/webassembly/heap_expand.mjs \
    /abs/path/to/ports/webassembly/build-standard/micropython.mjs

Expected behaviour

The port's own JS tests pass on the standard variant, same as they do on pyscript.

Observed behaviour

38 of 43 pass, 5 fail:

43 tests performed (460 individual testcases)
38 tests passed
5 tests failed: ports/webassembly/gc_behaviour.mjs ports/webassembly/heap_expand.mjs
ports/webassembly/js_proxy_reuse_free.mjs ports/webassembly/weakref_finalize_collect.mjs
ports/webassembly/weakref_ref_collect.mjs

heap_expand.mjs on its own:

Aborted(Assertion failed: The call to mp_js_do_exec is running asynchronously. If this was intended, add the async option to the ccall/cwrap call.)

All five call gc.collect(), the two weakref ones are even titled "requiring gc.collect()", so it looks like one fault rather than five.

Additional Information

As far as I can trace it: gc_collect() in ports/webassembly/main.c (the non-MICROPY_GC_SPLIT_HEAP_AUTO branch, ~line 216) calls emscripten_scan_stack() and emscripten_scan_registers(), and those suspend the wasm stack under Asyncify. Module.ccall() in api.js is called without { async: true }, so any collection taken while Python code is running is illegal and you get the assert above. The standard variant sets -s ASYNCIFY and leaves MICROPY_GC_SPLIT_HEAP_AUTO at 0, so it takes that branch.

The hang I started with is the same condition, it just doesn't assert when there are host objects on the heap, it spins instead. That's what made it slow to pin down, with no message and a live display in the picture it reads as an LVGL problem rather than glue. Heap size makes no difference either, I tried 8, 16, 32, 64 and 256 MB, any collection at all is enough and 16 KB allocations force one almost straight away.

CI doesn't see any of this because tools/ci.sh (ci_webassembly_build / ci_webassembly_run_tests) builds and tests VARIANT=pyscript only. pyscript sets MICROPY_GC_SPLIT_HEAP_AUTO, where gc_collect() just raises a flag and the real collection happens later from gc_collect_top_level() with no stack or register roots to scan, so nothing suspends and the fault can't occur there:

variant suite result
standard, as-is 38/43
pyscript, as-is 43/43
standard + GC_SPLIT_HEAP_AUTO + ALLOW_MEMORY_GROWTH 43/43

Two fixes I've tried, neither of which I'm confident is the one you'd want:

  1. Pass { async: true } to the mp_js_do_exec_async ccall in api.js (wrapped so the non-Asyncify builds still work, since ccall only returns a promise when Asyncify is on). That makes collection legal during runPythonAsync() and I've run a render loop with GC enabled through it happily. However it does nothing for runPython(), pyimport() or the node CLI, which are synchronous and so can't suspend by construction, and heap_expand.mjs goes through runPython(), so this doesn't fix the failing tests.

  2. Give the standard variant MICROPY_GC_SPLIT_HEAP_AUTO (plus ALLOW_MEMORY_GROWTH, since deferred collection grows the heap rather than reclaiming in place). That takes the suite to 43/43 and fixes every entry point, not just the async one. That being said it changes the default variant's GC strategy, which is your call rather than mine, and I did hit one thing I can't explain: a light-churn tick_inc/timer_handler loop that runs fine on the current config hangs under it, at 20 iterations as readily as at 1000. I haven't got to the bottom of that.

Happy to turn either into a PR if you've got a preference. Adding a standard-variant build/test to tools/ci.sh would want to go with whichever fix lands, otherwise the job just goes red on the five above.

I used generative AI tools while investigating and writing this up, the measurements and traces are mine and reproducible with the commands above.

Activity

  1. dpgeorge commented on Aug 31, 2026

    @dpgeorge
    Member

    Yep, the standard variant is lacking.

    Please see previous bug report and discussion at #19380, and proposals for improvements in #19427 and #19594 .

    Maybe you can test #19594 to see if it fits your use case?

  2. andrewleech commented on Sep 2, 2026

    @andrewleech
    SponsorContributorAuthor

    Thanks for the background info. Yes I've tested #19594 and it fixes my issues.

    Setup: the PR at fba8faa, emscripten 6.0.9, node 24.5.0 with
    --experimental-wasm-jspi, linux.

    VARIANT=jspi:

    • tests/ports/webassembly: 47 passed, 0 failed, 1 skipped
      (jspi_not_supported.mjs)
    • whole suite: 856 passed, 0 failed

    All five of the tests I listed above as failing on standard pass.

    lv_binding_micropython builds against VARIANT=jspi as a USER_C_MODULE
    unchanged. Sizes with the module linked in, same toolchain and source:

    variant .wasm
    standard (ASYNCIFY) 4,242,169 B
    pyscript 1,502,474 B
    jspi 1,492,719 B

    so JSPI costs nothing over pyscript, it comes out slightly smaller. Same
    direction @ntoll measured without a user module in the mix.

    Behaviour wise the two are indistinguishable across every LVGL script I ran. I
    do still have a busy spin in some display setup paths, however it reproduces
    identically on pyscript and jspi, so that one's mine to chase.

    run_sync() covers the one thing pyscript couldn't do for me. The harness needs
    to await fetch/WebSocket from python eventually, which was the only reason
    ASYNCIFY looked necessary, so dropping it costs me nothing.

    Happy for this issue to be closed as a duplicate of #19380.

    I used generative AI tools while running these tests, the numbers are all from
    real runs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions