Skip to content

extmod/modlwip: Fix TCP send and POLLOUT when out of memory. - #19705

Open
srgg wants to merge 3 commits into
micropython:masterfrom
srgg:modlwip-nonblocking-partial
Open

srgg wants to merge 3 commits into
micropython:masterfrom
srgg:modlwip-nonblocking-partial

Conversation

@srgg

@srgg srgg commented Sep 16, 2026 •

Copy link
Copy Markdown

Reworked after review: the PR now returns EAGAIN and carries three commits; the current description is this comment. The text below describes the first version.

Resolves #19704, resolves #19746.

Summary

A non-blocking socket.write() blocks for up to 10 s on ERR_MEM (analysis in #19704).

This change makes that path write the prefix that fits — halving write_len and retrying — and return ENOBUFS only when nothing fits, instead of blocking. EAGAIN is not used: POLLOUT reads tcp_sndbuf, so it would busy-spin a select caller; ENOBUFS is a resource error, not would-block. Blocking sockets are unchanged.

Testing

Measured on an OpenMV RT1062 (cyw43 Wi-Fi) streaming MJPEG to a reader throttled to 20 KB/s, on a v1.28.0-based tree: worst write() 1553086 us, 54 events over 100 ms per 320 s; the early return removes both.

Not built against master and run on no port here — the added block copies the adjacent tcp_sndbuf == 0 return, and CI compiles the MICROPY_PY_LWIP ports.

Trade-offs and Alternatives

Halving costs up to log2(tcp_sndbuf) tcp_write attempts on the exhausted pass; returning ENOBUFS with zero bytes is O(1) but makes no progress.

Generative AI

I used generative AI tools when creating this PR, but a human has checked the code and is responsible for the code and the description above.

@srgg srgg changed the title fix(modlwip): return a partial or ENOBUFS from a non-blocking send extmod/modlwip: Return a partial or ENOBUFS from a non-blocking send Sep 16, 2026
@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.59%. Comparing base (a129b2f) to head (84847f9).

Additional details and impacted files
@@            Coverage Diff             @@
##           master   #19705      +/-   ##
==========================================
+ Coverage   98.55%   98.59%   +0.03%     
==========================================
  Files         182      182              
  Lines       23346    23346              
  Branches        5        5              
==========================================
+ Hits        23009    23017       +8     
+ Misses        336      328       -8     
  Partials        1        1              
Flag Coverage Δ
unix-coverage-32bit 98.59% <ø> (+0.03%) ⬆️
unix-coverage-64bit 98.52% <ø> (ø)

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Code size report:

Reference:  tools/mpremote: Add tell() support for remote mounted files. [a129b2f]
Comparison: extmod/modlwip: Gate POLLOUT after ERR_MEM on a segment's allocation. [merge of 84847f9]
  mpy-cross:    +0 +0.000% 
   bare-arm:    +0 +0.000% 
minimal x86:    +0 +0.000% 
   unix x64:    +0 +0.000% standard
      stm32:    +0 +0.000% PYBV10
      esp32:    +0 +0.000% ESP32_GENERIC
     mimxrt:    +0 +0.000% TEENSY40
        rp2:  +120 +0.012% RPI_PICO_W
       samd:    +0 +0.000% ADAFRUIT_ITSYBITSY_M4_EXPRESS
  qemu rv32:    +0 +0.000% VIRT_RV32

Comment thread extmod/modlwip.c Outdated
@dpgeorge dpgeorge added the extmod Relates to extmod/ directory in source label Sep 19, 2026
Comment thread extmod/modlwip.c Outdated
@srgg

srgg commented Oct 2, 2026 •

Copy link
Copy Markdown
Author

@dpgeorge, thank you for the review. Both comments are taken:

The halving retry is gone too: with Nagle on, tcp_pbuf_prealloc allocates a full-MSS pbuf regardless of length, so it only reduced the segment count.

Prior art, read from source: lwIP's own netconn returns ERR_WOULDBLOCK from lwip_netconn_do_writemore on ERR_MEM, and Linux returns EAGAIN from sk_stream_wait_memory; both then gate socket writability with a per-socket rule (NETCONN_FLAG_CHECK_WRITESPACE low-water marks, sk_stream_moderate_sndbuf), which cannot see memory held by another socket.

PR reworked, three commits:

  1. A non-blocking send blocked on ERR_MEM: the loop now returns EAGAIN for a non-blocking socket; nothing is queued, and the caller retries. Previously, write() blocked up to 10 s and raised ENOMEM.

  2. The ERR_MEM loop ignored the socket timeout (modlwip: socket timeout ignored in send when out of memory #19746): a socket with settimeout(t) now gets ETIMEDOUT at t; the 10 s limit applies only to a socket with no timeout. Previously, the loop ended only at 10 s with ENOMEM. The bound is per send(): write() is write-all and loops over chunks with a timeout each, as before.

  3. POLLOUT reported writable after that EAGAIN: it now requires one segment's memory: a pbuf of mss bytes, one tcp_seg tried and released, and a free slot in snd_queuelen. The send retries once at a single MSS before returning EAGAIN. Previously, POLLOUT only checked tcp_sndbuf > 0, so the next write() failed again. Cost: one mem_malloc/mem_free and one memp_malloc/memp_free per poll on a socket that has seen ERR_MEM; nothing on any other poll. On lwIP 1.x (ESP8266) memp_malloc does not compile in modlwip.c, so the trial is one pbuf of mss + sizeof(struct tcp_seg) bytes; that variant is compiled by CI only.

Commit 1 stands alone. If a narrower PR is preferred, commits 2 and 3 can move to their own PRs.

Test

One script reproduces all three on any MICROPY_PY_LWIP board with Wi-Fi. The same file runs on both sides: on the host it sends itself to the board through mpremote and plays two slow readers, on the board it runs the test. Two slow readers fill the lwIP heap with unacked segments, so a socket's tcp_write meets ERR_MEM while its own tcp_sndbuf still reports room. Phase 1 covers commits 1 and 3: a non-blocking write() returns EAGAIN instead of blocking, and POLLOUT after that EAGAIN is truthful. Phase 2 covers commit 2: a send() on a socket with a timeout ends at that timeout.

  1. Save the script as repro.py and set WIFI_SSID / WIFI_PASS in it.
  2. From a host on the same network, with the board on USB: python3 repro.py. It needs mpremote or uvx on the host.
  3. Read the two DONE blocks, each followed by PASS or FAIL: <reason>. About 4 minutes in all.

Phase 1 passes when no non-blocking write() takes more than 100 ms, writes return EAGAIN, and no POLLOUT is followed by another EAGAIN. Phase 2 passes when every starved send() on a settimeout(0.5) socket ends by ETIMEDOUT within 600 ms and none by ENOMEM.

What was run: an OpenMV RT1062 (cyw43 Wi-Fi) on OpenMV's fork at MicroPython v1.28.0-49, where the retry loop waits in mp_hal_delay_ms(50) instead of poll_sockets(). Upstream has no board definition for this camera, so the same change was applied to that loop; the PR's own form has not been run, CI compiles it. The unpatched image is OpenMV's release build with an 8 KB lwIP heap; the patched image uses a 16 KB heap, so the two differ by more than this change.

repro.py
# Non-blocking and timed sends under lwIP memory starvation, one file for both sides.
# On the host (CPython) it sends itself to the board through mpremote and plays two slow
# readers; on the board (MicroPython) it runs the two phases. The readers fill the lwIP heap
# with unacked segments, so a socket's tcp_write meets ERR_MEM while its tcp_sndbuf still
# reports room.
#   unpatched:  phase 1 a non-blocking write() blocks for seconds; phase 2 a send() on a
#               settimeout(0.5) socket blocks past its timeout
#   patched:    both phases PASS
# Usage: set WIFI_SSID / WIFI_PASS, then python3 repro.py with the board on USB.
#        python3 repro.py <board-ip> plays the readers only, for a board started by hand.
import sys

WIFI_SSID = "..."
WIFI_PASS = "..."
PORT = 8080
PHASE_S = 120
TIMEOUT_S = 0.5
# TIMEOUT_S plus the 50 ms grain of the ERR_MEM loop and scheduling slack.
MAX_TIMED_SEND_MS = 600
# A non-blocking write copies one chunk; a stall sits far above 100 ms.
MAX_NONBLOCKING_WRITE_MS = 100
READER_RATE = 20000


def board_main():
    import network, socket, select, time, errno, os

    print("firmware: %s; %s" % (os.uname().version, sys.version))
    print("board:    %s" % os.uname().machine)
    print("modlwip send under lwIP memory starvation: 2 phases of %d s, each needs 2 slow readers:" % PHASE_S)
    print("   Phase 1 NON-BLOCKING - setblocking(False) sockets: write() with no memory must return EAGAIN at once,")
    print("                          and POLLOUT must not report writable while the next write still fails")
    print("   Phase 2 TIMEOUT      - settimeout(%.1f) sockets: send() with no memory must raise ETIMEDOUT at the timeout" % TIMEOUT_S)
    print()

    wlan = network.WLAN(network.STA_IF)

    def join():
        for attempt in (1, 2, 3):
            # A radio left associated or mid-join by an earlier run refuses a new join.
            wlan.active(False)
            time.sleep_ms(500)
            wlan.active(True)
            print("joining AP(%s), attempt %d of 3..." % (WIFI_SSID, attempt), end="")
            wlan.connect(WIFI_SSID, WIFI_PASS)
            t0 = time.ticks_ms()
            while not wlan.isconnected() and time.ticks_diff(time.ticks_ms(), t0) < 20000:
                time.sleep_ms(500)
                print(".", end="")
            if wlan.isconnected():
                print(" joined")
                return True
            print(" not joined, status %d" % wlan.status())
        return False

    try:
        joined = join()
    except KeyboardInterrupt:
        print()
        print("INTERRUPTED during the Wi-Fi join")
        return
    if not joined:
        seen = [s for s in wlan.scan() if s[0] == WIFI_SSID.encode()]
        if seen:
            print("FAILED: could not join AP(%s) in 3 attempts; a scan sees it on channel %d at %d dBm,"
                  " so check the password or the AP's access list" % (WIFI_SSID, seen[0][2], seen[0][3]))
        else:
            print("FAILED: could not join AP(%s) in 3 attempts; a scan does not see it,"
                  " so the AP is out of range or off" % WIFI_SSID)
        return
    ip = wlan.ifconfig()[0]

    # Every socket lands in open_socks the moment it exists; the finally below closes them all.
    open_socks = []
    payload = memoryview(bytearray(32768))

    def listen():
        server = socket.socket()
        open_socks.append(server)
        server.setsockopt(socket.SOL_SOCKET, socket.SO_REUSEADDR, 1)
        server.bind(("0.0.0.0", PORT))
        server.listen(2)
        server.settimeout(1.0)
        print("listening %s %d" % (ip, PORT))
        print()
        return server

    def accept_two(server, setup):
        print("waiting for 2 readers", end="")
        clients = []
        while len(clients) < 2:
            try:
                conn, addr = server.accept()
            except OSError:
                print(".", end="")
                continue
            open_socks.append(conn)
            if not clients:
                print()
            setup(conn)
            clients.append(conn)
            print("  Reader %d (%s:%d) connected" % (len(clients), addr[0], addr[1]))
        print()
        return clients

    def drop(clients):
        for c in clients:
            c.close()

    def verdict(failures):
        if failures:
            print("       FAIL: " + "; ".join(failures))
        else:
            print("       PASS")
        print()
        return not failures

    def phase_nonblocking(server):
        clients = accept_two(server, lambda c: c.setblocking(False))
        print("PHASE 1 NON-BLOCKING: setblocking(False) socket, write() with no memory;"
              " must return EAGAIN at once, and POLLOUT must not be followed by EAGAIN")
        pollers = {}
        for c in clients:
            pollers[c] = select.poll()
            pollers[c].register(c, select.POLLOUT)
        eagain = spins = enomem = 0
        maxwr = 0

        def nb_write(c):
            # Unpatched firmware blocks a non-blocking write in the ERR_MEM loop, then may raise ENOMEM.
            try:
                return c.write(payload)
            except OSError as e:
                if e.args[0] != errno.ENOMEM:
                    raise
                return -1

        end = time.ticks_add(time.ticks_ms(), PHASE_S * 1000)
        last = time.ticks_ms()
        while time.ticks_diff(end, time.ticks_ms()) > 0:
            if time.ticks_diff(time.ticks_ms(), last) >= 10000:
                last = time.ticks_ms()
                print("  t=%3ds  EAGAIN: %d  EAGAIN right after POLLOUT: %d  max(write) ms: %d" % (
                    time.ticks_diff(last, end) // 1000 + PHASE_S, eagain, spins, maxwr // 1000))
            for c in clients:
                t = time.ticks_us()
                n = nb_write(c)
                maxwr = max(maxwr, time.ticks_diff(time.ticks_us(), t))
                if n == -1:
                    enomem += 1
                    continue
                if n:
                    continue
                eagain += 1
                tp = time.ticks_us()
                if pollers[c].poll(1000):
                    n2 = nb_write(c)
                    if n2 == -1:
                        enomem += 1
                    elif not n2 and time.ticks_diff(time.ticks_us(), tp) < 2000:
                        spins += 1
        print("DONE:  EAGAIN: %d  EAGAIN right after POLLOUT: %d  max(write) ms: %d" % (eagain, spins, maxwr // 1000))
        drop(clients)
        failures = []
        if maxwr // 1000 > MAX_NONBLOCKING_WRITE_MS:
            failures.append("non-blocking write() blocked %d ms" % (maxwr // 1000))
        elif eagain == 0:
            failures.append("no EAGAIN, memory never ran out")
        if enomem:
            failures.append("%d writes raised ENOMEM" % enomem)
        if spins:
            failures.append("%d writes got EAGAIN right after POLLOUT" % spins)
        return verdict(failures)

    def phase_timeout(server):
        clients = accept_two(server, lambda c: c.settimeout(TIMEOUT_S))
        print("PHASE 2 TIMEOUT: settimeout(%.1f) socket, one send() with no memory;"
              " must raise ETIMEDOUT at the timeout" % TIMEOUT_S)
        timedout = enomem = 0
        maxsend = 0
        end = time.ticks_add(time.ticks_ms(), PHASE_S * 1000)
        last = time.ticks_ms()
        while time.ticks_diff(end, time.ticks_ms()) > 0:
            if time.ticks_diff(time.ticks_ms(), last) >= 10000:
                last = time.ticks_ms()
                print("  t=%3ds  ETIMEDOUT: %d  ENOMEM: %d  max(send) ms: %d" % (
                    time.ticks_diff(last, end) // 1000 + PHASE_S, timedout, enomem, maxsend // 1000))
            for c in clients:
                t = time.ticks_us()
                try:
                    # send() is one lwip_tcp_send call; write() loops over chunks, a clock each.
                    c.send(payload)
                except OSError as e:
                    if e.args[0] == errno.ETIMEDOUT:
                        timedout += 1
                    elif e.args[0] == errno.ENOMEM:
                        enomem += 1
                    else:
                        raise
                maxsend = max(maxsend, time.ticks_diff(time.ticks_us(), t))
        print("DONE:  ETIMEDOUT: %d  ENOMEM: %d  max(send) ms: %d" % (timedout, enomem, maxsend // 1000))
        drop(clients)
        failures = []
        if maxsend // 1000 > MAX_TIMED_SEND_MS:
            failures.append("send() blocked %d ms, socket timeout is %d ms"
                            % (maxsend // 1000, int(TIMEOUT_S * 1000)))
        elif timedout == 0:
            failures.append("no send blocked, memory never ran out")
        if enomem:
            failures.append("%d sends raised ENOMEM" % enomem)
        return verdict(failures)

    stage = "opening the server socket"
    try:
        server = listen()
        stage = "phase 1"
        ok1 = phase_nonblocking(server)
        print("Dropping readers, they reconnect for phase 2")
        print()
        stage = "phase 2"
        ok2 = phase_timeout(server)
        print("ALL DONE: phase 1 %s, phase 2 %s" % ("PASS" if ok1 else "FAIL", "PASS" if ok2 else "FAIL"))
    except KeyboardInterrupt:
        print()
        print("INTERRUPTED during %s: closing sockets" % stage)
    except OSError as e:
        print()
        print("FAILED: socket error %d during %s" % (e.args[0], stage))
    except Exception as e:
        print()
        print("FAILED during %s: %r" % (stage, e))
    finally:
        # One close failing must not leave the rest open.
        for s in open_socks:
            try:
                s.close()
            except OSError:
                pass


def host_main():
    import os, re, shutil, signal, socket, subprocess, threading, time

    # One line per failure mpremote reports as a traceback.
    causes = (
        ("no device found", "no board on USB: plug it in, or close the program holding its port"),
        ("could not enter raw repl", "the board does not answer on its serial port: reset or replug it"),
        ("failed to access", "the serial port is held by another program (an IDE, another mpremote)"),
        ("device disconnected", "the board dropped off USB during the run: replug it"),
    )
    stop = threading.Event()

    def reader(host, port):
        while not stop.is_set():
            s = None
            try:
                s = socket.create_connection((host, port), timeout=5)
                s.settimeout(1.0)
                while not stop.is_set():
                    try:
                        data = s.recv(4096)
                    except socket.timeout:
                        continue
                    if not data:
                        break
                    time.sleep(len(data) / READER_RATE)
            except OSError:
                pass
            finally:
                if s is not None:
                    s.close()
            time.sleep(0.5)

    def start_readers(host, port):
        for _ in range(2):
            threading.Thread(target=reader, args=(host, port), daemon=True).start()

    if len(sys.argv) > 1:
        print("2 readers to %s:%d at %d B/s each, for a board already running the test; Ctrl-C stops them"
              % (sys.argv[1], PORT, READER_RATE))
        start_readers(sys.argv[1], PORT)
        try:
            while True:
                time.sleep(1)
        except KeyboardInterrupt:
            stop.set()
            print("\nreaders stopped")
            return 0

    mpremote = ["mpremote"] if shutil.which("mpremote") else ["uvx", "mpremote"]
    try:
        # Own session: the terminal's Ctrl-C reaches this script alone, which then stops the
        # child once, instead of both dying mid-close.
        proc = subprocess.Popen(mpremote + ["run", os.path.abspath(__file__)], stdout=subprocess.PIPE,
                                stderr=subprocess.STDOUT, text=True, bufsize=1, start_new_session=True)
    except FileNotFoundError:
        print("FAILED: neither mpremote nor uvx is installed")
        return 2

    traceback_lines = []
    try:
        for line in proc.stdout:
            # mpremote prints the board's empty reply, b'', ahead of its raw-REPL traceback.
            if traceback_lines or line.startswith(("Traceback", "mpremote:")) or line.strip() == "b''":
                traceback_lines.append(line)
                continue
            sys.stdout.write(line)
            sys.stdout.flush()
            m = re.match(r"listening (\S+) (\d+)", line)
            if m:
                start_readers(m.group(1), int(m.group(2)))
        rc = proc.wait()
    except KeyboardInterrupt:
        print("\nINTERRUPTED: stopping the board")
        stop.set()
        proc.send_signal(signal.SIGINT)
        try:
            proc.wait(timeout=10)
        except subprocess.TimeoutExpired:
            proc.kill()
        return 130
    stop.set()

    if traceback_lines:
        text = "".join(traceback_lines)
        for needle, cause in causes:
            if needle in text:
                print("FAILED: " + cause)
                break
        else:
            print("FAILED: mpremote said: " + traceback_lines[-1].strip())
        return 1
    return rc


if sys.implementation.name == "micropython":
    board_main()
else:
    sys.exit(host_main())
Result on unpatched firmware (OpenMV v5.0.0, MicroPython v1.28.0-49, built 2026-07-02, OpenMV RT1062)
firmware: v1.28.0-49 on 2026-07-02; 3.4.0; OpenMV v5.0.0; MicroPython v1.28.0-49
board:    OpenMV IMXRT1060 with MIMXRT1062DVJ6A
modlwip send under lwIP memory starvation: 2 phases of 120 s, each needs 2 slow readers:
   Phase 1 NON-BLOCKING - setblocking(False) sockets: write() with no memory must return EAGAIN at once,
                          and POLLOUT must not report writable while the next write still fails
   Phase 2 TIMEOUT      - settimeout(0.5) sockets: send() with no memory must raise ETIMEDOUT at the timeout

joining AP(<ssid>), attempt 1 of 3............ joined
listening <board-ip> 8080

waiting for 2 readers
  Reader 1 (<host-ip>:54353) connected
.  Reader 2 (<host-ip>:54352) connected

PHASE 1 NON-BLOCKING: setblocking(False) socket, write() with no memory; must return EAGAIN at once, and POLLOUT must not be followed by EAGAIN
  t= 10s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1400
  t= 21s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1550
  t= 31s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1550
  t= 41s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1550
  t= 51s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1550
  t= 62s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1552
  t= 72s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1552
  t= 82s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1552
  t= 93s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 1552
  t=103s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 2400
  t=114s  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 2400
DONE:  EAGAIN: 0  EAGAIN right after POLLOUT: 0  max(write) ms: 2400
       FAIL: non-blocking write() blocked 2400 ms

Dropping readers, they reconnect for phase 2

waiting for 2 readers......
  Reader 1 (<host-ip>:54366) connected
  Reader 2 (<host-ip>:54367) connected

PHASE 2 TIMEOUT: settimeout(0.5) socket, one send() with no memory; must raise ETIMEDOUT at the timeout
  t= 10s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 1450
  t= 21s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 1550
  t= 32s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 1750
  t= 42s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t= 53s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t= 64s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t= 74s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t= 85s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t= 95s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3101
  t=105s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3401
  t=115s  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3401
DONE:  ETIMEDOUT: 0  ENOMEM: 0  max(send) ms: 3401
       FAIL: send() blocked 3401 ms, socket timeout is 500 ms

ALL DONE: phase 1 FAIL, phase 2 FAIL
Result with this PR's change (same board, OpenMV's fork at MicroPython v1.28.0-49 plus the change)
firmware: cf80cce8a0-dirty on 2026-10-02; 3.4.0; OpenMV 13d41e5381-dirty; MicroPython cf80cce8a0-dirty
board:    OpenMV IMXRT1060 with MIMXRT1062DVJ6A
modlwip send under lwIP memory starvation: 2 phases of 120 s, each needs 2 slow readers:
   Phase 1 NON-BLOCKING - setblocking(False) sockets: write() with no memory must return EAGAIN at once,
                          and POLLOUT must not report writable while the next write still fails
   Phase 2 TIMEOUT      - settimeout(0.5) sockets: send() with no memory must raise ETIMEDOUT at the timeout

joining AP(<ssid>), attempt 1 of 3......... joined
listening <board-ip> 8080

waiting for 2 readers
  Reader 1 (<host-ip>:54307) connected
.  Reader 2 (<host-ip>:54308) connected

PHASE 1 NON-BLOCKING: setblocking(False) socket, write() with no memory; must return EAGAIN at once, and POLLOUT must not be followed by EAGAIN
  t= 10s  EAGAIN: 161  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 21s  EAGAIN: 270  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 31s  EAGAIN: 376  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 42s  EAGAIN: 479  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 53s  EAGAIN: 612  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 64s  EAGAIN: 719  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 75s  EAGAIN: 827  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 85s  EAGAIN: 946  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t= 95s  EAGAIN: 1048  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t=105s  EAGAIN: 1132  EAGAIN right after POLLOUT: 0  max(write) ms: 4
  t=116s  EAGAIN: 1238  EAGAIN right after POLLOUT: 0  max(write) ms: 4
DONE:  EAGAIN: 1295  EAGAIN right after POLLOUT: 0  max(write) ms: 4
       PASS

Dropping readers, they reconnect for phase 2

waiting for 2 readers.......
  Reader 1 (<host-ip>:54327) connected
  Reader 2 (<host-ip>:54328) connected

PHASE 2 TIMEOUT: settimeout(0.5) socket, one send() with no memory; must raise ETIMEDOUT at the timeout
  t= 10s  ETIMEDOUT: 12  ENOMEM: 0  max(send) ms: 501
  t= 20s  ETIMEDOUT: 27  ENOMEM: 0  max(send) ms: 501
  t= 30s  ETIMEDOUT: 42  ENOMEM: 0  max(send) ms: 501
  t= 40s  ETIMEDOUT: 57  ENOMEM: 0  max(send) ms: 501
  t= 51s  ETIMEDOUT: 71  ENOMEM: 0  max(send) ms: 501
  t= 61s  ETIMEDOUT: 86  ENOMEM: 0  max(send) ms: 501
  t= 71s  ETIMEDOUT: 100  ENOMEM: 0  max(send) ms: 501
  t= 81s  ETIMEDOUT: 114  ENOMEM: 0  max(send) ms: 501
  t= 91s  ETIMEDOUT: 130  ENOMEM: 0  max(send) ms: 504
  t=102s  ETIMEDOUT: 145  ENOMEM: 0  max(send) ms: 504
  t=113s  ETIMEDOUT: 162  ENOMEM: 0  max(send) ms: 504
DONE:  ETIMEDOUT: 172  ENOMEM: 0  max(send) ms: 504
       PASS

ALL DONE: phase 1 PASS, phase 2 PASS

The defect was first seen in an application on this board: an MJPEG stream to one reader throttled to 20 KB/s stalled a non-blocking write() for up to 1553086 us, 54 times over 100 ms in 320 s.

@srgg
srgg marked this pull request as draft October 2, 2026 16:05
lwip_tcp_send sizes a write from tcp_sndbuf, a byte count, but tcp_write
allocates its segments from the heap and the MEMP_TCP_SEG pool and returns
ERR_MEM when those run out while the count still shows room.  The retry
loop then waited up to 10s, for a non-blocking socket too.

tcp_write queues nothing on ERR_MEM, so a non-blocking socket now returns
EAGAIN from that loop, as it already does when tcp_sndbuf is 0.  Blocking
sockets are unchanged.

Fixes issue micropython#19704.

Signed-off-by: srgg <[email protected]>
@srgg
srgg force-pushed the modlwip-nonblocking-partial branch 3 times, most recently from 9b4b80c to e0dfeb3 Compare October 4, 2026 17:27
@srgg
srgg marked this pull request as ready for review October 4, 2026 17:49
The tcp_sndbuf==0 wait in lwip_tcp_send ends with ETIMEDOUT at the socket's
timeout, but the ERR_MEM retry loop ended only when the write succeeded or
10s had passed, so a send on a socket with settimeout(t) could block far
past t.  Measured on an OpenMV RT1062 with t=0.5 and two slow readers:
sends of 1651ms to 3401ms, none raising ETIMEDOUT.

The loop now ends with ETIMEDOUT at the socket's timeout, counted from the
start of the call so the tcp_sndbuf wait and the loop share one clock, and
the 10s limit applies to a socket with no timeout only.  Same test: every
starved send raises ETIMEDOUT within 504ms.

Fixes issue micropython#19746.

Signed-off-by: srgg <[email protected]>
@srgg
srgg force-pushed the modlwip-nonblocking-partial branch from e0dfeb3 to a17b43b Compare October 4, 2026 18:16
@srgg srgg changed the title extmod/modlwip: Return a partial or ENOBUFS from a non-blocking send extmod/modlwip: Fix TCP send and POLLOUT when out of memory. Oct 4, 2026
POLLOUT reads tcp_sndbuf, a byte count, so after an ERR_MEM EAGAIN a
non-blocking socket polls writable and its next send gets EAGAIN again.
lwIP's sockets gate on per-pcb low-water marks
(NETCONN_FLAG_CHECK_WRITESPACE), which read one socket, while the memory
ERR_MEM lacks is shared by all of them.

After such an EAGAIN, POLLOUT now needs one segment's memory to be
allocatable: a pbuf of mss bytes and one tcp_seg, both tried and released,
plus a free slot in snd_queuelen.  Since tcp_write queues all of write_len
or nothing, the send retries once at a single mss before EAGAIN, so a
writable poll is followed by an accepted write.  Measured on an OpenMV
RT1062 with two slow readers: 1295 EAGAINs, none right after POLLOUT.

Signed-off-by: srgg <[email protected]>
@srgg
srgg force-pushed the modlwip-nonblocking-partial branch from a17b43b to 84847f9 Compare October 5, 2026 12:20
@dpgeorge

dpgeorge commented Oct 6, 2026

Copy link
Copy Markdown
Member

Commit 1 stands alone. If a narrower PR is preferred, commits 2 and 3 can move to their own PRs.

Right. This PR is now quite complicated and requires some effort to understand its impact.

Considering #19708 independently did exactly the same thing as Commit 1 here, I suggest we just merge that fix for now. And then follow up with a separate PR for the other changes.

@srgg

srgg commented Oct 6, 2026

Copy link
Copy Markdown
Author

I understand the concern. Commit 1 contains the original fix I needed, while Commits 2 and 3 address additional issues I discovered along the way.

I’d prefer to keep the changes together rather than spend time restructuring the PR. That said, if you’d prefer to keep this PR focused on the original fix, I’m fine with merging #19708 independently and leaving the other changes as they are for now.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

extmod Relates to extmod/ directory in source

Projects

None yet

Development

Successfully merging this pull request may close these issues.

modlwip: socket timeout ignored in send when out of memory modlwip: non-blocking socket blocks in send when out of memory

2 participants