Skip to content

qede: full-size (1514 byte) IPv6 frames are dropped on TX ("Packet drop as packet len 0x5ea > 0x5ee"), occasional #GP panic in qede_tx_bcopy #107

Description

@itilim

Summary
On OmniOS r151058 with a QLogic/Marvell FastLinQ QL45xxx (10GBASE-T) NIC, the qede driver silently drops outgoing IPv6 packets that result in a full-size Ethernet frame (1514 bytes, MTU 1500). The same frame size over IPv4 is not affected. The drop is logged as

qede: NOTICE: qede_send_tx_packet(0):Packet drop as packet len 0x5ea > 0x5ee
(0x5ea = 1514, 0x5ee = 1518 - the printed comparison is not true as printed, so the check seems to compare something other than the raw frame length.)

Practical impact: TCP streams that need a full-size IPv6 segment stall forever. On our file server a stalled LDAP/GC connection (idmapd -> domain controller, IPv6) blocked idmapd, which then blocked every new SMB logon in smbd. Existing SMB sessions kept working, new ones timed out; NFS/SMB "hang until reboot" over weeks.

The same code path (qede_send_tx_packet) produced a kernel panic once (stack below).

Environment
OmniOS r151058 (omnios-r151058-c1eded413b), i86pc, driver/network/qede 0.5.11-151058.0
qede 8.1.25, FW 8.18.19.0, MFW 8.22.6.0, chip "AH A2", PCI 1077:8070, subsystem pci1590,28d
Link: 10000 Mbps full duplex, 10GBASE-T, MTU 1500 (jumbo disabled), 4 RX/TX rings (num_fp=4)
Defaults otherwise: tx_copy_threshold=256, rx_copy_threshold=128, checksum=2, lro_enable=1
Single, non-VLAN, non-aggregated link (qede0), IPv4 + IPv6 (static + addrconf)
Only 1.2-2.4k drops/day under normal load, but each one can kill a TCP flow

Reproduction (deterministic)
Server 2001:db8::251, peer (domain controller) 2001:db8::10 - ICMPv6 echo, 10 packets each:

ping -s -n 2001:db8::10 1000 10   ->  0% loss      (frame 1042)
ping -s -n 2001:db8::10 1432 10   ->  0% loss      (IPv6 packet 1480, frame 1494)
ping -s -n 2001:db8::10 1452 10   -> 100% loss     (IPv6 packet 1500, frame 1514)
ping -s -n 192.0.2.10   1472 10   ->  0% loss      (IPv4 packet 1500, frame 1514)

Each lost full-size IPv6 packet increments stats:txTotalDiscards (+31 during the test, matching the lost packets plus TCP retransmits of the stuck flow). The peer is reachable, small packets and the same size over IPv4 work, so this is not a network/MTU problem on the path.

Stuck TCP flow observed at the same time (netstat -an -f inet6): ESTABLISHED to the peer's port 3268, Send-Q 1692 bytes, unchanged over minutes.

What was ruled out
LSO: qede0_lso_enable=0; in /kernel/drv/qede.conf, confirmed at boot ("qede_mac_get_capability(0):Disabling large segmentation offload"), after reboot. Drops continued unchanged (~1.2k-2k/day for 4 weeks) - LSO is not the trigger.
Driver version: newest qede package for r151058 already installed.
Switch/path: ping of smaller sizes and IPv4 full-size fine; kstat shows no CRC/align/oversize errors.

Workaround that works
ipadm set-ifprop -p mtu=1480 -m ipv6 qede0
With the IPv6 interface MTU at 1480 (no 1514-byte IPv6 frames), the 1452-byte test passes (fragmented) and stats:txTotalDiscards stayed flat at 43647 for 4 days (previously +1.2-2k/day); the service hang did not return.

Kernel panic, same code path (7 Aug)

BAD TRAP: type=d (#gp General protection) rp=fffffe004d5fd840 addr=42
sched: pid=0 ... #gp General protection addr=0x42
genunix:ddi_dma_sync+5b
qede:qede_tx_bcopy+97
qede:qede_send_tx_packet+17f
qede:qede_ring_tx+62
mac:mac_hwring_tx+26 / mac_ring_tx+27 / mac_provider_tx+6d
mac:mac_tx_send+2ad / mac_tx_soft_ring_process+92 / mac_tx_fanout_mode+176 / mac_tx+1a3
dld:str_mdata_fastpath_put+8e
ip:ip_xmit+8b7 / ip:ire_send_wire_v4+3c2 / ip:conn_ip_output+18f

Note this one is an IPv4 path (ire_send_wire_v4), so IPv4 can also reach the broken code, just not via the full-size-frame condition above. Seven crash dumps from this machine exist since 03/2025 (several with the same NIC); we can provide one if useful.

Not yet known / could be tested
Exact threshold between 1494 and 1514 byte frames for IPv6 (tested 1494 ok, 1514 lost)
Whether default_checksum=1 (IPv4-only checksum offload) or default_lro_enable=0 changes it
Whether this affects other QL45xxx boards / firmware versions

Ask
Does anybody recognise the "Packet drop as packet len X > Y" check in qede_send_tx_packet, and is this a known issue?
Happy to run dtrace/mdb on the live system (fbt::qede_send_tx_packet:entry etc.) or provide crash dumps.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions