Repository navigation
infinite loop of kevent returning EINVAL on macOS #1315
Description
Activity
- addedstatus: needs-triageThis issue needs to be triaged.This issue needs to be triaged.type: bugThis issue describes a bug.This issue describes a bug.
on Jul 8, 2021 ok, after writing all that, I realized since
statusis-1we should break out of the loop here so something a bit different must be happening (_do_pollcalled in a loop?)The 2.4.2 was released few days ago. Could you try it ?
I guess we should preferably search around the possibility of a negative time again :-/
@mr-salty how did you build haproxy ? It's not normal that your CFLAGS line is empty, and it's particularly lacking a few options such as
-fwrapvthat we've added a while ago to support wrapping arithmetic that the original code relies on, with recent compilers. And this one used to be responsible for some cases of invalid timestamps causing pollers to loop.I think I'll add a runtime check in the code to make sure this one was not disabled.
Could you please apply the attached patch ? It will detect the bugs and fail if your options are invalid.
build.diff.txtI am using the homebrew pre-built binary.
looking at https://github.com/Homebrew/homebrew-core/blame/master/Formula/haproxy.rb - it does muck with
CFLAGSbut that hasn't been changed in 9 years (!) so it's probably overdue for an update. I can play around with building it myself and send them a PR.I did update my machine to 2.4.2 the other day and haven't seen any issues yet, although AFAIK the problem only happened to me the one time in 5 months. I asked other devs at my company to keep an eye out as well.
Indeed, so that definitely justifies at least a sensitivity to integer wrapping. The thing is that compilers recently started to find it fun to break 40 years of properly working code by abusing undefined behaviors that used to work well for everyone, and now you can't get a properly working C program without at least a bit of options to calm some of their useless creativity. Thus I'll merge the proposed patch above to catch these earlier. Regarding the fact that it doesn't happen often, the most likely cause is that it will hit only after 49.7 days of uptime (when the internal time wraps).
By the way the simplest fix for the package above would be to pass their optimizations to CPU_CFLAGS instead of CFLAGS (the latter contains a concatenation of the various flags sources). It's worth noting that their variable was entirely empty, resulting in the loss of even -O2, and that ought to be addressed.
- added a commit that references this issue
on Jul 14, 2021 - added1.8This issue affects the HAProxy 1.8 stable branch. (EOL)This issue affects the HAProxy 1.8 stable branch. (EOL)1.9This issue affects the HAProxy 1.9 stable branch. (EOL)This issue affects the HAProxy 1.9 stable branch. (EOL)2.0This issue affects the HAProxy 2.0 stable branch. (EOL)This issue affects the HAProxy 2.0 stable branch. (EOL)and removedstatus: needs-triageThis issue needs to be triaged.This issue needs to be triaged.
on Jul 17, 2021 6 remaining items
- removed2.4This issue affects the HAProxy 2.4 stable branch. (EOL)This issue affects the HAProxy 2.4 stable branch. (EOL)2.3This issue affects the HAProxy 2.3 stable branch. (EOL)This issue affects the HAProxy 2.3 stable branch. (EOL)
on Jul 27, 2021 - added 2 commits that reference this issue
on Jul 28, 2021 - added 2 commits that reference this issue
on Aug 1, 2021 - added 2 commits that reference this issue
on Aug 13, 2021 - removed2.0This issue affects the HAProxy 2.0 stable branch. (EOL)This issue affects the HAProxy 2.0 stable branch. (EOL)2.1This issue affects the HAProxy 2.1 stable branch. (EOL)This issue affects the HAProxy 2.1 stable branch. (EOL)2.2This issue affects the HAProxy 2.2 stable branch. (EOL)This issue affects the HAProxy 2.2 stable branch. (EOL)1.9This issue affects the HAProxy 1.9 stable branch. (EOL)This issue affects the HAProxy 1.9 stable branch. (EOL)
on Aug 13, 2021 - removed1.8This issue affects the HAProxy 1.8 stable branch. (EOL)This issue affects the HAProxy 1.8 stable branch. (EOL)
on Aug 25, 2022 - added a commit that references this issue
on Aug 29, 2022
Detailed Description of the Problem
I noticed
haproxywas using ~100% CPU on my machine (Big Sur 11.4, HAProxy version 2.4.1-1ce7d49). It was not responding on port 19000 and was not writing anything to the log. Unfortunately the log got clobbered but I didn't note anything unusual, the last line was about a service going down and the mtime on the logfile was >24h old.so, I tried
dtrusson the process which showed an endless stream of:22 =
EINVAL; per the manpage: "The specified time limit or filter is invalid".After a restart it was back to normal.
Another dev at my company said that other people see this issue periodically, I will try to collect more data on how frequently it is happening.
Expected Behavior
haproxy should not get stuck in a hard loop...
Steps to Reproduce the Behavior
not sure, it just happened randomly while haproxy was running.
Do you have any idea what may have caused this?
looking at ev_kqueue.cc, either
timeoutorkev(whichdtrusshelpfully doesn't show us, although I guess we wouldn't see the contents anyways) must be malformed, so retrying doesn't help.I saw some similar looking issues: #635, "haproxy hang", "time wrap caused kevent infinite loop"
complete speculation on on my part is running out of fds may be part of it, since that has caused other issues for us in the past, but I didn't look for evidence of that at the time. even if that is the case, I'd argue it should be handled more gracefully.
Do you have an idea how to solve the issue?
could HAproxy handle
EINVALdifferently (also proposed in one of the earlier issues I linked to)?Or, perhaps more generally, one or both of:
while (1)put a limit on the number of retries before we give up (and perhaps exit?).What is your configuration?
Output of
haproxy -vvLast Outputs and Backtraces
No response
Additional Information
No response