Repository navigation
gracefulClose stops servers due to a lot of TCP states #438
Description
Activity
Please let me know what additional data would be helpful to diagnose the issue. I can provide larger TCP dumps and/or detailed netstat infos if needed. I can also provide direct ssh access to affected machines, if that's helpful for understanding or solving the issue.
Sorry for the delay. This is probably due to
gracefulClose. First of all, if you are usingwarp< 3.3.4, please upgrade towarp>= 3.3.5. Since 3.3.5,warpusesclose()for HTTP/1.1 by default.Even if the problem continues, please check if HTTP/2 connections are used. If so, please use
setGracefulCloseTimeout2 0to disablegracefulClosefor HTTP/2. If this fixes the problem, the source of bug is definitelygracefulClose. Timeout to close sockets is not handled correctly.We are definitely using the newest
warp, so we'll try it withsetGracefulCloseTimeout.This article would help your understanding: https://kazu-yamamoto.hatenablog.jp/entry/2019/09/20/165939
@larskuhtz Would you close this issue if already resolved?
- changed the title
[-]Regression with version >= 3.1.1.0[/-][+]gracefullclose stops servers due to a lot of TCP states[/+]on May 19, 2020 I had the same experience. I will try to fix.
Reacted by Colin WoodburyThank you! Should the
3.1.1.xseries be marked as deprecated on Hackage, once3.1.2.0is released?Yes. I will do so.
Reacted by Colin Woodbury@snoyberg I'm CC:ing to you here since the current approach of
gracefulClosewas suggested by you and this is relating to Warp.TCP and server:
CLOSE_WAITis the state that the TCP stack of the server received TCP FIN but not send TCP FIN. This means the server does notclose(2)the socket yet. The TCP stack waits forever in this situation.Warp: When TCP FIN is received,
connRecvreturns "".serveinforkreturns and thenconnCloseis called. So far, so good. ButgracefulClosecannot callclose. Why?https://github.com/haskell/network/blob/master/Network/Socket/Shutdown.hs#L61
I suspect two things:
- Asynchronous exception from Warp's time manager
- Callback is not fired by GHC TimerManager
But I don't have any clues yet so far. Could you suggest anything?
Note that
shutdowncan throw an exception.It might be wise if we call
recvBufNoWaitaftershutdownfor the case where FIN is already received. (Like C'sdo { } while ()loop).- changed the title
[-]gracefullclose stops servers due to a lot of TCP states[/-][+]gracefulClose stops servers due to a lot of TCP states[/+]on May 19, 2020 I think it was actually @nh2 who proposed the implementation of
gracefulClose. Maybe he has some thoughts, I unfortunately don't.Approach 4 in this article (https://kazu-yamamoto.hatenablog.jp/entry/2019/09/20/165939) was proposed by you. :-)
29 remaining items
@swamp-agr If you use Linux, please check
net.ipv4.tcp_fin_timeout. The default value, 60 (second), might be too long for your use case. Also, please checknet.ipv4.tcp_tw_reuse. This should be 1 for your case. The following settings for/etc/sysctl.confwould help:net.ipv4.tcp_tw_reuse = 1 net.ipv4.tcp_fin_timeout = 30Note that
tcp_tw_recyclewould be also related. But this is a bit dangerous and was removed Linux 4.1.2.
Anyway, please look into these parameters.Issue reproduced even with these values:
net.ipv4.tcp_tw_reuse=1 net.ipv4.tcp_fin_timeout=15tcp_tw_recycleis indeed dangerous since enabling it make server suffers and drops connections. I think it is incompatible with NGINXkeepalivesetting for upstream server.@swamp-agr If you believe this is a bug of Warp, please send this issue to Warp.
- Root cause found inside application code (as usual).
- Handlers were running indefinitely. It causes warp to wait endlessly.
- Client dropped connection.
- NGINX closed socket from its side.
- Warp still waits for application handler.
With a constant RPS server accumulates stalled sockets.
Expected Result: Application handler should finish its job. Warp should respond with
close.
Acutal Result: Application never stopped. Warp is waiting.You might close the issue.
@swamp-agr I close this issue. Please bring this issue to Warp.
The number of open file descriptors is moderate, but many of the TCP sockets are in a
CLOSE_WAITstate. Most of those sockets are not listed bylsof, but are only shown bynetstatwithout an associated process.I'm now hitting exactly that problem.
I asked a question, and provided an answer, on how and why these process-less, FD-less
CLOSE_WAITs exist:- https://stackoverflow.com/questions/77355636/close-wait-tcp-states-despite-closed-file-descriptors
- https://stackoverflow.com/questions/77355636/close-wait-tcp-states-despite-closed-file-descriptors/77420709#77420709
The answer is:
A
CLOSE_WAITstate without associated process occurs when a client waiting in the Linux kernel'slisten()backlog queue disconnects before the user-space applicationaccept()s it.I haven't figured out yet why my warp application stops
accept()ing for multiple minutes, creating theCLOSE_WAITs.@nh2 Thank you for bringing this answer!
And now I think I can answer your question.
Callbacks for the IO and Timer managers MUST NOT be blocked.
If they are blocked, the entire loops of the IO and Timer manager block are blocked.
In this situation, any bad things can happen (including the non-accepting).
If the callbacks can be blocked, we MUST useforkIO.
That's whytimeoutcallsforkIOif necessary.Originally, we tried to use the call back approach to avoid forking a new thread.
But it is not avoidable to callforkIOin this approach.
Thus, thetimeoutapproach is better becauseforkIOis called only when it is necessary.Thus, the
timeoutapproach is better becauseforkIOis called only when it is necessary.@kazu-yamamoto Just for me to get back into context:
Where are those callbacks / is this something that was recently changed, or that you plan to change (e.g. open or already-closed issue or PR)?
Because I'm still currently investigating what to do about those blocked accepts.
Your explanation seems to fit my symptoms ("the entire loops of the IO and Timer manager block are blocked") because the process really seems to stop doing almost everything for a while -- not very good for my web server when it happens :D
Where are those callbacks / is this something that was recently changed, or that you plan to change (e.g. open or already-closed issue or PR)?
I guess that you are talking about the graceful close.
Our final decision was to adopt thethreadDelayapproach.
So, our gracefull close is NOT suffering fromCLOSE_WAIT.See approach 3 in https://kazu-yamamoto.hatenablog.jp/entry/2019/09/20/165939
Probably, I should update this article.So, our gracefull close is NOT suffering from
CLOSE_WAIT.@kazu-yamamoto Because my server is suffering from 3000
CLOSE_WAITs a couple times per week; I'm on:network-3.1.2.9 wai-3.2.3 warp-3.3.23@nh2 Understood.
If you stop usinggracefulClose, doesCLOSE_WAITdisappear?I haven't figured out yet why my warp application stops
accept()ing for multiple minutes, creating theCLOSE_WAITs.An update on this:
My application was calling
unsafeFFI to process data coming from anmmap(to hash a 100 GB file). That is illegal, becauseunsafeforeign calls must only be used on extremely short-running functions. Otherwise the entire Haskell process blocks during GC. This happened to me, so my entire process blocked for 20 minutes (until the hashing of 100 GB completed).This of course caused my process to stop calling any function, including
accept(), and thus theCLOSE_WAITs accumulated (I have automated HTTP monitoring that queries my server every second to see if it's still up; if it doesn't reply within a 2 second timeout, my monitoring client disconnets, and so its un-accept()ed connection sits in the kernel's listen queue asCLOSE_WAIT, and every second a new one got added).You can read more about it here:
It was difficult to figure out because
mmapaccess does not show up instrace. Avoidmmapwhen you can, it's invisible and thus hard to debug!Also scrutinise any libraries for
unsafecalls, even if they only domemcpy; if they accept an arbitrary-sizedByteString, they will block your entire process until they are done.
This does not imply that there are no further buts in
networkorwarp, just that I found one certain cause of the Haskell process freezing that was a problem in my application.Reacted by Janus Troelsen and Andrey Prokopenko



We run a p2p network with Haskell nodes using
network+tls+warpfor the server andnetwork+tls+http-clientfor the client components.We observed that nodes that are using
networkversion< 3.1.1.0have been running without issues for weeks, while nodes that are usingnetwork >=3.1.1.0are stopping to make and serve requests after running for a few days.Bad nodes don't accept any incoming connections and fail to establish outgoing connections.
On the bad nodes there is no increase in memory consumption and CPU usage is low, since they are not doing anything useful without being able to make network connections. The number of open file descriptors is moderate, but many of the TCP sockets are in a
CLOSE_WAITstate. Most of those sockets are not listed bylsof, but are only shown bynetstatwithout an associated process.The following are two typical TCP sessions:
HTTP TCP sessions from other processes seem fine.