Repository navigation
No request timeout possibly means that sockets are being left open when queries fail #446
Description
Activity
Weird, doesn't the machine automatically time out after a while? The
requestdocs seem to suggest 20 seconds is the default timeout for connections on Linux. Any idea what is causing the callback to not fire? Is the reply just not coming back GCD servers or is there something wrong withgcloud-node? A snippet that could reproduce this issue would really help us investigate.Thanks for the quick response.
The timeout referred to in that doc is the timeout for TCP connections; I think my problem is with established connections that are open but not sending any data. If the TCP connection is healthy, there's no limit on how long it can be kept open which makes websockets and HTTP/2 possible.
I've looked at the gcloud-node code and couldn't see any places where dataset could "forget" to run the function, except that no timeout is being set for request.
The quickest solution is to see the agent's maxConnections to Infinity, which I should have done anyway, but if the connections never close the event loop will never be empty which means the process cannot gracefully exit. Therefore I think setting a timeout of something like 60 seconds just to be sure the socket is not orphaned would be a good thing to do.
I could be wrong about the open socket thing though. It would be cool if a nice Googler out there who knows more about how GCD's API servers and networks work could provide some insight into whether that's even possible.
Oh yeah it might help to know that I'm using a non-Google cloud server in Tokyo.
I'm writing on my phone now so I'll try to get some code up within the next few hours.
- addedapi: datastoreIssues related to the Datastore API.Issues related to the Datastore API.
on Mar 17, 2015 Sorry for taking so long to provide code.
There are two environments in which I've experienced this.
1: Getting some entities either by date or key
I have a class calledArticlewhose data I set with the contents of a Datastore entity.var Article = function(data){ data && extend(this,data) this.preview_image = data.preview_image || '' } Article.articleFromEntity = function(entity){ var article = new Article(entity.data) article.id = entity.key.path[1] return article }And I query some
articlesordered by date.Article.getArticlesBeforeDate = function(options,callback){ var date = options.date var query = dataset.createQuery(['Article']) query = query.filter('date <',date) query = query.order('-date') query = query.limit(5) dataset.runQuery(query,function(err,entities,endCursor){ var entities = entities || [] var articles = [] entities.forEach(function(entity){ var article = Article.articleFromEntity(entity) articles.push(article) }) callback(err,articles) }) }And by ID
Article.getArticleWithId = function(options,callback){ id = options.id var key = dataset.key(['Article',id]) dataset.get(key,function(err,entity){ if(err){ console.log(err) //very stupid retry setTimeout(function(){ Article.getArticleWithId(options,callback) },1000) return; } if(entity){ var article = Article.articleFromEntity(entity) callback(err,article) } else{ callback(err,null) } }) }When the socket pool gets full, both of those functions start to fail. I'm sure of this because in one environment, my API endpoints log when a request comes in successfully, but cannot complete because the
dataset.getandquery.runQueryfunctions never run the callback functions.2: Server discovery using GCD
I have a node module called Comrade which uses GCD to store information about servers in a cluster such as their IP addresses and to tell them when they should gracefully shut down. If you wouldn't mind looking at member.js#35, you can see that I am updating a single entity and I have a queue to prevent contention. That function is run about once every 10 seconds. Since the queue has a concurrency of 1, if any of the callbacks don't get called the queue will just keep filling up forever. The first time this happened, I thought it was a bug in
gcloud-nodethat's causing the callback to not be fired, so I added a 15 second timeout as a workaround, but that stops working after a while because eventually the callback stops being fired altogether no matter how many times I retry, so I think it can only be a network issue.I'm really not sure how to start approaching this issue. It would be most ideal if you provide a gcloud-node specific snippet that could be tested on its own that shows the connections are being left open, otherwise I'm inclined to say this is an issue that affects your code or is somewhere else along the pipeline. I don't see why the server would leave those connections open indefinitely, even if the callback was improperly not called.
It would be cool if a nice Googler out there who knows more about how GCD's API servers and networks work could provide some insight into whether that's even possible.
@jgeewax / @pcostell any insight?
@richardkazuomiller sorry for how long this has been outstanding. Are you able to run some tests against our latest version? A lot of how we handle making requests has changed (for one; since this issue, we have set maxSockets to infinity).
Thanks, and cool project!
/cc @eddavisson
I don't think I can provide any insight here. It seems like the client shouldn't depend on the specifics of how long the server might choose to keep idle sockets open.
@stephenplusplus Sorry I missed your last comment! Also thanks for saying it's a cool project. That means a lot (^_^)
As you suggested, I think this issue was (kinda) resolved when the socket pool size was increased to infinity (4188eb3). Even if the sockets were left open again, nothing bad would happen until the machine ran out of ports, which would be 10s of thousands of requests. Most deployments would probably restart the process by the time that many sockets get opened. You can probably leave it like it is and no one will die because of it.
However, I disagree with @eddavisson in this case because we know that there is an amount of time in which GCD should give a response. If a request takes more than 30 seconds, there's either something wrong with the network or something wrong with GCD, and the request is either never going to give a response or fail in some other way. If it was normal for a request to take several minutes to complete, a timeout may not be appropriate, but I do not think this is the case. GCD's servers may never keep the socket open for such a long time, but the Internet is a big place and there are environments in which a server can think it's connected to a certain thing, waiting for packets, but in reality some faulty router in the middle is keeping the client connection open and dropped the outgoing connection, resulting in a half-open connection (just one example). In any case, the client has no way of knowing when or if a response will arrive, so it needs some way to close the connection after some time has passed. Also I think it's important to keep in mind that packages that prevent the event loop from becoming empty - whether it's because of rogue
setTimeouts orsetIntervals or open sockets - are super annoying because the process will never exit unless you force it to.Since GCD is kind of a black box, whether or not to enable a socket timeout is up to you Googlers (Alphabetters?). Although it is not always the case, I think it is not unreasonable to work off the assumption that clients have a reliable connection to GCD because worrying about every edge case would be incredibly time consuming. However, adding a socket timeout would just require passing the
timeoutparameter torequest.Datasetwould just need to accept a default timeout as an option (EDIT: or don't allow an option, just hardcode some arbitrary number of milliseconds as a timeout). I know you're never supposed to tell an engineer something is easy but ... that's really easy! If you did that it would be super cool. I'd even be willing to implement it and submit a PR, but only if you're open to the idea of having a timeout.Infinity sockets is proving to be a pretty good bandaid (or I guess two birds one stone type of deal?), and since I posted this I've moved everything critical to shiny new GCE servers so I probably won't have weird network problems anyway. This is mostly an edge case and probably won't affect me ever again, so it's more of a philosophical issue than a technical one. At this point I personally don't care which way you decide to go. Feel free to close if you want.
Sorry for the long post but after six months I think it's best to get all of my feelings out there and let you guys decide what to do so we can all move on.
TL;DR version
Would adding a 60 second timeout to all requests break anything? If not, let's add a 60 second time out! If yes, let's close this issue. I don't need to hear any of the reasons either way; I feel like with the number of people involved in this now we're getting dangerously close to bikeshedding which I don't think would be doing the right thing (^_-)-☆
Even though it took six months, I'm glad it led to such an informative discussion! Feel free to chime in on any of our issues!
I'm totally open to a sixty second timeout. PR welcome 👍
Awesome! I'll start working on it soon.
@richardkazuomiller we still need you, buddy! :) No worries if you can't get around to it quickly, just a friendly ping if you're still interested.
19 remaining items
- added 4 commits that reference this issue
on Jan 27, 2026 - added a commit that references this issue
on Jan 28, 2026 - added a commit that references this issue
on Feb 5, 2026 - added a commit that references this issue
on Feb 5, 2026 - added a commit that references this issue
on Feb 17, 2026 - added a commit that references this issue
on Feb 23, 2026 - added a commit that references this issue
on Feb 25, 2026 - added a commit that references this issue
on Feb 26, 2026 - added a commit that references this issue
on Feb 26, 2026 - added a commit that references this issue
on Mar 5, 2026 - added a commit that references this issue
on Mar 5, 2026 - added a commit that references this issue
on Mar 18, 2026
I have a web app that running Node.js v0.10 and latest gcloud-node on multiple servers. I've run into two issues that occurred independent of one another but I think are symptoms of the same issue.
dataset.savewas never called, so the subsequent requests I had queued to prevent contention were never run.Before restarting the processes, I checked to see what sockets were open
In the output above, there are 15 connections to Google IPs, which remained open for several minutes before I restarted the Node processes. There are 15 because in Node 0.10 the default
maxSocketsfor the global HTTP agent is5, and there were three Node processes running on that machine.Leaving these sockets open indefinitely seems like a problem that should be fixed on the server side of GCD, but as long as I'm not missing something I propose that there should be a timeout of a few seconds for all requests sent by gcloud-node.
Anyone have any thoughts related to any of that?
Thanks in advance.