TCP Connection Life

Tags:

How long can I expect a client/server TCP connection to last in the wild?

I want it to stay permanently connected, but things happen, so the client will have to reconnect. At what point do I say that there's a problem in the code rather than there's a problem with some external equipment?

856

asked Oct 01 '08 17:10

Robert

2 Answers

I agree with Zan Lynx. There's no guarantee, but you can keep a connection alive almost indefinitely by sending data over it, assuming there are no connectivity or bandwidth issues.

Generally I've gone for the application level keep-alive approach, although this has usually because it's been in the client spec so I've had to do it. But just send some short piece of data every minute or two, to which you expect some sort of acknowledgement.

Whether you count one failure to acknowledge as the connection having failed is up to you. Generally this is what I have done in the past, although there was a case I had wait for three failed responses in a row to drop the connection because the app at the other end of the connection was extremely flaky about responding to "are you there?" requests.

If the connection fails, which at some point it probably will, even with machines on the same network, then just try to reestablish it. If that fails a set number of times then you have a problem. If your connection persistently fails after it's been connected for a while then again, you have a problem. Most likely in both cases it's probably some network issue, rather than your code, or maybe a problem with the TCP/IP stack on your machine (has been known: I encountered issues with this on an old version of QNX--it'd just randomly fall over). Having said that you might have a software problem, and the only way to know for sure is often to attach a debugger, or to get some logging in there. E.g. if you can always connect successfully, but after a time you stop getting ACKs, even after reconnect, then maybe your server is deadlocking, or getting stuck in a loop or something.

What's really useful is to set up a series of long-running tests under a variety of load conditions, from just sending the keep alive are you there?/ack requests and responses, to absolutely battering the server. This will generally give you more confidence about your software components, and can be really useful in shaking out some really weird problems which won't necessarily cause a problem with your connection, although they might result in problems with the transactions taking place. For example, I was once writing a telecoms application server that provided services such as number translation, and we'd just leave it running for days at a time. The thing was that when Saturday came round, for the whole day, it would reject every call request that came in, which amounted to millions of calls, and we had no idea why. It turned out to be because of a single typo in some date conversion code that only caused a problem on Saturdays.

Hope that helps.

139

answered Oct 19 '22 23:10

Bart Read

I think the most important idea here is theory vs. practice.

The original theory was that the connections had no lifetimes. If you had a connection, it stayed open forever, even if there was no traffic, until an event caused it to close.

The new theory is that most OS releases have turned on the keep-alive timer. This means that connections will last forever, as long as the system on the other end responds to an occasional TCP-level exchange.

In reality, many connections will be terminated after time, with a variety of criteria and situations.

Two really good examples are: The remote client is using DHCP, the lease expires, and the IP address changes.

Another example is firewalls, which seem to be increasingly intelligent, and can identify keep-alive traffic vs. real data, and close connections based on any high level criteria, especially idle time.

How you want to implement reconnect logic depends a lot on your architecture, the working environment, and your performance goals.

answered Oct 19 '22 23:10

benc

Related questions
                            
                                The best way to use a DB table as a job queue (a.k.a batch queue or message queue)
                            
                                Append an int to char*
                            
                                Escaping HTML in Rails
                            
                                What's the best technique for exiting from a constructor on an error condition in C++
                            
                                What is a good open source Java SE JTA TransactionManager implementation? [closed]
                            
                                VS2008 - Outputting a different file name for Debug/Release configurations
                            
                                ServletContext.getRequestDispatcher() vs ServletRequest.getRequestDispatcher()
                            
                                Basic Rails 404 Error Page
                            
                                Calculating elapsed time in a C program in milliseconds
                            
                                Creating "pretty" Qt Custom Widgets
                            
                                select count and other records in one single query
                            
                                PHP comments: # vs. //

Donate For Us

If you love us? You can donate to us via Paypal or buy me a coffee so we can maintain and grow! Thank you!

Donate Us With