Graceful Shutdown in Kubernetes: Drain HTTP Requests and Close Database Pools Cleanly

Every Kubernetes deployment ends pods constantly: rolling updates, autoscaling, node drains, spot instance reclaims. If your application treats each of those events as an instant death, users see failed requests, databases accumulate orphaned connections, and background work is silently lost.

A graceful shutdown is the opposite: the application stops taking new work, finishes what it already accepted, releases its resources in the right order, and exits on its own before Kubernetes has to force it.

This guide explains the full lifecycle. You will learn what Kubernetes actually does when it terminates a pod, why a naive SIGTERM handler is not enough, and how to build a shutdown sequence that drains HTTP requests and closes database connection pools. It includes a complete Node.js implementation, a Go equivalent, a matching Deployment manifest, a timeout budget, a test plan, and a troubleshooting table.

Why graceful shutdown matters

Consider a rolling update of an API with three replicas. Kubernetes starts a new pod, waits for it to become ready, then terminates an old one. During that termination, three things can go wrong:

  1. In-flight requests are cut off. A client waiting on a response gets a connection reset or a truncated body.
  2. New requests still arrive. The pod is shutting down, but load balancers and proxies may keep sending traffic for a few seconds, because the update to routing is not instant.