Distributed systems fail in partial, messy ways. A dependency doesn't crash cleanly. It slows down, returns intermittent 503s, or accepts connections and never answers. If your Go service responds to that with "just retry," you can turn a small slowdown into an outage.

This guide shows how to build a circuit breaker and a retry layer from scratch in Go, and how to combine them so they protect your services instead of hurting them. You'll get production-minded code, an explanation of each design decision, a testing strategy, and guidance on when to use an existing library instead.

What You Will Learn

  • Why naive retries cause retry storms and cascading failures
  • How a circuit breaker works as a state machine (closed, open, half-open)
  • How to build a rolling-window circuit breaker in Go with generics
  • How to implement exponential backoff with full jitter and a retry budget
  • How to compose both patterns around an HTTP client without causing extra load
  • How to test time-dependent resilience code deterministically
  • How to expose state changes through log/slog and Prometheus

The examples target Go 1.22 or newer, because they use math/rand/v2 and generics.

Why Resilience Needs Deliberate Design

In a monolith, a function call either returns or panics. In a microservice architecture, every network call adds a failure mode: latency spikes, connection resets, DNS hiccups, overloaded dependencies, and partial outages.