Distributed systems fail in partial, messy ways. A dependency doesn't crash cleanly. It slows down, returns intermittent 503s, or accepts connections and never answers. If your Go service responds to that with "just retry," you can turn a small slowdown into an outage.
This guide shows how to build a circuit breaker and a retry layer from scratch in Go, and how to combine them so they protect your services instead of hurting them. You'll get production-minded code, an explanation of each design decision, a testing strategy, and guidance on when to use an existing library instead.
What You Will Learn
Why naive retries cause retry storms and cascading failures
How a circuit breaker works as a state machine (closed, open, half-open)
How to build a rolling-window circuit breaker in Go with generics
How to implement exponential backoff with full jitter and a retry budget
How to compose both patterns around an HTTP client without causing extra load
How to test time-dependent resilience code deterministically
How to expose state changes through log/slog and Prometheus
The examples target Go 1.22 or newer, because they use math/rand/v2 and generics.
Why Resilience Needs Deliberate Design
In a monolith, a function call either returns or panics. In a microservice architecture, every network call adds a failure mode: latency spikes, connection resets, DNS hiccups, overloaded dependencies, and partial outages.
Cascading failure. Service A calls B, and B calls C. When C slows down, B's goroutines pile up waiting on it. B's latency rises, A's requests to B time out, and A's resources fill up too. One slow dependency spreads upstream until nothing responds.
Retry amplification. Retries multiply load on a struggling service. If a client, a gateway, and a service each make 3 attempts, one user request can become 3 × 3 × 3 = 27 requests against the bottom dependency. A dependency that was merely slow gets overwhelmed and cannot recover.
A circuit breaker addresses the first problem by failing fast once a dependency is clearly unhealthy. Disciplined retries address the second by recovering from brief faults without adding unbounded load. They only work well together if you design how they interact.
Understanding the Circuit Breaker State Machine
A circuit breaker wraps calls to a dependency and tracks their outcomes. It moves between three states:
State
Behavior
Transition
Closed
Calls pass through and outcomes are recorded
Opens when the failure ratio crosses a threshold (given enough traffic)
Open
Calls fail immediately without touching the dependency
Moves to half-open after a cooldown
Half-open
A limited number of probe calls pass through
Closes if probes succeed, reopens if any fail
The open state is the important one. It gives the dependency breathing room, and it frees your own goroutines and connections from waiting on something that won't answer.
Should You Write Your Own or Use a Library?
Mature libraries exist. For example, sony/gobreaker is a widely used Go implementation, and many service meshes and gRPC setups can handle some of this at the infrastructure layer.
Writing your own is worthwhile when:
You need to count failures by your own rules (for example, treat 429 as a failure but not 404)
You want the breaker and retry logic to share a policy and a retry budget
You need tight integration with your telemetry
You want to understand the behavior well enough to tune it under pressure
Prefer a library when your needs are standard and you don't want to own the maintenance. Either way, understanding the mechanics below will help you configure any implementation correctly.
Official Documentation Used in This Guide
All external references are official documentation:
A good breaker answers these questions explicitly:
What counts as a failure? A timeout does. A caller cancelling its own request usually does not. A 404 usually does not. A 503 does.
Over what window? A breaker based on "5 consecutive failures" is fragile at high traffic and slow to react at low traffic. A failure ratio over a rolling window, with a minimum request count, is more stable.
How long to stay open? Long enough for the dependency to recover, short enough to restore service quickly.
How many probes in half-open? Too many and you stampede a recovering service. Too few and recovery is slow.
The Implementation
This breaker uses a ring of time buckets for the rolling window and a generation counter so that results from calls started before a state change cannot corrupt the new state.
package resilience
import (
"context"
"errors"
"sync"
"time"
)
type State int
const (
StateClosed State = iota
StateOpen
StateHalfOpen
)
func (s State) String() string {
switch s {
case StateClosed:
return "closed"
case StateOpen:
return "open"
case StateHalfOpen:
return "half-open"
default:
return "unknown"
}
}
var (
ErrOpen = errors.New("circuit breaker is open")
ErrHalfOpenSaturated = errors.New("circuit breaker half-open probes in flight")
)
// Config controls breaker behavior. Zero values receive sensible defaults.
type Config struct {
Name string
Window time.Duration // rolling window length
Buckets int // number of buckets in the window
MinRequests int // minimum calls in window before tripping
FailureRatio float64 // 0..1; trip at or above this ratio
OpenTimeout time.Duration // cooldown before half-open
HalfOpenMaxCalls int // concurrent probes allowed
HalfOpenRequiredSuccesses int // successes needed to close
// IsFailure decides whether an error counts against the dependency.
IsFailure func(error) bool
// OnStateChange runs while the breaker's lock is held. Keep it fast
// and never call back into the breaker from it.
OnStateChange func(name string, from, to State)
// Now is injectable so tests can control time.
Now func() time.Time
}
func (c *Config) withDefaults() {
if c.Window <= 0 {
c.Window = 30 * time.Second
}
if c.Buckets <= 0 {
c.Buckets = 10
}
if c.MinRequests <= 0 {
c.MinRequests = 20
}
if c.FailureRatio <= 0 || c.FailureRatio > 1 {
c.FailureRatio = 0.5
}
if c.OpenTimeout <= 0 {
c.OpenTimeout = 15 * time.Second
}
if c.HalfOpenMaxCalls <= 0 {
c.HalfOpenMaxCalls = 3
}
if c.HalfOpenRequiredSuccesses <= 0 || c.HalfOpenRequiredSuccesses > c.HalfOpenMaxCalls {
c.HalfOpenRequiredSuccesses = c.HalfOpenMaxCalls
}
if c.IsFailure == nil {
c.IsFailure = defaultIsFailure
}
if c.Now == nil {
c.Now = time.Now
}
}
// By default, any error except the caller cancelling counts as a failure.
func defaultIsFailure(err error) bool {
return err != nil && !errors.Is(err, context.Canceled)
}
type bucket struct{ successes, failures int }
type Breaker struct {
cfg Config
mu sync.Mutex
state State
generation uint64
buckets []bucket
bucketDur time.Duration
head int
headStart time.Time
openUntil time.Time
halfOpenInFlight int
halfOpenSuccesses int
}
func NewBreaker(cfg Config) *Breaker {
cfg.withDefaults()
b := &Breaker{
cfg: cfg,
buckets: make([]bucket, cfg.Buckets),
bucketDur: cfg.Window / time.Duration(cfg.Buckets),
}
if b.bucketDur <= 0 {
b.bucketDur = time.Second
}
b.headStart = cfg.Now()
return b
}
func (b *Breaker) State() State {
b.mu.Lock()
defer b.mu.Unlock()
// Reflect an elapsed cooldown without requiring traffic.
if b.state == StateOpen && !b.cfg.Now().Before(b.openUntil) {
return StateHalfOpen
}
return b.state
}
// advance rotates buckets so the head bucket matches the current time.
func (b *Breaker) advance(now time.Time) {
elapsed := now.Sub(b.headStart)
if elapsed < b.bucketDur {
return
}
steps := int(elapsed / b.bucketDur)
if steps >= len(b.buckets) {
for i := range b.buckets {
b.buckets[i] = bucket{}
}
b.head = 0
b.headStart = now
return
}
for i := 0; i < steps; i++ {
b.head = (b.head + 1) % len(b.buckets)
b.buckets[b.head] = bucket{}
}
b.headStart = b.headStart.Add(time.Duration(steps) * b.bucketDur)
}
func (b *Breaker) totals() (successes, failures int) {
for _, bk := range b.buckets {
successes += bk.successes
failures += bk.failures
}
return
}
// transition must be called with b.mu held.
func (b *Breaker) transition(to State, now time.Time) {
from := b.state
if from == to {
return
}
b.state = to
b.generation++ // invalidates results from calls started earlier
for i := range b.buckets {
b.buckets[i] = bucket{}
}
b.head = 0
b.headStart = now
b.halfOpenInFlight = 0
b.halfOpenSuccesses = 0
if to == StateOpen {
b.openUntil = now.Add(b.cfg.OpenTimeout)
}
if b.cfg.OnStateChange != nil {
b.cfg.OnStateChange(b.cfg.Name, from, to)
}
}
// before decides whether a call may proceed and returns its generation.
func (b *Breaker) before() (uint64, error) {
b.mu.Lock()
defer b.mu.Unlock()
now := b.cfg.Now()
if b.state == StateOpen {
if now.Before(b.openUntil) {
return 0, ErrOpen
}
b.transition(StateHalfOpen, now)
}
switch b.state {
case StateClosed:
b.advance(now)
case StateHalfOpen:
if b.halfOpenInFlight >= b.cfg.HalfOpenMaxCalls {
return 0, ErrHalfOpenSaturated
}
b.halfOpenInFlight++
}
return b.generation, nil
}
// after records the outcome of a call admitted by before.
func (b *Breaker) after(gen uint64, err error) {
failure := b.cfg.IsFailure(err)
b.mu.Lock()
defer b.mu.Unlock()
if gen != b.generation {
return // the state changed while this call was in flight
}
now := b.cfg.Now()
switch b.state {
case StateClosed:
b.advance(now)
if failure {
b.buckets[b.head].failures++
} else {
b.buckets[b.head].successes++
}
s, f := b.totals()
total := s + f
if total >= b.cfg.MinRequests &&
float64(f)/float64(total) >= b.cfg.FailureRatio {
b.transition(StateOpen, now)
}
case StateHalfOpen:
b.halfOpenInFlight--
if failure {
b.transition(StateOpen, now)
return
}
b.halfOpenSuccesses++
if b.halfOpenSuccesses >= b.cfg.HalfOpenRequiredSuccesses {
b.transition(StateClosed, now)
}
}
}
var errPanic = errors.New("call panicked")
// Do runs fn through the breaker. It is a function rather than a method
// because Go methods cannot declare their own type parameters.
func Do[T any](ctx context.Context, b *Breaker, fn func(context.Context) (T, error)) (T, error) {
var zero T
if err := ctx.Err(); err != nil {
return zero, err
}
gen, err := b.before()
if err != nil {
return zero, err
}
completed := false
defer func() {
if !completed { // fn panicked; record a failure and let the panic continue
b.after(gen, errPanic)
}
}()
res, err := fn(ctx)
completed = true
b.after(gen, err)
return res, err
}
Design Decisions Worth Understanding
Rolling window over consecutive failures. Consecutive-failure counters trip too easily on noisy but healthy dependencies and react slowly in other cases. A ratio with MinRequests avoids tripping on two failures out of three calls during a quiet period.
Generation counter. Imagine 50 calls in flight when the breaker opens. When they finish, their results describe the old world. Without a generation check, late successes could count toward closing a half-open breaker, or late failures could immediately reopen it. Bumping the generation on every transition makes stale results harmless.
Caller cancellation is not a dependency failure. If your user closes the browser tab and the context is cancelled, the dependency did nothing wrong. Counting that would make breakers trip because of impatient clients. context.DeadlineExceeded, by contrast, usually does indicate slowness and counts as a failure by default. See the context package documentation for the semantics of cancellation versus deadlines.
Panic safety. The deferred handler records a failure if fn panics, so a half-open probe slot is never leaked.
Callback under lock.OnStateChange runs while the mutex is held, which keeps transitions atomic and ordered. The cost is that callbacks must be fast and must not re-enter the breaker.
Designing Retries That Help Instead of Hurt
A retry is only safe and useful when all of these hold:
The failure is transient. Retrying a 400 Bad Request or a validation error just repeats the failure.
The operation is idempotent. Retrying GET is generally safe. Retrying a payment POST can double-charge unless you use idempotency keys. A unique key per logical operation lets the server deduplicate; the Bulk UUID Generator is handy for producing sample keys while you develop and test.
Delays are randomized. Without jitter, thousands of clients that failed together retry together, producing synchronized spikes.
Total retry volume is bounded. Per-request attempt limits aren't enough. When every request is failing, a cap of 3 attempts still triples your traffic.
The caller's deadline is respected. Don't sleep through a backoff you cannot afford.
Backoff With Full Jitter
Exponential backoff doubles the delay cap on each attempt. "Full jitter" then picks a random delay between zero and that cap. This spreads retry load evenly instead of clustering it, and it is a widely recommended approach for contended systems.
The Retry Budget
A retry budget limits retries to a fraction of normal traffic. Each original request adds a small amount of credit, such as 0.1 tokens, and each retry spends one whole token. When the dependency is healthy, retries are rare and the budget stays full. When everything fails, the budget runs dry and retries stop, capping amplification at roughly 10% extra load instead of 200%.
package resilience
import (
"context"
"errors"
"fmt"
"math/rand/v2"
"sync"
"time"
)
// RetryBudget caps retries as a fraction of overall request volume.
type RetryBudget struct {
mu sync.Mutex
tokens float64
max float64
ratio float64
}
// NewRetryBudget creates a budget. ratio is tokens earned per request
// (0.1 means about one retry per ten requests); max is the burst ceiling.
func NewRetryBudget(ratio, max float64) *RetryBudget {
return &RetryBudget{tokens: max, max: max, ratio: ratio}
}
func (rb *RetryBudget) onRequest() {
rb.mu.Lock()
defer rb.mu.Unlock()
rb.tokens = min(rb.max, rb.tokens+rb.ratio)
}
func (rb *RetryBudget) tryAcquire() bool {
rb.mu.Lock()
defer rb.mu.Unlock()
if rb.tokens >= 1 {
rb.tokens--
return true
}
return false
}
// RetryAfterError lets a callee suggest a minimum wait, such as from a
// Retry-After header.
type RetryAfterError struct {
Err error
After time.Duration
}
func (e *RetryAfterError) Error() string { return e.Err.Error() }
func (e *RetryAfterError) Unwrap() error { return e.Err }
type RetryPolicy struct {
MaxAttempts int // total attempts including the first
BaseDelay time.Duration // first backoff cap
MaxDelay time.Duration // ceiling for any single wait
Retryable func(error) bool
Budget *RetryBudget // optional
}
func (p RetryPolicy) normalized() RetryPolicy {
if p.MaxAttempts <= 0 {
p.MaxAttempts = 3
}
if p.BaseDelay <= 0 {
p.BaseDelay = 100 * time.Millisecond
}
if p.MaxDelay <= 0 {
p.MaxDelay = 5 * time.Second
}
if p.Retryable == nil {
p.Retryable = func(err error) bool {
return !errors.Is(err, ErrOpen) && !errors.Is(err, context.Canceled)
}
}
return p
}
// backoff returns a full-jitter delay: random in [0, min(max, base*2^attempt)).
func (p RetryPolicy) backoff(attempt int) time.Duration {
if attempt > 30 { // avoid shift overflow
attempt = 30
}
ceiling := min(p.MaxDelay, p.BaseDelay<<attempt)
if ceiling <= 0 {
return 0
}
return rand.N(ceiling)
}
func sleep(ctx context.Context, d time.Duration) error {
t := time.NewTimer(d)
defer t.Stop()
select {
case <-ctx.Done():
return ctx.Err()
case <-t.C:
return nil
}
}
// Retry runs fn up to MaxAttempts times, honoring the context deadline,
// the retry budget, and any server-provided Retry-After hint.
func Retry[T any](ctx context.Context, p RetryPolicy, fn func(ctx context.Context, attempt int) (T, error)) (T, error) {
p = p.normalized()
var zero T
if p.Budget != nil {
p.Budget.onRequest()
}
var lastErr error
for attempt := 0; attempt < p.MaxAttempts; attempt++ {
res, err := fn(ctx, attempt)
if err == nil {
return res, nil
}
lastErr = err
if attempt == p.MaxAttempts-1 || !p.Retryable(err) {
break
}
if p.Budget != nil && !p.Budget.tryAcquire() {
return zero, fmt.Errorf("retry budget exhausted after %d attempt(s): %w", attempt+1, lastErr)
}
delay := p.backoff(attempt)
var ra *RetryAfterError
if errors.As(err, &ra) && ra.After > delay {
delay = min(ra.After, p.MaxDelay)
}
// Don't sleep past the caller's deadline.
if dl, ok := ctx.Deadline(); ok && time.Until(dl) <= delay {
return zero, fmt.Errorf("deadline too close to retry after %d attempt(s): %w", attempt+1, lastErr)
}
if err := sleep(ctx, delay); err != nil {
return zero, fmt.Errorf("retry interrupted after %d attempt(s): %w", attempt+1, errors.Join(err, lastErr))
}
}
return zero, fmt.Errorf("giving up after %d attempt(s): %w", p.MaxAttempts, lastErr)
}
Details to notice:
time.NewTimer with Stop avoids the leak you would get from time.After in a loop on older Go versions. See the time package documentation for timer behavior.
Errors are wrapped with %w, so callers can still use errors.Is and errors.As. See the errors package documentation.
Composing the Breaker and the Retrier
The order in which you nest these patterns changes system behavior.
Retry outside, breaker inside (recommended here). Each attempt passes through the breaker. The breaker sees the true load and failure rate, and once it opens, ErrOpen is marked non-retryable, so the retry loop stops immediately instead of sleeping and trying again against a dependency known to be unhealthy.
Breaker outside, retry inside. The breaker counts one logical failure after all retries are exhausted. This hides real attempt volume from the breaker and delays tripping, and the dependency absorbs all the retries before the breaker reacts.
For most services, the first arrangement is safer. The code below uses it.
A Resilient HTTP Client
This client combines both layers, applies a per-attempt timeout, treats only the right status codes as failures, and honors Retry-After.
package resilience
import (
"context"
"encoding/json"
"errors"
"fmt"
"io"
"net"
"net/http"
"strconv"
"time"
)
// StatusError carries an HTTP status for classification decisions.
type StatusError struct {
Code int
RetryAfter time.Duration
}
func (e *StatusError) Error() string { return fmt.Sprintf("unexpected status %d", e.Code) }
type Client struct {
HTTP *http.Client
Breaker *Breaker
Policy RetryPolicy
AttemptTimeout time.Duration
}
// NewClient wires sensible classification rules into the breaker and retrier.
func NewClient(name string, onChange func(string, State, State)) *Client {
breaker := NewBreaker(Config{
Name: name,
IsFailure: isDependencyFailure,
OnStateChange: onChange,
})
return &Client{
HTTP: &http.Client{},
Breaker: breaker,
Policy: RetryPolicy{
MaxAttempts: 3,
BaseDelay: 100 * time.Millisecond,
MaxDelay: 2 * time.Second,
Retryable: isRetryable,
Budget: NewRetryBudget(0.1, 10),
},
AttemptTimeout: 2 * time.Second,
}
}
// 5xx and 429 indicate dependency trouble; other 4xx are caller errors.
func isDependencyFailure(err error) bool {
if err == nil || errors.Is(err, context.Canceled) {
return false
}
var se *StatusError
if errors.As(err, &se) {
return se.Code >= 500 || se.Code == http.StatusTooManyRequests
}
return true // network errors, timeouts, decode failures
}
func isRetryable(err error) bool {
if err == nil || errors.Is(err, ErrOpen) || errors.Is(err, ErrHalfOpenSaturated) ||
errors.Is(err, context.Canceled) {
return false
}
var se *StatusError
if errors.As(err, &se) {
switch se.Code {
case http.StatusTooManyRequests, http.StatusBadGateway,
http.StatusServiceUnavailable, http.StatusGatewayTimeout:
return true
}
return false
}
var ne net.Error
if errors.As(err, &ne) {
return true
}
return errors.Is(err, context.DeadlineExceeded)
}
func parseRetryAfter(v string) time.Duration {
if v == "" {
return 0
}
if secs, err := strconv.Atoi(v); err == nil && secs >= 0 {
return time.Duration(secs) * time.Second
}
if t, err := http.ParseTime(v); err == nil {
if d := time.Until(t); d > 0 {
return d
}
}
return 0
}
// GetJSON performs an idempotent GET and decodes the JSON body into out.
func (c *Client) GetJSON(ctx context.Context, url string, out any) error {
_, err := Retry(ctx, c.Policy, func(ctx context.Context, _ int) (struct{}, error) {
return Do(ctx, c.Breaker, func(ctx context.Context) (struct{}, error) {
return struct{}{}, c.attempt(ctx, url, out)
})
})
return err
}
func (c *Client) attempt(ctx context.Context, url string, out any) error {
ctx, cancel := context.WithTimeout(ctx, c.AttemptTimeout)
defer cancel()
req, err := http.NewRequestWithContext(ctx, http.MethodGet, url, nil)
if err != nil {
return err
}
resp, err := c.HTTP.Do(req)
if err != nil {
return err
}
defer resp.Body.Close()
if resp.StatusCode >= 400 {
// Drain a bounded amount so the connection can be reused.
_, _ = io.Copy(io.Discard, io.LimitReader(resp.Body, 1<<20))
se := &StatusError{Code: resp.StatusCode}
if resp.StatusCode == http.StatusTooManyRequests || resp.StatusCode == http.StatusServiceUnavailable {
se.RetryAfter = parseRetryAfter(resp.Header.Get("Retry-After"))
}
if se.RetryAfter > 0 {
return &RetryAfterError{Err: se, After: se.RetryAfter}
}
return se
}
return json.NewDecoder(resp.Body).Decode(out)
}
Why These Choices
The per-attempt timeout is separate from the overall deadline. A single hung attempt cannot consume the caller's whole budget, so retries remain possible.
404 and 400 don't trip the breaker. They describe the request, not the dependency's health.
Response bodies are always closed and drained. Skipping this leaks connections and quietly degrades throughput. See the net/http documentation for client and body-handling guidance.
Only GET is wrapped. Retrying non-idempotent operations requires idempotency keys and server support, which is a separate design decision.
ErrOpen is not retryable. This is the line that stops the two patterns from fighting each other.
If you debug upstream error payloads while building this, the JSON Formatter & Validator is a quick way to inspect and validate responses.
Usage Example
package main
import (
"context"
"log/slog"
"os"
"time"
"example.com/yourapp/resilience"
)
type Profile struct {
ID string `json:"id"`
Name string `json:"name"`
}
func main() {
logger := slog.New(slog.NewJSONHandler(os.Stdout, nil))
client := resilience.NewClient("profile-service",
func(name string, from, to resilience.State) {
logger.Warn("circuit breaker state change",
slog.String("breaker", name),
slog.String("from", from.String()),
slog.String("to", to.String()))
})
ctx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
defer cancel()
var p Profile
if err := client.GetJSON(ctx, "https://profiles.internal/v1/users/42", &p); err != nil {
logger.Error("profile lookup failed", slog.Any("error", err))
return // or serve a degraded fallback
}
logger.Info("profile loaded", slog.String("name", p.Name))
}
Structured logging with log/slog keeps state transitions searchable. See the log/slog documentation.
Fallbacks: What To Do When the Breaker Is Open
An open breaker is only useful if the caller handles it well. Options, in rough order of preference:
Serve cached or stale data when freshness isn't critical.
Return a degraded response, such as a profile page without recommendations.
Queue the work for later if the operation is asynchronous by nature.
Return a clear error with a suitable status code, such as 503, and let the caller decide.
Run with the race detector. Breakers are shared mutable state, so run go test -race ./.... See the Go race detector documentation.
Use httptest for end-to-end behavior. The net/http/httptest package lets you build a server that fails a configurable number of times, then recovers, to verify retries and the breaker together.
Test the unhappy paths: panics in fn, context cancellation mid-backoff, and a half-open breaker receiving concurrent calls.
Load test the combined behavior to confirm your retry budget actually caps amplification. A methodology for measuring latency and bottlenecks under load is covered in Load Testing LLM Streaming Endpoints: TTFT & Memory, and the principles carry over to Go services.
Observability: You Can't Tune What You Can't See
Resilience patterns that fail silently are dangerous. At minimum, expose:
Breaker state (as a gauge) and transition counts
Rejected calls (ErrOpen count)
Retry attempts and budget exhaustion events
Per-attempt latency and outcome
Here's a small example using the official Prometheus Go client:
Alert on sustained open state and on retry-budget exhaustion rates, not on every transition. Brief trips are the system doing its job.
Tuning Guidance
There are no universal numbers. Treat the defaults above as starting points and adjust based on measurement.
Parameter
Too low
Too high
FailureRatio
Trips on normal noise
Reacts too slowly to real outages
MinRequests
Trips on tiny samples
Never trips on low-traffic paths
Window
Jumpy, over-sensitive
Slow to notice change
OpenTimeout
Hammers a recovering service
Prolongs the outage for users
HalfOpenMaxCalls
Slow recovery
Probe stampede
MaxAttempts
Fragile to brief blips
Higher amplification risk
Retry budget ratio
Few recoveries from blips
Weak protection during outages
A practical approach: derive per-attempt timeouts from the dependency's observed p99 latency, set the overall deadline from your own SLO, and size MaxAttempts so that attempts × per-attempt timeout fits inside that deadline.
Common Mistakes
Retrying at every layer. Pick one layer, usually the one closest to the call, to own retries. Layered retries multiply (3 × 3 × 3 = 27).
Retrying non-idempotent operations without idempotency keys.
Using fixed delays with no jitter.
Counting 4xx client errors as dependency failures.
Having no per-attempt timeout, so a single hung call burns the entire deadline.
One global breaker for everything. Use a breaker per dependency, or even per route, so one bad endpoint doesn't block healthy ones.
Forgetting to close or drain response bodies.
Ignoring Retry-After from an overloaded dependency that is asking you to slow down.
Not testing the open and half-open paths until production forces you to.
Related Patterns That Complete the Picture
Circuit breakers and retries are two parts of a broader resilience toolkit:
Timeouts and deadlines are the foundation. Every outbound call needs one, and Go's context makes propagating them idiomatic. The Go team's context blog post is the canonical introduction.
Bulkheads isolate resources, for example a bounded semaphore per dependency so one slow service can't consume all your goroutines or connections.
Service-level retry policies are supported natively by gRPC. If you use gRPC, review the official gRPC retry guide before layering your own retries on top, to avoid double retries.
Frequently Asked Questions
What is a circuit breaker in Go?
A circuit breaker is a wrapper around calls to a dependency that tracks failures and temporarily stops sending requests once the dependency looks unhealthy. In Go it is usually implemented as a small state machine guarded by a mutex, with closed, open, and half-open states.
Should retries be inside or outside the circuit breaker?
Generally outside. Each attempt passes through the breaker so it sees real load, and an open circuit stops the retry loop immediately. Placing retries inside hides attempt volume from the breaker and delays tripping.
What is a retry budget?
A retry budget limits retries to a fraction of total traffic, for instance 10%. It prevents retry storms by ensuring that when most calls fail, retries cannot multiply the load on an already struggling dependency.
Why use jitter in exponential backoff?
Without jitter, all clients that failed at the same moment retry at the same moments, creating synchronized load spikes. Randomizing delays spreads retries over time.
Should a timeout count as a circuit breaker failure?
Usually yes. Deadline-exceeded errors typically indicate a slow dependency. A caller cancelling its own request, however, is not the dependency's fault and shouldn't count.
When should I use a library instead of building my own?
Use a library when your requirements are standard and you want to avoid maintaining the code. Build your own when you need custom failure classification, shared retry budgets, or deep telemetry integration. Either way, review the implementation's defaults, since poorly tuned defaults are a common cause of surprises.
Do I still need these patterns with a service mesh?
A mesh can provide retries, timeouts, and outlier detection at the infrastructure layer, which is valuable. Application-level patterns still matter for semantic decisions, such as which errors are retryable and what fallback to serve. If you use both, make sure retries aren't configured at both layers.
Key Takeaways
Retries without limits amplify failures, and breakers without careful failure classification trip for the wrong reasons.
Use a rolling-window failure ratio with a minimum request count, and a generation counter to ignore stale results.
Retry only transient errors on idempotent operations, with exponential backoff, full jitter, a retry budget, and deadline awareness.
Nest retry outside, breaker inside, and make ErrOpen non-retryable.
Inject a clock so you can test state transitions deterministically, and run tests with -race.
Instrument everything. Resilience you cannot observe is resilience you cannot trust.
Treat all numeric settings as starting points and tune them from real measurements in your own environment.
The code in this article is educational and intended as a starting point. Review, test, and load test it against your own workloads and failure modes before using it in production.
3ResizeObserver Guide: Watch Element Size in JS & React