LLM Streaming Error Handling in Python: Retries & Failures
Learn how to handle LLM streaming errors in Python: classify failures, retry safely, survive mid-stream drops, and ship FastAPI code that fails cleanly.
Handling LLM Streaming Errors, Retries, and Mid-Stream Failures in Python
Streaming makes an LLM feature feel fast. It also changes how failures work. With a normal API call, a request either succeeds or fails, and you get a clear HTTP status either way. With a streamed response, the server can send a 200 OK, deliver forty tokens, and then fail. By that point the status line is already on the wire, the user is already reading, and your retry logic has nothing clean to retry.
This guide covers how to handle that in a Python backend: how to classify failures, what is safe to retry, how to detect stalled streams, how to tell the client what happened, and how to test all of it. The examples use FastAPI and the Anthropic Python SDK, but the patterns apply to any provider and any ASGI framework.
If you haven't read the first article in this series, Optimizing LLM Streaming Payloads in Python Backends covers protocol choice, buffering, and backpressure. This one picks up where it stops: what happens when the stream breaks.
Key takeaways
Streaming has two failure phases: before the first token and after it. They need different handling.
Retry automatically only while nothing has been sent to the user.
Once tokens are out, fail cleanly with a typed error event instead of retrying silently.
Add an idle timeout between chunks. A stalled stream is a failure that never raises an exception.
Always clean up the upstream connection when the client leaves.
Test failure paths with a fake upstream. You can't reliably reproduce these bugs against a real provider.
Why streaming errors are different
In a non-streaming call, the HTTP status tells you what happened. A rate limit is a 429, an overloaded service is a 529 or 503, and a bad request is a 400. Your error handler maps each status to a response.
Streaming breaks that model. Once the server starts sending the event stream, the HTTP status is committed. Anthropic's streaming documentation states that the API may occasionally send errors inside the event stream itself, and gives overloaded_error as an example. In a non-streaming call that error would normally be an HTTP 529. Anthropic's docs also warn that new event types may be added over time, so your code should tolerate event types it doesn't recognize.
There are three practical consequences:
A 200 doesn't mean success. You have to watch the stream to know the outcome.
Failures arrive in-band. They show up as error events, dropped connections, or silence.
Partial output creates state. Once a user has seen half a paragraph, "just retry" is no longer a neutral action.
A taxonomy of streaming failures
Before writing code, name the failures. Each one has a different cause and a different right response.
Failure
When it happens
Typical signal
Safe to auto-retry?
Rejected request
Before the stream opens
HTTP 400/401/403/404
No. Fix the request.
Rate limited
Before the stream opens
HTTP 429, sometimes Retry-After
Yes, after a delay.
Provider overloaded or unavailable
Before or during the stream
HTTP 5xx, or an in-stream error event
Only if no tokens were sent.
Connection failure
Before or during the stream
Connection reset, DNS or TLS error
Only if no tokens were sent.
Stall
During the stream
No events for a long time
Only if no tokens were sent.
Malformed event
During the stream
Parse error in the SDK
Treat as a failure.
Client disconnect
During the stream
Cancelled task, failed write
No. Stop and clean up.
Truncation
End of stream
Stop reason such as max_tokens
Not an error. Report it.
The last row matters. Hitting the output limit is a normal end to a generation, not a failure. Anthropic documents stop reasons in its API guides, and your done event should carry the finish reason so the client can show a "response was cut off" hint instead of an error.
The core rule: only retry what nobody has seen
The most important design decision is where you draw the retry line.
Before the first token reaches the client, retrying is safe. The user has seen nothing, so a second attempt is invisible. You can also retry against a different region or model.
After the first token, an automatic retry means one of two bad outcomes. Either the user sees the answer restart from the top, or you stitch two different generations together and get text that contradicts itself.
So the policy is simple: retry silently until the first chunk is emitted, then stop retrying and report failures explicitly. Everything below implements that rule.
SDK-level retries help with the first half. Both the Anthropic and OpenAI Python SDKs retry certain failed requests by default and expose a max_retries option. In general, that covers opening the connection. It can't unsend tokens you have already delivered, so mid-stream failures still need your own handling.
Building the resilient stream in Python
The code below needs Python 3.11 or newer, because it uses asyncio.timeout. The design has four small parts: a failure classifier, a backoff function, a wrapper around the provider SDK, and a resilient generator that enforces the retry rule.
Classify failures once
Put the "is this retryable?" decision in one function. Scattering it through your code is how retries end up on 400 errors.
# app/streaming.py
import asyncio
import random
from dataclasses import dataclass
from typing import AsyncIterator, Callable
import anthropic
@dataclass(frozen=True)
class Failure:
code: str # safe, stable identifier sent to clients
retryable: bool
retry_after: float | None = None
class StreamFailure(Exception):
"""Raised when a stream cannot be completed."""
def __init__(self, code: str, partial: bool):
super().__init__(code)
self.code = code
self.partial = partial # True if some tokens were already yielded
def _retry_after(exc: anthropic.APIStatusError) -> float | None:
raw = exc.response.headers.get("retry-after")
try:
return float(raw) if raw else None
except ValueError: # Retry-After may also be an HTTP date; ignore it here
return None
def classify(exc: Exception) -> Failure:
if isinstance(exc, TimeoutError): # our idle watchdog
return Failure("upstream_timeout", True)
if isinstance(exc, anthropic.APIStatusError):
status = exc.status_code
if status == 429:
return Failure("rate_limited", True, _retry_after(exc))
if status >= 500: # includes 529 overloaded
return Failure("upstream_unavailable", True, _retry_after(exc))
if status in (408, 409):
return Failure("upstream_timeout", True)
return Failure("upstream_rejected", False)
if isinstance(exc, anthropic.APIConnectionError): # includes SDK timeouts
return Failure("upstream_unreachable", True)
return Failure("internal_error", False)
Anthropic's error documentation lists the status codes and error types, including overloaded_error for 529 and api_error for 500. It also recommends catching the SDK's typed exception classes rather than string-matching messages, which is what the classifier does. The SDK typically surfaces an in-stream error event as an exception too, so one classifier covers both phases. Verify that against the SDK version you pin.
The classifier deliberately returns a short code string rather than the raw exception message. Provider error text can contain request details you don't want to forward to browsers.
Back off with jitter, and respect Retry-After
Fixed delays cause synchronized retry storms: every client that failed together retries together. Exponential backoff with full jitter spreads them out. If the provider sends a Retry-After header, treat it as a floor. The header is defined in RFC 9110, and OpenAI's rate-limit guide also recommends exponential backoff with randomness.
def backoff_delay(attempt: int, retry_after: float | None = None,
base: float = 0.5, cap: float = 8.0) -> float:
delay = random.uniform(0, min(cap, base * 2 ** attempt))
if retry_after:
delay = max(delay, retry_after)
return min(delay, 30.0) # never make a user wait longer than this
Wrap the provider call
Keep the SDK behind a tiny async generator. This isolates the provider-specific code, and it lets tests swap in a fake.
client = anthropic.AsyncAnthropic() # reads ANTHROPIC_API_KEY from the environment
MODEL = "your-model-name" # load from config, not hard-coded
async def anthropic_tokens(prompt: str) -> AsyncIterator[str]:
async with client.messages.stream(
model=MODEL,
max_tokens=1024,
messages=[{"role": "user", "content": prompt}],
) as stream:
async for text in stream.text_stream:
yield text
The resilient generator
This function enforces the retry rule and adds the idle watchdog.
async def resilient_stream(
open_upstream: Callable[[], AsyncIterator[str]],
*,
max_attempts: int = 3,
idle_timeout: float = 30.0,
) -> AsyncIterator[str]:
attempt = 0
while True:
emitted = False
upstream = open_upstream()
try:
while True:
try:
async with asyncio.timeout(idle_timeout):
chunk = await anext(upstream)
except StopAsyncIteration:
return
emitted = True
yield chunk
except Exception as exc: # CancelledError is not an Exception
info = classify(exc)
attempt += 1
if emitted or not info.retryable or attempt >= max_attempts:
raise StreamFailure(info.code, partial=emitted) from exc
await asyncio.sleep(backoff_delay(attempt, info.retry_after))
finally:
await upstream.aclose()
A few details are worth understanding:
emitted flips to True on the first chunk. From then on, every failure becomes a StreamFailure with partial=True instead of a retry.
The yield sits outside the timeout block. The idle timer measures how long the upstream takes to produce the next chunk, not how long your client takes to consume it.
asyncio.CancelledError derives from BaseException in modern Python, so except Exception doesn't swallow it. Cancellation propagates, which is what you want when a client disconnects. The finally block still closes the upstream stream.
Pick idle_timeout with your model in mind. Models that reason before answering can pause for a while between the request and the first token, so a value that is right for a fast chat model may be too tight for a reasoning model. Measure your real gaps before tightening it.
The Python documentation for asyncio.timeout explains how the context manager converts a cancellation into a TimeoutError, which is why the classifier can treat it like any other failure.
Turn pre-stream failures into real HTTP errors
Here is a useful trick. Streaming only commits the HTTP status once you return the response and start writing. If you pull the first chunk before returning, then any failure that happens before the first token can still be a proper HTTP error: a 429, a 503, a 504. Load balancers, API clients, and monitoring all understand those. OpenRouter's documentation describes the same split between pre-stream and mid-stream errors and notes that only the pre-stream phase allows normal error statuses.
# app/main.py
import json
import logging
import uuid
from fastapi import FastAPI, HTTPException
from fastapi.responses import StreamingResponse
from pydantic import BaseModel, Field
from app.streaming import StreamFailure, anthropic_tokens, resilient_stream
app = FastAPI()
log = logging.getLogger("llm.stream")
HTTP_STATUS = {
"rate_limited": 429,
"upstream_timeout": 504,
"upstream_unavailable": 503,
"upstream_unreachable": 503,
}
class ChatRequest(BaseModel):
prompt: str = Field(min_length=1, max_length=8000)
def sse(event: str, data: dict, event_id: str | None = None) -> str:
head = f"id: {event_id}\n" if event_id else ""
return f"{head}event: {event}\ndata: {json.dumps(data)}\n\n"
async def events(upstream, first, request_id: str):
sent = 0
try:
if first is not None:
sent += 1
yield sse("token", {"text": first}, f"{request_id}:{sent}")
async for chunk in upstream:
sent += 1
yield sse("token", {"text": chunk}, f"{request_id}:{sent}")
yield sse("done", {"request_id": request_id, "chunks": sent})
except StreamFailure as failure:
log.warning("stream_failed", extra={"request_id": request_id,
"code": failure.code, "chunks": sent})
yield sse("error", {"code": failure.code, "request_id": request_id,
"partial": failure.partial})
except Exception:
log.exception("stream_crashed", extra={"request_id": request_id})
yield sse("error", {"code": "internal_error", "request_id": request_id,
"partial": sent > 0})
finally:
await upstream.aclose()
@app.post("/v1/chat/stream")
async def chat_stream(body: ChatRequest):
request_id = uuid.uuid4().hex
upstream = resilient_stream(lambda: anthropic_tokens(body.prompt))
try:
first = await anext(upstream) # retries happen inside, invisibly
except StopAsyncIteration:
first = None
except StreamFailure as failure:
raise HTTPException(
status_code=HTTP_STATUS.get(failure.code, 502),
detail={"code": failure.code, "request_id": request_id},
)
return StreamingResponse(
events(upstream, first, request_id),
media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"},
)
Every stream ends with exactly one terminal event, either done or error. A client that sees the connection close without either knows something went wrong at the network layer.
The error event never carries raw exception text. It carries a stable code and a request ID. Your logs hold the detail.
Design the error contract for your frontend
The frontend can only behave well if the backend's contract is explicit. Here is a small, stable one:
Event
Payload
Meaning
token
{"text": "..."}
A piece of the response.
done
{"request_id", "chunks"}
The stream completed. Add a finish reason if you track one.
error
{"code", "request_id", "partial"}
The stream failed. partial says whether text was already shown.
The partial flag drives the UI. If it is false, the client can safely offer "Try again" or retry automatically. If it is true, the client should keep the text it has, mark it as incomplete, and let the user decide.
SSE has some built-in reconnection behavior you should know about. The WHATWG specification defines the id: field and the Last-Event-ID request header, and MDN's guide to server-sent events explains that the browser's EventSource reconnects on its own after a dropped connection. That's convenient for feeds, but it's a trap for one-shot LLM generations. A silent reconnect would start a brand-new generation unless your server deliberately supports resuming. Many teams use a fetch-based stream reader, which gives explicit control over reconnection, and use EventSource only where resume is implemented.
When you debug payloads during development, pasting a captured data: line into the JSON Formatter & Validator is a quick way to check that your events are valid JSON.
Detect stalls and keep idle connections alive
Some failures never raise an exception. A network path silently drops packets, a provider hangs, or a proxy stops forwarding, and your code waits forever. The idle timeout in resilient_stream handles the upstream side. There are two more places to guard.
HTTP client timeouts. Set explicit connect, read, and write timeouts on the client your SDK uses. The HTTPX timeout documentation explains the four timeout types. Note that a read timeout measures the gap between received chunks, not the total request duration, which suits streaming.
Downstream heartbeats. Reverse proxies and load balancers often close connections that carry no bytes for a while. If a generation has a long pause before the first token, send an SSE comment line such as : keep-alive every 15 seconds or so. Comment lines start with a colon and are ignored by SSE parsers, so they cost almost nothing. Heartbeats also help the server notice a dead client sooner, because a failed write raises an error.
Handle client disconnects and stop paying for tokens
When a user closes the tab, your server should stop generating. A stream nobody reads still consumes provider capacity, and in most cases it still costs money.
In Starlette-based frameworks, a disconnect during a streaming response typically shows up as the response task being cancelled or as a failed write. Behavior differs across server and framework versions, so test it on your own stack. Two habits make cleanup reliable:
Use async with around the upstream stream, as anthropic_tokens does, so leaving the block closes the HTTP connection and ends generation upstream.
Put aclose() in a finally block, as both resilient_stream and events do. Never catch CancelledError and continue. If you must catch it to run cleanup, re-raise it afterward.
Don't rely on polling request.is_disconnected() inside your generator as your only mechanism. Polling adds latency and CPU cost, and it can be unreliable while a response body is being written. Structured cleanup is more dependable.
Recovering from mid-stream failures
Once tokens have reached the user and the stream dies, you have four options. None is perfect.
1. Fail clearly (the default). Send an error event with partial: true. The client keeps the visible text, marks it incomplete, and offers a retry. This is honest, simple, and the right choice for most products.
2. Continue from the partial text. For plain prose, you can start a new request that includes what was already generated and asks the model to carry on:
This works reasonably for essays or summaries, but it can produce a visible seam or slight repetition, and the continuation may drift in tone. Use it only for low-stakes text, and keep the cut point clean by trimming any half-finished sentence.
3. Resume from a server-side buffer. If the failure was the client's connection rather than the provider's, the generation may still be running. Store emitted chunks under the request ID, and when the client reconnects with Last-Event-ID, replay what it missed and keep streaming. This gives the best user experience but needs shared storage and lifecycle management. A later article in this series covers it in the context of queues and Redis.
4. Regenerate and replace. The client discards the partial answer and starts over. It is simple and consistent, but it costs extra tokens and can look jarring.
Note that a failed generation may still be billed for the tokens produced before the failure. Check your provider's billing documentation rather than assuming otherwise, and count retries in your cost dashboards.
Special cases: tool calls, structured output, and stop reasons
Tool calls. Never resume a half-streamed tool call. A partial JSON argument string is not a valid call, and executing a guess can cause real side effects. Discard the incomplete call, and retry the whole turn only if the tool is idempotent.
Structured output. Streamed JSON is not valid until it is complete. If the stream dies partway, treat the whole object as failed. Do not try to repair truncated JSON and hand it to downstream code. A later article covers this topic in depth.
Truncation. A stop reason such as max_tokens means the model hit the limit you set. Report it in done with a finish reason so the UI can say the answer was cut short, and consider raising the limit rather than retrying.
Refusals and moderation. A model declining a request or a provider filter ending a stream is not a transient error. Retrying it burns tokens and won't change the outcome.
Test failure paths with a fake upstream
You can't reliably trigger a mid-stream overloaded_error on demand, so test against a fake. Because resilient_stream takes a factory, this is easy.
# tests/test_streaming.py
import asyncio
import pytest
from app import streaming as mod
from app.streaming import StreamFailure, resilient_stream
async def flaky(tokens, fail_at, exc):
for i, token in enumerate(tokens):
if i == fail_at:
raise exc
yield token
def collect(factory):
async def run():
return [chunk async for chunk in resilient_stream(factory)]
return asyncio.run(run())
def test_retries_before_first_token(monkeypatch):
monkeypatch.setattr(mod, "backoff_delay", lambda *a, **k: 0)
calls = 0
def factory():
nonlocal calls
calls += 1
return flaky(["a", "b"], 0 if calls == 1 else None, TimeoutError())
assert collect(factory) == ["a", "b"]
assert calls == 2
def test_no_retry_after_first_token(monkeypatch):
monkeypatch.setattr(mod, "backoff_delay", lambda *a, **k: 0)
calls = 0
def factory():
nonlocal calls
calls += 1
return flaky(["a", "b", "c"], 1, TimeoutError())
with pytest.raises(StreamFailure) as info:
collect(factory)
assert info.value.partial is True
assert calls == 1
These two tests pin down the most important behavior: retry before the first token, never after. Add more for the cases you care about, such as a non-retryable rejection, exhausted retries, and a stalled upstream that sleeps past the idle timeout. For load-oriented testing of these paths, generating unique request IDs in bulk with the Bulk UUID/GUID Generator is handy for building test fixtures. Full load testing and observability get their own article next in this series.
What to log and alert on
Every stream should produce one structured log line when it ends, whatever the outcome. Include:
request_id, model name, and the number of attempts used
outcome: done, error, or client_disconnect
error code and whether the failure was partial
chunks sent and total duration
From those fields you can build the alerts that matter: a rising share of partial failures, a spike in rate_limited, or many retries per request. Anthropic returns a request ID with each response, and its error documentation recommends logging it because support can look up a specific request with it. Store it next to your own ID.
Production checklist
Retries happen only before the first token is emitted.
Backoff uses jitter and honors Retry-After.
Retryability is decided in one classifier, using typed SDK exceptions.
An idle timeout guards every upstream read.
HTTP client timeouts are set explicitly.
Heartbeats keep proxies from closing quiet connections.
Pre-stream failures return real HTTP status codes.
Every stream ends with exactly one done or error event.
Error events carry codes and request IDs, never raw exception text.
Upstream streams are closed on completion, failure, and cancellation.
Tool calls and structured output are never resumed mid-way.
Failure paths are covered by tests with a fake upstream.
Each stream logs one structured summary line.
Frequently asked questions
Should I retry an LLM stream that fails halfway?
Not automatically. Retry silently only if no tokens have reached the user. After that, send a clear error event and let the client decide, or use a deliberate resume strategy.
Why do I get a 200 status but the response is an error?
Streaming commits the HTTP status when the response starts. If the provider fails afterward, the error arrives as an event inside the stream. Anthropic documents this behavior for overloaded_error, and it's why you need to read the stream to know the outcome.
Do the OpenAI and Anthropic SDKs already retry for me?
They retry certain failed requests by default and let you configure max_retries. In general this covers opening the request, not recovering a stream that fails after tokens have been delivered. You still need your own mid-stream handling.
How long should my idle timeout be?
Measure the real gaps between chunks for your model under load, then set the timeout comfortably above the worst normal gap. Reasoning models can pause before the first token, so give them more room than fast chat models.
Does EventSource handle reconnection for me?
It reconnects automatically after a dropped connection and sends Last-Event-ID. But without server-side resume support, a reconnect starts a new generation. For one-shot LLM responses, many teams prefer a fetch-based reader with explicit retry control.
What should the client show when a stream fails midway?
Keep the text already received, label it as incomplete, and offer a retry. Never silently replace or discard what the user has already read.
Conclusion
Reliable LLM streaming comes down to a few clear rules. Retry only before the first token. After that, fail loudly with a typed error and enough context for the UI to respond well. Guard every read with a timeout, close every upstream connection you open, and keep provider error text out of the browser. None of it is exotic, but all of it has to be deliberate, because streaming errors don't announce themselves through status codes.
The next articles in this series look at load testing and observability for streaming endpoints, scaling with queues and Redis fan-out, and streaming structured output and tool calls.