Handling LLM Streaming Errors, Retries, and Mid-Stream Failures in Python

Streaming makes an LLM feature feel fast. It also changes how failures work. With a normal API call, a request either succeeds or fails, and you get a clear HTTP status either way. With a streamed response, the server can send a 200 OK, deliver forty tokens, and then fail. By that point the status line is already on the wire, the user is already reading, and your retry logic has nothing clean to retry.

This guide covers how to handle that in a Python backend: how to classify failures, what is safe to retry, how to detect stalled streams, how to tell the client what happened, and how to test all of it. The examples use FastAPI and the Anthropic Python SDK, but the patterns apply to any provider and any ASGI framework.

If you haven't read the first article in this series, Optimizing LLM Streaming Payloads in Python Backends covers protocol choice, buffering, and backpressure. This one picks up where it stops: what happens when the stream breaks.

Key takeaways

  • Streaming has two failure phases: before the first token and after it. They need different handling.
  • Retry automatically only while nothing has been sent to the user.
  • Once tokens are out, fail cleanly with a typed error event instead of retrying silently.