Nobody is watching when an agent runs, so every failure has to say what to do next in the status code itself — stop, or try again.
Ask for a model that doesn't exist and the answer is 400 model_not_found — a typo in the request, not an outage. Spend past a key's cap and the answer is 402 request_cap_exhausted, and an empty balance gives 402 insufficient_credits. Neither clears by waiting, so an agent stops instead of retrying into its budget.
The two worth retrying look different on purpose: 429 when you are rate limited, and 5xx when a provider fails. And if a provider dies mid-answer, the stream ends with an explicit error event rather than simply stopping, so a half-finished reply is never mistaken for a complete one.