AI API errors: 401, 429, 5xx and timeouts
Separate authentication, quota, rate limits and transient failures for OpenAI, Claude and Gemini APIs. Use error details before deciding to retry.
Quick answer
An API error does not by itself establish a provider outage. Read the provider’s error body and request ID, then check the matching status component. Authentication and billing problems need configuration changes; transient failures may justify limited retries.
Start with the error category
Use this table to choose a first check. Exact meanings and remedies depend on the provider, endpoint and error body.
| What you see | First check | What it does not establish |
|---|---|---|
| 400 / 404 / 413 | Request format, endpoint, resource and size; read the provider error | A general service outage |
| 401 / 403 | Credentials, access permissions and provider-specific restrictions | That all users are blocked |
| 429 | Rate-limit headers, quota, credits and spend limits | That retrying immediately will work |
| 500 / 503 / 529 | Provider error details and the relevant official component | A confirmed global incident |
| Timeout / connection error | Client deadline, DNS, TLS, proxy and request duration | Which side caused the failure |
OpenAI: a 429 can mean different limits
Inspect error.code, not just HTTP 429. OpenAI distinguishes request rate limits from exhausted credits and organization or project limits. Billing and quota failures will not be fixed by repeated retries. For transient rate limits or overload, honor Retry-After when present and limit retries.
Use the OpenAI API entry for API components. The ChatGPT application entry cannot confirm the result of an API request or the health of a specific GPT model.
Claude: overload and spend caps are different
Claude documents 529 as overload, 401 as authentication failure and 403 as permission failure. A 429 can reflect rate limits or a spend cap. A usage-tier spend-cap 429 has no Retry-After header and keeps failing until access resumes. Read the full error before selecting a retry policy.
For streaming responses, Claude documents that an error can occur after HTTP 200. Check stream completion and error events; receiving response headers alone does not establish a successful generation.
Gemini: verify the developer endpoint
Google recommends bounded exponential backoff with jitter for transient failures. Check API version, model and supported parameters as well. Consult the developer API’s own report; the Gemini app incident feed does not cover Gemini API or AI Studio requests.
Give retries a budget
As an application design choice, set both a maximum attempt count and a total time budget. Account for retries already performed by your SDK. Use a delay with exponential backoff and jitter for errors the provider identifies as transient, and respect applicable retry headers.
Before replaying an interrupted request, check whether it may already have produced output, a charge or a downstream action. Retrying a model request and repeating an agent’s tool action are separate decisions. Preserve request IDs for troubleshooting and stop automatic retries when the permitted budget is exhausted.
Troubleshoot a specific response
Use a dedicated guide for the error returned by your request. These pages explain request failures; they do not provide account-limit monitoring.
Related service pages
Each entry shows its own collection coverage. A linked service is not necessarily automatically collected.
Editorial review date applies to this guide. Live report collection times appear on the service pages.