Requesty
Back|SEP '26BEST PRACTICES / ROUTING
21 MIN READ|

429 Too Many Requests: Causes and Fixes for LLM APIs

Last updated

HTTP 429 Too Many Requests means the server is limiting requests made in a period of time. A Retry-After header may tell you when to try again. RFC 6585 defines the status, but leaves the counting method and client identity to the server.

On LLM APIs, read the error body too. A 429 can indicate request, token or concurrency pressure, exhausted credits, a spend cap or shared-capacity contention. The fix depends on which layer rejected the request and which limit it enforced.

What does 429 Too Many Requests mean?

An API has identified the caller and restricted its request rate. LLM providers usually identify the caller by API key, then enforce the limit at the organization, project, account or deployment level. The limit can cover one model or several; HTTP does not prescribe a universal scope.

An illustrative response:

http
HTTP/1.1 429 Too Many Requests
Retry-After: 120
Content-Type: application/json
 
{"error":{"type":"rate_limit_error","message":"Request limit reached. Try again later."}}

Here, the server asks for a 120-second wait. The header is optional and can instead contain an HTTP date under RFC 9110's Retry-After definition. There is no universal one-minute reset. Caches must not store the 429 response itself.

Capture the failed API call

Capture the failed response's status, headers, body and request ID, plus the endpoint, model and timestamp. If your backend calls the provider, inspect its logs: your client sees only your backend's response.

For the OpenAI SDK, Python exposes error.status_code, error.body and error.response.headers; Node exposes error.status, error.error and error.headers. These Python exception fields and Node fields let you diagnose without parsing message text. Redact error bodies before sharing them because messages can repeat sensitive input.

Diagnose the 429 before you retry

A chart showing unused requests per minute cannot explain every rejection. Find the exhausted dimension and its scope before changing retry settings.

CauseImmediate fixPrevention
Temporary request or token throttleHonor the wait hintPace aggregate requests and tokens
Credits, billing or spend capStop short retriesRestore credits, adjust a cap or await reset
Daily quota exhaustedWait for reset or request an increaseSchedule deferred work
Concurrent-call limitWait for accepted calls to finishBound parallelism
Shared-capacity contentionBack off or use an independent pathSmooth bursts and review capacity

Sources: OpenAI errors, Claude limits, Gemini limits, DeepSeek concurrency and Vertex capacity.

A 429 diagnosis flow maps credits or spend caps to account changes, daily quotas to their reset, in-flight limits to reduced concurrency, request and token limits to pacing, and shared capacity to backoff or independent fallback.
The response body and rejecting layer determine the recovery action. Based on HTTP and provider documentation checked September 29, 2026.

Diagnose one failed request

0 of 6 done

Similar codes need different recovery

Provider and codeMeaningRecovery
OpenAI 429 billing codeCredits, spend or approved usage exhaustedFix the identified account condition
Claude 429 enforced_spend_limit_reachedTier monthly spend capHigher limit or monthly reset
Claude 529 overloaded_errorProvider overloadBounded backoff or fallback
OpenAI 503 server_is_overloadedServing capacityHonor hint and retry within budget
Bedrock 429 ModelNotReadyExceptionModel readinessBounded readiness retry
Groq 498Flex Tier capacity exceededReview capacity or change path
OpenRouter 402Credits or in-flight spending budgetInspect metadata before retrying

Sources: OpenAI errors, Claude errors, Bedrock Converse exceptions, Groq errors and OpenRouter errors.

Provider-specific 429 errors and rate-limit headers

Verified September 29, 2026. Numeric allowances depend on your account, model and endpoint. Use your current console for those values; use the panels below to identify the rejecting limit and the field that explains it.

RPM means requests per minute; RPS means requests per second. TPM means tokens per minute. RPD and TPD are the corresponding daily dimensions. Input and output tokens can have separate pools, while concurrency measures calls still in progress.

Provider or pathStart here
OpenAIerror.code: traffic or billing?
AnthropicNested spend-cap code and retry-after
Gemini APIEndpoint-specific error and project limits
Vertex AIConsumption framework and capacity message
Azure OpenAIDeployment allocation and retry-after-ms
AWS BedrockSDK exception; Runtime or Mantle?
MistralRPS, TPM and top-level error fields
GroqResponse headers ending in requests measure daily allowance
DeepSeekAccount connections still in progress
xAISecond-level bursts and all counted tokens
OpenRouterPlatform counter and provider metadata

OpenAI

Limits: Organization/project limits cover model-specific RPM, RPD, TPM, TPD and modality dimensions. Model families can share buckets; rapid traffic increases also trigger slow_down.

Read: error.code and error.type, plus these headers:

text
Retry-After
x-ratelimit-limit-requests / x-ratelimit-limit-tokens
x-ratelimit-remaining-requests / x-ratelimit-remaining-tokens
x-ratelimit-reset-requests / x-ratelimit-reset-tokens

Resets are durations such as 6m0s. Conditional project headers are x-ratelimit-limit-project-tokens, x-ratelimit-remaining-project-tokens and x-ratelimit-reset-project-tokens.

Fix: Retry transient pressure with the hint. Billing codes credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded and organization_usage_limit_exceeded require account action; the type can remain insufficient_quota. Inspect Organization Limits, project settings and billing. Overload uses 503 server_is_overloaded; older endpoints also used 503 slow_down.

Honor Retry-After with backoff and jitter

Retries recover from temporary rejection. They do not create throughput. If aggregate traffic still exceeds the provider's acceptance rate, the next attempt reaches the same bottleneck.

First exclude billing, daily-quota and policy stops. Then use the server's minimum wait, add random delay, and bound total attempts and elapsed time. OpenAI's rate-limit guidance and the AWS Builders Library explain backoff and jitter; OpenAI also states that unsuccessful requests contribute to its per-minute limit.

Parse the hint before sleeping

HintMeaningCommon mistake
Retry-After: 120At least 120 secondsReading milliseconds
Retry-After: Wed, 21 Oct 2015 07:28:00 GMTWait until the HTTP dateUsing parseInt
Azure retry-after-ms: 2000At least 2 secondsWaiting 2,000 seconds
Google RetryInfo.retryDelayBody-level minimum, when presentChecking only headers

Sources: RFC 9110, Azure headers and Google RetryInfo. The millisecond value is an example.

Use wall-clock time for HTTP dates and a monotonic clock for elapsed budgets. A past date adds no minimum wait; a missing or malformed hint falls back to jittered exponential backoff. When a valid hint exceeds the remaining deadline, return the error or defer the job instead of shortening the wait.

Full jitter draws a wait uniformly from zero to the current exponential cap. The examples below add that draw to a valid server minimum. This keeps clients from retrying together without undercutting the provider's hint; when hints conflict, they use the larger valid delay.

Illustrative retry timeline: attempt one at zero seconds returns Retry-After 3, attempt two follows at 3.12 seconds, and a full-jitter draw of 2.18 seconds places attempt three at 5.30 seconds.
Assume zero call duration: 3 + 0.12 = 3.12 seconds, then 3.12 + 2.18 = 5.30 seconds. The sampled jitter stays inside each attempt's exponential cap.

Choose one retry owner

The Python and Node SDKs default to two retries, or three attempts; five wrapper attempts could therefore produce 5 × 3 = 15 transport attempts before gateway retries. If a wrapper owns recovery, set max_retries=0 or maxRetries: 0 and include gateway attempts in the budget. AWS recommends one retry point to avoid that multiplication.

Compact Python and TypeScript retry examples

These examples target direct, non-streaming OpenAI Responses text generation. They stop on billing and policy codes, honor x-should-retry: false, check completion, and propagate connection failures and timeouts without replaying an ambiguously accepted request.

Assume five total attempts, a 60-second operation deadline, a 20-second attempt deadline and a 1,024-token output cap. Exponential jitter caps grow from 2 to 16 seconds. Python uses nested asyncio deadlines, separate from SDK network-phase timeouts; Node's referenced timers abort the request and independently reject the caller, including a stalled custom error body.

Set OPENAI_API_KEY. Python requires 3.11+ and OpenAI 3.22.1; Node requires 22+ and OpenAI 7.25.0. Calls use documented Responses parameters and the SDK exception contracts linked above.

Install with python -m pip install openai==3.22.1, save as retry.py, and run python retry.py.

retry.py
import asyncio
import logging
import math
import random
import re
import time
from datetime import timezone
from email.utils import parsedate_to_datetime
 
from openai import APIStatusError, AsyncOpenAI
 
STOP = {
    "insufficient_quota",
    "credit_balance_exhausted",
    "organization_spend_limit_exceeded",
    "project_spend_limit_exceeded",
    "organization_usage_limit_exceeded",
    "cyber_policy",
    "misalignment_policy_violation",
}
 
 
def server_wait(headers):
    waits = []
    for name, divisor in (("retry-after-ms", 1000), ("retry-after", 1)):
        value = headers.get(name, "").strip()
        try:
            pattern = (
                r"[0-9]+(?:\.[0-9]+)?" if divisor == 1000 else r"[0-9]+"
            )
            if re.fullmatch(pattern, value):
                seconds = float(value) / divisor
            elif name == "retry-after" and value:
                date = parsedate_to_datetime(value)
                seconds = (
                    date.replace(tzinfo=date.tzinfo or timezone.utc).timestamp()
                    - time.time()
                )
            else:
                continue
            waits.append(max(0.0, seconds))
        except (TypeError, ValueError, OverflowError, OSError):
            continue
    return max(waits) if waits else 0.0
 
 
async def generate(client, prompt):
    deadline = time.monotonic() + 60
    async with asyncio.timeout(60):
        for attempt in range(5):
            try:
                async with asyncio.timeout(
                    min(20, deadline - time.monotonic())
                ):
                    result = await client.responses.create(
                        model="gpt-6-sol",
                        input=prompt,
                        max_output_tokens=1024,
                        timeout=20,
                    )
                    if result.status != "completed":
                        raise RuntimeError(
                            f"Response not completed: {result.status}"
                        )
                    return result
            except APIStatusError as e:
                if (
                    e.code in STOP
                    or e.type == "insufficient_quota"
                    or attempt == 4
                    or e.response.headers.get("x-should-retry") == "false"
                    or e.status_code not in {429, 500, 502, 503, 504}
                ):
                    raise
                delay = server_wait(e.response.headers) + random.uniform(
                    0, min(16, 2 ** (attempt + 1))
                )
                if (
                    not math.isfinite(delay)
                    or delay >= deadline - time.monotonic()
                ):
                    raise
                logging.warning(
                    "retry status=%s request_id=%s wait=%.2fs",
                    e.status_code,
                    e.request_id,
                    delay,
                )
                await asyncio.sleep(delay)
 
 
async def main():
    async with AsyncOpenAI(max_retries=0) as client:
        response = await generate(client, "Explain HTTP 429 in one sentence.")
        print(response.output_text)
 
 
asyncio.run(main())

Built-in retry settings are useful for prototypes, but these wrappers own classification and minimum waits. Python 3.22.1 declines server delays above 120 seconds; Node 7.25.0 substitutes its own backoff above 60 seconds. Disabling them prevents an early SDK retry before the wrapper can defer a long-wait request.

Test through the real SDK by pointing its base_url or baseURL at a local server that returns scripted 429s, wait headers and successful responses. Verify transport counts, future/past dates, milliseconds, conflicting/overflowing hints, billing/policy stops, cancellation, slow bodies, five-attempt exhaustion and failed/incomplete responses; run the stalled Node error-body case in a standalone process and check timer cleanup.

Stop policy failures and validate the result

OpenAI's cybersecurity guide documents cyber_policy without an HTTP-status mapping. Misalignment monitoring documents pre-stream 403 misalignment_policy_violation and instructs you to stop the affected workflow. Treat these as review conditions before retrying or choosing a fallback.

Adapt error classification for each provider: Claude's spend-cap code is nested, Google can supply a body-level delay, and gateways can normalize native fields. An OpenAI-compatible API reduces client changes but does not make error semantics identical. The Responses output cap includes visible and reasoning tokens; check refusals and application validity as well as completion status.

Prevent the next burst of 429s

Budget requests, tokens and concurrency separately

Assume 120 RPM, 60,000 TPM and 3,000 quota-counted tokens per call: 60,000 ÷ 3,000 = 20 calls/minute, so sustained throughput is min(120, 20) = 20 calls/minute. A semaphore bounds simultaneous calls but does not enforce RPM or token budgets; calculate separate input/output pools and admission reservations where applicable. Coordinate workers sharing organization, project or model limits, include retry traffic, and smooth short-window bursts instead of spending the minute's allowance at once.

Right-size output reservations

Azure, Bedrock Runtime and Mantle include the output cap in admission pressure; direct Claude OTPM counts generated tokens in real time and excludes max_tokens. Runtime also weights output: assume 5× burndown, 1,000 ordinary input tokens, 200 output tokens and no cache writes, giving 1,000 + (200 × 5) = 2,000 quota tokens versus 1,200 unweighted tokens. Set caps from observed output needs and inspect truncation so lower reservations do not create more failed tasks.

Cache eligible prompt prefixes

Prompt caching reuses processing while retaining the API call; Claude, Bedrock Runtime, Mantle and Groq exclude specified cache-read tokens from quota accounting, with host/model exceptions. xAI still counts cached prompt tokens toward TPM; do not assume a universal exemption for OpenAI or Gemini. Keep response reuse separate: a stored answer avoids a call only when freshness, authorization and tenant isolation permit reuse.

Move deferred work to a Batch API

OpenAI, Claude, Gemini, Mistral and Groq document separate batch limits or separation from real-time limits. Use those asynchronous paths for evaluations and extraction that do not need an immediate answer, while respecting their queue and processing constraints. Combining tasks into one synchronous prompt reduces request count but retains token workload and requires reliable output-to-input mapping.

Distribute traffic across independent capacity

A new key in the same OpenAI organization/project, Gemini project or DeepSeek account does not add headroom. Choose separately allocated deployments, approved regions or providers, and test context, schema, tools and input modalities before failover. Regional/global endpoints and cross-Region paths change capacity access and inference location, so preserve residency requirements throughout the chain.

Fail over to independent capacity with Requesty

Our routing gateway uses fallback policies to move a request to another model or provider when an upstream path returns 429. When the backup succeeds, your application receives the response instead of the primary provider's error. Weighted load balancing spreads traffic across the providers and weights you choose before failures occur.

Requesty itself limits in-flight requests, not requests per minute. Its concurrency check runs before provider routing, so a gateway rejection needs a free slot; an upstream rejection can trigger your configured fallback chain. If every path fails, the request returns an error.

  1. 1
    Create a fallback policy

    Open Routing Policies in the console, click Create Policy and choose Fallback. Follow our setup guide and name this example rate-limit-safe.

  2. 2
    Add compatible capacity

    Order approved models/providers and fit per-model retries to your operation budget. Preserve schema, tools and residency; gateway region and model inference location are separate selections.

  3. 3
    Call your policy

    Use policy/<name> with our OpenAI-compatible endpoint:

    Python
    import os
    from openai import OpenAI
     
    client = OpenAI(
        api_key=os.environ["REQUESTY_API_KEY"],
        base_url="https://router.requesty.ai/v1",
        max_retries=0,
    )
    response = client.chat.completions.create(
        model="policy/rate-limit-safe",
        messages=[{"role": "user", "content": "Draft a release note."}],
    )
    print(response.choices[0].message.content)

Policies support managed credentials and BYOK credential paths, including managed-first or own-key-first fallback. BYOK supports at most one own key per provider. Use regional routing to select the gateway and approved inference paths; a policy name alone does not pin a region.

Distinguish gateway and upstream 429s

Read error.origin in our error reference:

JSON
{"error":{"origin":"router","message":"Your organization reached its limit of concurrent requests"}}

Reduce parallelism or wait for accepted calls to finish. This in-flight rejection is neither queued nor retried on your behalf.

An application request first reaches Requesty's concurrency gate. A router 429 stops before any provider attempt; an accepted request can reach provider A, receive an upstream 429, and try compatible provider B through its fallback policy.
A router-origin concurrency rejection happens before the fallback chain starts. Source: Requesty documentation.

In Requesty's observability view, use request logs to filter status and inspect requested/used model and provider, request ID, key, user and trace. Performance monitoring groups errors and latency by model/provider. Track attempts alongside completed tasks so successful fallback does not conceal a throttled primary path.

Requesty
Give your primary provider a backup

Choose approved capacity, configure recovery once, and inspect which provider served the request. Create a Requesty account and set up your first fallback chain.

A deployment checklist for 429 recovery

429 recovery deployment checklist

0 of 12 done

Classification
Retries
Traffic
Fallback
Observability

Start with one failed response. Identify the rejecting layer and exhausted dimension, then apply its recovery. A temporary throttle needs pacing and a wait; an account cap needs account action or reset; fallback needs independent capacity.

Sources

Provider behavior and SDK contracts checked September 29, 2026. Inline links support individual claims.

Frequently asked questions
What does 429 Too Many Requests mean?
HTTP 429 Too Many Requests means a server is limiting requests made in a period of time. On LLM APIs, the same status can also indicate token limits, concurrency limits, exhausted credits, spend caps or shared-capacity contention. Read the error body to identify the cause.
How do I fix a 429 Too Many Requests error?
Stop repeated requests and follow a valid Retry-After hint. API developers should check the error code, quota scope and billing status, then pace traffic, reduce the relevant workload or use approved independent capacity. Billing and spend-cap errors need an account change or reset.
How long should I wait after a 429 error?
Wait at least as long as the server's valid Retry-After value, which can be seconds or an HTTP date. If it exceeds your application's deadline, defer the request rather than retrying early. Daily quotas and monthly spend caps follow their own reset schedules.
What if a 429 response has no Retry-After header?
First check for billing, daily-quota or spend-cap errors that short retries cannot fix. For a temporary throttle without a valid wait hint, use bounded exponential backoff with jitter and reduce the arrival rate. Some APIs provide structured retry information in the response body.
Why do I get 429 errors when I am below my API limit?
You may be below one limit while exceeding another, such as tokens, concurrent requests or a short-window burst allowance. Other workers can share the same quota, and admission estimates can exceed billed token usage. Some providers return 429 for shared-capacity contention or rapid traffic increases.
Does a new API key fix rate limits?
Not when the keys share the same organization, project, account or model quota. OpenAI, Claude, Gemini, DeepSeek and other providers enforce limits above the individual-key level. Request authorized capacity or choose a path with an independent quota instead.
Is 429 the same as 503 or 529?
No. HTTP 429 is the rate-limiting status, while 503 indicates service unavailability and Anthropic uses 529 for overload. Provider implementations can use 429 for capacity or account limits too, so classify the response using its body and API path.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders