HTTP 429 Too Many Requests means the server is limiting requests made in a period of time. A Retry-After header may tell you when to try again. RFC 6585 defines the status, but leaves the counting method and client identity to the server.
On LLM APIs, read the error body too. A 429 can indicate request, token or concurrency pressure, exhausted credits, a spend cap or shared-capacity contention. The fix depends on which layer rejected the request and which limit it enforced.
What does 429 Too Many Requests mean?
An API has identified the caller and restricted its request rate. LLM providers usually identify the caller by API key, then enforce the limit at the organization, project, account or deployment level. The limit can cover one model or several; HTTP does not prescribe a universal scope.
An illustrative response:
HTTP/1.1 429 Too Many Requests
Retry-After: 120
Content-Type: application/json
{"error":{"type":"rate_limit_error","message":"Request limit reached. Try again later."}}Here, the server asks for a 120-second wait. The header is optional and can instead contain an HTTP date under RFC 9110's Retry-After definition. There is no universal one-minute reset. Caches must not store the 429 response itself.
Capture the failed API call
Capture the failed response's status, headers, body and request ID, plus the endpoint, model and timestamp. If your backend calls the provider, inspect its logs: your client sees only your backend's response.
For the OpenAI SDK, Python exposes error.status_code, error.body and error.response.headers; Node exposes error.status, error.error and error.headers. These Python exception fields and Node fields let you diagnose without parsing message text. Redact error bodies before sharing them because messages can repeat sensitive input.
Diagnose the 429 before you retry
A chart showing unused requests per minute cannot explain every rejection. Find the exhausted dimension and its scope before changing retry settings.
| Cause | Immediate fix | Prevention |
|---|---|---|
| Temporary request or token throttle | Honor the wait hint | Pace aggregate requests and tokens |
| Credits, billing or spend cap | Stop short retries | Restore credits, adjust a cap or await reset |
| Daily quota exhausted | Wait for reset or request an increase | Schedule deferred work |
| Concurrent-call limit | Wait for accepted calls to finish | Bound parallelism |
| Shared-capacity contention | Back off or use an independent path | Smooth bursts and review capacity |
Sources: OpenAI errors, Claude limits, Gemini limits, DeepSeek concurrency and Vertex capacity.

Diagnose one failed request
0 of 6 done
Similar codes need different recovery
| Provider and code | Meaning | Recovery |
|---|---|---|
| OpenAI 429 billing code | Credits, spend or approved usage exhausted | Fix the identified account condition |
Claude 429 enforced_spend_limit_reached | Tier monthly spend cap | Higher limit or monthly reset |
Claude 529 overloaded_error | Provider overload | Bounded backoff or fallback |
OpenAI 503 server_is_overloaded | Serving capacity | Honor hint and retry within budget |
Bedrock 429 ModelNotReadyException | Model readiness | Bounded readiness retry |
| Groq 498 | Flex Tier capacity exceeded | Review capacity or change path |
| OpenRouter 402 | Credits or in-flight spending budget | Inspect metadata before retrying |
Sources: OpenAI errors, Claude errors, Bedrock Converse exceptions, Groq errors and OpenRouter errors.
Provider-specific 429 errors and rate-limit headers
Verified September 29, 2026. Numeric allowances depend on your account, model and endpoint. Use your current console for those values; use the panels below to identify the rejecting limit and the field that explains it.
RPM means requests per minute; RPS means requests per second. TPM means tokens per minute. RPD and TPD are the corresponding daily dimensions. Input and output tokens can have separate pools, while concurrency measures calls still in progress.
| Provider or path | Start here |
|---|---|
| OpenAI | error.code: traffic or billing? |
| Anthropic | Nested spend-cap code and retry-after |
| Gemini API | Endpoint-specific error and project limits |
| Vertex AI | Consumption framework and capacity message |
| Azure OpenAI | Deployment allocation and retry-after-ms |
| AWS Bedrock | SDK exception; Runtime or Mantle? |
| Mistral | RPS, TPM and top-level error fields |
| Groq | Response headers ending in requests measure daily allowance |
| DeepSeek | Account connections still in progress |
| xAI | Second-level bursts and all counted tokens |
| OpenRouter | Platform counter and provider metadata |
OpenAI
Limits: Organization/project limits cover model-specific RPM, RPD, TPM, TPD and modality dimensions. Model families can share buckets; rapid traffic increases also trigger slow_down.
Read: error.code and error.type, plus these headers:
Retry-After
x-ratelimit-limit-requests / x-ratelimit-limit-tokens
x-ratelimit-remaining-requests / x-ratelimit-remaining-tokens
x-ratelimit-reset-requests / x-ratelimit-reset-tokensResets are durations such as 6m0s. Conditional project headers are x-ratelimit-limit-project-tokens, x-ratelimit-remaining-project-tokens and x-ratelimit-reset-project-tokens.
Fix: Retry transient pressure with the hint. Billing codes credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded and organization_usage_limit_exceeded require account action; the type can remain insufficient_quota. Inspect Organization Limits, project settings and billing. Overload uses 503 server_is_overloaded; older endpoints also used 503 slow_down.
Anthropic direct API
Limits: Claude Messages uses model-class RPM, input TPM and output TPM, with organization/workspace constraints and short-window enforcement.
Read: error.type: rate_limit_error and nested error.details.error_code. Standard headers expand as follows:
| Dimension | Limit / remaining / reset headers |
|---|---|
| Requests | anthropic-ratelimit-requests-limit, anthropic-ratelimit-requests-remaining, anthropic-ratelimit-requests-reset |
| Combined tokens | anthropic-ratelimit-tokens-limit, anthropic-ratelimit-tokens-remaining, anthropic-ratelimit-tokens-reset |
| Input | anthropic-ratelimit-input-tokens-limit, anthropic-ratelimit-input-tokens-remaining, anthropic-ratelimit-input-tokens-reset |
| Output | anthropic-ratelimit-output-tokens-limit, anthropic-ratelimit-output-tokens-remaining, anthropic-ratelimit-output-tokens-reset |
retry-after is seconds; resets are RFC 3339 timestamps. Remaining token counts are rounded to thousands; combined headers show the most restrictive token limit.
Fix: Honor the normal rate-error hint. enforced_spend_limit_reached has no hint: raise the tier limit or wait until 00:00 UTC on the month's first day. Check Rate limits and Billing. User-set organization/workspace spend caps return 400, with a Claude Code workspace exception. 529 means overload; Fast mode and Priority have additional pools.
Gemini Developer API
Limits: Per-project limits include RPM, input TPM, RPD and model-specific image/daily-token dimensions. Some accounts also have rolling 10-minute spend-rate limits. Keys share the project's quota.
Read: GenerateContent-style errors use numeric error.code: 429 and error.status: RESOURCE_EXHAUSTED. Structured details can include quota violations or google.rpc.RetryInfo.retryDelay, as shown in Google's SDK issue tracker. RetryInfo defines a minimum wait. The Interactions API instead uses string codes such as rate_limit_exceeded, quota_exceeded and too_many_requests. Use those details and the console; the rate guide specifies no numeric header family.
Fix: Check AI Studio limits and its upgrade/increase route. RPD resets at midnight Pacific time. Exhausted Prepay credits return 402. Vertex uses a separate quota system.
Vertex AI
Limits: Consumption framework matters. Current Standard PayGo sets per-model baseline TPM from the organization's eligible spend over a rolling 30-day period. Model baselines are independent, with no separate tier RPM limit. Its 429 indicates shared-resource contention; other paths have project/model or feature quotas.
Read: 429 RESOURCE_EXHAUSTED and the framework message:
PayGo: Resource exhausted, please try again later.
Provisioned: Too many requests. Exceeded the Provisioned Throughput.The guide specifies no numeric remaining-quota header family. Inspect structured details, Cloud Monitoring and applicable quotas.
Fix: Smooth bursts, ramp gradually and back off. A global endpoint accesses multi-region capacity when locality requirements permit. Request increases for quota-governed paths; Standard PayGo contention needs available capacity. Provisioned Throughput changes capacity handling but does not guarantee every request succeeds.
Azure OpenAI
Limits: Azure quota is per subscription, region, model and deployment type. TPM is allocated to deployments; model-specific ratios derive RPM. Admission estimates include prompt content and the output cap, and short windows are typically 1 or 10 seconds.
Read:
retry-after-ms
x-ratelimit-limit-requests / x-ratelimit-limit-tokens
x-ratelimit-remaining-requests / x-ratelimit-remaining-tokens
x-ratelimit-reset-requests / x-ratelimit-reset-tokensUse retry-after-ms as milliseconds; reset units are unspecified in the guide. Provisioned utilization appears in the error message; a documented PTU example used RateLimitReached.
Fix: Pace traffic, right-size output reservations and compare returned limits with deployment allocation. Reallocate TPM or request quota. Temporary capacity adjustments also affect limits; spare subscription quota and successful billed-token charts do not establish deployment headroom.
AWS Bedrock
Limits: Runtime has account/Region/model TPM, applicable RPM and daily pools. Mantle has independent input/output TPM allocations and no RPM quota.
Read: Converse documents 429 ThrottlingException and ModelNotReadyException. Use SDK exceptions, Service Quotas and CloudWatch; no numeric remaining-quota header family is documented there.
Fix: Right-size output reservations. Runtime reserves input plus output cap and applies output burndown; Converse uses inferenceConfig.maxTokens. Mantle reserves input plus cap against input TPM, returns unused reservation after completion, and can cut off generation at output TPM. Only Opus 4.7 has published Mantle defaults as of September 2026; other models use internal capacity. Request adjustable Runtime increases through Service Quotas, Mantle through AWS Support. Account for SDK readiness retries. Estimated TPM metrics exclude reservation pressure.
Mistral
Limits: Organization/model RPS and TPM are enforced independently. The error glossary also lists concurrency as a 429 cause.
Read: Documented headers are Retry-After and X-RateLimit-Remaining; the sources do not specify the latter's units or a full Limit/Reset family. Error fields are top-level:
{
"object": "error",
"message": "Rate limit exceeded.",
"type": "rate_limit_error"
}This example uses an illustrative message. Other fields include param and code; inspect the returned values.
Fix: Honor the wait hint and pace requests and tokens separately. Check organization limits. The increase process asks for target RPS, expected token volume and use case. Batch processing does not count against real-time limits.
Groq
Limits: Organization limits include RPM, RPD, TPM, TPD and audio usage; some accounts have separate input/output dimensions.
Read:
retry-after
x-ratelimit-limit-requests / x-ratelimit-limit-tokens
x-ratelimit-remaining-requests / x-ratelimit-remaining-tokens
x-ratelimit-reset-requests / x-ratelimit-reset-tokensThe *-requests response headers always mean RPD; *-tokens headers mean TPM. Resets are durations, including fractional seconds. retry-after gives seconds on rate-limit 429s. Error fields include error.message and error.type; do not reuse a subtype from an unrelated example.
Fix: Check Console Limits for the exhausted dimension. Retry temporary pressure with the hint; daily exhaustion needs reset or a higher allowance. The Developer plan and requested capacity raise limits. Batch has no impact on standard limits. Distinguish 498 Flex Tier capacity exhaustion and 503 unavailability.
DeepSeek
Limits: Account concurrency counts each connection from submission until response completion, regardless of API key. Excess concurrency returns 429.
Read: Track unfinished calls, including waiting connections. The rate/error references specify no header family or stable 429 subtype. Documented statuses are 429 Rate Limit Reached, 402 Insufficient Balance and 503 Server Overloaded.
Fix: Bound parallelism and retain slots through completion. Waiting requests receive empty lines or streaming : keep-alive comments; these do not release concurrency. The server closes a connection if inference has not started after 10 minutes. Use the capacity-expansion route linked from the official rate page. Expanded accounts also have user_id constraints; ordinary user IDs and additional keys do not create independent account capacity.
xAI
Limits: Team/model RPS and TPM apply independently. RPS is derived from RPM divided by 60, so a minute's request budget cannot be spent in one second. Voice/audio also has concurrent-session limits.
Read: TPM includes prompt, completion, reasoning and cached prompt tokens. The rate guide specifies no numeric header family, guaranteed Retry-After or stable 429 subtype. Inspect the returned body and your team/model limits rather than assuming OpenAI-compatible responses use OpenAI's headers.
Fix: Pace at the second-level interval and budget every counted token category. Check personalized limits on the xAI Models page. Tiers scale with cumulative spend; use the console increase route or sales for enterprise capacity. Cache discounts lower billing without releasing TPM. Another key under the same team/model does not create another pool.
OpenRouter
Limits: Platform caps, upstream limits and DDoS protection are separate from credit/in-flight spending gates. Additional keys/accounts do not raise globally governed capacity.
Read: error.metadata.error_type identifies the category; provider_code carries upstream codes when available. Platform 429s include X-RateLimit-Limit, X-RateLimit-Remaining, X-RateLimit-Reset; successful inference responses omit them. Retry-After is conditional on attempted providers' hints. Inspect account counters:
curl https://openrouter.ai/api/v1/key \
-H "Authorization: Bearer $OPENROUTER_API_KEY"Free daily counters use UTC days; the legacy rate_limit object is deprecated.
Fix: Check platform counters and routing restrictions; OpenRouter already supports provider/model fallback. 402 metadata distinguishes credits from the transient openrouter_in_flight_budget case. Inspect body/stream errors after HTTP 200 as well. Reset-header encoding is unspecified; use an explicit retry hint for timing.
Honor Retry-After with backoff and jitter
Retries recover from temporary rejection. They do not create throughput. If aggregate traffic still exceeds the provider's acceptance rate, the next attempt reaches the same bottleneck.
First exclude billing, daily-quota and policy stops. Then use the server's minimum wait, add random delay, and bound total attempts and elapsed time. OpenAI's rate-limit guidance and the AWS Builders Library explain backoff and jitter; OpenAI also states that unsuccessful requests contribute to its per-minute limit.
Parse the hint before sleeping
| Hint | Meaning | Common mistake |
|---|---|---|
Retry-After: 120 | At least 120 seconds | Reading milliseconds |
Retry-After: Wed, 21 Oct 2015 07:28:00 GMT | Wait until the HTTP date | Using parseInt |
Azure retry-after-ms: 2000 | At least 2 seconds | Waiting 2,000 seconds |
Google RetryInfo.retryDelay | Body-level minimum, when present | Checking only headers |
Sources: RFC 9110, Azure headers and Google RetryInfo. The millisecond value is an example.
Use wall-clock time for HTTP dates and a monotonic clock for elapsed budgets. A past date adds no minimum wait; a missing or malformed hint falls back to jittered exponential backoff. When a valid hint exceeds the remaining deadline, return the error or defer the job instead of shortening the wait.
Full jitter draws a wait uniformly from zero to the current exponential cap. The examples below add that draw to a valid server minimum. This keeps clients from retrying together without undercutting the provider's hint; when hints conflict, they use the larger valid delay.

Choose one retry owner
The Python and Node SDKs default to two retries, or three attempts; five wrapper attempts could therefore produce 5 × 3 = 15 transport attempts before gateway retries. If a wrapper owns recovery, set max_retries=0 or maxRetries: 0 and include gateway attempts in the budget. AWS recommends one retry point to avoid that multiplication.
Compact Python and TypeScript retry examples
These examples target direct, non-streaming OpenAI Responses text generation. They stop on billing and policy codes, honor x-should-retry: false, check completion, and propagate connection failures and timeouts without replaying an ambiguously accepted request.
Assume five total attempts, a 60-second operation deadline, a 20-second attempt deadline and a 1,024-token output cap. Exponential jitter caps grow from 2 to 16 seconds. Python uses nested asyncio deadlines, separate from SDK network-phase timeouts; Node's referenced timers abort the request and independently reject the caller, including a stalled custom error body.
Set OPENAI_API_KEY. Python requires 3.11+ and OpenAI 3.22.1; Node requires 22+ and OpenAI 7.25.0. Calls use documented Responses parameters and the SDK exception contracts linked above.
Install with python -m pip install openai==3.22.1, save as retry.py, and run python retry.py.
import asyncio
import logging
import math
import random
import re
import time
from datetime import timezone
from email.utils import parsedate_to_datetime
from openai import APIStatusError, AsyncOpenAI
STOP = {
"insufficient_quota",
"credit_balance_exhausted",
"organization_spend_limit_exceeded",
"project_spend_limit_exceeded",
"organization_usage_limit_exceeded",
"cyber_policy",
"misalignment_policy_violation",
}
def server_wait(headers):
waits = []
for name, divisor in (("retry-after-ms", 1000), ("retry-after", 1)):
value = headers.get(name, "").strip()
try:
pattern = (
r"[0-9]+(?:\.[0-9]+)?" if divisor == 1000 else r"[0-9]+"
)
if re.fullmatch(pattern, value):
seconds = float(value) / divisor
elif name == "retry-after" and value:
date = parsedate_to_datetime(value)
seconds = (
date.replace(tzinfo=date.tzinfo or timezone.utc).timestamp()
- time.time()
)
else:
continue
waits.append(max(0.0, seconds))
except (TypeError, ValueError, OverflowError, OSError):
continue
return max(waits) if waits else 0.0
async def generate(client, prompt):
deadline = time.monotonic() + 60
async with asyncio.timeout(60):
for attempt in range(5):
try:
async with asyncio.timeout(
min(20, deadline - time.monotonic())
):
result = await client.responses.create(
model="gpt-6-sol",
input=prompt,
max_output_tokens=1024,
timeout=20,
)
if result.status != "completed":
raise RuntimeError(
f"Response not completed: {result.status}"
)
return result
except APIStatusError as e:
if (
e.code in STOP
or e.type == "insufficient_quota"
or attempt == 4
or e.response.headers.get("x-should-retry") == "false"
or e.status_code not in {429, 500, 502, 503, 504}
):
raise
delay = server_wait(e.response.headers) + random.uniform(
0, min(16, 2 ** (attempt + 1))
)
if (
not math.isfinite(delay)
or delay >= deadline - time.monotonic()
):
raise
logging.warning(
"retry status=%s request_id=%s wait=%.2fs",
e.status_code,
e.request_id,
delay,
)
await asyncio.sleep(delay)
async def main():
async with AsyncOpenAI(max_retries=0) as client:
response = await generate(client, "Explain HTTP 429 in one sentence.")
print(response.output_text)
asyncio.run(main())Install with npm install --save-exact openai@7.25.0. Save as retry.mts and compile/run with your project's NodeNext TypeScript setup.
import OpenAI from 'openai';
import { setTimeout as sleep } from 'node:timers/promises';
const stop = new Set([
'insufficient_quota',
'credit_balance_exhausted',
'organization_spend_limit_exceeded',
'project_spend_limit_exceeded',
'organization_usage_limit_exceeded',
'cyber_policy',
'misalignment_policy_violation',
]);
function serverWait(headers?: Headers): number {
const waits: number[] = [];
const ms = headers?.get('retry-after-ms')?.trim();
if (ms && /^[0-9]+(?:\.[0-9]+)?$/.test(ms)) {
waits.push(Number(ms));
}
const value = headers?.get('retry-after')?.trim();
if (value && /^[0-9]+$/.test(value)) {
waits.push(Number(value) * 1000);
} else if (
value && /[A-Za-z]/.test(value) && Number.isFinite(Date.parse(value))
) {
waits.push(Math.max(0, Date.parse(value) - Date.now()));
}
return waits.length ? Math.max(...waits) : 0;
}
async function bounded<T>(
parent: AbortSignal,
ms: number,
run: (s: AbortSignal) => PromiseLike<T>,
) {
parent.throwIfAborted();
const controller = new AbortController();
const signal = AbortSignal.any([parent, controller.signal]);
let rejectAbort: (reason: unknown) => void = () => {};
const stopped = new Promise<never>((_, reject) => {
rejectAbort = reject;
});
const abort = () => rejectAbort(signal.reason);
signal.addEventListener('abort', abort, { once: true });
const timer = setTimeout(
() => controller.abort(new Error('LLM deadline expired')),
ms,
);
try {
signal.throwIfAborted();
return await Promise.race([run(signal), stopped]);
} finally {
clearTimeout(timer);
signal.removeEventListener('abort', abort);
}
}
async function generate(
client: OpenAI,
prompt: string,
caller = new AbortController().signal,
) {
const deadline = performance.now() + 60_000;
return bounded(caller, 60_000, async (signal) => {
for (let attempt = 0; attempt < 5; attempt++) {
signal.throwIfAborted();
try {
const result = await bounded(
signal,
Math.max(1, Math.min(20_000, deadline - performance.now())),
(s) => client.responses.create(
{
model: 'gpt-6-sol',
input: prompt,
max_output_tokens: 1024,
},
{ signal: s, timeout: 20_000 },
),
);
if (result.status !== 'completed') {
throw new Error(`Response not completed: ${result.status}`);
}
return result;
} catch (e) {
signal.throwIfAborted();
if (
!(e instanceof OpenAI.APIError)
|| stop.has(e.code ?? '')
|| e.type === 'insufficient_quota'
|| e.headers?.get('x-should-retry') === 'false'
|| attempt === 4
|| ![429, 500, 502, 503, 504].includes(e.status ?? 0)
) {
throw e;
}
const delay = serverWait(e.headers)
+ Math.random() * Math.min(16_000, 1000 * 2 ** (attempt + 1));
if (
!Number.isFinite(delay)
|| delay >= deadline - performance.now()
) {
throw e;
}
console.warn({
status: e.status,
requestID: e.requestID,
waitMs: Math.ceil(delay),
});
await sleep(Math.ceil(delay), undefined, { signal });
}
}
throw new Error('Attempt limit reached');
});
}
const response = await generate(
new OpenAI({ maxRetries: 0 }),
'Explain HTTP 429 in one sentence.',
);
console.log(response.output_text);Built-in retry settings are useful for prototypes, but these wrappers own classification and minimum waits. Python 3.22.1 declines server delays above 120 seconds; Node 7.25.0 substitutes its own backoff above 60 seconds. Disabling them prevents an early SDK retry before the wrapper can defer a long-wait request.
Test through the real SDK by pointing its base_url or baseURL at a local server that returns scripted 429s, wait headers and successful responses. Verify transport counts, future/past dates, milliseconds, conflicting/overflowing hints, billing/policy stops, cancellation, slow bodies, five-attempt exhaustion and failed/incomplete responses; run the stalled Node error-body case in a standalone process and check timer cleanup.
Stop policy failures and validate the result
OpenAI's cybersecurity guide documents cyber_policy without an HTTP-status mapping. Misalignment monitoring documents pre-stream 403 misalignment_policy_violation and instructs you to stop the affected workflow. Treat these as review conditions before retrying or choosing a fallback.
Adapt error classification for each provider: Claude's spend-cap code is nested, Google can supply a body-level delay, and gateways can normalize native fields. An OpenAI-compatible API reduces client changes but does not make error semantics identical. The Responses output cap includes visible and reasoning tokens; check refusals and application validity as well as completion status.
These examples do not consume streams or execute tools. OpenRouter stops failover after output reaches you; a timeout or client cancellation does not undo earlier effects. Use endpoint-supported idempotency and application reconciliation for side-effecting workflows.
Prevent the next burst of 429s
Budget requests, tokens and concurrency separately
Assume 120 RPM, 60,000 TPM and 3,000 quota-counted tokens per call: 60,000 ÷ 3,000 = 20 calls/minute, so sustained throughput is min(120, 20) = 20 calls/minute. A semaphore bounds simultaneous calls but does not enforce RPM or token budgets; calculate separate input/output pools and admission reservations where applicable. Coordinate workers sharing organization, project or model limits, include retry traffic, and smooth short-window bursts instead of spending the minute's allowance at once.
Right-size output reservations
Azure, Bedrock Runtime and Mantle include the output cap in admission pressure; direct Claude OTPM counts generated tokens in real time and excludes max_tokens. Runtime also weights output: assume 5× burndown, 1,000 ordinary input tokens, 200 output tokens and no cache writes, giving 1,000 + (200 × 5) = 2,000 quota tokens versus 1,200 unweighted tokens. Set caps from observed output needs and inspect truncation so lower reservations do not create more failed tasks.
Cache eligible prompt prefixes
Prompt caching reuses processing while retaining the API call; Claude, Bedrock Runtime, Mantle and Groq exclude specified cache-read tokens from quota accounting, with host/model exceptions. xAI still counts cached prompt tokens toward TPM; do not assume a universal exemption for OpenAI or Gemini. Keep response reuse separate: a stored answer avoids a call only when freshness, authorization and tenant isolation permit reuse.
Move deferred work to a Batch API
OpenAI, Claude, Gemini, Mistral and Groq document separate batch limits or separation from real-time limits. Use those asynchronous paths for evaluations and extraction that do not need an immediate answer, while respecting their queue and processing constraints. Combining tasks into one synchronous prompt reduces request count but retains token workload and requires reliable output-to-input mapping.
Distribute traffic across independent capacity
A new key in the same OpenAI organization/project, Gemini project or DeepSeek account does not add headroom. Choose separately allocated deployments, approved regions or providers, and test context, schema, tools and input modalities before failover. Regional/global endpoints and cross-Region paths change capacity access and inference location, so preserve residency requirements throughout the chain.
Fail over to independent capacity with Requesty
Our routing gateway uses fallback policies to move a request to another model or provider when an upstream path returns 429. When the backup succeeds, your application receives the response instead of the primary provider's error. Weighted load balancing spreads traffic across the providers and weights you choose before failures occur.
Requesty itself limits in-flight requests, not requests per minute. Its concurrency check runs before provider routing, so a gateway rejection needs a free slot; an upstream rejection can trigger your configured fallback chain. If every path fails, the request returns an error.
- 1Create a fallback policy
Open Routing Policies in the console, click Create Policy and choose Fallback. Follow our setup guide and name this example
rate-limit-safe. - 2Add compatible capacity
Order approved models/providers and fit per-model retries to your operation budget. Preserve schema, tools and residency; gateway region and model inference location are separate selections.
- 3Call your policy
Use
policy/<name>with our OpenAI-compatible endpoint:Pythonimport os from openai import OpenAI client = OpenAI( api_key=os.environ["REQUESTY_API_KEY"], base_url="https://router.requesty.ai/v1", max_retries=0, ) response = client.chat.completions.create( model="policy/rate-limit-safe", messages=[{"role": "user", "content": "Draft a release note."}], ) print(response.choices[0].message.content)
Policies support managed credentials and BYOK credential paths, including managed-first or own-key-first fallback. BYOK supports at most one own key per provider. Use regional routing to select the gateway and approved inference paths; a policy name alone does not pin a region.
Distinguish gateway and upstream 429s
Read error.origin in our error reference:
{"error":{"origin":"router","message":"Your organization reached its limit of concurrent requests"}}Reduce parallelism or wait for accepted calls to finish. This in-flight rejection is neither queued nor retried on your behalf.
{"error":{"origin":"provider","message":"Too many requests"}}Inspect the attempted provider and configured policy. Fallback tries the next compatible path after its configured retries.

In Requesty's observability view, use request logs to filter status and inspect requested/used model and provider, request ID, key, user and trace. Performance monitoring groups errors and latency by model/provider. Track attempts alongside completed tasks so successful fallback does not conceal a throttled primary path.
Choose approved capacity, configure recovery once, and inspect which provider served the request. Create a Requesty account and set up your first fallback chain.
A deployment checklist for 429 recovery
429 recovery deployment checklist
0 of 12 done
Start with one failed response. Identify the rejecting layer and exhausted dimension, then apply its recovery. A temporary throttle needs pacing and a wait; an account cap needs account action or reset; fallback needs independent capacity.
Sources
Provider behavior and SDK contracts checked September 29, 2026. Inline links support individual claims.
- HTTP: RFC 6585 and RFC 9110.
- Provider references: OpenAI, Claude, Gemini, Vertex AI, Azure, Bedrock Runtime, Mantle, Mistral, Groq, DeepSeek, xAI and OpenRouter.
- Retry mechanics: AWS Builders Library, Python SDK source and Node SDK source.
- Requesty: Fallback policies, in-flight limits and error origins.
Frequently asked questions
- What does 429 Too Many Requests mean?
- HTTP 429 Too Many Requests means a server is limiting requests made in a period of time. On LLM APIs, the same status can also indicate token limits, concurrency limits, exhausted credits, spend caps or shared-capacity contention. Read the error body to identify the cause.
- How do I fix a 429 Too Many Requests error?
- Stop repeated requests and follow a valid Retry-After hint. API developers should check the error code, quota scope and billing status, then pace traffic, reduce the relevant workload or use approved independent capacity. Billing and spend-cap errors need an account change or reset.
- How long should I wait after a 429 error?
- Wait at least as long as the server's valid Retry-After value, which can be seconds or an HTTP date. If it exceeds your application's deadline, defer the request rather than retrying early. Daily quotas and monthly spend caps follow their own reset schedules.
- What if a 429 response has no Retry-After header?
- First check for billing, daily-quota or spend-cap errors that short retries cannot fix. For a temporary throttle without a valid wait hint, use bounded exponential backoff with jitter and reduce the arrival rate. Some APIs provide structured retry information in the response body.
- Why do I get 429 errors when I am below my API limit?
- You may be below one limit while exceeding another, such as tokens, concurrent requests or a short-window burst allowance. Other workers can share the same quota, and admission estimates can exceed billed token usage. Some providers return 429 for shared-capacity contention or rapid traffic increases.
- Does a new API key fix rate limits?
- Not when the keys share the same organization, project, account or model quota. OpenAI, Claude, Gemini, DeepSeek and other providers enforce limits above the individual-key level. Request authorized capacity or choose a path with an independent quota instead.
- Is 429 the same as 503 or 529?
- No. HTTP 429 is the rate-limiting status, while 503 indicates service unavailability and Anthropic uses 529 for overload. Provider implementations can use 429 for capacity or account limits too, so classify the response using its body and API path.
- SEP '26
OpenAI-Compatible APIs: Meaning, Providers and Examples
An OpenAI compatible API lets you reuse your SDK, but features vary. Compare provider endpoints, tools, streaming, JSON, vision and Responses support.
- SEP '26
How to Get a Claude API Key (and Manage Credits and Billing)
Get a Claude API key in the Console, test it with curl or an SDK, buy credits, set monthly spending limits, and fix common billing errors safely.
- SEP '26
Claude Code API Key: Setup Without a Subscription
Use a Claude Code API key instead of Pro or Max. Follow macOS, Linux and Windows setup, verify billing, compare costs and switch back without hidden overrides.
