Run your coding agents on any model.
Point Claude Code, Cline, Roo and Copilot at one endpoint. Switch models per task, cap what an agent can spend, and see every session it ran.
One environment variable away
Claude Code reads the Anthropic variables. Everything else takes an OpenAI compatible base URL. The agent does not know it moved.
export ANTHROPIC_BASE_URL="https://router.requesty.ai" export ANTHROPIC_AUTH_TOKEN="your_requesty_api_key" export ANTHROPIC_MODEL="anthropic/claude-fable-5" claude
export OPENAI_BASE_URL="https://router.requesty.ai/v1" export OPENAI_API_KEY="your_requesty_api_key" # or paste the same two values into the # settings panel in Cline, Roo or VS Code
Agents fail on rate limits, not on availability
Two thirds of provider failures on the gateway are 429s. An agent that gives up on the first throttle looks broken to the person waiting, so the fallback chain matters more than any one provider's uptime.
Share of April 2026 requests where the upstream provider returned a failure. Throttling, not downtime, is what stops an agent mid-task. See the data note.
prompt-cache hit rate on Claude Code traffic through the gateway, April 2026. It is what keeps a long session affordable.
calls in a P95 Claude Code session. One unretried throttle anywhere in the run is a dead task.
eventual success on a routing policy in April 2026, against 85.01% for callers pinned to one provider.
Cap what an agent can burn
Give each agent its own key with its own ceiling. An agent loop becomes a capped experiment, not an incident.
- A key per agentCap it, alert on it, revoke it on its own
The loop stopped at the ceiling, not at the invoice.
Replay what the agent did
Session reconstruction stitches a run back together, request by request. You find the call that went wrong instead of guessing at it.
- Tool call analyticsWhich tools ran, how often, what they cost
- Finish reasonsTruncation and refusals, not just errors
Works with the agents you already run
Two minute setup for each. Most are a base URL and a key.
Building your own agent?
The same endpoint works for LangChain, CrewAI, AutoGen, LlamaIndex, PydanticAI and the Vercel AI SDK.
Start without a bill: 200 requests per day on free models, enough to point an agent at Requesty tonight.
Browse the free modelsPoint your agent at one endpoint
600+ models, budgets per key, and every session on the record.

