Requesty
AI coding agents

Run your coding agents on any model.

Point Claude Code, Cline, Roo and Copilot at one endpoint. Switch models per task, cap what an agent can spend, and see every session it ran.

Speak to founders
600+ modelsone base URL200 requests a day on free models
~/acme/api / claude
via requesty
$ export ANTHROPIC_BASE_URL="https://router.requesty.ai"
$ claude
> refactor the billing webhook and add tests
  reading src/billing/webhook.ts
  editing src/billing/webhook.ts +64 −18
  running pnpm test billing 24 passed
provider returned 429, retried on the fallback chain
    anthropic/claude-opus-5 → anthropic/claude-fable-5 ok
session 4m 12stokens 184.2Kcache hits 71%cost $0.42
Setup

One environment variable away

Claude Code reads the Anthropic variables. Everything else takes an OpenAI compatible base URL. The agent does not know it moved.

claude code / zsh
anthropic vars
export ANTHROPIC_BASE_URL="https://router.requesty.ai"
export ANTHROPIC_AUTH_TOKEN="your_requesty_api_key"
export ANTHROPIC_MODEL="anthropic/claude-fable-5"

claude
everything else / zsh
openai vars
export OPENAI_BASE_URL="https://router.requesty.ai/v1"
export OPENAI_API_KEY="your_requesty_api_key"

# or paste the same two values into the
# settings panel in Cline, Roo or VS Code
Reliability

Agents fail on rate limits, not on availability

Two thirds of provider failures on the gateway are 429s. An agent that gives up on the first throttle looks broken to the person waiting, so the fallback chain matters more than any one provider's uptime.

provider failures by status code, April 2026
open data
429Rate limited
65.8%
400Bad request
19.4%
403Forbidden
9.4%
5xxProvider unavailable
4.8%

Share of April 2026 requests where the upstream provider returned a failure. Throttling, not downtime, is what stops an agent mid-task. See the data note.

92%

prompt-cache hit rate on Claude Code traffic through the gateway, April 2026. It is what keeps a long session affordable.

Open data note

209

calls in a P95 Claude Code session. One unretried throttle anywhere in the run is a dead task.

Open data note

99.25%

eventual success on a routing policy in April 2026, against 85.01% for callers pinned to one provider.

Open data note

Spend controls

Cap what an agent can burn

Give each agent its own key with its own ceiling. An agent loop becomes a capped experiment, not an incident.

  • A key per agentCap it, alert on it, revoke it on its own
keys / sk-agent-eval
capped
monthly cap$500 / $500
requests refused until the cap is raised
80% of the cap
email to the key owner
14 Apr
95% of the cap
webhook to #ai-spend
16 Apr
cap reached
key paused, other keys unaffected
17 Apr

The loop stopped at the ceiling, not at the invoice.

Debugging

Replay what the agent did

Session reconstruction stitches a run back together, request by request. You find the call that went wrong instead of guessing at it.

  • Tool call analyticsWhich tools ran, how often, what they cost
  • Finish reasonsTruncation and refusals, not just errors
session / agent-run-4f21c
reconstructed
Six calls, in order
01read repository layout
02search for webhook handler
03edit webhook.ts
04run test suite
05patch assertion
06retry with larger budget
step 05 finish_reason: length, the model ran out of output budget
Frameworks

Building your own agent?

The same endpoint works for LangChain, CrewAI, AutoGen, LlamaIndex, PydanticAI and the Vercel AI SDK.

Start without a bill: 200 requests per day on free models, enough to point an agent at Requesty tonight.

Browse the free models

Point your agent at one endpoint

600+ models, budgets per key, and every session on the record.

Speak to founders