AI Gateway and LLM Router

The AI Gateway for

Also: up to 60% savings on model spend.

600+ models. Real-time analytics. Intelligent routing. Change one line and get auto-caching, failover, governance and observability.

Or email us at sales@requesty.ai

your appPOST /v1/chat/completions
Gatewayrouter.requesty.ai/v1
  • authok
  • policypass
  • cachewarm
  • failoverarmed
  • Claude Opus 5live ยท 310ms
  • gpt-5.6-solfallback
  • gemini-3.6-flashready
  • deepseek-v4-proready
  • kimi-k3ready
  • +595 more

Trusted by teams at

ZoomInfoShopifySiemensPfizerCapgeminiPwCAmadeusSageChargebeeDemandbaseRelevance AIAppnovation
600+
Models
99.99%
Uptime SLA
14ms
Failover
225B+
Tokens per day

Dashboard

Your AI ops, at a glance

Heatmaps, cost breakdowns and cache gauges across all your AI providers, on a dashboard that refreshes as the requests land.

requestyDashboard / Overview7d30d90d30D
Total requests
143.2K
18.7%
Total cost
$1,247
12.3%
Cached input
87.3%
4.2 pts
Success rate
99.85%
30 days, this account
Latency distribution24H UTC
<50ms
<200ms
<500ms
<1s
<2s
>2s
00061218
<50ms3.6%
<200ms32.4%
<500ms43.1%
<1s17.2%
<2s3.4%
Share of requests, low to high. Peak load at 14:00 UTC.
Auto-cachingPROVIDER PREFIX
87.3%Cached input
2.63B
Tokens cached
78.6%
Off input
$1,029
Saved
Daily cost by model7D
M
T
W
T
F
S
S
opus-4.6gpt-5.4gemini-3.6-flashdeepseek-v4-prokimi-k3
By model30D
opus-4.6Anthropic34.2%$427
gpt-5.4OpenAI25.5%$318
gemini-3.6-flashGoogle19.6%$244
deepseek-v4-proDeepSeek12.7%$158
kimi-k3Moonshot8.0%$100

Plug and play

Integrate in a minute

Point your base URL at Requesty and keep everything else. One line of your code changes.

  • OpenAI-compatible API, works with any SDK
  • Your prompts stay yours: switch model or leave with one line
  • Automatic failover and load balancing included
Read the quickstart
from openai import OpenAI
client = OpenAI(
    api_key="YOUR_REQUESTY_KEY",
    base_url="https://router.requesty.ai/v1",
)
resp = client.chat.completions.create(
    model="anthropic/claude-fable-5",
    messages=[{"role": "user", "content": "Hello"}],
)

EU customers swap the host for router.eu.requesty.ai. Same API, same key, and the request terminates in Frankfurt against EU-hosted model endpoints.

That one base URL reaches:

  • Anthropic
  • OpenAI
  • Google
  • Mistral
  • DeepSeek
  • Moonshot AI
  • Z.ai
  • xAI
  • Alibaba
  • AWS Bedrock
  • Azure OpenAI
  • Meta
  • and every other major provider

01Route

Intelligent infrastructure

Where a request goes, what it is allowed to cost, and what happens when the provider behind it fails.

Region routingONE ENDPOINT
Gateway regions and what the routing policy on the key does with a request from each
RegionServesPolicy
eu-central-1pinnedEU, FrankfurtEU models only
us-east-1US, Virginianearest healthy
ap-southeast-1APAC, Singaporenearest healthy
Request from an EU key for a US-only modelrefused, not forwarded
A pin is not a preference. A call that would leave eu-central-1 returns an error rather than crossing the border.

02See

Real-time observability

Cost, latency and usage across every provider, current enough that you find a problem while it is still small.

Cost attributionLAST 30D
30d spend
$1,247
Prev 30d
$1,422
Change
-12.3%
Per 1K req
$8.71
The six largest spend lines this month, by team, key and model
TeamKeyModelSpend
platformprod-apiopus-4.6$101
supporttriage-svcopus-4.6$79
platformprod-apigpt-5.4$76
supporttriage-svcgpt-5.4$59
platformprod-apigemini-3.6-flash$58
supporttriage-svcgemini-3.6-flash$45

03GovernSecurity

Governance and guardrails

Sensitive data scrubbed before the outbound call, policy enforced per team, user and key.

PII detection and scrubbing

Personal data redacted before it reaches the model. Emails, phone numbers, national IDs and cards, among eight entity types.

PII redaction8 ENTITY TYPES
Your application sends
Customer sarah.chen@northwind.co called about invoice 8841. Callback +44 7700 900412, card ending 4242.
Scrubbed at the gateway, 3 of 8 types matched
The model receives
Customer [EMAIL] called about invoice 8841. Callback [PHONE], card ending [CARD].
Detection runs in your region. 8 types: email, phone, credit card, national ID, IBAN, IP address, postal address and date of birth.

Content guardrails

Enforce content policies, block prompt injections and filter harmful outputs, on the way in and on the way out.

GuardrailsLAST 24H
Prompt injectionBlocking4,773 checked
Content policyBlocking4,773 checked
PII redactionRedacting318 scrubbed
Topic filterBlocking4,773 checked
Toxicity shieldBlocking12 blocked
Rate limiterThrottling0 throttled
Every guardrail is enforcing. Two of them had something to act on in the last day.

Team management

Role-based access with Owner, Admin, Billing, Developer and Viewer roles. Set per-team budgets, model allowlists and usage quotas.

Roles and permissions37 MEMBERS
Permissions by role
RoleKeysBudgetsBilling
OwnerYesYesYes
AdminYesYesNo
BillingNoYesYes
DeveloperYesNoNo
ViewerNoNoNo
Five roles, plus per-team budgets, model allowlists and usage quotas on top.

Audit logs

Complete audit trail of every action. Track who did what, when and from where. Export logs for compliance and review.

Audit logEXPORTABLE
Recent audit events
TimeActionOrigin
14:22j.okafor raised the research budget to $260/moSSO
14:09ci-eval-runner key rotatedAPI
13:54p.nowak pinned acme-prod to eu-central-1SSO
13:31d.mensah added gpt-5.4 to the support allowlistSSO
12:47Guardrail toxicity-shield blocked 4 requestsSYSTEM
Every action, actor, timestamp and origin. Streamed to your SIEM or exported as JSON.

In production

What it looks like when the gateway stops being your problem

We evaluated over 7 different gateways and ended up with Requesty for their enterprise readiness and configurability.
Arkady Landes
Senior Cloud DevOps Manager, ZoomInfo
Read the ZoomInfo story
2.45M
Requests, 100% EU
77B
Tokens, ~31K per request
Requesty became our default AI control plane. We control where data goes, which models run on which tasks, and what we spend, from one place. For a regulated business, that is exactly the layer we needed.
Rechtsanwalt Michael Friedmann
Vorstand, CLIO (Chief Legal Innovation Officer), anwalt.de services AG
Read the anwalt.de story
100%
Traffic, EU residency
~50%
Tokens served from cache
Requesty sits in the critical path of every phone call we answer, and we simply stopped thinking about it. Calls are faster, costs are lower and the data stays in Europe.
Ignacio Ruisรกnchez
Founder and CEO, NotarBot
Read the NotarBot story
Also in productionAll case studies
99.95%
success rate or better, every month so far, Mozart AI
40%
lower AI cost per meeting, Amberscript
3x
faster rollout of new models, LOUPZ
330+
football clubs on the platform, Soccerment

Review and residency

Built to survive procurement

The questions security and legal ask, answered before they ask them, including the one where the answer is not yet.

In progress
SOC 2 Type II

Observation period under way with an independent auditor. Controls in place and documented; the target date is on the trust page.

Contractual
GDPR, DPA on request

An Article 28 DPA, signed on request at any spend, with sub-processors listed.

Regional
EU residency

Frankfurt, on AWS eu-central-1, via router.eu.requesty.ai.

Retention
Zero data retention, pinnable

Pin your traffic to endpoints that retain nothing. 131 of the EU model endpoints qualify.

Pricing

One number, and it sits on top of the model cost

No per-seat pricing, no feature gates, no minimum. You pay for what your application spends, plus 5%.

Pay as you go
5%
markup on model cost

Every feature included: routing, caching, failover, guardrails, analytics and audit logs. No seat fees, no feature tiers.

See pricing
0%Bring your own keys
on your own provider contracts

Keep the pricing you have negotiated with each provider and still get routing, limits and observability across all of them.

How BYOK works
VolumeEnterprise
discounts and committed spend

SSO, SCIM, private regions, a signed DPA and a named contact. Procurement paperwork handled by people, not a form.

Speak to a founder

Questions and answers

What is Requesty?

An AI gateway between your app and 600+ models from every major provider. Change your base URL to router.requesty.ai and get intelligent routing, fallbacks, cost optimisation, auto-caching, governance and observability.

How do I integrate?

One line of code: client = OpenAI(base_url='https://router.requesty.ai/v1', api_key='your-key'). Works with all major SDKs.

How does it reduce costs?

Smart routing to cheaper equivalent models, auto-caching, automatic fallback from expensive providers, per-user spending limits, and real-time cost analytics.

Does it work with Cursor, Cline and Continue?

Yes. Native OAuth integrations, any model, and no rate limit from us. The only ceilings are the ones you set on the key.

How does pricing work?

5% markup on model costs. All features included. Enterprise plans available with volume discounts.

Can I use my own API keys?

Yes. Bring your own keys for any provider while getting Requesty's routing and observability. Or use our unified key.

Pricing and plan comparison ย ยทย  Docs ย ยทย  EU residency

Start building with Requesty

One line of code. 600+ models. One key.

Free tier, no card. Or email sales@requesty.ai.