Built for enterprise teams

AI Gateway for Enterprise Teams

Secure, compliant, and scalable AI infrastructure. Deploy 600+ models with enterprise-grade governance, security, and observability.

View documentation

Or email us at sales@requesty.ai

600+
AI models
225B+
Tokens per day
99.99%
Uptime SLA
131
EU ZDR endpoints

Trusted by leading companies

ShopifyAmadeusChargebeeDemandbasePfizerPwCCapgeminiSageSiemensRelevance AIAppnovation

00The control plane

Every control, on every request

Five checks between your SDK call and the provider, and a record after it.

  1. Stage 1 of 8, CallerSDKOne base URL, unchanged application coden/a
  2. Stage 2 of 8, IdentityAuthSSO via Okta, key mapped to a team0.4ms
  3. Stage 3 of 8, PolicyBudgetsupport, $1,800/mo cap, 67% used0.3ms
  4. Stage 4 of 8, PolicyAllowlist3 of 600+ models cleared for this team0.2ms
  5. Stage 5 of 8, GuardrailRedact3 of 8 PII types matched and scrubbed1.8ms
  6. Stage 6 of 8, RoutingRegionPinned eu-central-1, Frankfurt0.6ms
  7. Stage 7 of 8, UpstreamProvideropus-4.6, warm prefix at cache-read raten/a
  8. Stage 8 of 8, RecordAuditActor, action, origin, streamed to your SIEMasync

Five checks, 3.3ms in front of a call that takes seconds.Measured at the Frankfurt gateway, p50

01GovernGovernance

Enterprise governance

Manage your entire AI infrastructure with precision. Set policies, control access, and maintain compliance across every team.

Permissions by role4 ROLES
What each role may do
RoleAPIBudgetsBillingPolicies
OwnerYesYesYesYes
AdminYesYesNoYes
MemberYesNoNoNo
ViewerNoNoNoNo
Four roles, and the matrix is the whole definition: there is no fifth column hidden in a plan tier.

02SecureSecurity

Security and compliance

Built with enterprise security at the core. Your data stays private, protected, and under your control at all times.

Compliance register4 CONTROLS
SOC 2 Type IIIn progressObservation period under way with an independent auditor
GDPRCompliantFull EU data protection regulation compliance
DPAOn requestArticle 28 processor agreement, signed at any spend
ISO 27001In progressTarget date published on the trust page
SOC 2 Type II is an observation period, not a certificate. The target date is on the trust page.
Data handling8 ENTITY TYPES
Your application sends
Customer sarah.chen@northwind.co called about invoice 8841. Callback +44 7700 900412, card ending 4242.
Scrubbed at the gateway, 3 of 8 types matched
The model receives
Customer [EMAIL] called about invoice 8841. Callback [PHONE], card ending [CARD].
Encryption and privacyEnd-to-end encryption in transit and at rest. Zero data retention policy. We never train on your data or store prompts beyond processing.
Data residencyChoose where your data lives. EU requests stay in Frankfurt, US in Virginia. A call that would leave the pinned region is refused, not forwarded.
Automatic detection and redaction of personal data before it reaches any model. Detection runs in your region: email, phone, credit card, national ID, IBAN, IP address, postal address and date of birth.

At scale

The gateway is the boring part

Three numbers decide whether an AI layer stops being a topic.

03RouteInfrastructure

Intelligent infrastructure

Enterprise-grade routing, caching, and reliability. Built for production workloads at any scale.

Failover, last event14MS TOTAL
StepTook
Endpoint degraded2ms
Health check6ms
Pick next provider3ms
Switch traffic3ms
Total14ms
The retry happens inside the gateway, so the caller sees one request. Warm prefixes bill at the provider's cache-read rate.
Region routingONE ENDPOINT
Gateway regions and the policy applied to each
RegionServesPolicy
eu-central-1pinnedEU, FrankfurtEU models only
us-east-1US, Virginianearest healthy
ap-southeast-1APAC, Singaporenearest healthy
Route to the nearest region automatically, or pin one. The routing policy on the key decides, not the caller.

04SeeObservability

Enterprise observability

Complete visibility into your AI infrastructure at scale. Monitor costs, performance, and usage across all teams and providers.

Enterprise cost and performanceLAST 30D
30d spend
$12,470
Requests
1.94M
p95 latency
950ms
Success rate
99.85%
Cost / day
$415
p95 latency
950ms
Requests / hr
2.7K peak
Errors, 24h
11
Track spending by model, team, user and project in real time, and set alerts before budgets are exceeded. Alerts fire on the p95, not the mean.
By agent92.3K CALLS, 30D
support-triageHealthy42.1K calls$182
p50 165msp95 500ms
code-reviewerHealthy18.8K calls$126
p50 545msp95 1.4s
data-extractorHealthy31.4K calls$259
p50 285msp95 835ms
Track latency, cost and success rates per agent. Identify bottlenecks, optimise routing, and hold an SLA for every agent in the fleet.
Usage by teamLAST 30D
platform41%795K
support26%504K
research19%368K
growth14%271K
Understand which models, teams and users drive consumption, so model selection and capacity planning start from data rather than from an argument.

One line to integrate

OpenAI SDK compatible

Switch in seconds. Change one line of code and you're running on Requesty with full enterprise governance.

  • Works with existing code
  • Existing SDK calls unchanged
  • Instant governance layer
integration.py
# Before: OpenAI
client = OpenAI(
    api_key="sk-..."
)

# After: Requesty
client = OpenAI(
    api_key="req-...",
    base_url="https://router.requesty.ai/v1",
)

# That's it. Full enterprise governance applied.

EU customers swap the host for router.eu.requesty.ai. Same API, same key, and the request terminates in Frankfurt.

Simple process

How it works

Get started with enterprise-grade AI infrastructure in three straightforward steps.

Step 01
Talk to our team

Schedule a call to discuss your requirements, compliance needs, and integration points. We’ll design a solution tailored to your organization.

Step 02
Custom onboarding

We’ll set up your workspace with your policies, connect your IdP, configure team permissions, and migrate your existing infrastructure.

Step 03
Go live

Deploy with one line of code. We provide hands-on support during migration, optimization, and scaling to ensure a smooth transition.

Enterprise FAQ

Common questions from enterprise teams evaluating Requesty.

Does Requesty support SSO?

Yes. Requesty supports SSO via Okta, Azure AD, and Google Workspace on enterprise plans, with full SAML and OIDC support.

Is there a free trial?

Yes. You can start on the pay-as-you-go plan with no commitment, then upgrade to enterprise when you need SSO, RBAC, dedicated support, or custom SLAs.

Can Requesty be self-hosted?

Not today. Requesty is a fully managed cloud platform with EU data residency available in Frankfurt. Self-hosting is not offered at this time.

How does RBAC work?

Role-based access is fully managed on the platform. Admins see all activity, spend, and keys across the organization. Users only see their own keys, usage, and logs. Groups and team-level budgets let you delegate further.

What models can our team use?

All 600+ models by default. Enterprise plans let admins restrict access to an approved list of models and providers, so users can only call models your organization has vetted.

How is enterprise pricing calculated?

Enterprise plans are priced on request. Contact our team to discuss your volume, required features (SSO, RBAC, guardrails, audit logs), and support needs.

Provenance. Model count (600+), tokens per day (225B+), EU zero-retention endpoints (131) and the 14ms failover are platform measurements. Gateway overhead of 3.3ms is the sum of the five checks shown above, measured at the Frankfurt gateway at p50. Uptime SLA of 99.99% is a contractual commitment, not an observed result. SOC 2 Type II is an observation period under way, not a certification. Customer logos are the marks of Requesty customers as published on /customers.

Ready to deploy AI at scale?

Get in touch with our team to discuss your requirements and see how Requesty can transform your organization's AI infrastructure.

View documentation

Or email sales@requesty.ai.