Requesty
Case Study|Enterprise AI Gateway

How Ethira Scaled an EU-First AI Platform Across 13 Providers with Requesty

Ethira · Information Security, Third Party Risk Management

4.4x
Request growth in seven months
61
Models from 13 providers, one API
23
API keys, one pane of glass
93%
Of traffic on EU regional routing
41.65%
Cache hit rate after optimization

Our volume can triple in a month when big customers onboard. Requesty just absorbs it. We have never had to think about whether the model layer keeps up.

Duncan Mills
Duncan Mills
Customer Success, Ethira

About Ethira

Ethira is a Third Party Risk Management platform that understands risk from how third party providers are used and integrated into systems, processes and services, in relation to threat intelligence, vendor security posture and markers of good and bad practice.

The product is AI-native from the ground up: conversational assistants, deep document analysis over security evidence, research workflows, web search and autonomous agents all run on large language models. Ethira serves European customers with European expectations on data residency, so every one of those workloads has to stay inside the EU.

Growth on Requesty

Ethira came to Requesty in November 2025 and scaled fast. Monthly request volume grew 4.4x in seven months, from 115K requests in the first month to over 500K by June 2026, and the road there was anything but linear: in their second month alone, volume jumped nearly 7x as the product took off. The gateway absorbed every step of that curve on the same single integration.

  • Built for spikes. When Ethira's heaviest document processing waves hit, monthly volume swung from 200K to over 500K requests within weeks. No re-architecture, no capacity planning calls, no new integrations. The gateway scaled with them.
  • One endpoint, every provider. Ethira integrated once through the OpenAI-compatible API and immediately had 19 models from 7 providers in play.
  • Constant model experimentation. 61 distinct models from 13 providers used all time, peaking at 35 models from 11 providers in a single month. New frontier models go into production testing the week they ship.
  • Caching optimization. In June 2026 the team restructured prompts for cacheability and pushed their cache hit rate from single digits to 41.65% in one month, a major efficiency win on document-heavy workloads.

Switch Any Model, Any Provider, Anytime

Ethira's product spans very different workloads: high-volume conversational AI, deep document analysis over security evidence, lightweight completions, and tool-calling agents. Each gets the model that fits, and the mix changes constantly as the model landscape evolves.

  • The right model per workload. Fast, cost-efficient models carry the high-volume conversational traffic, while frontier models with large context windows handle heavy document analysis. Lightweight tasks run on near-free models.
  • Switching is a config change. Moving a workload from one provider's model to another's takes minutes through the gateway. No new SDKs, no new auth, no redeploys. Ethira has moved workloads freely across OpenAI, Anthropic, Google, Bedrock, Azure, and Vertex.
  • Same-week frontier adoption. When a new model generation ships, Ethira benchmarks it against the incumbent on real production traffic and promotes it if it wins. Up to 35 models from 11 providers have been in active use in a single month.
  • Agents included. Roughly one in ten completions involves tool calls. Agentic workloads route through the same gateway with the same flexibility.

Changing the model behind a feature used to be an engineering decision. Now it's a product decision. We pick whatever is best this month and move on.

Fredrik Gustafson
Fredrik Gustafson
Head of Product, Ethira

Full Visibility, From Production to Every Preview Environment

Ethira runs a sophisticated engineering operation: production plus more than 11 color-coded staging and preview environments, with 23 scoped API keys across them. Requesty gives the team one pane of glass over all of it.

  • Per-key, per-environment observability. Every environment has its own keys, so usage, latency, and errors are attributable to the exact service and branch that generated them. Production traffic and preview experiments never blur together.
  • Per-model and per-provider analytics. The team sees exactly how every model performs on their real workloads: latency, success rates, and usage, side by side, making model decisions data-driven instead of anecdotal.
  • Request-level traceability. Every request is logged at the gateway. For a company whose product is AI accountability, having a complete audit trail of its own AI usage is not optional, and with Requesty it comes built in.

Every environment, every key, every model, one dashboard. When something looks off in a preview branch, we see it in seconds, not after it ships.

Rafael Correia
Rafael Correia
Founding Engineer, Ethira

EU-First by Design

Ethira serves European customers with European expectations on data residency. 93% of their traffic runs through Requesty's EU deployment, with workloads pinned to specific European regions: Claude on eu-central-1, Kimi on eu-west-2, GPT models on francecentral. Requesty's regional routing keeps inference inside the EU without giving up the flexibility to use every provider.

  • EU data residency. Inference stays in European regions across providers, aligned with the compliance posture Ethira's own customers demand.
  • Regional pinning per model. Each model routes to its designated EU region, configured at the gateway, not scattered across application code.
  • Redundancy without leaving Europe. Multiple providers and regions are available behind the same EU-resident endpoint, so flexibility never comes at the cost of residency.

Our customers ask where their data goes. With Requesty we answer in one sentence: it stays in the EU, across every provider we use.

Lucas de Araujo
Lucas de Araujo
CTO, Ethira

The Results

4.4x · Request growth in seven months

From 115K monthly requests in November 2025 to over 500K by June 2026, including a near-7x single-month spike, all on the same single gateway integration.

61 · Models from 13 providers, one API

OpenAI, Anthropic, Google, Bedrock, Azure, Vertex, and more, with up to 35 models in production use in a single month and switching done in minutes, not sprints.

23 · API keys, one pane of glass

Production and 11+ preview environments each run on scoped keys with full per-key observability: usage, latency, and errors attributable to the exact service that generated them.

93% · Of traffic on EU regional routing

Region-pinned deployments across providers keep inference inside the EU, matching the data residency story Ethira sells to its own customers.

41.65% · Cache hit rate after optimization

Up from single digits in one month, after the team restructured document-analysis prompts for cacheability through Requesty's automatic prompt caching.

In Their Words

Requesty grew with us. We started with one endpoint, and seven months later we run an EU-resident, multi-provider platform with full observability on exactly the same integration.

Lucas de Araujo
Lucas de Araujo
CTO, Ethira

EU-resident AI infrastructure that scales with you

One line of code. 600+ models. Full control.

Speak to founders