Requesty
Case Study|Enterprise AI Gateway

How anwalt.de, Germany's Largest Legal Marketplace, Powers Multi-Model AI on EU Sovereign Infrastructure with Requesty

anwalt.de services AG · Legal Tech, Online Legal Services

2.45M
Requests routed, 100% EU
77B
Tokens processed
9
Routing policies across 58 models
100%
Requesty uptime, zero downtime

Switching a workload from one model to another used to be a project. With Requesty policies, it is a configuration change. That is the productivity unlock for a small engineering team shipping real legal AI features.

Rechtsanwalt Michael Friedmann
Rechtsanwalt Michael Friedmann
Vorstand, CLIO, anwalt.de services AG

About anwalt.de

anwalt.de services AG operates anwalt.de and Frag-einen-Anwalt.de, two of the most established legal platforms in the German-speaking market. The company connects consumers and businesses with qualified attorneys, publishes legal content at scale, and powers digital workflows for thousands of law firms.

The team is building a new generation of AI-assisted legal products: from automated case intake and document summarization to attorney-facing reasoning tools. Every workload must run inside the EU, under strict data protection rules, and stay compatible with the confidentiality expectations of regulated legal practice.

The Challenge

anwalt.de wanted to ship AI features fast without compromising on jurisdiction, cost control, or model choice. Going direct to each AI provider would have meant managing separate contracts, separate regional deployments, and separate billing for every model family the product needed.

  • EU-only data residency, no exceptions. Legal content cannot leave EU regions, which rules out any default cloud routing that may fall back to US endpoints.
  • Multi-provider by design. Different legal tasks need different models. Cheap, high-volume summarization. Deep reasoning for complex matters. Tool use for structured extraction. No single provider covers all three well.
  • Cost predictability at scale. Token spend on frontier models can balloon overnight. The team needed hard guardrails on which workloads are allowed to hit the expensive models.
  • Operational simplicity. One integration, one key, one analytics view. Not five SDKs, five dashboards, and five invoices.

We needed an AI layer that respected EU rules by default and let us pick the right model per task without rewriting our stack every quarter. Requesty gave us both on day one.

Rechtsanwalt Michael Friedmann
Rechtsanwalt Michael Friedmann
Vorstand, CLIO, anwalt.de services AG

Why Requesty

EU sovereign routing

Every request is pinned to EU regions across Google, Anthropic, OpenAI, and open-source providers. Zero US fallback in production.

Policy-based virtual models

Workloads call stable policy IDs like summarizer1m or claude-4.6 instead of raw model names, so the team swaps models behind the scenes without redeploying.

Cost shaping per workload

High-volume traffic is steered to fast, low-cost models. Premium reasoning models are reserved for the small slice of requests that justify them.

Single integration

One API key, one SDK, full observability over tokens, latency, success rate, and policy mix across every provider.

The Results

2.45M · Requests routed, 100% EU

Between late January and late June 2026, the platform served 2,454,707 requests through Requesty. 99.9998% landed on EU regions across Google (europe-west, europe-north, europe-central), Anthropic (eu-west, eu-central, eu-north), and Azure OpenAI (France Central, Sweden Central).

77B · Tokens processed

66.9 billion input tokens, 7.0 billion output tokens, and 3.85 billion reasoning tokens processed. 3.7 billion input tokens served from cache, reducing cost and latency on repeated context without any application changes.

9 · Routing policies across 58 models

Nine routing policies map distinct workload types to the right model. The summarizer1m policy alone handles 81% of traffic on fast Gemini-class models, while the claude-4.6 policy is reserved for high-value reasoning tasks. 58 unique models exercised, fully abstracted behind stable policy IDs.

100% · Requesty uptime, zero downtime

Zero Requesty downtime across the entire deployment. Even when individual providers experienced outages, Requesty's automatic failover kept anwalt.de's AI features online without interruption.

In Their Words

Requesty became our default AI control plane. We control where data goes, which models run on which tasks, and what we spend, from one place. For a regulated business, that is exactly the layer we needed.

Rechtsanwalt Michael Friedmann
Rechtsanwalt Michael Friedmann
Vorstand, CLIO, anwalt.de services AG