2026

22 posts

Claude in the EU: 47 region deployments, one quota wall, and how to fail over without…

DeepSeek V4.1 Flash retires V4 Pro by redirect: on 14 September your pinned model id…

ChatGPT, Claude and Grok went down on the same day: correlated failure is now the default…

The $20 tier is being hollowed out: what to do when the best model lives above your plan

Record on Terminal-Bench, catastrophic on a private eval: Fable 5.1 and the case for your…

Fable 5.1 cut cache reads 75% to $0.25 per million: the launch line nobody screenshotted

Five open weight releases in nine days: GLM-5.3-Flash, Qwen3.8-Flash, Hy4 and the…

20 trillion tokens in 6 days: what the Ox Alpha stealth launch taught us about model IDs

Free model IDs now ship with expiry dates: how to use promotional inference without…

36x the price for 22% more quality: the Pareto data that should decide your model mix

Why AI gateways should be built in Go: a language audit of every major gateway

OpenRouter Is Down? How to Fail Over to a Backup Provider in 2 Minutes

OpenRouter Rate Limits: Why They Happen and How Multi Provider Fallback Fixes Them

The supply side is exploding: 136 providers, 502 models, and 2,084 client apps in a…

Why Traditional Latency Routing Fails and How We Fixed It With One Formula

No model stays #1: the leader's share of gateway traffic fell from 29% to 8% in eight…

Inside Sakana Fugu Ultra: We Reverse Engineered Its Multi Agent Architecture

How to route LLM requests by cost and latency

What the gateway saw in April 2026: agents live on Anthropic, open-source models got…

Agentic routing, benchmarked: Requesty adds 16ms of overhead, OpenRouter adds 55ms

Designing fallback retries: why Requesty uses 500ms → 4s with jitter

Routing policies 101: fallback, load balancing, and latency in production

2025

17 posts

Case Study: How E-commerce Chatbots Scale to Black Friday Traffic with Requesty

Cross-Provider Caching Deep Dive: Maximize Performance Across Your Stack

LLM Gateway vs Direct API Calls: Benchmarking Latency & Uptime

Rate-Limiting, Retries & 429s: Bullet-Proofing Your AI Pipeline

Smart Routing Demystified: Choosing the Fastest-Cheapest Model per Request

Solving Provider Outages: Real-World Failover War Stories

The Future of LLM Routing: On-device, Edge AI, and Federated Models

Top 7 Smart-Routing Strategies (with YAML/JSON Examples)

Smarter-Than-Human Model Picking: Introducing Requesty Smart Routing

Intelligent LLM Routing in Enterprise AI: Uptime, Cost Efficiency, and Model Selection

Introducing Smart Routing: Smart AI Model Selection!

Supercharging Cline with Requesty: Models, Fallbacks, and Optimizations

Handling LLM Platform Outages: What to Do When OpenAI, Anthropic, DeepSeek, or Others Go…

Implementing Zero-Downtime LLM Architecture: Beyond Basic Fallbacks

Claude-3-5-Sonnet: Save Over 50% on AI Costs with Cline & Requesty Router

Switching LLM Providers: Why It’s Harder Than It Seems

What is LLM Routing?