Requesty
Live from production

Watch our router pick the fastest provider, in real time.

Requesty routes every latency-optimized request with Thompson Sampling: one draw from each provider's live latency distribution, lowest draw wins. Sending everything to one provider is the naive strategy this replaces; the traffic split you see is the router exploiting the fastest option while auditing the rest. Refreshed every minute from production traffic. This applies to requests using our latency routing policy: enable it on your API key to get routed like this.

131model families
271provider candidates
46live matchups
1.8ssaved per request vs provider-blind picking

Live traffic share by model family46 matchups

Where our router sends each model's traffic right now. The share that doesn't go to the leader is deliberate: continuous probing of every provider, so degradations are caught within minutes.

gpt-4.1-nano

1azure @francecentral80.5%481ms713ms$0.112openai17.0%596ms901ms$0.013azure @swedencentral2.4%1.19s2.77s7%$0.144azure0.1%1.02s1.72s$0.14
fastest provider, last 24h2 lead changes

gemini-2.5-flash-lite

1vertex @europe-west113.6% err20.1%1.29s8.12s4%$0.192vertex @europe-west81.8% err19.0%1.19s5.52s2%$0.133vertex @europe-central21.5% err15.7%1.30s6.58s$0.134vertex @europe-west48.1% err15.5%1.28s6.94s5%$0.175vertex @us-east113.1%1.35s3.65s10%$0.126vertex @europe-north112.5%1.34s5.91s$0.137vertex15.6% err4.0%1.85s4.84s$0.128vertex @us-west10.1%2.41s8.99s23%$0.129google21.9% err0.0%18.5s35.4s$0.3610vertex @us-central10.0%4.24s11.8s54%$0.08
fastest provider, last 24h9 lead changes

deepseek-v4-flash

1sference78.1%1.41s8.14s89%$0.082fireworks13.2%3.58s11.6s15%$0.173deepseek3.8%4.75s15.0s50%$0.094tensorx3.2%7.82s31.7s30%$0.125novita1.7%5.38s15.2s92%$0.14
fastest provider, last 24h6 lead changes

gemini-3.1-flash-lite

1google38.0%1.84s2.88s16%$0.292vertex @us29.7%1.82s3.27s29%$0.283vertex @eu26.2%2.06s4.20s7%$0.344vertex2.7% err6.1%2.39s4.19s2%$0.44
fastest provider, last 24h14 lead changes

claude-haiku-4-5

1vertex83.4%860ms1.21s$1.732vertex @europe-west11.1% err12.5%1.76s6.72s32%$1.193bedrock @eu-central-12.0%1.89s4.31s30%$1.024bedrock @eu-west-11.1%1.79s4.12s2%$1.395bedrock @eu-north-10.7%2.12s4.83s2%$1.516anthropic0.2%12.6s31.2s16%$1.697bedrock @eu-west-30.1%2.81s6.25s58%$0.76
fastest provider, last 24h16 lead changes

gemini-3-flash-preview

1vertex71.5%4.79s13.1s2%$1.062google2% err28.5%6.38s16.1s15%$1.06
fastest provider, last 24h1 lead change

gpt-4.1-mini

1openai69.1%1.01s1.80s4%$0.112azure @francecentral29.5%1.27s2.39s7%$0.393azure @westus30.8%3.48s15.3s$0.774azure @eastus20.5%8.89s63.5s$0.86
fastest provider, last 24h5 lead changes

gemini-2.5-flash

1coding22.6%2.99s6.00s$1.042coding @europe-west819.4%2.60s4.66s$1.173coding @europe-north117.8%2.96s5.89s$1.224vertex @us-east117.3%4.16s4.19s$1.825coding @europe-central210.2%4.07s12.9s$1.496coding @europe-west410.1%3.23s9.89s$1.177coding @europe-west11.6%4.01s6.12s$1.358vertex7.3% err0.6%6.55s18.2s2%$1.369vertex @europe-west10.2%10.4s32.6s12%$1.4110google0.0%14.8s38.9s$1.6211vertex @europe-west414.6% err0.0%13.3s60.7s5%$1.3912vertex @europe-north18.1% err0.0%15.8s27.2s2%$1.4713vertex @europe-west818.2% err0.0%18.7s31.7s4%$1.3814vertex @us-south10.0%11.0s22.9s$1.77
fastest provider, last 24h22 lead changes

gpt-5.6-luna

1azure @swedencentral0.5% err86.5%2.14s5.43s22%$1.212openai10.5%6.43s26.6s71%$0.463azure @eastus23.0%7.67s22.5s90%$0.22
fastest provider, last 24hno lead changes

gpt-4.1

1azure @francecentral56.2%1.05s1.77s63%$1.132azure @swedencentral30.0%1.34s5.09s38%$1.633openai13.9%1.52s3.57s14%$2.294azure0.0%18.8s33.6s20%$3.09
fastest provider, last 24h11 lead changes

gpt-4o-mini

1azure @swedencentral88.0%888ms2.35s1%$0.162openai12.0%1.29s2.09s$0.21
fastest provider, last 24h9 lead changes

gpt-5-mini

1azure @francecentral83.2%2.28s5.21s5%$0.572azure @swedencentral15.2%4.19s11.0s26%$0.973azure1.5%6.81s16.5s15%$0.524azure @uksouth0.0%10.7s22.8s5%$0.475openai0.0%10.3s16.5s4%$1.05
fastest provider, last 24h15 lead changes

glm-5.2

1sference75.4%1.85s21.3s92%$0.382tensorx13.4%5.63s32.0s87%$0.583fireworks7.0%7.11s50.2s87%$0.364zai2.9%8.33s69.7s50%$1.935inceptron0.6%18.3s106.3s66%$0.886nebius0.6%23.8s118.9s$2.24
fastest provider, last 24h2 lead changes

claude-opus-4-8

1bedrock @eu-west-338.5%3.45s18.1s41%$3.922anthropic21.2%4.68s22.2s30%$5.723vertex @eu13.9%5.96s40.7s84%$1.844vertex12.3%5.23s25.9s73%$4.185bedrock @eu-north-15.5%7.89s46.9s80%$1.966bedrock @eu-central-14.4%7.68s33.4s81%$2.007bedrock @eu-west-13.3%9.68s47.6s60%$3.188bedrock1.0%9.93s34.3s34%$4.66
fastest provider, last 24h14 lead changes

claude-opus-4-5

1vertex85.7%3.66s7.04s7%$6.542bedrock11.1%6.99s17.8s12%$6.023anthropic3.2%7.65s14.8s6%$6.03
fastest provider, last 24h11 lead changes

claude-sonnet-4-6

1vertex @europe-west118.4%4.71s35.1s56%$2.102bedrock @eu-west-116.2%4.57s29.1s42%$2.943bedrock @eu-central-114.0%5.25s35.3s64%$2.054bedrock @us-east-113.9%6.98s48.8s$5.915bedrock @eu-west-312.5%4.99s46.1s84%$1.146vertex11.0%4.37s8.07s17%$3.887anthropic7.4%7.59s76.0s56%$2.518bedrock3.4%6.36s31.1s$5.839bedrock @eu-north-13.1%6.62s25.2s87%$0.94
fastest provider, last 24h17 lead changes

gpt-5.4

1azure @swedencentral34.1%2.54s20.1s38%$3.052azure @francecentral29.4%2.54s17.9s60%$1.633azure @eastus226.1%2.53s14.5s15%$3.094openai1.8% err10.4%3.65s22.7s49%$2.08
fastest provider, last 24h15 lead changes

gemini-3.5-flash

1vertex48.9%4.75s27.5s36%$2.002vertex @eu5.1% err29.4%5.52s21.9s20%$2.453vertex @us21.8%6.99s29.1s21%$4.40
fastest provider, last 24h9 lead changes

claude-sonnet-5

1vertex @eu40.3%5.34s37.9s78%$0.852bedrock @eu-central-115.2%7.87s47.3s35%$2.693bedrock13.6%6.87s22.1s58%$1.144anthropic13.6%8.00s43.2s66%$1.085vertex6.3%12.5s41.3s86%$0.636bedrock @eu-west-26.1%11.7s35.9s79%$0.827bedrock @eu-north-15.0%10.2s34.6s65%$1.14
fastest provider, last 24h13 lead changes

gpt-5.4-mini

1azure83.8%721ms1.33s1%$1.072openai15.2%1.27s5.33s18%$0.803azure @eastus21.0%1.19s1.65s54%$0.43
fastest provider, last 24h6 lead changes

gemini-3.5-flash-lite

1vertex @eu87.7%964ms1.78s21%$0.332vertex12.3%1.79s2.76s$0.34
fastest provider, last 24h2 lead changes

minimax-m3

1fireworks67.4%3.04s8.21s11%$0.332minimaxi28.5%4.43s40.0s59%$0.313tensorx4.1%10.4s51.1s90%$0.19
fastest provider, last 24h6 lead changes

gpt-5.1

1azure @francecentral51.6%2.38s9.74s2%$2.602azure @swedencentral47.0%2.55s10.8s9%$2.013openai1.4%5.57s8.74s80%$1.10
fastest provider, last 24h8 lead changes

deepseek-v4-pro

1tensorx85.3%2.65s10.4s90%$0.602fireworks10.9%9.36s81.1s23%$2.203deepseek3.2%15.3s83.0s81%$0.164nebius0.6%16.0s49.5s$1.76
fastest provider, last 24h8 lead changes

gpt-5-nano

1openai46.7%6.53s21.4s9%$0.292azure45.9%6.44s13.9s$0.143azure @francecentral7.4%14.4s49.9s29%$0.234azure @swedencentral0.0%31.3s50.4s19%$0.31
fastest provider, last 24h9 lead changes

gemini-3.1-pro-preview

1google76.3%17.6s69.1s30%$2.732vertex16.7% err23.7%34.0s106.0s4%$4.75
fastest provider, last 24h3 lead changes

gpt-5.6-terra

1openai56.5%3.75s13.6s87%$0.802azure @swedencentral1% err26.4%5.78s36.9s72%$1.203azure @eastus217.1%5.52s13.8s67%$1.17
fastest provider, last 24h12 lead changes

claude-opus-4-7

1vertex29.6%3.22s10.8s53%$4.082vertex @eu25.1%3.71s18.0s88%$1.463bedrock @eu-north-119.4%3.65s11.4s65%$3.694bedrock9.6%10.7s113.7s35%$6.705bedrock @eu-central-18.4%5.07s22.5s88%$1.666bedrock @eu-west-14.0%10.7s45.6s39%$5.147anthropic3.9%6.67s26.3s91%$1.11
fastest provider, last 24h13 lead changes

gpt-5.6-sol

1azure @swedencentral1.3% err50.0%6.90s37.4s82%$1.982openai42.6%7.51s24.1s92%$1.113azure @eastus27.4%15.4s61.0s91%$1.11
fastest provider, last 24h7 lead changes

claude-opus-4-6

1vertex @europe-west153.1%6.70s28.9s72%$3.262anthropic23.2%12.4s139.7s63%$3.903bedrock22.3%12.9s103.9s$23.34vertex1.5%30.1s142.6s22%$5.77
fastest provider, last 24h9 lead changes

gpt-5.5

1azure @swedencentral80.3%3.20s15.6s27%$4.942openai19.7%7.78s84.5s37%$5.69
fastest provider, last 24h9 lead changes

gemma-4-26B-A4B-it

1deepinfra52.9%2.13s5.65s$0.082parasail47.1%2.20s6.98s14%$0.13
fastest provider, last 24h5 lead changes

qwen3.7-plus

1fireworks93.3%6.96s21.4s13%$0.982alibaba6.7%17.7s32.4s30%$0.25
fastest provider, last 24h4 lead changes

claude-sonnet-4-5

1vertex26.0%4.24s16.8s26%$3.732anthropic18.3%4.88s19.9s39%$2.923bedrock @eu-central-116.4%3.89s5.82s4%$3.874bedrock @eu-west-314.7%4.09s6.51s$3.885bedrock @eu-north-111.3%4.08s5.87s$3.826vertex @europe-west110.6%8.95s60.8s29%$3.237bedrock @eu-west-12.8%9.21s39.9s15%$7.87
fastest provider, last 24h14 lead changes

gemini-2.5-pro

1coding56.5%7.94s12.6s$5.032vertex @europe-west121.7%10.2s23.2s3%$4.123coding @europe-central212.2%15.5s35.5s3%$2.704google6.9%10.8s17.4s$6.035vertex2.7%14.0s25.4s2%$6.666vertex @europe-central20.0%27.4s57.7s14%$2.43
fastest provider, last 24h10 lead changes

gpt-oss-120b

1groq66.3%936ms2.06s46%$0.252nebius33.7%1.31s4.72s70%$0.20
fastest provider, last 24h10 lead changes

claude-fable-5

1anthropic48.0%10.0s50.2s50%$10.22vertex35.5%14.4s327.9s38%$12.33vertex @eu16.5%24.7s113.7s44%$9.71
fastest provider, last 24h9 lead changes

kimi-k2.7-code

1moonshot54.0%2.69s21.3s95%$0.692tensorx33.6%3.55s9.76s67%$0.833fireworks11.9%5.39s21.4s67%$0.854parasail0.5%24.6s108.6s59%$0.51
fastest provider, last 24h6 lead changes

nemotron-3-ultra-550b-a55b

1nebius96.1%3.00s6.58s$2.022nvidia3.9%19.5s86.4s37%
fastest provider, last 24h12 lead changes

kimi-k2.6

1moonshot77.1%12.9s106.8s84%$0.372inceptron22.9%35.7s297.7s32%$1.07
fastest provider, last 24h7 lead changes

qwen3-30b-a3b-instruct-2507

1alibaba63.4%2.09s7.09s$0.292nebius36.6%2.81s15.2s13%$0.13
fastest provider, last 24h4 lead changes

minimax-m2.5

1bedrock @eu-west-161.3%4.60s16.8s24%$0.402inceptron38.7%5.85s19.5s83%$0.11
fastest provider, last 24h6 lead changes

gemini-2.5-flash-image

1vertex @europe-west199.8%1.44s6.60s$25.22vertex0.2%9.04s10.5s$13.3
fastest provider, last 24h10 lead changes

kimi-k2.5

1bedrock @eu-north-186.0%2.65s14.1s16%$0.812bedrock @us-west-214.0%7.02s21.3s6%$0.86
fastest provider, last 24h5 lead changes

GLM-5.1

1deepinfra56.7%37.0s65.0s50%$0.732zai43.3%45.4s86.6s20%$1.41
fastest provider, last 24h3 lead changes

glm-5.1

1nebius97.9%8.56s25.6s$2.712zai2.1%29.3s47.8s22%$1.29
fastest provider, last 24h2 lead changes

All models with live data85 models

single provider, no routing contest
gpt-5.4-nanoopenai2.10sdeepseek-chatdeepseek2.71sgpt-4oopenai3.85sgpt-5.2openai5.38sQwen3-235B-A22Bdeepinfra5.07sgrok-4-1-fast-reasoningxai12.2sdeepseek-v3.2novita4.74sgrok-4-1-fast-non-reasoningxai1.87skimi-k3moonshot26.5shy3novita16.9smistral-small-latestmistral1.37smistral-large-latestmistral3.65sgrok-4.5xai11.9sdeepseek-reasonerdeepseek4.52sclaude-haiku-4-5-20251001anthropic1.76sDeepSeek-V4-Flashdeepinfra7.37smistral-medium-latestmistral5.08smistral-small-2603mistral960msopen-mistral-7bmistral1.74sgemma-4-31b-itgoogle13.2sqwen-turboalibaba1.75sgemini-3.6-flashvertex6.13sqwen3-32bnebius2.63snemotron-3-super-120b-a12bnvidia8.50sqwen3-maxalibaba1.19sgpt-oss-20bgroq419mslaguna-m.1poolside7.78sgemma-4-31B-itdeepinfra38.0smimo-v2.5xiaomi3.58sMiniMax-M2.5minimaxi2.40skimi-k2.7-Codeinceptron1.45sgrok-4.3xai54.8sQwen3.5-27Bdeepinfra38.7sgpt-5openai10.1sQwen3.5-2Bdeepinfra11.6sgpt-5-chat-latestopenai463msgpt-5-chatopenai1.99smimo-v2.5-proxiaomi8.36sgemini-3-pro-imagevertex45.2sgpt-4o-2024-05-13openai1.56sleanstral-1-5mistral1.50snemotron-3-nano-30b-a3bnvidia5.71sparasail-deepseek-v4-flashparasail13.3smistral-small-2503mistral911mssonarperplexity8.89smistral-medium-3-5mistral2.82sMiniMax-M2.7minimaxi26.5sgemini-3.1-flash-imagevertex29.3sparasail-qwen25-vl-72b-instructparasail27.0sgpt-5.3-codexopenai17.9sgpt-5.3-chatopenai3.10snemotron-3-nano-omni-30b-a3b-reasoningnvidia4.47sgemini-3-pro-previewgoogle13.0sqwen3-235b-a22b-instruct-2507nebius3.42sQwen3.5-35B-A3Bdeepinfra70.9ssonar-properplexity5.89sMiniMax-M2.7-highspeedminimaxi19.6sQwen3-235B-A22B-Instruct-2507deepinfra6.34sNemotron-3-Nano-30B-A3Bdeepinfra9.56sgpt-4o-2024-08-06openai5.74sqwen3.6-plusalibaba5.71sgpt-5.2-codexopenai4.45sDeepSeek-R1deepinfra13.4sKimi-K2.6deepinfra6.85sglm-5.2-fastfireworks6.16sSeed-2.0-prodeepinfra8.88sGLM-4.6zai15.6sQwen2.5-72B-Instructdeepinfra9.48sgrok-4-fastxai8.59so4-miniopenai5.11sgrok-3-minixai10.2sNVIDIA-Nemotron-3-Super-120B-A12Bdeepinfra3.20sDeepSeek-V3deepinfra20.1sdevstral-latestmistral118.1snemotron-3.5-content-safetynvidia317msqwen-2.5-72b-instructnovita16.8sGLM-5.2zai108.2sMiMo-V2.5-Prodeepinfra3.33sqwen3-next-80b-a3b-thinkingnebius20.9sgrok-4xai6.02sqwen-maxalibaba2.46sMeta-Llama-3.1-8B-Instruct-Turbodeepinfra542msgemini-2.5-flash-image-previewvertex6.36sthinkingcap-qwen3.6-27bsference12.8sqwen3.7-maxalibaba18.4s

How the router decides

01

Measure every request

Each production request records time to first token and total latency per provider, model, region, stream mode, and input size. Stats are outlier-trimmed and recency-weighted with a 10 minute half-life, so a provider that degrades shows up within minutes.

02

Sample, don't average

For each request, the router draws one latency sample from every candidate's live distribution (Thompson Sampling) and routes to the lowest draw. Providers with little data get wider distributions, so new candidates keep getting explored without dedicated canary traffic.

03

Penalize unreliability

Candidates returning capacity errors (429s and 5xxs) over the last 15 minutes have their samples inflated by a penalty factor, up to 10x. An unreliable provider loses routing share immediately and wins it back as soon as it recovers.

Reading the board

Win share is the percentage of simulated routing decisions each candidate wins against the others in its group, using the same sampler the router runs in production. Latency is the geometric mean of total response time over the last hour. The winner is not always the lowest average: a candidate with volatile latency or recent capacity errors loses share to a steadier one. Groups split by streaming mode and input size because provider rankings flip between them.

Route with RequestyRead the docsThis data is public: fetch it yourself from /public/v1/routing-leaderboard