2026

19 posts

Which of the smartest models can you run inside the EU, and what is the residency premium?

Claude in the EU: 47 region deployments, one quota wall, and how to fail over without…

DeepSeek V4.1 Flash retires V4 Pro by redirect: on 14 September your pinned model id…

GLM-5.3 Flash vs DeepSeek V4 Flash 0731: which flash model fits your API workload?

GPT-6 Astra is the best model you may not be allowed to call: gated access, the EU gap…

The list price is now a range: peak hours, promo windows and host floors

GPT-6 Astra scores 61 on the independent index, the same as Sol, at 2.5x the price

Record on Terminal-Bench, catastrophic on a private eval: Fable 5.1 and the case for your…

Fable 5.1 cut cache reads 75% to $0.25 per million: the launch line nobody screenshotted

Five open weight releases in nine days: GLM-5.3-Flash, Qwen3.8-Flash, Hy4 and the…

A 320B model trained on Chinese AI chips at 1/100 frontier price: the inference supply…

20 trillion tokens in 6 days: what the Ox Alpha stealth launch taught us about model IDs

Claude Opus 5 vs GPT-5.6 Sol: API speed, coding and cost

Free model IDs now ship with expiry dates: how to use promotional inference without…

36x the price for 22% more quality: the Pareto data that should decide your model mix

The supply side is exploding: 136 providers, 502 models, and 2,084 client apps in a…

No model stays #1: the leader's share of gateway traffic fell from 29% to 8% in eight…

Inside Sakana Fugu Ultra: We Reverse Engineered Its Multi Agent Architecture

Best AI Coding Model (2026): Benchmarks, Cost, and Real World Performance