Requesty
Back|SEP '26AI MODELS / PRICING
7 MIN READ|

Grok 4.7 and MiMo-V2.6 Pro launched four hours apart with the same benchmark score: one costs a fifth per token and a twenty fourth per eval

Last updated

Two flagship launches landed on the evening of 21 September, four hours apart, and our tracker watched the conversation move from one to the other in real time.

At 16:17 UTC, SpaceXAI posted Grok 4.7 is here, "a notable improvement over Grok 4.6 at the same price and speed." By our snapshot the next midday it had 27,088 likes, 2,375 bookmarks and 12.3 million impressions. Vercel had it 40% off on its gateway within the hour, and we listed xai/grok-4.7 in the Requesty catalog at 16:48 UTC.

At around 20:00 UTC, Xiaomi released MiMo-V2.6 Pro and Flash with open weights. Nathan Lambert posted the shapes: Pro at 1.02 trillion parameters with 42 billion active, Flash at 310 billion with 15 billion active, "top open model on Artificial Analysis", with an RL dashboard and a tech report. OpenRouter had all three variants live at 21:08, all taking text, image, video and audio with a 1M context. TestingCatalog reported a sixth place on the Intelligence Index, "on par with Claude Opus 5 and GPT-5.6 Sol."

Then the verdict from the account that had spent the afternoon hyping Grok. Chubby, 145,000 followers, at 22:14 UTC that the real deal was "MiMo-2.6! Better than Grok 4.7, much cheaper and holy moly is china back." 892 likes, 138,000 impressions. Deedy Das ran the cost math the next morning: under agentic coding assumptions, MiMo-V2.6 Pro comes out 15 times cheaper than Kimi K3, 6 times cheaper than GLM-5.3 and twice as expensive as DeepSeek V4.1 Flash.

Hourly tracked posts naming Grok 4.7 and MiMo-V2.6, 21 September 12:00 to 22 September 12:00 UTC
Hourly tracked posts naming Grok 4.7 and MiMo-V2.6, 21 September 12:00 to 22 September 12:00 UTC

The attention curve

In the 24 hours from 12:00 UTC on 21 September, our tracker logged 88 posts from 75 distinct authors naming Grok 4.7 and 110 posts from 86 authors naming MiMo. Grok's curve is the classic API launch shape we have measured on every closed flagship this year: a spike of 21 posts in the announcement hour, 9 the next, then single digits by 19:00 UTC. MiMo's curve is 22 and 23 posts in its first two hours, 13 in the third, and it was still producing 9 posts an hour at 08:00 the next morning as Asia and Europe woke up to it.

The launch tweets tell the same story from the other direction. SpaceXAI's post reached 12.3 million impressions on the back of a two million follower account. The MiMo conversation had no single anchor; the largest post about it, Chubby's reversal, did 138,000 impressions. The open weight model won the second day on distributed word of mouth, which is the pattern we documented when five open weight releases landed in nine days in August.

Same score, different bill

Artificial Analysis had both flagships scored within a day. Its Grok 4.7 page (evaluated at the xhigh reasoning setting) and its MiMo-V2.6 Pro page report an Intelligence Index of 46 for each, which it describes as well above average among comparable models. That is the headline the community ran with, and it is fair as far as it goes.

The two pages also report what it cost to produce that score, and those numbers are not close:

Grok 4.7 (xhigh)MiMo-V2.6 Pro
Intelligence Index4646
Input, $ per 1M2.000.435
Output, $ per 1M6.000.87
Cache read, $ per 1M0.500.0036
Tokens generated to complete the index240M140M
Cost to evaluate$4,967$207
Context window500K1M
Weightsclosedopen

Per token, MiMo-V2.6 Pro is about a fifth of the price: 4.6 times cheaper on input, 6.9 times cheaper on output. That alone would make it a compelling swap for a model at the same index score. But the eval bill came in at 24 times less, and the extra multiplier is verbosity. Grok 4.7 generated 240 million tokens to get through the same test set that MiMo covered in 140 million. Artificial Analysis calls it "very verbose" against a median of 94 million for its peer group, and "notably slow" on top.

We flagged the same effect on GPT-6 Astra's independent scores: a reasoning model that writes more to reach an answer costs more per task than its output price suggests, and the list price hides it. If you are comparing models on price per million tokens, you are comparing the wrong number. Tokens per task times price per token is the bill, and Grok 4.7 loses on both factors here.

List prices per million tokens for nine models in the Requesty catalog, 22 September 2026, log scale
List prices per million tokens for nine models in the Requesty catalog, 22 September 2026, log scale

Where each model sits in the catalog

The chart above places both launches against the models a production team is choosing between this week, from the Requesty catalog as of 22 September. A few things stand out.

MiMo-V2.6 Flash is the new floor. At $0.14 in, $0.28 out and $0.0028 per cached token, it lists below DeepSeek V4.1 Flash on every line and below GLM-5.3 Flash on output and cache. The cache read price matters more than it looks: on our platform, open weight Flash tier models run cache hit rates between 81% and 96%, so most of an agent's input bill is the cache line. MiMo's $0.0028 is a tenth of GLM-5.3 Flash's $0.03.

Grok 4.7 is priced as a frontier model with a mid tier index score. Its $2 and $6 sit between the open weight group and GPT-5.6 Sol at $4 and $20. Kimi K3, the one open weight model priced like a frontier model at $3 and $15, is now more expensive than Grok 4.7 on both lines, which is going to be uncomfortable for Kimi K3's hosts.

The cache read spread is 350 to 1. From MiMo-V2.6 Flash at $0.0028 to GPT-6 Astra at $1.00. For a coding agent that reads the same repository context again on every turn, that line dominates the invoice, and we walked through the mechanics in why cache hit rate decides your agent bill.

Both MiMo variants are in the Requesty catalog at Xiaomi's first party price as of 22 September, at xiaomi/mimo-v2.6-pro and xiaomi/mimo-v2.6-flash. Third party hosts will follow, and their prices will diverge; for the previous DeepSeek Flash release the host spread on the cache read line alone reached 13 times.

What we do not know yet

An index score is one number. Nobody outside Xiaomi has run MiMo-V2.6 Pro for a week on real agent traffic, and the first reports are mixed in the way they always are on day one. One OpenRouter user reported the model looping the same tool calls three times in a session. Grok 4.7 has a 500K context against MiMo's 1M, and Grok is text and image only where MiMo takes video and audio. Reasoning settings differ: the Grok score is at xhigh, and a lower setting would be cheaper and probably lower scoring.

The honest position for a production team is the one we take with every launch. Do not swap on the index score. Run both through the eval harness you already have on your own prompts, measure tokens per task as well as pass rate, and let a routing policy send the requests that MiMo handles to MiMo, with Grok or a closed frontier model as the fallback for the ones it does not. A Pareto view of price against quality is the tool for deciding where that line sits.

What we can say with confidence is narrower and still useful: two models scored identically this week, one of them costs a fifth per token and writes 40% less to get there, and the cheaper one has open weights, which means its price will fall further as hosts compete. That is the third time this pattern has held in six weeks. The cheapest models ranking updates daily if you want to watch it happen.

Methodology

Launch timing and post counts come from our public post tracker, which polls X and Reddit for model and gateway conversation every ten minutes; counts cover 21 September 12:00 to 22 September 12:00 UTC and match posts naming each model by regular expression. Engagement figures for quoted posts were refetched from the X API at 12:30 UTC on 22 September so they reflect at least twelve hours of maturity. Prices are first party list prices from the Requesty catalog on 22 September, cross checked against SpaceXAI's published model page for Grok 4.7. Benchmark figures, token counts and evaluation costs are Artificial Analysis's own, read from its model pages on 22 September; we have not independently reproduced them.

Frequently asked questions
How much does Grok 4.7 cost per million tokens?
SpaceXAI lists Grok 4.7 at $2.00 per million input tokens and $6.00 per million output tokens, with cached input at $0.50, the same as Grok 4.6. Input above the standard context threshold is billed at double. The Requesty catalog carries the same first party price at xai/grok-4.7 with a 500K context window.
How much does Xiaomi MiMo-V2.6 Pro cost?
On first party routes MiMo-V2.6 Pro lists at $0.435 per million input tokens, $0.87 per million output and $0.0036 per million cached input, with a 1M token context. MiMo-V2.6 Flash lists at $0.14 in, $0.28 out and $0.0028 cached. Xiaomi also offers a Pro Ultraspeed variant at $4.35 in and $8.70 out. Third party hosts will set their own prices as the weights spread.
Is MiMo-V2.6 Pro as good as Grok 4.7?
On the Artificial Analysis Intelligence Index both score 46, which places them well above the median comparable model. Grok 4.7 was evaluated at its xhigh reasoning setting. A shared index score does not mean identical behaviour on your workload: the two models differ in verbosity, speed, modality support and context length, and the only way to know which one handles your prompts is to run them through your own eval harness.
Why did the Grok 4.7 benchmark run cost 24 times more than MiMo-V2.6 Pro?
Two effects compound. Grok 4.7 is roughly 4.6 times more expensive on input and 6.9 times on output. It also generated 240 million tokens to complete the index against 140 million for MiMo-V2.6 Pro, so it wrote 1.7 times as much to reach the same score. Artificial Analysis reports $4,967 to evaluate Grok 4.7 and $207 for MiMo-V2.6 Pro. Output verbosity is a cost multiplier that the price list does not show.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders