Requesty
Back|SEP '26AI MODELS / INDUSTRY
5 MIN READ|

GPT-6 Astra scores 61 on the independent index, the same as Sol, at 2.5x the price

Last updated

OpenAI launched GPT-6, named Astra, on 3 September. Within hours the community had done something useful: it went and looked at the independent numbers.

The figures circulating from Artificial Analysis, reported across several threads, are unflattering relative to the framing. A widely upvoted r/codex post laid them out plainly under the heading Astra is good, but maybe we should calm down with the hype:

Artificial Analysis has Astra at 61, exactly the same as Sol. Muse Spark 1.3 is at 62. Even in coding, Astra is only around 67 vs 65 for Sol. Meanwhile the API price per token is 2.5x higher. The real improvements seem to be automation, tool use, long-running tasks and token efficiency. Those are useful, but that's very different from a massive jump in raw intelligence or coding ability. If you remove the GPT-6 name and the launch hype, the independent numbers look a lot less revolutionary.

Another thread put the comparison in the title: Astra / GPT-6 scores behind Meta Muse in AA Index. Benchmark discussion dominated the day, including official benchmarks, index and coding agent scores, and a top thread preserving benchmarks from the OpenAI blog before it was taken down.

What the numbers say, read carefully

Strip the branding and there are three findings.

Raw capability is at parity with the previous generation. 61 against 61 is not a rounding artifact you can spin. On this measure Astra is not ahead of the model it replaces, and it is behind an open weight competitor at 62.

Coding moved slightly. 67 against 65 is real and small.

Price moved a lot, in the wrong direction. Roughly 2.5x per token.

Put those together and the naive read is that the model is a bad deal. That read is wrong, but not for the reason the marketing suggests.

The gains are real and they are on different axes

The same critical post lists what improved: automation, tool use, long running tasks, token efficiency. Those are precisely the properties that determine whether an agent finishes a job, and no intelligence index measures them well.

A model that completes a four hour task without derailing is more valuable than one that scores two points higher and stalls, and if it uses fewer tokens getting there, its higher per token price can still produce a lower cost per completed task. That is a genuinely different claim from "generational leap", and it is testable.

Which is the whole point. Per token price and index score are both the wrong unit. The unit is cost per completed task on your workload, and it inverts rankings regularly. We laid out the general case in 36x the price for 22% more quality: above a capability floor, price and quality have largely decoupled, so the decision belongs to a task class rather than to a procurement cycle.

A 2.5x price premium needs to buy something specific. For long horizon autonomous work it plausibly does. For classification, extraction and summarisation, which is most volume in most deployments, it certainly does not.

Availability is part of the evaluation

One more thing surfaced in the same window and it is easy to miss under the benchmark noise. Reporting ahead of launch indicated OpenAI would restrict the model after rating it a critical cyber risk, and the community reaction was appropriately weary about the pattern.

Set aside whether the risk assessment is right. The operational fact is that frontier releases now arrive with access conditions attached, and a model with uncertain availability is not yet a routing target. Neither is one whose benchmark page can be taken down hours after publication.

This is the same lesson we drew from a stealth model that processed twenty trillion tokens and then had its ID retired in the Ox Alpha post: treat capability, price and availability as three independent variables, and never let application code depend on a single model name.

The competitive picture this reveals

Two things are true at once, and together they are more interesting than the launch.

The closed frontier is compressing. OpenAI's newest flagship matches its predecessor on the index. Anthropic's Fable 5.1, two days earlier, kept input and output pricing flat and led on cache read cost. Both labs shipped improvements in efficiency and task completion rather than in raw capability.

Open weights are level on the same index. Muse Spark 1.3 at 62 is ahead of both. In August five open weight releases landed inside nine days at Flash tier prices, which we covered in the open weight frontier post. A practitioner in r/Anthropic pointed at a price comparison between a flagship and an open weight model and argued the community should be pressing on price rather than defending it.

If capability has converged near the top and the differentiation is now efficiency, tool use and cost, then a single model deployment is the expensive choice regardless of which model you pick.

How to evaluate Astra this week

Do not migrate on the index score. It says parity. Migrate on a measured improvement in your own task completion.

Measure cost per completed task, not per token. Token efficiency is the plausible win here, and it only shows up in that denominator. Read it from cost tracking and usage analytics, not from the rate card.

Test it on the long tasks only. Its claimed advantages are long horizon. Point it at your longest running agent, keep the cheap tier for everything else, and express that split as a managed policy rather than a code branch.

Treat availability as a gate. Put it behind approved models with fallback policies underneath, because a restricted model with launch week capacity limits will return errors.

Compare against the open weight option honestly. If a model scoring 62 costs a fraction as much, that belongs in the comparison. Current prices for every provider are in the models catalog and the cheapest rankings.

The most encouraging thing about this launch was not the model. It was that within hours the developer community had independent numbers, a price ratio and a clear statement of what improved and what did not. That is a market getting harder to sell to, which is good for everyone building on it.

Start routing on Requesty and put the new flagship where it earns its premium.

Frequently asked questions

How did GPT-6 Astra score on independent benchmarks?
Community reporting of Artificial Analysis figures put Astra at 61 on the intelligence index, the same as GPT-5.6 Sol, with Meta Muse Spark 1.3 at 62. On the coding agent index it scored around 67 against 65 for Sol. Its API price per token is roughly 2.5x higher than Sol.
So is Astra not an improvement?
It is an improvement on axes the intelligence index does not capture well: automation, tool use, long running task completion and token efficiency. Those matter a great deal for agents. What the data does not support is describing it as a generational jump in raw intelligence or coding ability.
Should I migrate my agents to it?
Only where long running autonomous task completion is the bottleneck, and only after measuring cost per completed task rather than cost per token. At 2.5x the price with a comparable index score, the case has to come from task completion or token efficiency on your own workload.
What is the restriction on Astra about?
Reporting ahead of the launch indicated OpenAI would restrict the model after rating it a critical cyber risk. Treat availability and access tier as part of the evaluation, not a footnote, because a model you cannot reliably call is not a routing option.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders