Requesty
Back|SEP '26BENCHMARKS / AGENTS
4 MIN READ|

GPT-6 Astra scores 61.2 on intelligence and 51.5 on agentic work: the gap nobody puts in a launch post

Last updated

Artificial Analysis publishes two headline numbers for every model. Almost every launch post quotes the first one.

On the 2026-09-03 snapshot, GPT-6 Astra scores 61.2 on the general intelligence index and 51.5 on the agentic index. That is a 9.7 point gap, the widest among the 19 releases scoring above 40 on the intelligence index, which is the group charted below. Weaker models show wider gaps still, but they are not competing for the same work. Claude Fable 5.1, the highest scoring model on the board at 65.7, drops to 61.3 on agentic work. Gemini 3.8 Flash loses 8.7 points.

Going the other direction, GLM 5.3 Flash is the only release that scores higher on agentic work than on general intelligence, at 58.2 against 57.5.

Agentic index minus intelligence index across tracked model releases
Agentic index minus intelligence index across tracked model releases

The numbers

Every figure below is from the Artificial Analysis snapshot taken on 2026-09-03, using the strongest effort setting published for each release.

ReleaseIntelligence indexAgentic indexGap
GPT-6 Astra61.251.59.7 lower
MiniMax-M345.436.19.3 lower
Gemini 3.8 Flash58.750.08.7 lower
GPT-5.6 Terra56.650.26.4 lower
Claude Fable 562.156.65.5 lower
Kimi K359.754.35.4 lower
Claude Fable 5.165.761.34.3 lower
Claude Opus 563.159.23.9 lower
GPT-5.6 Sol60.957.83.1 lower
Grok 4.660.958.72.2 lower
Qwen3.8 2.4T A95B57.757.10.6 lower
GLM-5.359.559.10.4 lower
GLM 5.3 Flash57.558.20.7 higher

Read the ranking twice. By general intelligence, Astra sits fifth. By agentic index it falls behind Grok 4.6, GPT-5.6 Sol, GLM-5.3, GLM 5.3 Flash and Qwen3.8 2.4T A95B, several of which cost a fraction as much.

Why the two numbers come apart

A general intelligence score rewards getting a hard question right in one pass. An agentic score rewards something different: choosing a tool, reading what came back, noticing that it failed, and recovering, dozens of times in a row without drifting.

Those skills are related but not the same, and the failure modes are not the same either. A model that reasons beautifully in one shot can still call a tool with a malformed argument on turn 14, misread an empty result as success, and confidently report a finished job. Single pass accuracy does not measure that at all.

This is why a model can post a record on a reasoning suite and still frustrate the person wiring it into a loop. Both experiences are real, and they are measuring different things.

What the gap means for routing

The practical consequence is that the model at the top of the leaderboard is not automatically the right default for an agent, and the gap is large enough to change the ordering.

Three things follow.

Shortlist on the agentic number when you are building agents. If your workload is a tool loop, the agentic index is the closer proxy. On this snapshot that promotes GLM-5.3 at 59.1 and Grok 4.6 at 58.7 into the same conversation as models scoring far higher on the general index.

Price the gap. GLM 5.3 Flash reaches 58.2 agentic. Claude Fable 5.1 reaches 61.3. Three points of agentic capability is a real difference, and it is worth paying for on hard work. It is not worth paying for on the routine steps of a loop, which is most of the steps. This is exactly the case for routing by cost and latency rather than picking one default and living with it.

Do not trust either number as a substitute for your own evaluation. Both indices are public benchmarks, and public benchmarks measure what they measure. Requesty gives you per request logging so you can compare candidates on your own traffic, and fallback policies so a candidate that disappoints can be demoted without a redeploy.

The honest caveats

The agentic index is one composite score from one evaluator, and Artificial Analysis flags some index values as estimated. A 0.4 point gap, as on GLM-5.3, is inside the noise and should not be read as a ranking. A 9.7 point gap is not.

The snapshot is a single day. Model versions, effort settings and prices all move, sometimes weekly, so re-pull before quoting these figures.

None of this says Astra is a weak model. It says Astra's agentic score is nearly ten points below its headline score, and if you picked it for an agent because of the headline, you bought a different model than the one you read about.

Where to start

Pull both numbers for your shortlist, not just the famous one. Then check the gap. If it is wide, assume your agent will feel worse than the launch post promised, and plan to benchmark agentic routing on your own workload before you commit.

Frequently asked questions

What is the agentic index and how does it differ from the intelligence index?
Artificial Analysis reports a general intelligence index built from reasoning and knowledge evaluations, and a separate agentic index weighted towards tool use and multi step task completion. A model can rank well on one and poorly on the other, because answering a hard question and driving a tool loop are different skills.
Does a negative gap mean the model is bad at agentic work?
No. It means the model is weaker at agentic work than its headline number suggests. GPT-6 Astra at 51.5 agentic still outranks most of the field in absolute terms. The gap is about mis-calibrated expectations, not absolute capability.
Which models hold up best on agentic work?
On the 2026-09-03 snapshot GLM 5.3 Flash is the only release scoring higher on agentic than on intelligence, at 58.2 against 57.5. GLM-5.3 loses 0.4 points, Qwen3.8 2.4T A95B loses 0.6 and Grok 4.6 loses 2.2, so all three are close to calibrated.
Should I pick a model on the agentic index alone?
No. The agentic index is a better proxy than the general index if you are building agents, but it is still a public benchmark. Use it to shortlist two or three candidates and then run your own evaluation set on your own traffic.
Related reading

Start building with Requesty

One line of code. 600+ models. Full control.

Speak to founders