Opus 5.5 leads AA Intelligence; GPT-6 Astra leads FrontierSWE v2
Two September releases lead different measures: Anthropic's Opus 5.5 tops the current AA composite, while OpenAI's Astra leads Proximal Labs' long-horizon coding evaluation.
Event date Published
The two September leaders
Claude Opus 5.5, released September 22, leads the current Artificial Analysis Intelligence Index at 57.6. GPT-6 Astra, released September 3, leads FrontierSWE v2 at 65.5%. The measures answer different questions; their scores are not on a shared scale.
- Opus 5.5 · AA Intelligence Index
- 57.6 #1 of 364 Artificial Analysis, September 25: first among 364 catalogued models with an index score. This index's composition changed since August; the old and current scores are not directly comparable.
- Astra · FrontierSWE v2
- 65.5% #1 of 16 Proximal Labs: first among 16 evaluated models on the v2 overall Mean@5 score. Five trials per task, 34 tasks, Proximus harness.
Three peers, two different orderings
AA Intelligence Index (points)
Claude Opus 5.5 57.6
Claude Fable 5.1 53.4
GPT-6 Astra 52.7
FrontierSWE v2 Mean@5 (%)
GPT-6 Astra 65.5
Claude Opus 5.5 62.3
Claude Fable 5.1 56.3
| Panel | Item | Score in the panel's units |
|---|---|---|
| AA Intelligence Index (points) | Claude Opus 5.5 | 57.6 |
| AA Intelligence Index (points) | Claude Fable 5.1 | 53.4 |
| AA Intelligence Index (points) | GPT-6 Astra | 52.7 |
| FrontierSWE v2 Mean@5 (%) | GPT-6 Astra | 65.5 |
| FrontierSWE v2 Mean@5 (%) | Claude Opus 5.5 | 62.3 |
| FrontierSWE v2 Mean@5 (%) | Claude Fable 5.1 | 56.3 |
AA Intelligence is Artificial Analysis's current composite; its composition changed since August. FrontierSWE v2 is Proximal Labs' 34-task Mean@5 percentage. Each panel starts at zero and uses its own units. Compare models within a panel, not values between panels.