Two September releases lead different measures: Anthropic's Opus 5.5 tops the current AA composite, while OpenAI's Astra leads Proximal Labs' long-horizon coding evaluation.
- Opus 5.5 · AA Intelligence Index
- 57.6 #1 of 364 Artificial Analysis, September 25: first among 364 catalogued models with an index score. This index's composition changed since August; the old and current scores are not directly comparable.
- Astra · FrontierSWE v2
- 65.5% #1 of 16 Proximal Labs: first among 16 evaluated models on the v2 overall Mean@5 score. Five trials per task, 34 tasks, Proximus harness.