Kimi K3 debuts with seven top-three finishes
Across 15 first observations, Kimi K3 places in the top three seven times and leads the newly added AutomationBench-AA board.
Event date Published
A strong but uneven first profile
Kimi K3 enters EveryBench with 15 first observations. Seven are top-three placements, including one first place. The other eight range from fourth to a tie for fifteenth among models with a reported score on each benchmark.
- First observations
- 15 Kimi K3 was absent from the earlier EveryBench data.
- Top-three placements
- 7 Across the 15 benchmark results shown by EveryBench.
- First places
- 1 AutomationBench-AA, a newly added rolling board.
The strongest results cluster around agentic and composite work
Kimi K3 leads AutomationBench-AA at 52.7%, places second on AA-Briefcase at 1547 Elo and Vals Index at 74.7, and places third on the AA Intelligence Index at 57.1 and GDPval at 58.4%. AutomationBench-AA was newly added, so Kimi did not displace an earlier leader.
- AutomationBench-AA
- 52.7% #1 of 34 Ahead of Grok 4.5 at 51.4% and GPT-5.6 Sol at 51.2%.
- AA-Briefcase
- 1547 Elo #2 of 41 Behind Claude Fable 5 at 1583 Elo.
- Vals Index
- 74.7 #2 of 37 Behind Claude Fable 5 at 75.1.
- AA Intelligence Index
- 57.1 #3 of 318 Behind Claude Fable 5 and GPT-5.6 Sol.
- GDPval
- 58.4% #3 of 117 This version of GDPval is fixed rather than continually updated.
Where Kimi K3 sits against nearby leaders
AutomationBench-AA (%)
AA-Briefcase (Elo)
Vals Index
AA Intelligence Index
GDPval (%)
| Panel | Item | Benchmark score |
|---|---|---|
| AutomationBench-AA (%) | Kimi K3 | 52.7 |
| AutomationBench-AA (%) | Grok 4.5 | 51.4 |
| AutomationBench-AA (%) | GPT-5.6 Sol | 51.2 |
| AA-Briefcase (Elo) | Claude Fable 5 | 1,583 |
| AA-Briefcase (Elo) | Kimi K3 | 1,547 |
| AA-Briefcase (Elo) | GPT-5.6 Sol | 1,495 |
| Vals Index | Claude Fable 5 | 75.1 |
| Vals Index | Kimi K3 | 74.7 |
| Vals Index | GPT-5.6 Sol | 73.1 |
| AA Intelligence Index | Claude Fable 5 | 59.9 |
| AA Intelligence Index | GPT-5.6 Sol | 58.9 |
| AA Intelligence Index | Kimi K3 | 57.1 |
| AA Intelligence Index | Claude Opus 4.8 | 55.7 |
| GDPval (%) | Claude Fable 5 | 63 |
| GDPval (%) | GPT-5.6 Sol | 62.4 |
| GDPval (%) | Kimi K3 | 58.4 |
| GDPval (%) | Claude Sonnet 5 | 55.3 |
Kimi K3 is highlighted in gold. Each panel uses a labeled local scale to make nearby gaps readable; compare positions only within the same benchmark.
Eight new boards split their winners
Six distinct models lead the eight newly added benchmarks. Claude Fable 5 and GPT-5.4 lead two each; GLM-5.2, Claude Opus 4.8, GPT-5.5, and Kimi K3 lead one each. No model wins more than two.
- New benchmarks
- 8 Three coding, three agentic, and two multimodal benchmarks.
- Distinct leaders
- 6 No model leads more than two of the eight.