Claude Fable 5 leads eight new practical benchmark views
Fable leads more of the 19 added coding and agentic measures than any other model, while GPT-5.6 Terra and Claude Opus 4.8 lead different kinds of work.
Event date Published
Fable leads the broadest share of the expansion
EveryBench added 19 coding and agentic measures across 10 benchmark families. Claude Fable 5 leads eight of them: LiveBench Coding, four OpenHands views, Toloka Arena, TextQuests, and Vals Vibe Code. Claude Opus 4.8 wins all three FrontierCode task sets and ties Fable on OpenHands Greenfield.
- Added practical measures
- 19 Coding and agentic results across 10 benchmark families.
- Claude Fable 5 leads
- 8 of 19 More of the added measures than any other model.
- Claude Opus 4.8
- 3 wins + 1 tie Three FrontierCode wins and an OpenHands Greenfield tie.
The leaders change with the task
Claude Fable 5 leads LiveBench Coding at 85.99, ahead of GPT-5.6 Sol at 83.94. GPT-5.6 Terra leads LiveBench Agentic Coding at 67.98, with Sol second at 65.61. Fable also leads the OpenHands Index at 81.0 and Vals Vibe Code at 90.352, while Claude Opus 4.8 leads FrontierCode Extended at 51.79%.
- LiveBench Coding
- Fable 85.99 GPT-5.6 Sol follows at 83.94.
- LiveBench Agentic Coding
- Terra 67.98 GPT-5.6 Sol follows at 65.61.
- OpenHands Index
- Fable 81.0 Claude Opus 4.8 follows at 71.88.
- Vals Vibe Code
- Fable 90.352 Claude Opus 4.8 follows at 82.725.
- FrontierCode Extended
- Opus 4.8 51.79% GPT-5.5 follows at 44.76%.
Five practical tasks, four different leader patterns
LiveBench Coding
LiveBench Agentic Coding
OpenHands Index
Vals Vibe Code
FrontierCode Extended
| Panel | Item | Score (%) |
|---|---|---|
| LiveBench Coding | Claude Fable 5 | 85.99 |
| LiveBench Coding | GPT-5.6 Sol | 83.94 |
| LiveBench Coding | GPT-5.2-Codex | 83.62 |
| LiveBench Agentic Coding | GPT-5.6 Terra | 67.98 |
| LiveBench Agentic Coding | GPT-5.6 Sol | 65.61 |
| LiveBench Agentic Coding | Grok 4.5 | 59.8 |
| OpenHands Index | Claude Fable 5 | 81 |
| OpenHands Index | Claude Opus 4.8 | 71.88 |
| OpenHands Index | Claude Opus 4.7 | 69.66 |
| Vals Vibe Code | Claude Fable 5 | 90.35 |
| Vals Vibe Code | Claude Opus 4.8 | 82.73 |
| Vals Vibe Code | Claude Sonnet 5 | 81.33 |
| FrontierCode Extended | Claude Opus 4.8 | 51.79 |
| FrontierCode Extended | GPT-5.5 | 44.76 |
| FrontierCode Extended | Claude Opus 4.7 | 43.24 |
Gold marks the leader in each panel. The same model does not lead every kind of practical work.