Claude Fable 5 leads the new Omniscience Index; 231 of 272 models score below zero
Artificial Analysis's hallucination-penalized knowledge index lands with first results for 272 models. Claude Fable 5 leads at 40.1, 7.2 points ahead of Gemini 3.1 Pro — and the models least likely to hallucinate are different ones entirely.
Event date Published
The Omniscience Index joins EveryBench
Artificial Analysis's Omniscience Index now has first results on EveryBench for 272 models, alongside a companion column carrying each model's non-hallucination rate. The index rewards correct answers and penalizes confidently wrong ones, so a model that abstains outscores one that fabricates, and scores run on a roughly -100 to +100 range. That penalty bites in practice: 231 of the 272 models score below zero, meaning they hallucinate more than they answer correctly on this question set.
- Models with first results
- 272 The Omniscience Index and the companion non-hallucination rate cover the same 272 models.
- Models above zero
- 41 of 272 The other 231 lose more to hallucination penalties than they earn from correct answers.
- Leader's margin
- 7.2 points Claude Fable 5 at 40.1 over Gemini 3.1 Pro at 32.9 — the widest gap anywhere in the top five.
Claude Fable 5 leads the first results
Claude Fable 5 tops the debut Omniscience Index leaderboard at 40.1, with Gemini 3.1 Pro second at 32.9 and Claude Opus 4.8 third at 27.4. Grok 4.5 and Claude Opus 4.7 complete the top five, separated by just 0.2 points. These are first results from a single first-hand evaluation, not before/after movements.
- Claude Fable 5
- 40.1 #1 of 272 Leads the debut Omniscience Index results.
- Gemini 3.1 Pro
- 32.9 #2 of 272 7.2 points behind the leader.
- Claude Opus 4.8
- 27.4 #3 of 272 5.5 points behind second place.
- Grok 4.5
- 26.4 #4 of 272 0.2 points ahead of fifth place.
- Claude Opus 4.7
- 26.2 #5 of 272 Completes the top five.
Knowing the most is not hallucinating the least
Omniscience Index (knowledge net of hallucination penalty)
Non-hallucination rate, %
| Panel | Item | Score |
|---|---|---|
| Omniscience Index (knowledge net of hallucination penalty) | Claude Fable 5 | 40.15 |
| Omniscience Index (knowledge net of hallucination penalty) | Gemini 3.1 Pro | 32.93 |
| Omniscience Index (knowledge net of hallucination penalty) | Claude Opus 4.8 | 27.43 |
| Omniscience Index (knowledge net of hallucination penalty) | Grok 4.20 | 15.35 |
| Omniscience Index (knowledge net of hallucination penalty) | MiniMax M3 | 1.37 |
| Omniscience Index (knowledge net of hallucination penalty) | Command A+ | -3.98 |
| Non-hallucination rate, % | Claude Fable 5 | 45.15 |
| Non-hallucination rate, % | Gemini 3.1 Pro | 50.13 |
| Non-hallucination rate, % | Claude Opus 4.8 | 64.11 |
| Non-hallucination rate, % | Grok 4.20 | 83.27 |
| Non-hallucination rate, % | MiniMax M3 | 83.92 |
| Non-hallucination rate, % | Command A+ | 85.92 |
The same six models on both new columns, sharing one first-hand evaluation. Index leaders earn the most net credit for correct answers; non-hallucination leaders rarely fabricate but land near or below zero on the index. No model beats Grok 4.20 on both measures at once.
Knowledge leaders and reliability leaders barely overlap
The two columns rank models by different virtues. Claude Fable 5 earns the most credit net of penalties, but 49 of the 272 models fabricate less often than it does. Command A+ fabricates least of all — an 85.9% non-hallucination rate — yet lands slightly below zero on the index because its correct answers do not outweigh its penalties. If a wrong answer costs you more than no answer, read the non-hallucination column; no model beats Grok 4.20 on both measures at once. These are first results from a single first-hand evaluation, not a settled ranking.
Also new on the board
The same update surfaces seven more first-hand Artificial Analysis evaluations and reworks the LMArena coverage.
- Terminal-Bench Hard
- 265 models First-hand Artificial Analysis results on the harder agentic terminal tasks; GPT-5.6 Sol leads at 65.9%. Existing OpenRouter-cited figures stay visible as a second reporter.
- IFBench
- 274 models MiniMax M3 leads instruction following at 82.9%.
- LiveCodeBench and AIME 2025
- 196 and 166 models Gemini 3 Pro leads LiveCodeBench at 91.7%; AIME 2025 tops out at 99% (GPT-5.2), so the ceiling is close.
- τ³-Banking
- 118 models Kimi K3 leads at 33.4% — even the best models complete about a third of these fintech support tasks.
- Harvey LAB and EnterpriseOps-Gym
- 33 and 27 models Kimi K3 leads the private legal tasks at 94.6%; Claude Fable 5 leads the enterprise-operations tasks at 51.1%.
- LMArena boards
- 5 added, 1 rescored WebDev, Vision, Image-to-WebDev, Document, and Search arenas join; the Agent arena now carries LMArena's Net-Improvement score instead of an Elo.