← Back to New
BENCHMARKS

Claude Fable 5 leads the new Omniscience Index; 231 of 272 models score below zero

Artificial Analysis's hallucination-penalized knowledge index lands with first results for 272 models. Claude Fable 5 leads at 40.1, 7.2 points ahead of Gemini 3.1 Pro — and the models least likely to hallucinate are different ones entirely.

Event date Published

The Omniscience Index joins EveryBench

Artificial Analysis's Omniscience Index now has first results on EveryBench for 272 models, alongside a companion column carrying each model's non-hallucination rate. The index rewards correct answers and penalizes confidently wrong ones, so a model that abstains outscores one that fabricates, and scores run on a roughly -100 to +100 range. That penalty bites in practice: 231 of the 272 models score below zero, meaning they hallucinate more than they answer correctly on this question set.

Models with first results
272
The Omniscience Index and the companion non-hallucination rate cover the same 272 models.
Models above zero
41 of 272
The other 231 lose more to hallucination penalties than they earn from correct answers.
Leader's margin
7.2 points
Claude Fable 5 at 40.1 over Gemini 3.1 Pro at 32.9 — the widest gap anywhere in the top five.

Claude Fable 5 leads the first results

Claude Fable 5 tops the debut Omniscience Index leaderboard at 40.1, with Gemini 3.1 Pro second at 32.9 and Claude Opus 4.8 third at 27.4. Grok 4.5 and Claude Opus 4.7 complete the top five, separated by just 0.2 points. These are first results from a single first-hand evaluation, not before/after movements.

Claude Fable 5
40.1 #1 of 272
Leads the debut Omniscience Index results.
Gemini 3.1 Pro
32.9 #2 of 272
7.2 points behind the leader.
Claude Opus 4.8
27.4 #3 of 272
5.5 points behind second place.
Grok 4.5
26.4 #4 of 272
0.2 points ahead of fifth place.
Claude Opus 4.7
26.2 #5 of 272
Completes the top five.

Knowing the most is not hallucinating the least

Two dot plots compare the same six models. On the Omniscience Index, Claude Fable 5 leads at 40.15, Gemini 3.1 Pro scores 32.93, Claude Opus 4.8 scores 27.43, Grok 4.20 scores 15.35, MiniMax M3 scores 1.37, and Command A+ sits below zero at -3.98. On non-hallucination rate the order reverses: Command A+ leads at 85.9%, MiniMax M3 at 83.9%, and Grok 4.20 at 83.3%, while Claude Opus 4.8 sits at 64.1%, Gemini 3.1 Pro at 50.1%, and Claude Fable 5 at 45.1%.
PanelItemScore
Omniscience Index (knowledge net of hallucination penalty)Claude Fable 540.15
Omniscience Index (knowledge net of hallucination penalty)Gemini 3.1 Pro32.93
Omniscience Index (knowledge net of hallucination penalty)Claude Opus 4.827.43
Omniscience Index (knowledge net of hallucination penalty)Grok 4.2015.35
Omniscience Index (knowledge net of hallucination penalty)MiniMax M31.37
Omniscience Index (knowledge net of hallucination penalty)Command A+-3.98
Non-hallucination rate, %Claude Fable 545.15
Non-hallucination rate, %Gemini 3.1 Pro50.13
Non-hallucination rate, %Claude Opus 4.864.11
Non-hallucination rate, %Grok 4.2083.27
Non-hallucination rate, %MiniMax M383.92
Non-hallucination rate, %Command A+85.92

The same six models on both new columns, sharing one first-hand evaluation. Index leaders earn the most net credit for correct answers; non-hallucination leaders rarely fabricate but land near or below zero on the index. No model beats Grok 4.20 on both measures at once.

Knowledge leaders and reliability leaders barely overlap

The two columns rank models by different virtues. Claude Fable 5 earns the most credit net of penalties, but 49 of the 272 models fabricate less often than it does. Command A+ fabricates least of all — an 85.9% non-hallucination rate — yet lands slightly below zero on the index because its correct answers do not outweigh its penalties. If a wrong answer costs you more than no answer, read the non-hallucination column; no model beats Grok 4.20 on both measures at once. These are first results from a single first-hand evaluation, not a settled ranking.

Also new on the board

The same update surfaces seven more first-hand Artificial Analysis evaluations and reworks the LMArena coverage.

Terminal-Bench Hard
265 models
First-hand Artificial Analysis results on the harder agentic terminal tasks; GPT-5.6 Sol leads at 65.9%. Existing OpenRouter-cited figures stay visible as a second reporter.
IFBench
274 models
MiniMax M3 leads instruction following at 82.9%.
LiveCodeBench and AIME 2025
196 and 166 models
Gemini 3 Pro leads LiveCodeBench at 91.7%; AIME 2025 tops out at 99% (GPT-5.2), so the ceiling is close.
τ³-Banking
118 models
Kimi K3 leads at 33.4% — even the best models complete about a third of these fintech support tasks.
Harvey LAB and EnterpriseOps-Gym
33 and 27 models
Kimi K3 leads the private legal tasks at 94.6%; Claude Fable 5 leads the enterprise-operations tasks at 51.1%.
LMArena boards
5 added, 1 rescored
WebDev, Vision, Image-to-WebDev, Document, and Search arenas join; the Agent arena now carries LMArena's Net-Improvement score instead of an Elo.