DeepSeek V4 Flash's July update lifts results across benchmark views
LiveBench rises 8.7 points. Artificial Analysis's stable DeepSeek V4 Flash entry gains 9.7 points on its Intelligence Index, and Epoch AI's WeirdML results rise by 13 to 17 points at matched reasoning settings.
Event date Published
A large LiveBench improvement
DeepSeek V4 Flash's July update scores 74.2 on LiveBench's current 37-model board, ranking 18th. The earlier result scores 65.5 and ranks 34th: an 8.7-point improvement. The largest gains are in reasoning, data analysis, and agentic coding.
- LiveBench
- 74.2 #18 of 37 Up 8.7 points from the earlier DeepSeek V4 Flash result, 65.5 and #34.
- LiveBench Reasoning
- 86.6 #15 of 37 Up 16.1 points from the earlier result, 70.6 and #35.
- LiveBench Data Analysis
- 79.3 #4 of 37 Up 11.3 points from the earlier result, 68.0 and #29.
- LiveBench Agentic Coding
- 46.8 #22 of 37 Up 9.1 points from the earlier result, 37.6 and #36.
Four LiveBench views of the July update
LiveBench
LiveBench Reasoning
LiveBench Data Analysis
LiveBench Agentic Coding
| Panel | Item | LiveBench score (%) |
|---|---|---|
| LiveBench | Earlier result | 65.5 |
| LiveBench | July update | 74.2 |
| LiveBench Reasoning | Earlier result | 70.6 |
| LiveBench Reasoning | July update | 86.6 |
| LiveBench Data Analysis | Earlier result | 68.0 |
| LiveBench Data Analysis | July update | 79.3 |
| LiveBench Agentic Coding | Earlier result | 37.6 |
| LiveBench Agentic Coding | July update | 46.8 |
Each panel uses the same 0 to 100 LiveBench score (%) scale. The July update is highlighted; values are comparable within each panel on LiveBench's current board.
Other matched results point in the same direction
Artificial Analysis's stable DeepSeek V4 Flash entry records a gain in its Intelligence Index at maximum reasoning effort. The source does not name the evaluated revision, so this is reported as an observed movement. Epoch AI's explicitly labelled 0731 WeirdML External result also rises at both matched effort settings.
The improvement is not uniform on every task
Gert Labs' combined score rises from 51.8% to 55.6%, but its components diverge: one-shot coding rises from 29.7% to 40.6%, decision-making falls from 50.2% to 40.5%, and agentic coding is nearly unchanged. These comparisons support a material update, not a claim that every task or configuration improved by the same amount.