← Back to New
MODEL RELEASE

DeepSeek V4 Flash's July update lifts results across benchmark views

LiveBench rises 8.7 points. Artificial Analysis's stable DeepSeek V4 Flash entry gains 9.7 points on its Intelligence Index, and Epoch AI's WeirdML results rise by 13 to 17 points at matched reasoning settings.

Event date Published

A large LiveBench improvement

DeepSeek V4 Flash's July update scores 74.2 on LiveBench's current 37-model board, ranking 18th. The earlier result scores 65.5 and ranks 34th: an 8.7-point improvement. The largest gains are in reasoning, data analysis, and agentic coding.

LiveBench
74.2 #18 of 37
Up 8.7 points from the earlier DeepSeek V4 Flash result, 65.5 and #34.
LiveBench Reasoning
86.6 #15 of 37
Up 16.1 points from the earlier result, 70.6 and #35.
LiveBench Data Analysis
79.3 #4 of 37
Up 11.3 points from the earlier result, 68.0 and #29.
LiveBench Agentic Coding
46.8 #22 of 37
Up 9.1 points from the earlier result, 37.6 and #36.

Four LiveBench views of the July update

Four 0 to 100 percentage-score panels compare an earlier DeepSeek V4 Flash result with the July update. LiveBench is 65.5 versus 74.2, LiveBench Reasoning 70.6 versus 86.6, LiveBench Data Analysis 68.0 versus 79.3, and LiveBench Agentic Coding 37.6 versus 46.8.
PanelItemLiveBench score (%)
LiveBenchEarlier result65.5
LiveBenchJuly update74.2
LiveBench ReasoningEarlier result70.6
LiveBench ReasoningJuly update86.6
LiveBench Data AnalysisEarlier result68.0
LiveBench Data AnalysisJuly update79.3
LiveBench Agentic CodingEarlier result37.6
LiveBench Agentic CodingJuly update46.8

Each panel uses the same 0 to 100 LiveBench score (%) scale. The July update is highlighted; values are comparable within each panel on LiveBench's current board.

Other matched results point in the same direction

Artificial Analysis's stable DeepSeek V4 Flash entry records a gain in its Intelligence Index at maximum reasoning effort. The source does not name the evaluated revision, so this is reported as an observed movement. Epoch AI's explicitly labelled 0731 WeirdML External result also rises at both matched effort settings.

AA Intelligence Index
42.1to51.8+9.7
Artificial Analysis, maximum reasoning effort.
Epoch AI WeirdML, high effort
43.8to57.1+13.3
WeirdML External at high effort.
Epoch AI WeirdML, max effort
45.6to63+17.3
WeirdML External at max effort.

The improvement is not uniform on every task

Gert Labs' combined score rises from 51.8% to 55.6%, but its components diverge: one-shot coding rises from 29.7% to 40.6%, decision-making falls from 50.2% to 40.5%, and agentic coding is nearly unchanged. These comparisons support a material update, not a claim that every task or configuration improved by the same amount.