← Back to New
MODEL RELEASE

Gemini 3.7 Flash takes GPQA Diamond at 94.8%, Grok 4.6 pushes Terminal-Bench to 88.4 — a day apart at the frontier

Google released Gemini 3.7 Flash on 2026-08-13 and xAI released Grok 4.6 on 2026-08-12. One now leads GPQA Diamond outright; the other sits top five on Terminal-Bench 2.1 and the AA Coding and Intelligence indices, at very different speeds and prices.

Event date Published

Two August frontier models arrive a day apart

Google released Gemini 3.7 Flash on 2026-08-13 and xAI released Grok 4.6 on 2026-08-12. Both are frontier models. Gemini leads GPQA Diamond outright; both sit top ten on Terminal-Bench 2.1 — Grok also top five on the AA Intelligence Index — at very different prices and speeds.

Gemini 3.7 Flash — GPQA Diamond
94.8% #1 of 354
Epoch AI reported, 94.8% (0.948). Board leader, ahead of GPT-5.4 Pro at 94.6% and Gemini 3.1 Pro at 94.4%.
Grok 4.6 — GPQA Diamond
94.0% #5 of 354
Epoch AI reported, 94.0% (0.940). Fifth, 0.8 points behind Gemini 3.7 Flash.
Gemini 3.7 Flash — median speed
361 tokens/sec #6 of 190
Artificial Analysis median across 190 models with reported speed, 361.2 tokens/sec, against 59.2 tokens/sec for Grok 4.6 — about 6.1 times faster.
Grok 4.6 — median speed
59 tokens/sec #149 of 190
Artificial Analysis median, 59.2 tokens/sec. Gemini 3.7 Flash at 361.2 tokens/sec is the faster of the two.
Output price
$3.75 vs $6.00 per 1M tokens
Gemini 3.7 Flash $0.75 in / $3.75 out; Grok 4.6 $2.00 in / $6.00 out — Gemini is 37.5% less per output token.

Where they separate: four benchmarks for the two new models

Four panels plotting Gemini 3.7 Flash and Grok 4.6. GPQA Diamond: Gemini 94.8, Grok 94.0. Terminal-Bench 2.1: Grok 88.4, Gemini 85.8. AA Intelligence Index: Grok 60.9, Gemini 56.0. HLE: Gemini 47.9, Grok 44.1.
PanelItemScore
GPQA Diamond (%)Gemini 3.7 Flash94.8
GPQA Diamond (%)Grok 4.694.0
Terminal-Bench 2.1Gemini 3.7 Flash85.8
Terminal-Bench 2.1Grok 4.688.4
AA Intelligence IndexGemini 3.7 Flash56.0
AA Intelligence IndexGrok 4.660.9
HLE (%)Gemini 3.7 Flash47.9
HLE (%)Grok 4.644.1

Same benchmarks, same reporters for both models — Epoch AI for GPQA Diamond, Artificial Analysis for HLE, AA Intelligence Index, and Terminal-Bench 2.1 (T-Bench fallback is off). Each panel is on its own domain; higher is better. The two are effectively tied on GPQA, Grok leads on Terminal and Intelligence, Gemini leads on HLE.

What else the current data says

In the original run-102 snapshot, both models reported broadly and their edges differed. LiveBench overall was a virtual tie (Gemini 78.8, Grok 78.0), and on tau3-Banking Grok was second at 50.7 while Gemini was twenty-first at 35.5, with Qwen3.8 Max leading at 51.3. These are historical ranks, not a current coverage count; the price and speed gaps were not small.

HLE — Gemini 3.7 Flash
47.9% #5 of 337
Against Grok 44.1% (#11 of 337). Board leader Claude Fable 5 at 55.5%. Gemini is top five on HLE.
LiveBench overall
78.8 vs 78.0
Gemini 78.8 (#6 of 45), Grok 78.0 (#8 of 45) — 0.8 points apart, behind Claude Fable 5 at 83.0.
Context window
Not reported
Neither Gemini 3.7 Flash nor Grok 4.6 reports context_window_tokens in this data.

Which to try first

If GPQA Diamond or HLE at the lowest latency and price matters, Gemini 3.7 Flash is the obvious first test — it leads GPQA outright, is about six times faster, and costs less per token. If Terminal-Bench 2.1, the AA Intelligence Index, or the broader agentic benchmarks (Gert Labs agentic coding 78.7% vs 66.9%) matter more, Grok 4.6 is ahead on those, with the cost that implies. Both are frontier; neither aggregate index alone settles which helps on your workload.