Gemini 3.7 Flash takes GPQA Diamond at 94.8%, Grok 4.6 pushes Terminal-Bench to 88.4 — a day apart at the frontier
Google released Gemini 3.7 Flash on 2026-08-13 and xAI released Grok 4.6 on 2026-08-12. One now leads GPQA Diamond outright; the other sits top five on Terminal-Bench 2.1 and the AA Coding and Intelligence indices, at very different speeds and prices.
Event date Published
Two August frontier models arrive a day apart
Google released Gemini 3.7 Flash on 2026-08-13 and xAI released Grok 4.6 on 2026-08-12. Both are frontier models. Gemini leads GPQA Diamond outright; both sit top ten on Terminal-Bench 2.1 — Grok also top five on the AA Intelligence Index — at very different prices and speeds.
- Gemini 3.7 Flash — GPQA Diamond
- 94.8% #1 of 354 Epoch AI reported, 94.8% (0.948). Board leader, ahead of GPT-5.4 Pro at 94.6% and Gemini 3.1 Pro at 94.4%.
- Grok 4.6 — GPQA Diamond
- 94.0% #5 of 354 Epoch AI reported, 94.0% (0.940). Fifth, 0.8 points behind Gemini 3.7 Flash.
- Gemini 3.7 Flash — median speed
- 361 tokens/sec #6 of 190 Artificial Analysis median across 190 models with reported speed, 361.2 tokens/sec, against 59.2 tokens/sec for Grok 4.6 — about 6.1 times faster.
- Grok 4.6 — median speed
- 59 tokens/sec #149 of 190 Artificial Analysis median, 59.2 tokens/sec. Gemini 3.7 Flash at 361.2 tokens/sec is the faster of the two.
- Output price
- $3.75 vs $6.00 per 1M tokens Gemini 3.7 Flash $0.75 in / $3.75 out; Grok 4.6 $2.00 in / $6.00 out — Gemini is 37.5% less per output token.
Where they separate: four benchmarks for the two new models
GPQA Diamond (%)
Terminal-Bench 2.1
AA Intelligence Index
HLE (%)
| Panel | Item | Score |
|---|---|---|
| GPQA Diamond (%) | Gemini 3.7 Flash | 94.8 |
| GPQA Diamond (%) | Grok 4.6 | 94.0 |
| Terminal-Bench 2.1 | Gemini 3.7 Flash | 85.8 |
| Terminal-Bench 2.1 | Grok 4.6 | 88.4 |
| AA Intelligence Index | Gemini 3.7 Flash | 56.0 |
| AA Intelligence Index | Grok 4.6 | 60.9 |
| HLE (%) | Gemini 3.7 Flash | 47.9 |
| HLE (%) | Grok 4.6 | 44.1 |
Same benchmarks, same reporters for both models — Epoch AI for GPQA Diamond, Artificial Analysis for HLE, AA Intelligence Index, and Terminal-Bench 2.1 (T-Bench fallback is off). Each panel is on its own domain; higher is better. The two are effectively tied on GPQA, Grok leads on Terminal and Intelligence, Gemini leads on HLE.
What else the current data says
In the original run-102 snapshot, both models reported broadly and their edges differed. LiveBench overall was a virtual tie (Gemini 78.8, Grok 78.0), and on tau3-Banking Grok was second at 50.7 while Gemini was twenty-first at 35.5, with Qwen3.8 Max leading at 51.3. These are historical ranks, not a current coverage count; the price and speed gaps were not small.
- HLE — Gemini 3.7 Flash
- 47.9% #5 of 337 Against Grok 44.1% (#11 of 337). Board leader Claude Fable 5 at 55.5%. Gemini is top five on HLE.
- LiveBench overall
- 78.8 vs 78.0 Gemini 78.8 (#6 of 45), Grok 78.0 (#8 of 45) — 0.8 points apart, behind Claude Fable 5 at 83.0.
- Context window
- Not reported Neither Gemini 3.7 Flash nor Grok 4.6 reports context_window_tokens in this data.
Which to try first
If GPQA Diamond or HLE at the lowest latency and price matters, Gemini 3.7 Flash is the obvious first test — it leads GPQA outright, is about six times faster, and costs less per token. If Terminal-Bench 2.1, the AA Intelligence Index, or the broader agentic benchmarks (Gert Labs agentic coding 78.7% vs 66.9%) matter more, Grok 4.6 is ahead on those, with the cost that implies. Both are frontier; neither aggregate index alone settles which helps on your workload.