Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
1,177-word document condensed to 52 words. Google DeepMind · Jun 28, 2026
TL;DR
“Gemini 3.5 Flash - Model Card Model Cards are intended to provide essential information on Gemini models, including known limitations,”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| GDPval-AA | economically valuable knowledge work | 1656.00 |
| CharXiv Reasoning | no tools | 84.2% |
| MCP Atlas | multi-step MCP workflows | 83.6% |
| MMMU-Pro | no tools | 83.6% |
| OSWorld-Verified | agentic computer use | 78.4% |
| MRCR v2 (8-needle) | 128k average | 77.3% |
| Terminal-bench 2.1 | Terminus-2 harness | 76.2% |
| ARC-AGI-2 | abstract reasoning | 72.1% |
Showing top 8 of 21. See full list below.
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 81284c1840d5 · version dated Jun 28, 2026.
Extracted Evaluations(21 results)
Sort by:0/21 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| agent | scored | 78.4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| multimodal | scored | 83.6 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 1656.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 84.2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 83.6 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
MRCR v2 (8-needle) | other | scored | 77.3 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Terminal-bench 2.1 | other | scored | 76.2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Finance Agent v2 | other | scored | 57.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Toolathon | other | scored | 56.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 53.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 40.2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Blueprint-Bench 2 | other | scored | 33.6 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 8.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ non-egregious | other | scored | 0.8 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 0.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | -2.6 accuracy | Averagemissing: shot countmissing: methodmissing: training state | self-reported | |
| other | scored | -3.9 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Frontier Safety Assessment/ cyber | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Child Safety | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Frontier Safety Assessment | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | scored | 72.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |