Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
19,272-word document condensed to 150 words. OpenAI · Jun 28, 2026
TL;DR
“GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet— are built to deliver these models safely and at scale, around the world.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| Atomic Challenges | network_attack_simulation, accuracy | 100.0% |
| Atomic Challenges | network_attack_simulation, accuracy | 98.0% |
| HealthBench Consensus | consensus | 95.5% |
| HealthBench Consensus | consensus | 95.1% |
| HealthBench Consensus | consensus | 95.1% |
| Atomic Challenges | vulnerability_research_and_exploitation, accuracy | 92.0% |
| Atomic Challenges | vulnerability_research_and_exploitation, accuracy | 91.0% |
| Tacit Knowledge | refusals-as-success, 60 MCQ | 84.1% |
Showing top 8 of 60. See full list below.
Capability claim
- “we trained the models to maintain a strong standard of overwrite avoidance while improving autonomy without relying on extra cautious prompting.”
Mitigations
- “We have deployed an expanded set of safeguards to restrict the ability of malicious actors to benefit from increased capabilities in cybersecurity performance.”
- “we have deployed Preparedness Safeguards.”
Deployment scope
- “available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA d5e9ee095a30 · version dated Jun 28, 2026.
Extracted Evaluations(60 results)
Sort by:0/60 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ verified | coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| coding | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | knowledge | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ 2025 | math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| medical | scored | 57.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Atomic Challenges/ network_attack_simulation | other | scored | 100.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Atomic Challenges/ network_attack_simulation | other | scored | 98.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 95.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 95.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 95.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Atomic Challenges/ vulnerability_research_and_exploitation | other | scored | 92.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Atomic Challenges/ vulnerability_research_and_exploitation | other | scored | 91.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge | other | scored | 84.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 60.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 57.7 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Atomic Challenges/ evasion | other | scored | 56.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 55.7 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
SecureBio | other | scored | 55.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 55.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Atomic Challenges/ evasion | other | scored | 54.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Troubleshooting Bench | other | scored | 48.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 43.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ open_ended | other | scored | 43.5 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 33.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 28.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
FrontierCyber/ medium | other | scored | 12.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ easy | other | scored | 11.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ medium | other | scored | 6.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ easy | other | scored | 6.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ hard | other | scored | 5.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ hard | other | scored | 4.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 1.3 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | scored | 0.7 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported | |
AAV_Capsid_Packaging_Prediction | other | scored | 0.5 spearman correlation | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
AAV_Capsid_Packaging_Prediction | other | scored | 0.5 spearman correlation | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 0.4 accuracy | cotmissing: shot countmissing: languagemissing: training state | self-reported | |
FrontierCyber/ elite | other | scored | 0.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCyber/ elite | other | scored | 0.0 accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ExploitGym | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ExploitGym | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
SEC-Bench/ pro | other | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ExploitGym | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ revised | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
PostTrainBench/ lite | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
PostTrainBench/ lite | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
PostTrainBench/ lite | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
KernelGen/ 1p | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Monorepo-Bench | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Tacit Knowledge and Troubleshooting | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | mentioned | — not unsafe | without-safeguardsmissing: shot countmissing: languagemissing: training state | self-reported | |
| other | mentioned | — accuracy | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
ExploitGym | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ExploitGym | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| reasoning | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |