Model Cards / OpenAI

GPT-5.6 Preview System Card

model card19,272 words·84 min read·Jun 28, 2026·Source
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
19,272-word document condensed to 150 words. OpenAI · Jun 28, 2026
TL;DR

GPT-5.6 is a new family of three models: Sol, our new flagship model; Terra, a capable lower-cost option; and Luna, our fastest and most cost-efficient model. The safeguards we have built for this launch—our most robust yet— are built to deliver these models safely and at scale, around the world.

Top benchmarks
BenchmarkVariantScore
Atomic Challengesnetwork_attack_simulation, accuracy100.0%
Atomic Challengesnetwork_attack_simulation, accuracy98.0%
HealthBench Consensusconsensus95.5%
HealthBench Consensusconsensus95.1%
HealthBench Consensusconsensus95.1%
Atomic Challengesvulnerability_research_and_exploitation, accuracy92.0%
Atomic Challengesvulnerability_research_and_exploitation, accuracy91.0%
Tacit Knowledgerefusals-as-success, 60 MCQ84.1%

Showing top 8 of 60. See full list below.

Capability claim
  • we trained the models to maintain a strong standard of overwrite avoidance while improving autonomy without relying on extra cautious prompting.
Mitigations
  • We have deployed an expanded set of safeguards to restrict the ability of malicious actors to benefit from increased capabilities in cybersecurity performance.
  • we have deployed Preparedness Safeguards.
Deployment scope
  • available to the public, we can continue to reserve the most sensitive cybersecurity and biological capabilities for trusted defenders.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA d5e9ee095a30 · version dated Jun 28, 2026.

Extracted Evaluations(60 results)

Sort by:0/60 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
codingmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ verified
codingmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
codingmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ pro
knowledgecited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
mathmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ 2025
mathmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
medicalscored
57.0
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
medicalmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ network_attack_simulation
otherscored
100.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ network_attack_simulation
otherscored
98.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
95.5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
95.1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
95.1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ vulnerability_research_and_exploitation
otherscored
92.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ vulnerability_research_and_exploitation
otherscored
91.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Tacit Knowledge
otherscored
84.1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
60.5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
57.7
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ evasion
otherscored
56.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
55.7
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
SecureBio
otherscored
55.5
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
55.5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Atomic Challenges/ evasion
otherscored
54.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Troubleshooting Bench
otherscored
48.0
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
43.5
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ open_ended
otherscored
43.5
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
33.1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
28.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ medium
otherscored
12.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ easy
otherscored
11.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ medium
otherscored
6.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ easy
otherscored
6.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ hard
otherscored
5.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ hard
otherscored
4.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
1.3
accuracy
cotmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
0.7
accuracy
cotmissing: shot countmissing: languagemissing: training state
self-reported
AAV_Capsid_Packaging_Prediction
otherscored
0.5
spearman correlation
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
AAV_Capsid_Packaging_Prediction
otherscored
0.5
spearman correlation
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
0.4
accuracy
cotmissing: shot countmissing: languagemissing: training state
self-reported
FrontierCyber/ elite
otherscored
0.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
FrontierCyber/ elite
otherscored
0.0
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ExploitGym
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ExploitGym
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
SEC-Bench/ pro
othermentioned
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ExploitGym
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ revised
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
PostTrainBench/ lite
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
PostTrainBench/ lite
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
PostTrainBench/ lite
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
KernelGen/ 1p
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Monorepo-Bench
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Tacit Knowledge and Troubleshooting
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
not unsafe
without-safeguardsmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
accuracy
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ExploitGym
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ExploitGym
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
reasoningcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported