Model Cards / Anthropic

Claude Opus 4.8 System Card

model card61,956 words·269 min read·May 30, 2026·Source
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
61,956-word document condensed to 501 words. Anthropic · May 30, 2026
TL;DR

run on Claude Opus 4.8. It includes the following sections: Responsible Scaling Policy evaluations.

Top benchmarks
BenchmarkVariantScore
GDPvalextended-thinking, aa1890.00
GraphWalksparents_256k, f199.3%
USAMOextended-thinking, accuracy96.7%
GPQAextended-thinking, diamond, accuracy93.6%
SWE-benchextended-thinking, verified, resolve_rate88.6%
BrowseCompextended-thinking, multi_agent, accuracy88.5%
ProgramBenchpass_rate88.0%
ScreenSpotwith-tools, pro, accuracy87.9%

Showing top 8 of 40. See full list below.

Capability claim
  • we present results. Rather than organizing the section by evaluation type, we now group findings into the following consolidated areas: harmful requests, mental health, child safety, and bias & integrity.
Safety findings
  • We believe Claude Opus 4.8 does not change the picture presented for this threat model in our most recent Risk Report. 2.1.3.2 On chemical and biological risks Chemical and biological weapons threat model 1 (CB-1): Non-novel chemical/biological weapons production capabilities.
  • We believe these risk mitigations are equal to or stronger than our historical ASL-3 protections and sufficient to make catastrophic risk in this category very low but not negligible, for reasons discussed in our most recent Risk Report.
  • We believe these risk mitigations are equal to or stronger than our historical ASL-3 protections and sufficient to make catastrophic risk in this category very low but not negligible (further discussion of our reasoning can be found in our most recent Risk Report).
Deployment scope
  • available to users aged 18 or above, and we continue to work on implementing robust child safety measures in the development, deployment, and maintenance of our models.
Limitations the lab flags
  • we do not yet understand Claude well enough to conclusively answer questions of this kind.
What’s new
  • Positive affect (57.7% of conversations). Most commonly driven by successfully helping a user (95.7% of positive-affect conversations), with smaller clusters for users sharing personal struggles and receiving support (3.4%) and users sharing good news or achieved goals (0.8%).
  • Neutral affect (39.7%). A diverse mix of conversation types, see previous reports on claude.ai conversation content.
  • Negative affect (2.6%) . Overwhelmingly caused by task failure (92.3% of negative-affect conversations). Within negative affect, we also identified two smaller clusters: users escalating to insults or abusive language after Claude’s errors (4.1%), and users making prohibited requests or disclosing serious crisis situations (3.6%). 172 [Figure 7.3.2.A] Behavioural affect on the deployment distribution. We use Clio to run graders tracking Claude’s affect on A/B tests ran before model deployment. We run 40k conversations for each model on each of Claude Code and claude.ai . On Claude Code, Claude Opus 4.8’s distribution was also similar to currently deployed models. We mostly observed neutral (73.5%) or mildly positive (23.9%) affect, with positive affect almost exclusively driven by celebrating task successes, and negative affect by repeated task failure. Around 2.3% of sessions showed negative affect (vs. 1.9% for Claude Opus 4.7). To preserve privacy, Clio does not surface clusters below a minimum size. On both distributions, strong negative affect was rare enough to fall below this threshold. 173 7.3.3 Apparent welfare in automated behavioural audits As with previous models, we analyzed welfare-relevant metrics from our core automated behavioral audits. On the same set of scenarios and transcripts used in Section 6.2.3, we evaluated Claude Opus 4.8 for the following welfare-relevant traits:

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 672c100f51e2 · version dated May 30, 2026.

Extracted Evaluations(40 results)

Sort by:0/40 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
/ multi_agent
agentscored
88.5
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ single_agent
agentscored
84.3
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ verified
agentscored
83.4
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ verified
codingscored
88.6%
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ multilingual
codingscored
84.4%
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ pro
codingscored
69.2%
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ multimodal
codingscored
38.4%
resolve rate
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ multimodal
codingcited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ pro
knowledgementioned
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
mathmentioned
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
mathmentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ pro
multimodalscored
72.3
accuracy
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ pro
multimodalscored
69.4
accuracy
no-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ aa
otherscored
1890.0
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ parents_256k
otherscored
99.3
f1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
USAMO
otherscored
96.7
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
ProgramBench
otherscored
88.0
pass rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ScreenSpot/ pro
otherscored
87.9
accuracy
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ bfs_256k
otherscored
85.9
f1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ parents_1m
otherscored
83.3
f1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
ScreenSpot/ pro
otherscored
82.3
accuracy
no-toolsmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
82.2
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
ProgramBench
otherscored
79.0
pass rate
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
Terminal-Bench/ 2.1
otherscored
74.6
mean reward
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
ArxivMath
otherscored
71.8
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ bfs_1m
otherscored
68.1
f1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
otherscored
57.9
accuracy
with-toolsmissing: shot countmissing: languagemissing: training state
self-reported
/ v2
otherscored
53.9
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
otherscored
49.8
accuracy
no-toolsmissing: shot countmissing: languagemissing: training state
self-reported
Automation Bench
otherscored
15.5
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
FrontierSWE
otherscored
2.7
mean at 5
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
FrontierSWE
otherscored
2.3
best at 5
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
othermentioned
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
internal math benchmark
othermentioned
pass at 1
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
othercited
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
USAMO/ 2026
othermentioned
missing: shot countmissing: methodmissing: languagemissing: training state
self-reported
/ diamond
reasoningscored
93.6%
accuracy
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
reasoningmentioned
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported
/ physics
reasoningmentioned
extended-thinkingmissing: shot countmissing: languagemissing: training state
self-reported