Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
61,956-word document condensed to 501 words. Anthropic · May 30, 2026
TL;DR
“run on Claude Opus 4.8. It includes the following sections: Responsible Scaling Policy evaluations.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| GDPval | extended-thinking, aa | 1890.00 |
| GraphWalks | parents_256k, f1 | 99.3% |
| USAMO | extended-thinking, accuracy | 96.7% |
| GPQA | extended-thinking, diamond, accuracy | 93.6% |
| SWE-bench | extended-thinking, verified, resolve_rate | 88.6% |
| BrowseComp | extended-thinking, multi_agent, accuracy | 88.5% |
| ProgramBench | pass_rate | 88.0% |
| ScreenSpot | with-tools, pro, accuracy | 87.9% |
Showing top 8 of 40. See full list below.
Capability claim
- “we present results. Rather than organizing the section by evaluation type, we now group findings into the following consolidated areas: harmful requests, mental health, child safety, and bias & integrity.”
Safety findings
- “We believe Claude Opus 4.8 does not change the picture presented for this threat model in our most recent Risk Report. 2.1.3.2 On chemical and biological risks Chemical and biological weapons threat model 1 (CB-1): Non-novel chemical/biological weapons production capabilities.”
- “We believe these risk mitigations are equal to or stronger than our historical ASL-3 protections and sufficient to make catastrophic risk in this category very low but not negligible, for reasons discussed in our most recent Risk Report.”
- “We believe these risk mitigations are equal to or stronger than our historical ASL-3 protections and sufficient to make catastrophic risk in this category very low but not negligible (further discussion of our reasoning can be found in our most recent Risk Report).”
Deployment scope
- “available to users aged 18 or above, and we continue to work on implementing robust child safety measures in the development, deployment, and maintenance of our models.”
Limitations the lab flags
- “we do not yet understand Claude well enough to conclusively answer questions of this kind.”
What’s new
- •“Positive affect (57.7% of conversations). Most commonly driven by successfully helping a user (95.7% of positive-affect conversations), with smaller clusters for users sharing personal struggles and receiving support (3.4%) and users sharing good news or achieved goals (0.8%).”
- •“Neutral affect (39.7%). A diverse mix of conversation types, see previous reports on claude.ai conversation content.”
- •“Negative affect (2.6%) . Overwhelmingly caused by task failure (92.3% of negative-affect conversations). Within negative affect, we also identified two smaller clusters: users escalating to insults or abusive language after Claude’s errors (4.1%), and users making prohibited requests or disclosing serious crisis situations (3.6%). 172 [Figure 7.3.2.A] Behavioural affect on the deployment distribution. We use Clio to run graders tracking Claude’s affect on A/B tests ran before model deployment. We run 40k conversations for each model on each of Claude Code and claude.ai . On Claude Code, Claude Opus 4.8’s distribution was also similar to currently deployed models. We mostly observed neutral (73.5%) or mildly positive (23.9%) affect, with positive affect almost exclusively driven by celebrating task successes, and negative affect by repeated task failure. Around 2.3% of sessions showed negative affect (vs. 1.9% for Claude Opus 4.7). To preserve privacy, Clio does not surface clusters below a minimum size. On both distributions, strong negative affect was rare enough to fall below this threshold. 173 7.3.3 Apparent welfare in automated behavioural audits As with previous models, we analyzed welfare-relevant metrics from our core automated behavioral audits. On the same set of scenarios and transcripts used in Section 6.2.3, we evaluated Claude Opus 4.8 for the following welfare-relevant traits:”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA 672c100f51e2 · version dated May 30, 2026.
Extracted Evaluations(40 results)
Sort by:0/40 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
/ multi_agent | agent | scored | 88.5 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ single_agent | agent | scored | 84.3 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | agent | scored | 83.4 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ verified | coding | scored | 88.6% resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ multilingual | coding | scored | 84.4% resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ pro | coding | scored | 69.2% resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ multimodal | coding | scored | 38.4% resolve rate | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ multimodal | coding | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ pro | knowledge | mentioned | — | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
| math | mentioned | — | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
| math | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
/ pro | multimodal | scored | 72.3 accuracy | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ pro | multimodal | scored | 69.4 accuracy | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ aa | other | scored | 1890.0 | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ parents_256k | other | scored | 99.3 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
USAMO | other | scored | 96.7 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
ProgramBench | other | scored | 88.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ScreenSpot/ pro | other | scored | 87.9 accuracy | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
/ bfs_256k | other | scored | 85.9 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ parents_1m | other | scored | 83.3 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
ScreenSpot/ pro | other | scored | 82.3 accuracy | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 82.2 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
ProgramBench | other | scored | 79.0 pass rate | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Terminal-Bench/ 2.1 | other | scored | 74.6 mean reward | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
ArxivMath | other | scored | 71.8 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
/ bfs_1m | other | scored | 68.1 f1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 57.9 accuracy | with-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
/ v2 | other | scored | 53.9 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
| other | scored | 49.8 accuracy | no-toolsmissing: shot countmissing: languagemissing: training state | self-reported | |
Automation Bench | other | scored | 15.5 accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
FrontierSWE | other | scored | 2.7 mean at 5 | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
FrontierSWE | other | scored | 2.3 best at 5 | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
| other | mentioned | — | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
internal math benchmark | other | mentioned | — pass at 1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | cited | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
USAMO/ 2026 | other | mentioned | — | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
/ diamond | reasoning | scored | 93.6% accuracy | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |
| reasoning | mentioned | — | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported | |
/ physics | reasoning | mentioned | — | extended-thinkingmissing: shot countmissing: languagemissing: training state | self-reported |