Model Cards / Anthropic

Claude Opus 4.7 System Card

model card63,024 words·274 min read·May 15, 2026·Source
Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
63,024-word document condensed to 242 words. Anthropic · May 15, 2026
TL;DR

This system card describes Claude Opus 4.7, a large language model from Anthropic. Overall, the model shows superior capabilities to those of its predecessor, Claude Opus 4.6, but weaker capabilities than those of our most powerful model, Claude Mythos Preview.

Capability claim
  • We are releasing Opus 4.7 with a new set of cybersecurity safeguards.
Safety findings
  • We believe our risk mitigations are sufficient to make catastrophic risk from non-novel chemical/biological weapons production very low but not negligible.
  • We believe that catastrophic risk from novel chemical/biological weapons remains low (with substantial uncertainty). The overall picture is similar to the one from our most recent Risk Report.
  • We believe that the overall risk is very low, and that this model in particular adds little to the risk picture we previously laid out for Claude Mythos Preview .
Deployment scope
  • accessible to experts, we interpret a model’s performance on this task primarily based on the expert’s assessment of uplift.
Limitations the lab flags
  • limitations included sycophantic agreement under pushback, verbose responses that buried actionable content, degraded reference accuracy, and overconfidence in the feasibility of synthesis steps.
  • open questions — particularly around fully explaining the evaluation- awareness results — that they would have preferred more time to resolve; and that the internal-usage evidence base for this model was thinner than for some prior releases.

Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA f055e7ef9acc · version dated May 15, 2026.

Extracted Evaluations(48 results)

Sort by:0/48 rows fully reproducible (0%)
BenchmarkCategoryStateScoreSetupSource
agentmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
agentmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
agentmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
agentmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ verified
codingmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ pro
codingmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ multilingual
codingmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ multimodal
codingmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
multilingualmentioned
Averageinstruction-tunedmissing: shot countmissing: method
self-reported
Firefox 147
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
EARL
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
SHADE-Arena
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Minimal-LinuxBench
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Terminal-Bench
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
USAMO/ 2026
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
MRCR/ v2
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Lab-Bench/ figqa
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
CharXiv/ reasoning
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
ScreenSpot/ pro
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
VendingBench
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ aa
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
GMMLU
othermentioned
Averageinstruction-tunedmissing: shot countmissing: method
self-reported
MILU
othermentioned
Averageinstruction-tunedmissing: shot countmissing: method
self-reported
INCLUDE
othermentioned
Averageinstruction-tunedmissing: shot countmissing: method
self-reported
BioPipelineBench/ verified
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
BioMysteryBench/ verified
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Long-form virology tasks
othermentioned
without-safeguardsinstruction-tunedmissing: shot countmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Sequence-to-function/ modeling
othermentioned
extended-thinkinginstruction-tunedmissing: shot countmissing: language
self-reported
Sequence-to-function/ design
othermentioned
extended-thinkinginstruction-tunedmissing: shot countmissing: language
self-reported
AECI
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Petri
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Reward hacking evaluations
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Automated Behavioral Audit
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Claude self-preference evaluation
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
Decision theory evaluation
othermentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
/ diamond
reasoningmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
reasoningmentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported
safetymentioned
instruction-tunedmissing: shot countmissing: methodmissing: language
self-reported