Chaptered summary is still being generated for this document. Showing a heuristic brief in the meantime.
Summary
66,240-word document condensed to 287 words. Anthropic · Jun 28, 2026
TL;DR
“throughout the system card. We preserved model welfare context about the initial safeguards, noting where the prior version is described.”
Top benchmarks
| Benchmark | Variant | Score |
|---|---|---|
| GDPval-AA | economically valuable knowledge work | 1932.00 |
| SWE-bench Verified | max effort | 95.5% |
| SWE-bench Verified | production safeguards | 95.0% |
| SWE-bench Multilingual | 9 languages | 92.2% |
| GraphWalks BFS 256K | long context | 91.1% |
| CharXiv Reasoning | no tools | 88.9% |
| Terminal-bench 2.1 | mini-SWE harness | 88.0% |
| BrowseComp | single-agent | 88.0% |
Showing top 8 of 27. See full list below.
Capability claim
- “we are releasing it in these two forms: Fable 5, which is for general use but comes with additional safeguards that block its ability to perform tasks in high-risk domains such as biology and cybersecurity; and Mythos 5, which has relevant safeguards lifted but is only made available to a small number of trusted partners (beginning with those in Project Glasswing).”
Safety findings
- “We believe these mitigations make catastrophic risk in this category low but still not negligible, for reasons discussed in our most recent Risk Report.”
- “We believe that Mythos 5 falls short of the specific threshold in version 3.3 of our RSP and in our FCF. But we are nonetheless concerned about the risks it poses in this category, and we think that world-class human expert substitution may now be possible in a few areas.”
- “we believe to be grader-incentivized. Rates of three behaviors in response to steering against grader awareness with three different vectors, on training environments with high-risk of grader exploitation.”
Mitigations
- “we have deployed Mythos 5 to the general public with additional safeguards as Claude Fable 5.”
Deployment scope
- “available to a small number of trusted partners (beginning with those in Project Glasswing).”
Limitations the lab flags
- “We do not yet measure this thoroughly, and signals of affect may be more directly tied to, and conditional on, experiential states.”
What’s new
- •“Noted CAIS’s contribution to the VCT CB-1 evaluation in Section 2.2.4.1.”
- •“Corrected a minor error in our description of the”
Every italicized passage is a verbatim substring of the source document (checked deterministically after extraction). Field selection is heuristic — some quotes may lack surrounding context and some claims may be absent if no matching pattern appeared. For citation, open the source: original model card · source SHA ac781cf43bf7 · version dated Jun 28, 2026.
Extracted Evaluations(27 results)
Sort by:0/27 rows fully reproducible (0%)
| Benchmark | Category | State | Score | Setup | Source |
|---|---|---|---|---|---|
| agent | scored | 88.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| agent | scored | 85.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| agent | scored | 85.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 95.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| coding | scored | 95.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| medical | scored | 62.7 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 1932.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 92.2 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
GraphWalks BFS 256K | other | scored | 91.1 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 88.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Terminal-bench 2.1 | other | scored | 88.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Terminal-bench 2.1 | other | scored | 84.3 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
BioMysteryBench | other | scored | 83.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 80.3 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 80.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
ArxivMath | other | scored | 78.5 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 66.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 59.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
| other | scored | 57.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
RiemannBench | other | scored | 55.0 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
| other | scored | 54.9 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported | |
Blueprint-Bench 2 | other | scored | 38.6 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
GDP.pdf | other | scored | 29.8 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
FrontierCode | other | scored | 29.3 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
CritPt | other | scored | 28.6 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Automation Bench | other | scored | 17.4 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |
Legal Agent Benchmark | other | scored | 13.3 | missing: shot countmissing: methodmissing: languagemissing: training state | self-reported |