Claude Opus 5

Claude Opus 5

anthropic· claudehealthy

Claude Opus 5: curated reference profile, Proprietary license

Claude Opus 5: curated reference profile drawn from the model's published reports. Not a Supermagin measurement. Run the official eval in the Studio to earn a verified score.

Anthropic+4% net improvement

Curated reference profile · not a Supermagin measurement

Model card & documentation

Claude Opus 5

Curated reference profile · Proprietary license · v2026.07.28 · 2.5s latency

These figures are drawn from the model's published reports and are NOT a Supermagin measurement. Run the official eval in the Studio to earn a verified, re-verifiable score.

Capabilities

  • agent: 96%
  • reasoning: 94%
  • cybersecurity: 93%
  • multimodal: 92%

Notes

  • Low hallucination rate on audited evals.
  • Reference evals: SWE-Bench Verified, GPQA Diamond, MMLU-Pro, MATH-500, ARC-AGI-2, BigCodeBench, τ-bench.

Pooled score

0%

credibility 100% · consistency 0.96

Score trend across attested runs

51990 total samples

More attested runs are needed to show a trend.

Category scores

reasoning
94%
coding
91%
agent
96%
multimodal
92%
cybersecurity
93%

Progress by year

2023
82%
2024
88%
2025
92%
2026
96%

Success rate

0%

Failure rate

0%

Hallucination

0%

Avg latency

2500ms

per task

Avg cost

$0.0200

per task

Calibration gap

+0.05

positive = overconfident

Top failure signals

hallucination ×572logic_flaw ×1248instruction_failure ×936

Pooled brain diagnostics

Output Entropy
healthy
Top-Token Confidence
healthy
Participation Ratio
healthy
Condition Number
healthy
Gradient Chain Health
healthy

Failure rate by task type

code
7%18197
qa
5%15597
reasoning
6%10398
classification
5%7799

Curated reference profile, descriptive data, not a Supermagin measurement.