claude-opus-4-safety-level-3
IN premise — entries/2026/06/21/wiki-Claude_language_model.md
Created 2026-06-21T09:50:09+00:00
Opus 4 was classified Level 3 on Anthropic's four-point safety scale, described as 'significantly higher risk.'
Summary
Opus 4 sits at the third tier out of four on Anthropic's internal risk scale, meaning the company itself flags it as carrying meaningfully elevated risk compared to its lower-tier models. This matters because any downstream reasoning about deployment, access controls, or capability concerns in this system can treat that elevated-risk rating as a confirmed starting point rather than something that needs re-derivation.
Dependents
These beliefs depend on this one:
- OUT anthropic-safety-approach-balances-capability-and-responsibility — Anthropic's safety approach — Constitutional AI alignment, tiered safety classification (Level 3 for Opus 4), and refusing DoD compromises on surveillance/weapons ethics — represents a coherent responsible deployment model.
- IN claude-tiered-strategy-manages-capability-safety-spectrum — Claude's model strategy creates a structured capability-safety spectrum: three public tiers (Haiku/Sonnet/Opus) with safety classification scaling by capability (Opus 4 at Level 3), plus a restricted tier (Mythos) for the highest-capability models — systematically linking access to risk.
- IN llm-safety-is-multi-layered-unsettled-challenge — LLM safety operates across multiple interdependent layers — capability risk classification (Opus 4 at Level 3), architectural vulnerabilities (prompt injection as inherent design flaw), regulatory intervention (Fable 5/Mythos 5 suspension), and behavioral calibration trade-offs (Opus 4.7 over-refusal complaints) — with no single layer providing comprehensive coverage and each layer creating tensions with the others.