Anthropic Responsible Scaling Policy (v1.0 2023 -> v2.0 2024 -> v2.1/2.2 2025 -> v3.0 2026)
Anthropic · 2023-09-19 (v1.0); 2024-10-15 (v2.0); 2026-02-24 (v3.0)
Anthropic's rulebook for when it will add safeguards as its models get more capable. The 2026 rewrite replaced the unconditional promise to pause if safeguards fall behind with a conditional one, and added more public reporting.
What it calls for
- Safety evaluations
- Pause/moratorium
- Transparency
- Compute governance
Scope
Single company; version history at https://www.anthropic.com/rsp-updates
What actually happened
Anthropic activated ASL-3 protections for Claude Opus 4 (May 2025) under v2.x. RSP v3.0 (effective 2026-02-24, https://www.anthropic.com/responsible-scaling-policy/rsp-v3-0) replaced the unconditional commitment not to train or deploy without adequate safeguards with conditional delay commitments (Appendix A: 'We will delay AI development and deployment as needed' — but only if Anthropic has a clear lead, or until it matches competitors' risk posture), alongside periodic Risk Reports and a non-binding Frontier Safety Roadmap; GovAI's analysis calls this an honest but weakened commitment (https://www.governance.ai/analysis/anthropics-rsp-v3-0-how-it-works-whats-changed-and-some-reflections). Anthropic did hold a separate red line: it refused the Pentagon's Feb 2026 demand to drop weapons/surveillance restrictions and won an Aug 28 2026 ruling that the retaliatory blacklist was unlawful (https://techcrunch.com/2026/08/28/anthropic-gets-its-first-court-win-over-the-pentagons-supply-chain-risk-label/).