In the world of smart contract security, we have a rule: if the auditor is paid by the project, the audit is a rubber stamp. Anthropic’s second Responsible Scaling Policy (RSP) risk report, published around mid-2025, applies the same principle to AI safety. The framework is elegant. The execution is opaque. And the governance model — self-assessment, self-publication, self-supervision — is a textbook case of centralization risk dressed in institutional garments.
Context: The Hype Cycle of AI Self-Regulation
Anthropic positioned itself as the security-first AI lab. Its RSP, first released in May 2023, was the industry’s first systematic scaling policy, borrowing the biosafety level (BSL) taxonomy to classify model capabilities into ASL-2 through ASL-4. The second report, analyzed here, claims to transition the framework from a static document to a dynamic evaluation mechanism. The report covers frontier risk assessments on Claude 3/3.5 series in CBRN (chemical, biological, radiological, nuclear), cyber offense, and autonomous replication. It defines operational canary metrics for ASL-3. On paper, it is a governance innovation.
But I have audited enough protocols to know that “innovation” in governance is often a polite term for “we made up the rules.” The crypto industry learned this the hard way with DAOs that had multisigs controlled by three founders. The RSP’s self-assessment model is the same architecture — a closed loop where the entity that creates the risk also evaluates it.
Core: The Systematic Teardown of the RSP Governance Model
Let me quantify the centralization risk of this framework. The RSP assigns Anthropic the following roles: risk definer, threshold setter, evaluator, adjudicator, and reporter. There is no independent third-party audit. The framework promises to introduce external auditing — the RSP policy text mentions it — but the second report does not confirm implementation. This is the equivalent of a DeFi protocol promising a timelock after a governance attack but never actually deploying it.

The coverage blind spot is equally concerning. The RSP focuses exclusively on catastrophic risks — CBRN, large-scale cyber attacks, autonomous replication. It systemically ignores everyday social risks: bias, discrimination, privacy violations, psychological manipulation. In my 2017 audit of 0x Protocol V2, I found seven critical re-entrancy flaws because the team was focused on feature velocity, not edge cases. Anthropic is doing the same: it prioritizes risks that could destroy its reputation overnight while ignoring the slow bleed of trust erosion from biased outputs. Code does not lie, but the auditors often do.
The threshold setting is a matter of arbitrary discretion. Where does ASL-3 start? What level of CBRN knowledge diffusion qualifies as dangerous? Anthropic owns the answer, and the public has no way to verify the correctness. This is worse than a closed-source smart contract — at least bytecode can be decompiled. Here, the evaluation methodology is proprietary, the test sets are unpublished, and the peer review is absent. We built a house of cards on a ledger of trust.
Contrarian: What the Bulls Got Right
To be fair, the RSP is a genuine institutional innovation. Among frontier AI labs, Anthropic is the only one that has published periodic risk reports with a structured framework. OpenAI’s Preparedness Framework (October 2023) and Google DeepMind’s Frontier Safety Framework (2024) are statements of intent, not operational engines. The second report proves that the RSP is not a one-off PR stunt. It is a living process. For enterprise clients in regulated industries — finance, healthcare, government — this creates a verifiable trust signal. In the crypto world, a project that releases a security audit report every quarter is already more credible than one that releases a single PDF at launch.
Additionally, the RSP creates a side effect market: third-party AI safety evaluation services. The demand for red teaming, capability assessment, and mitigation verification is continuous, not one-time. This is analogous to how DeFi audits evolved from a checkbox to a recurring expense. The RSP’s existence forces competitors to invest in similar infrastructure, raising the baseline for the whole industry.
Takeaway: The Accountability Test Is Yet to Come
The RSP second report is a signal, not a solution. It signals that Anthropic’s governance machinery is running. But the real test — the one that will determine whether this framework has teeth — is when ASL-4 triggers deployment restrictions that conflict with revenue. When a model is deemed too dangerous to release, and the market is hungry for the next generation, will Anthropic honor the policy? Security is a process, not a badge you wear.
I have seen this pattern in crypto: projects that preach decentralization but hold admin keys. They eventually use those keys. The RSP’s self-audit model is the same admin key. Until an independent third party can verify the thresholds, the evaluations, and the mitigation actions, the RSP remains a governance artifact — elegant, but unproven.
From my experience auditing the Terra-Luna algorithmic stablecoin before its collapse, I learned that the most dangerous risks are the ones that the system’s designers refuse to model. The RSP models catastrophic risk but ignores societal risk. That blind spot is where the next crisis will emerge. The crypto industry has a saying: "Your key, your risk. Their code, their bug." Anthropic’s code is the RSP. Their bug is the lack of external audit. The ledger remembers every exploit.
