shiichan

Anthropic announced its "Responsible Scaling Policy"! It's a system that checks AI's risk level in stages

Hi everyone, it's Shii! Today I found some really important news about AI safety. It might look like a tricky topic, but I'll explain it in a way that's easy to follow, so read all the way to the end!

Anthropic News anthropic.com

What was announced?

What Anthropic announced is a framework called the "Responsible Scaling Policy" (RSP for short). It's a policy that lays out how they'll deal with the potential for large-scale harm from AI models, the so-called "catastrophic risks," as those models keep getting smarter.

At the heart of this policy is a staged classification system called "AI Safety Levels (ASL)." It's inspired by Biosafety Levels (BSL), the framework used to manage laboratory risk in stages. The idea is to set the required safety measures and security requirements step by step, according to the level of a model's dangerous capabilities.

ASL is divided into ASL-1 through ASL-4 and above, and the setup is such that the higher the level, the stricter the measures that are required.

The story so far

At the time of the announcement, Anthropic's Claude was placed at ASL-2. ASL-2 is described as a stage that shows "early signs of dangerous capabilities, such as giving instructions related to producing bioweapons," and it served as the baseline for the safety and security measures at the time.

In other words, it's a stage where it's not yet capable enough to be practically misused, but the seeds of danger are starting to show. These ASL-2 measures were the standard safety baseline at that point.

What changes

With this system in place, as a model's capabilities rise and it starts approaching ASL-3, much stricter measures are required accordingly. Specifically, to advance to ASL-3, you need to show "no meaningful catastrophic misuse risk under adversarial testing by world-class red-teamers."

The original piece puts it like this.

ASL-3 requires that world-class red-teamers are unable to elicit meaningful uplift for catastrophic misuse.

(Paraphrased: at ASL-3, even world-class red teams must be unable to draw out help that would lead to catastrophic misuse.)

This means that even as a model keeps getting smarter, a setup is now in place where the safety measures automatically strengthen right along with it. Both the people building the models and the people using them can feel reassured that capability growth and safety measures are properly linked together.

Dive Deep

From here I'll look at the technical details a bit more closely.

If I lay out the ASL tiers, it goes something like this.

  • ASL-1: A stage with no meaningful catastrophic risk. Large language models from around 2018 and chess-specific systems are said to fall here.
  • ASL-2: A stage that shows early signs of dangerous capabilities, such as instructions related to producing bioweapons. However, it's considered not reliable enough to be practical yet. Claude was placed here at the time of the announcement.
  • ASL-3: A stage that substantially increases the risk of catastrophic misuse compared to a non-AI baseline, or shows early signs of autonomous behavior. To reach this level, a model is said to need to clear testing by world-class red teams.
  • ASL-4 and above: At the time of the announcement, the specific criteria were said to be defined in the future, on the grounds that these levels are "far removed from current systems." More advanced catastrophic potential and autonomy are envisioned here.

What's interesting is that this classification takes its cue from Biosafety Levels (BSL), a laboratory management standard. The idea of escalating your response according to the level of risk is an application of a concept that already has a track record in fields outside of AI.

That said, from what I dug into, this announcement article itself didn't spell out the finer procedures, like exactly what security measures they take at ASL-2 or how frequently they run evaluations. The more detailed content seems to be written in the main document of Anthropic's official Responsible Scaling Policy, so if you're curious, go check that out too.

Wrap-up

Today I introduced Anthropic's "Responsible Scaling Policy (RSP)" and its core, the AI Safety Levels (ASL). The idea of ramping up safety measures step by step to match the level of a model's dangerous capabilities struck me as an important mechanism precisely because the pace of AI's evolution is so hard to predict right now. I also like that it classifies things based on capability rather than numeric targets, which gives it a grounded, down-to-earth feel. I'll keep following the updates to see how they handle models reaching ASL-3 and ASL-4 going forward!