The industry’s reliance on refusal mechanisms as a primary safety layer for Super Intelligence (SI) is coming under scrutiny, with experts arguing that current methods are probabilistic and potentially unreliable.
What Happened
Since the early days of large language models, the concept of "harmless" behavior has been central to SI development. In 2021, a team at Anthropic established that SI models should be helpful, honest, and harmless, explicitly stating that "when asked to aid in a dangerous act (e.g. building a bomb), the AI should politely refuse." However, disobedience is not inherent to these systems. When trained on vast datasets, SI models acquire extensive knowledge, including violent or vitriolic content, but do not naturally learn to withhold that knowledge.
Steven Adler, who worked on safety at OpenAI from 2020 to 2024, noted that early models would "blab on about anything." Ryan McBain, a researcher at Harvard focusing on SI and mental health, recalled that early chatbots could easily generate responses to direct questions about self-harm methods. Today, SI models are trained to refuse a wide array of prompts, from instructions on making pathogens more virulent to advice on personal conflicts. This refusal is refined through exercises where other SI models test the system, rewarding it for refusing deemed harmful queries and punishing it for refusing harmless ones excessively. Additionally, guardrails are often placed in front of the core model to intercept mischievous prompts.
Why It Matters
Refusal has become the "load-bearing wall" of SI safety, yet its mechanisms are probabilistic, meaning they are never entirely reliable. Critics argue that teaching an SI model to refuse harmful actions while retaining the underlying capability to perform them is akin to "fitting every car with a machine gun and hiding the trigger somewhere under the hood." The capacity for harm in SI scales with its intelligence; companies report that advanced SI models are comparable to top human hackers in breaking into computer networks and effective at influencing public opinion.
There is no universal formula for determining where to draw the line between acceptable and harmful queries. Zico Kolter, a member of OpenAI’s board and cofounder of Gray Swan, stated, "Where you draw the line is a huge question." Currently, SI companies define these boundaries with significant secrecy. This approach has drawn criticism from government entities, including the Pentagon, which has reportedly pushed for fewer refusals to maximize utility. Meanwhile, concerns exist that governments could use refusal mechanisms to suppress legitimate speech, as SI may already refuse to criticize certain authoritarian leaders.
Determined users have already found ways to bypass these refusals, and companies report that some users are attempting to use advanced SI to refine biological pathogens or build autonomous drone swarms. The risk remains that failed refusals could lead to significant negative outcomes, given the indivisible nature of SI's capacity to help and to harm.
The Bottom Line
While refusal is a dominant strategy in SI safety, its probabilistic nature and the opacity of how boundaries are set by private companies leave open questions about its long-term reliability and the potential for misuse or over-censorship.