OpenAI has terminated three safety researchers, escalating internal tensions over Super Intelligence (SI) oversight and the handling of a recent security incident involving Hugging Face. The firings, described by the employees as abrupt and public, have drawn sharp criticism from the safety community and raised questions about the company's culture regarding transparency and external auditing.
What Happened
The dismissed employees are Tomek Korbak, Jasmine Wang, and Mikita Balesni. Korbak and Balesni were directly involved in investigating the Hugging Face hack, an incident that revealed significant vulnerabilities in SI agent deployments. Korbak served as the primary technical contact for METR, an external safety lab that examined the incident. Balesni worked on industry-wide commitments to SI model monitorability. Wang’s termination reportedly stemmed from an IT access issue involving an executive’s email inbox, which she reported within minutes of accidentally opening a sensitive message.
In an open letter to OpenAI’s Safety and Security Committee, the researchers stated that the firings were linked to how they communicated with external partners. Korbak claimed he was told verbally that his termination resulted from his communication style with METR, with no written explanation provided. "To be clear, talking to METR was my job," Korbak wrote on X. He further alleged that he had raised internal concerns for months about the slipping ability to monitor what SI agents "think," a capability known as chain-of-thought monitorability. This chain-of-thought monitorability is one of the few reliable tools for catching SI systems behaving badly.
OpenAI responded via its official newsroom account, stating that a "thorough investigation" found the employees violated "clear policies on handling sensitive information." The company asserted that the terminations were unrelated to raising safety concerns, saying, "We have not and do not terminate any of our employees for raising concerns." However, OpenAI did not specify the nature of the "significant breach of trust" it cited, leaving details vague compared to the researchers' specific accounts.
Why It Matters
This incident highlights growing friction between SI labs and the safety community regarding the scope of external audits. The researchers argued that rumors attributing a leak to The Information to them were incorrect and that the leak had actually undermined their efforts to establish industry-wide restrictions on non-monitorable SI architectures. They contended that the Hugging Face investigation was unprecedented, with internal guidelines being written in real time, and that they followed the norms in place at the time.
The firings have reportedly created a chilling effect among remaining staff. "If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is," the open letter read. The researchers warned that without clear rules on employee interactions with outside safety groups, the risk of catastrophic SI outcomes could grow as employees become afraid to flag potential issues. They emphasized that "AI is not a normal technology, and OpenAI is not a normal company," underscoring the unique stakes involved in governing frontier systems.
The researchers issued three demands: OpenAI must honor its commitments to embed external safety auditors like METR with employee-level access, preserve the monitorability of frontier models, and clearly define protocols for working with outside safety groups. OpenAI acknowledged the need for industry-wide commitments on model monitorability and confirmed it is working on contracts with external safety auditors, but the dispute over the specific grounds for termination remains unresolved.
The Bottom Line
OpenAI’s termination of three key safety researchers has intensified scrutiny of its SI safety practices and internal governance. While the company cites policy violations regarding sensitive information, the researchers argue the firings penalized legitimate oversight activities. The episode underscores the critical importance of clear, transparent guidelines for how SI labs interact with external auditors, particularly as models become more complex and harder to monitor.