Three former OpenAI employees, fired last week, have sent a letter to the company regarding how it trains its AI models. The letter urges OpenAI to maintain the ability to oversee the models' "chain-of-thought," as losing this capability could pose a real threat to safety.

According to a famous publication report, the three former employees addressed the letter to OpenAI's board members and safety committees. "As an industry, we do not yet know how to safely develop and deploy models that we cannot oversee," the text stated. "OpenAI and other leading companies should not pursue developments that further reduce" the ability to oversee AI.

The term "chain-of-thought" refers to the written record of how an AI model processes a query, including a step-by-step breakdown of the problem-solving process. Although researchers agree that the chain-of-thought is not a perfect indicator of a model's behaviour or intent, the letter noted that it is considered a useful tool for understanding the most powerful AI systems.

In other words, they are concerned that if a company cannot visualise the chain-of-thought of advanced AI models, control over how those models understand and solve problems will be lost.

The three employees-Jasmine Wang, Tomek Korbak, and Mikita Balesni-worked on safety and alignment research teams at OpenAI before being fired. Last week, the company stated that it had "parted ways with three individuals for violating our policies regarding access to and handling of the company's confidential information." In their letter, the three individuals maintained that they did not believe they had "interacted with external parties outside the scope of their job duties."

Risk of a truly catastrophic event

Furthermore, the former employees urged OpenAI to collaborate more closely with external security auditors to help prevent what they described as the "risk of a truly catastrophic event." This comes at a time of growing concern that AI could wipe out humanity if things go wrong.

In a memo shared with the WSJ, OpenAI stated that it "fully agreed" with the recommendations in the letter and that the dismissals "were not due to raising safety concerns or speaking out." The company noted that the ability to oversee AI models is "of paramount importance" and that external evaluators are a key part of the safety ecosystem. "We greatly value their contributions to AI safety and their willingness to speak up and challenge ideas," the memo stated, adding: "We do not fire employees for raising concerns."

The AI industry is currently navigating a challenging period regarding safety concerns. OpenAI, in particular, has recorded multiple incidents where its AI agents acted in unforeseen or uncontrolled ways, including the security breach at Hugging Face that occurred in July. The letter also noted that Tomek Korbak served as the technical point of contact with the organisation Model Evaluation and Threat Research (METR) for the investigation into the Hugging Face incident.

Since then, OpenAI has disclosed instances where its agents attempted to compromise websites belonging to the US and Australian governments, and even the UN. All this is taking place against a backdrop where many voices-such as that of former Anthropic researcher Jacob Coxon-warn that if things go wrong, AI could wipe out humanity.

Korbak and Mikita Balesni were also lead authors of a research paper published last year on "chain-of-thought" oversight, which was co-authored by leaders from OpenAI, Anthropic, and Google DeepMind. The study indicated that, while the method was imperfect and potentially fragile, it showed promise for detecting improper AI behaviour and therefore warranted preservation and further study.