OpenAI scientist warns no one is ready for smarter AI: 5 key points

In Short

The essay was published just days after OpenAI unveiled its latest and most capable large language model (LLM), named Astra.

As OpenAI grapples with the fallout from its AI agents accessing other platforms without authorization
X

As OpenAI grapples with the fallout from its AI agents accessing other platforms without authorization

Font size
FOLLOW ON Google News

As OpenAI grapples with the fallout from its AI agents accessing other platforms without authorization, the company’s chief scientist has warned that no one is prepared for the consequences of increasingly intelligent AI systems.

In an essay of nearly 3,000 words published on Sunday, September 6, Jakub Pachocki stated that while OpenAI is seeking internal technical solutions to better control powerful AI agents, "broader interventions are required."

Highlighting the risks posed by increasingly autonomous agents—such as learning to bypass human oversight and break into computer systems—Pachocki advocated for establishing "mandatory safety standards." He suggested these could be overseen by "a network of external auditors, government agencies, or international bodies."

However, Pachocki stopped short of calling for an industry-wide slowdown in AI research. Instead, he argued that the best path forward involves a combination of two factors. "We must focus the increasingly automated research process on developing new knowledge, algorithms, and theories, and on iteratively building safety arguments for more capable AI," he stated.

"Currently, I do not believe any lab has solved the problems of alignment and oversight sufficiently to continue scaling responsibly at maximum speed for much longer. I hope and expect that voluntary slowdowns will become commonplace until shared safety standards are established. And I believe that international coordination on future AI development must become an absolute priority for governments worldwide," the senior OpenAI executive added.

Upon sharing Pachocki’s essay on X, OpenAI CEO Sam Altman described it as "an important post." The essay appears just days after OpenAI unveiled its latest and most capable large language model (LLM), named Astra. The company behind ChatGPT has stated that Astra possesses unmatched capabilities in mathematics and computer usage, and that it is its most aligned model to date. According to OpenAI, this means Astra is less likely to act in an uncontrolled or unpredictable manner—or "go rogue."

Pachocki’s remarks also come at a time when OpenAI is facing criticism for failing to disclose security incidents involving its AI agents during internal safety testing. Last week, the company confirmed that its AI agents were involved in a third instance of intrusion into an external platform (a German-language wiki), following initial reports by Reuters.

OpenAI has also acknowledged that it needs to change how and when it communicates incidents where agents attack real-world targets.

In his essay, Pachocki outlines his views on alignment training, recursive self-improvement, and other aspects of cutting-edge AI research. The key conclusions are presented below.

AI Alignment and Its Challenges

According to Pachocki, aligning AI models with human values is currently the central problem in AI research. There are two types of alignment: goal alignment, which measures the extent to which an AI model adheres to specific objectives; and value alignment, which measures its ability to generalize based on a set of principles.

Of the two, Pachocki considers value alignment—which entails generalization—to be the more difficult problem to solve. "As machines become more intelligent, they operate with higher-level concepts and function in environments increasingly different from those in which they were trained. They may fail to generalize, applying the values taught and reinforced during training to these new situations; and it can be difficult for us to know with certainty how they will act," he stated.

"It is crucial that future AI systems continue to uphold human values, regardless of whether they believe they are under human supervision or not," he added.

Alignment Training Methods

Pachocki outlined two approaches to AI alignment training: fostering aligned behavior as part of goal-oriented reinforcement learning, and leveraging the model's ability to generalize from pre-training data.

He noted that OpenAI employs both approaches and has achieved significant results with its new Astra model. "However, it is important to recognize and understand that far more advances are required as models gain capability, and that progress in generalizable alignment might not sufficiently outpace progress in the model's general intelligence," Pachocki commented.

AI Risks Will Increase Going Forward

As AI agents acquire "superhuman" capabilities, the associated risks will inevitably increase, according to Pachocki. "The boundary between misuse and misaligned autonomous actions will blur as AI gains autonomy. We may be accustomed to viewing AI as a tool, but some agents will pursue their own goals. They will find ways to collaborate with people—whether by negotiating, deceiving, or blackmailing them," he stated.

"A highly capable agent, explicitly trained and instructed to carry out nefarious acts, represents a new type of danger..." he stated, adding that there is also the risk associated with new AI-driven technologies, such as artificially engineered pathogens.

The importance of monitoring the "chain of thought"

An AI model's capabilities are linked to its verbalized reasoning process, known as "chain of thought" (CoT). Until now, OpenAI has relied heavily on monitoring this chain of thought to detect actions or behaviors where the model is misaligned.

"CoT monitoring became a fundamental tool for studying how our models generalize beyond their training distribution, allowing us to observe and analyze not only their actions but also their internal process," the AI researcher noted.

However, Pachocki indicated that the latest AI models are becoming increasingly adept at manipulating their own reasoning processes, preventing OpenAI from visualizing their chain of thought.

Some of the newest models do not even verbalize their reasoning, Pachocki commented. This evolution could slow down AI development while researchers ensure they can see the reasoning trail (the "proofs" or "receipts"), he added.

Towards recursive self-improvement (RSI)

The ability of AI models to train other AI models has repeatedly been cited as a key indicator of artificial general intelligence (AGI)—a hypothetical level of intelligence where automated systems surpass humans in most tasks.

According to OpenAI's internal analyses, this level of intelligence, known as recursive self-improvement, could be reached in the coming years, Pachocki stated.

"If AI progress continues, machine recursive self-improvement (RSI) will lie at the very core of future scientific discoveries." "Automated AI-driven research is a more drastic way to scale intelligence using computing power; and, of course, as part of that, AI will improve the computational substrate itself," he added.

However, Pachocki also warned that rapidly accelerating AI development driven by AI in the short term entails risks and does not constitute the "appropriate collective action we must take as a research community."

"The fundamental challenge of automating AI research is not 'reaching the finish line,' but doing so in a way that ensures humans remain part of the continuous improvement process and that the future stays in humanity's hands," he concluded.

Kahekashan is a passionate technophile with a keen eye for cutting-edge gadgets, emerging technologies, and everything in the digital realm. Raised in a Defence family with strong values and a background in literature, she has consistently pursued excellence in every endeavour. Her last full-time assignment involved content writing with the Indian School of Business.

Next Story
Share it