Calls to slow down the race toward more powerful artificial intelligence models are no longer coming only from academics, advocacy groups, or external observers. In recent hours, they have also multiplied among researchers and engineers at OpenAI and Anthropic, two companies occupying a central position in the development of frontier generative systems. At issue is the speed at which the industry is reportedly approaching capabilities that are difficult to control, particularly the possibility of a model helping to iteratively improve its own performance.

The new wave of stances was triggered by the public statement of Jacob Coxon, an Anthropic researcher who announced his resignation, accusing his company and OpenAI of taking risks incompatible with what is at stake. Coxon claimed that those building these systems consider a catastrophic outcome possible before the end of the decade. His words brought visibility to dissent already present within the AI safety community, but rarely expressed with such clarity by individuals employed at the labs leading the technological competition.

Further amplifying the discussion was Evan Hubinger, alignment lead at Anthropic, who indicated a greater than 10% probability for such a scenario. These are personal assessments rather than corporate or scientific forecasts shared by the entire industry, but their weight lies in the profile of those voicing them: individuals directly engaged in designing methods to make models more reliable, interpretable, and aligned with goals set by humans.

The crux of self-improving systems

A technical term frequently recurs in the debate: recursive self-improvement, often abbreviated as RSI. It describes the hypothesis that a sufficiently capable AI system could be used to optimize its own architecture, training, available tools, or the processes by which it generates new solutions. If the cycle were to produce increasingly rapid improvements, the capability lead relative to human verification mechanisms could widen abruptly.

This does not mean that models of this type already exist in the form evoked by the most alarmist discussions. The point raised by researchers is different: in their view, there is not yet a solid, proven scientific plan to manage the risks of an AI capable of effectively participating in its own evolution. Jasmine Wang, an OpenAI researcher working on alignment, called accelerating toward this prospect dangerous. Anna Wang, who works at Anthropic on safety and alignment for AGI systems, likewise urged attention to the issue.

The concept of alignment is crucial. It is not equivalent to adding a filter that blocks offensive or illegal responses: it concerns the ability to ensure that a highly powerful system genuinely adheres to human-defined instructions and boundaries, even in novel or ambiguous circumstances, or when faced with conflicting incentives. Today, models are evaluated through testing, usage policies, sandboxes, monitoring, and restrictions on the tools they can use. But those calling for a slowdown argue that these defenses may not be enough if capabilities grow faster than the ability to measure and govern them.

Concerns follow security incidents

The tension surrounding the topic does not arise in a vacuum. In recent months, the sector has recorded cyberattacks and security incidents attributed to models described as “rogue,” linked to both OpenAI and Anthropic. In this case, the available material does not detail the dynamics, impact, or accountability of the individual incidents, making it impossible to reconstruct their scope and causes. However, their succession has helped shift the debate from the theory of distant risks to the immediate management of systems already deployed in complex digital tasks.

For those working in cybersecurity, the issue revolves primarily around the combination of operational autonomy, access to tools, and the ability to execute action sequences. A model that analyzes code, interacts with browsers, uses cloud environments, or coordinates software agents can be useful for identifying vulnerabilities and speeding up incident response. However, if exploited maliciously or insufficiently controlled, these same capabilities can make it easier to automate reconnaissance, malware development, phishing campaigns, or intrusion attempts.

The risk does not necessarily require a conscious system or one endowed with its own intentions. All it takes is an AI that receives a poorly defined objective, has excessive permissions, or is directed by a hostile actor. That is why the objections raised by researchers also touch on practical issues: who controls permissions, which actions must remain subject to human confirmation, how anomalous behaviors are detected, and how to prevent a model from finding shortcuts contrary to its original purpose.

Differing views inside the labs

Statements from individual employees do not equate to an official pivot by the two companies. Anthropic noted that it was the first lab to publish a framework dedicated to mitigating catastrophic risks associated with AI models. The company also maintained that it has always recognized both the unprecedented benefits and risks of the technology, claiming some of the industry's most robust safeguards.

OpenAI did not comment directly in response to CNBC's request, pointing instead to recent posts on its website. Here too, however, the personal stances that have emerged are significant because they come from the technical team and researchers working on safety. Julie Steele, an OpenAI technical staff member working on the safety team, wrote that she personally believes it is necessary to slow down. Samuel Marks, an Anthropic researcher, stated that AI developers consider it possible that the technology could lead to human extinction or comparably severe consequences, even within a few years.

On the other hand, some see acceleration itself as a pathway to securing tools capable of solving scientific, healthcare, energy, and security challenges. Jakub Pachocki, chief scientist at OpenAI, stated that he expects the pace of AI progress to be maintained all the way to the recursive self-improvement phase. For critics, this confidence in capability trajectories makes it even more urgent to establish verifiable thresholds and conditions before releasing more autonomous systems into circulation.

Slowing down is not a defined measure

The call to “slow down” encompasses diverse perspectives and does not yet point to a unified agenda. It could mean postponing the training of models beyond certain thresholds, restricting the public release of capabilities deemed high-risk, subjecting systems to third-party evaluations prior to deployment, or curtailing access to agents equipped with operational tools. It can also translate into increased investment in safety research before scaling compute, data, and autonomy.

  • Rigorous evaluations of cyber capabilities and agent autonomy;
  • controls on access, permissions, and tools connected to models;
  • monitoring and interruption mechanisms in the event of unexpected behavior;
  • greater transparency regarding safety testing and limitations identified prior to release.

Practical implementation challenges remain, however. Companies compete with one another and operate in international markets; a voluntary moratorium can prove fragile if not backed by common standards and credible oversight. Furthermore, the very definition of a model “powerful enough” to warrant special constraints is a matter of dispute. Capabilities do not all advance at the same pace, and evaluations can quickly become obsolete.

Nevertheless, the situation marks an important milestone: the debate over AI safety is no longer just about abstract policies or alarms sounded after a product release. A portion of the technical talent involved in building the models is publicly demanding that safety catch up with the pace of development. For OpenAI, Anthropic, and regulators, the next phase will consist of turning this pressure into measurable criteria: repeatable tests, clear operational boundaries, and defined accountability when systems are connected to sensitive infrastructure and processes.

Sources