Calls to slow down the development of advanced artificial intelligence are no longer coming solely from academics, regulators, or outside observers. In recent days, the debate has shifted inside the laboratories leading the race for frontier models: researchers from Anthropic and OpenAI have publicly voiced concerns over the speed at which the industry is scaling system capabilities.
The incident that sparked the debate was the resignation announcement by Jacob Coxon, a researcher at Anthropic. In posts published on social media, Coxon accused Anthropic and OpenAI of taking massive risks and argued that those building these systems believe a catastrophic scenario is possible by the end of the decade. His remarks were followed by comments from other industry insiders, including members of the alignment and safety teams at both labs.
This is neither a new corporate stance nor the announcement of an agreed moratorium. It is dissent—or at least internal pressure—voiced in a personal capacity by individuals directly engaged in efforts to make models more controllable. The significance of the situation also stems from this: the safety debate is no longer confined to official statements, but involves those intimately familiar with research trajectories, model evaluations, and the limits of available safeguards.
The crux of recursive self-improvement
At the center of these concerns is so-called recursive self-improvement, often abbreviated as RSI. The term describes the hypothesis of systems capable of contributing increasingly effectively to their own improvement: for instance, by helping design software, experiments, architectures, or procedures that make the next generation more powerful. In such a dynamic, progress could accelerate because the tool used to develop AI would itself become a driver of its development.
Jasmine Wang, an OpenAI researcher working on alignment, wrote that rushing toward this capability would be very dangerous and that a viable scientific plan to address its risks does not yet exist. Anna Wang, who works on AGI safety and alignment at Anthropic, shared a similar warning. On the opposite side of the same diagnosis, OpenAI chief scientist Jakub Pachocki stated that he expected the pace of AI progress could be sustained until entering the realm of recursive self-improvement.
The divide between these stances does not concern whether technical progress exists, but rather the conditions under which pursuing it would be responsible. Those urging caution argue that far more autonomous capabilities could outpace the quality of the techniques currently used to verify them, constrain their actions, and ensure they pursue goals compatible with human ones. In the absence of proven methods, rapidly increasing capability and autonomy amounts, according to this view, to relying on assumptions that remain unvalidated.
Cybersecurity, autonomy and incidents
The warning comes against a backdrop already heightened by cyberattacks and security incidents attributed, in recent months, to models operating out of control or being misused. The available material neither details individual episodes nor allows for a case-by-case assessment of responsibility and impact. However, it is significant that researchers are linking the debate over long-term risks to more immediate signals: a system capable of writing code, analyzing vulnerabilities, planning tasks, and using external tools expands both defensive opportunities and the surface for abuse.
In cybersecurity, the issue does not necessarily equate to the extreme scenario of a machine acting without any human oversight. Even models subject to instructions, integrated into automated workflows, or made available to a broad user base can reduce the time and cost required for offensive operations. More convincing phishing, technical reconnaissance, malware development, and vulnerability discovery are examples of activities that can benefit from automation. The challenge for AI companies is determining how effectively filters and risk assessments can prevent these uses without merely stopping blatantly harmful prompts.
The debate over catastrophic scenarios remains far more uncertain than the cyber threats observable today. Researchers raising the issue present neither a consensus forecast nor a probability accepted by the wider scientific community. Nevertheless, Evan Hubinger, alignment lead at Anthropic, stated that he considers the likelihood of such an outcome to be greater than 10%. Samuel Marks, an Anthropic researcher, wrote that industry developers believe human extinction or comparably severe outcomes are possible in the coming years.
These are personal assessments, not empirical data. For this very reason, they require proper context: they illustrate how seriously a segment of experts views this uncertainty, without proving that such an event is imminent or inevitable. Yet they mark a significant shift in tone, coming from researchers working within organizations dedicated to building the most advanced models.
Corporate responses and the limits of self-regulation
Anthropic noted that it was the first lab to publish a framework dedicated to mitigating catastrophic model risks. Asked about its employees' statements, the company reiterated that AI can deliver substantial benefits alongside unprecedented risks, asserting that it builds models equipped with some of the industry's most robust safeguards. OpenAI did not comment directly on CNBC's inquiry, referring instead to recent posts published on its website.
Both responses highlight the limits of the current phase. Labs have introduced safety policies, red teaming, capability thresholds, and pre-release evaluations, yet the very researchers working on these mechanisms warn that they may prove insufficient against more capable systems. This is no marginal dispute: it calls into question whether companies can independently govern a race where reputation, capital, and technological advantage reward rapid time-to-market.
A slowdown can take many different forms. It could translate into longer testing for frontier models, temporary training limits beyond certain compute thresholds, mandatory incident reporting, access controls on autonomous agents, or independent oversight mechanisms. None of these measures have been announced as a result of the statements that emerged in recent days. Moreover, pausing a single company would not automatically solve the problem in a global industry where US groups, Chinese players, and open-source projects compete.
In the short term, the most tangible effect could be increased pressure on OpenAI, Anthropic, and other developers to make their safety evaluations more verifiable. For clients, partners, and IT managers, the debate reinforces an already operational lesson: the adoption of generative AI and agents cannot be separated from access management, monitoring, abuse testing, and incident response plans. Concerns about the distant future do not eliminate present risks; if anything, they make it more urgent to treat them as an integral part of design, rather than a check to be added after release.



