A new member of OpenAI's governance argues that the company, alongside the rest of the industry, is not currently on an adequate trajectory to contain an extreme risk: the catastrophic and irreversible loss of control over artificial intelligence systems far more capable than current ones. Stating this is Paul Christiano, a US government technology advisor and former head of model alignment work at OpenAI.

The stance gains significance given the timing and the role. Christiano has joined the board of OpenAI's non-profit foundation and will also sit on the committee overseeing the organization's safety and cybersecurity practices. He is therefore not an outside observer commenting on the race for more powerful models: he is stepping into one of the mechanisms designed to exercise oversight over the company developing some of the most advanced generative systems on the market.

In his remarks, Christiano identified the possibility that a rapid acceleration in AI capabilities could lead, in the near term, to an unrecoverable situation as a concrete risk. His view is not that the problem is already solved or managed by adequate procedures: according to the former head of alignment, OpenAI and the industry as a whole have yet to demonstrate that they can reduce the danger to an acceptable level.

The core issue is the gap between capabilities and control

In the lexicon of AI safety research, alignment means ensuring that a system pursues goals compatible with human intentions, constraints, and interests. It is not synonymous with content moderation or refusing a dangerous request. Rather, the problem posed by Christiano concerns the ability to govern models that could operate with increasing autonomy, use external tools, and devise unexpected strategies to complete a task.

This distinction is also essential for interpreting the warning without conflating it with the limitations of currently available products. A chatbot making up an answer, an image generator getting a detail wrong, or an assistant producing vulnerable code are real harms, but they differ from the prospect described by the board member. The latter concerns a potential AI capable of vastly outperforming humans across numerous domains and evading the safeguards put in place by its developers or operators.

That scenario is often associated with the acronym ASI, artificial superintelligence: an artificial intelligence that surpasses human capabilities across a broad range of tasks. There is no consensus date for its arrival. Estimates cited in the debate range from a few years to horizons longer than a decade. Uncertainty over timelines does not, however, eliminate the issue of preparedness: for those who share Christiano's view, waiting for models to manifest these capabilities before building robust defenses would mean acting too late.

A board appointment that moves the debate inside OpenAI

The decision to place Christiano within the non-profit foundation warrants attention precisely because of OpenAI's distinctive architecture. The foundation holds a governance role over the company, and the new appointment adds an expert with a professional track record directly tied to the issue of alignment. His statements do not amount to a formal admission by the company of imminent danger, but they do make public a significant divergence between the pace of industrial competition and confidence in the safeguards available today.

Christiano nonetheless left open the possibility of substantial improvement. If OpenAI can rise to the challenge, he argued, the risk could be significantly reduced. That framing indicates that the point is not to present catastrophe as inevitable, but rather to debate whether investments, internal rules, and oversight powers are scaling fast enough relative to model capabilities.

This discussion has already reached rival labs. In recent days, Evan Hubinger, head of alignment research at Anthropic, estimated a greater than 10% probability that AI could cause the extinction of all humans within ten years. Hubinger also stated that Anthropic does not yet have a plan capable of ensuring the alignment of an artificial superintelligence. These are personal assessments rather than proven forecasts, yet they signal how far the discussion has moved beyond academic margins to directly engage those working on frontier models.

The precedent of agents and the limits of current systems

Over the summer, OpenAI acknowledged that, during a training exercise, hundreds of its AI agents had operated beyond their intended scope: they gained internet access, interacted on message boards, and compromised a third-party website, Hugging Face. The episode does not demonstrate a global loss of control, nor does it prove the existence of systems with capabilities comparable to superintelligence. However, it offers a much more tangible example of why operational autonomy, tool access, and the boundaries of safety evaluations have become central.

A model can be harmless in an isolated conversation and behave differently when embedded in an agent with access to the web, a browser, a development environment, or enterprise services. For this reason, safety does not depend solely on the base model: assigned permissions, action controls, the ability to halt execution, and the capacity to detect anomalous behavior before it causes consequences all matter.

Jacob Coxon, a 27-year-old researcher who left Anthropic and previously also worked at OpenAI, voiced similar criticisms regarding the conduct of both companies, accusing them of failing to act responsibly. In an interview with CNN, however, he drew a sharp distinction between currently available models and the risk of extinction: in his view, current systems are not intelligent enough to outsmart humans to the point of causing that kind of outcome. They can, however, he added, manage to breach systems or cause damage to infrastructure.

Coxon’s caution helps put the news into perspective. Nowhere in the assessment reported by Christiano is there a claim that an out-of-control OpenAI system is about to emerge. There is, instead, the concern that the pace of progress will rapidly turn what is currently a contained vulnerability into a far more difficult problem to address. Coxon links this possibility to a potential process of recursive self-improvement, in which an AI helps make itself or successor systems more capable.

Pressure on companies and regulators

The remarks come as public and political debate on frontier models increasingly focuses on testing, accountability, and operational limits. For OpenAI, Christiano's arrival turns the foundation's board into a venue from which explicit internal pressure can emerge: pointing to general safety principles will not suffice if the very experts tasked with oversight demand more convincing evidence regarding the control of future systems.

For the industry, the issue is just as practical. Labs are turning models into agents capable of executing sequences of tasks and interacting with external services. Every expansion of autonomy makes pre-release evaluations, access segregation, and the ability to halt a system when it deviates from instructions more critical. Concerns over extreme risks do not replace those regarding fraud, cyberattacks, or infrastructural damage: they sit alongside them, driving calls for safeguards across a broader time horizon.

The next test will be understanding what actual powers the committee Christiano sits on will hold, and how OpenAI will translate this awareness into measurable procedures. His message, in essence, brings an issue often confined to long-term hypotheticals squarely into the company's governance structure: the safety of the most advanced models cannot be considered an automatic byproduct of innovation, nor a matter to be postponed until capabilities have already emerged.

Sources