For years, generative artificial intelligence has mostly been described as an answering machine: writing text, summarizing a document, producing code, suggesting ideas. With GPT-6 Astra, unveiled by OpenAI on September 3, the center of gravity shifts. The point is no longer just how well a model can formulate an answer, but how much work it can complete within a real digital environment.
OpenAI describes Astra as its most capable model for computer use, browsing, software engineering, cybersecurity, science, and professional work. The rollout begins with a limited group of organizations, with a gradual expansion to ChatGPT and the APIs. Behind the marketing statements, however, lies a concrete product transformation: AI is increasingly being trained to navigate between applications, files, websites, and tools, maintaining context across long-running tasks and making operational decisions along the way.
From chatbot to executor
The difference is visible in the use cases highlighted by OpenAI in its presentation. Astra can fill out forms, update records in a CRM, organize calendars, conduct online research, work on spreadsheets, generate presentations following a template, build websites and test them in a browser, install software, and verify that a workflow actually works.
It is an important list not because every single feature is entirely new, but because it indicates the direction in which frontier models are heading: less isolated conversation, more end-to-end execution. In other words, AI is attempting to fill the space that currently separates a request from a finished result. It is that space—made of clicks, checks, edits, switching between applications, and verifications—that absorbs a massive share of office work.
According to data published by OpenAI, Astra improves significantly over GPT-5.6 Sol in benchmarks dedicated to computer use and automation. On OSWorld 2.0, for instance, the company reports a score of 72.6% compared to 65.7% for the previous model; on AutomationBench, the reported leap is from 18.1% to 41.4%. These are useful figures to gauge the trajectory, but they should not be confused with a guarantee of reliability in day-to-day work: a controlled benchmark does not replicate the complexity of an enterprise management system, a website with a changing interface, or a process involving incomplete data.
The real game is delegation
The most interesting shift is therefore organizational. If a model can execute a sequence of tasks for forty minutes, an hour, or more, the relationship between humans and software changes. AI is no longer used merely to accelerate a single phase: an objective is delegated, and the outcome is reviewed.
This explains why model makers are investing so heavily in working memory, navigation capabilities, permission management, and instruction comprehension. To be genuinely useful as an agent, a system cannot just be “smart”: it must remember what it was asked, understand what it can do without asking for confirmation, pause when a decision is consequential, and not lose the thread when the user modifies part of the task.
OpenAI claims that Astra was specifically trained to better maintain its bearings during long tasks and handle ambiguity more sensibly. It is a less spectacular aspect than a new benchmark record in mathematics, but arguably more important for enterprise adoption. An agent that less frequently misinterprets the “why” of a task is worth more than a model that gains a few points on a leaderboard and then requires constant supervision.
The most sensitive side: cybersecurity
However, this leap in capability comes at a cost in terms of risk. OpenAI states that GPT-6 Astra is its first broadly deployed model to reach the “Critical” level for cybersecurity capabilities in its Preparedness Framework. In internal testing, the model demonstrated vastly superior capabilities in finding and exploiting vulnerabilities; the company says that during evaluations, Astra even discovered two previously unknown zero-day vulnerabilities, which were subsequently reported to the maintainers.
This is probably the most significant part of the entire release. The more capable a model becomes of acting autonomously on a computer, the more necessary it becomes to monitor not just what it says, but what it does. OpenAI has therefore introduced trajectory monitoring, automated controls, and specific limitations on more advanced cyber activities. Here too, the issue goes beyond OpenAI: it is a structural challenge for the entire new generation of agents.
What changes for workers
The temptation is to translate every advance into a binary question: how many jobs will it replace? That question is too simplistic. In the short term, models like Astra appear better suited to compressing entire sequences of tasks than replacing whole professions. Someone who previously spent half a day juggling research, spreadsheets, documents, and presentations may find themselves defining the problem, delegating part of the execution, and spending more time on review.
This shifts value toward three skill sets: formulating the goal clearly, providing access to the right data, and verifying the output. This does not necessarily mean less work in absolute terms; it could mean that the same person manages multiple processes simultaneously. That is why the impact of AI on work may show up in productivity and team structure before it appears in aggregate employment figures.
The model market enters a new phase
With Astra, the competition among major labs becomes even less of a race over “who answers better” and increasingly a contest over who can turn general capabilities into reliable work. The competitive edge is shifting from pure intelligence to integration: browsers, computer use, APIs, enterprise tools, permissions, memory, monitoring, and costs.
This is also why the lines between model and product are beginning to blur. An outstanding model lacking a sound operating environment can be less useful than a slightly less capable model embedded in a system that knows how to access files, request permissions, use applications, and verify results.
The real test for GPT-6 Astra, then, will not be the highest score on a benchmark. It will be seeing how many real-world tasks people begin handing over to it from start to finish and, above all, how often they can accept the outcome without having to redo the work from scratch. That is where the shift from chatbot to digital coworker is measured.



