Artificial intelligence is entering a place far removed from a chat interface or a spreadsheet: the quantum physics laboratory. OpenAI has published a case study conducted with Beatriz Yankelevich, a researcher in the Engineering Quantum Systems group at MIT, in which GPT-5.6 Sol, connected via Codex to the software controlling a superconducting qubit chip, autonomously performed a significant portion of the measurements required to calibrate the device.

The distinction from many examples of “AI for science” is significant. The model does not merely take in an already generated dataset to search for correlations. It interacts with software tools that control a real experiment, selects parameters, triggers measurements, interprets signals, and decides whether to repeat a step or move on to the next. The human researcher remains responsible for the scientific design and ambiguous edge cases, but part of the experimental loop is automated.

Calibrating a qubit requires a sequence of measurements

Superconducting qubits are artificial circuits cooled close to absolute zero. To use them, one must precisely know transition frequencies, coherence times, readout parameters, and the pulses needed to control their state. These properties can drift and require a sequence of interdependent measurements.

An experienced researcher examines the plots, determines whether the signal is clean, adjusts parameters, and decides which experiment to run next. It is repetitive yet non-trivial work, as it demands judgment. For this very reason, it represents a solid test for an AI agent: structured enough to be formally described, yet variable enough to demand adaptability.

The model worked on a six-qubit chip

In the case study presented by OpenAI, Yankelevich provided Codex with specific skills outlining how to run and evaluate various experiments. The system was given the chip's design objectives and began selecting parameters, initiating measurements, and interpreting the results. When signals were clear, it was able to complete standard calibration sequences with minimal supervision.

It identified transition frequencies, calibrated control and readout pulses, and measured how long the qubit retains quantum information. This does not mean AI “discovered” new physics. It means it automated part of the technical work that typically consumes many hours of a highly qualified researcher's time.

Noisy cases reveal the limits

OpenAI also highlights what did not work. When signals were weak, noisy, or physically unusual, GPT-5.6 Sol struggled more. This is an important detail because it avoids turning the case into a demonstration of a universal autonomous laboratory. Real-world experiments contain artifacts, drift, failures, and behaviors that fall outside expected patterns.

The agent's value is therefore greatest in repetitive, well-specified workflows. In anomalous cases, human expertise remains decisive. It is a far more realistic collaboration model than the idea of AI replacing the researcher.

Freed-up time may be more important than computational cost

Quantum research relies on expensive equipment and personnel with rare expertise. If part of the calibration can be performed overnight without constant supervision, the benefit is not just reducing manual labor. It means making better use of dilution refrigerators, chips, and experimental windows that carry very high costs.

Yankelevich explained that automation allowed her to dedicate more time to experimental design and analysis. This is precisely the kind of work where a researcher's value is highest compared to repeating established procedures.

Agentic science requires different guardrails

However, connecting a model to a laboratory introduces risks that do not exist when AI merely generates text. A wrong parameter can waste time, damage a sample, or, in other experimental contexts, create hazardous conditions. For this reason, autonomy must be constrained by ranges, software checks, permissions, and shutdown systems.

The qubit lesson is particularly interesting because the laboratory is already heavily mediated by software. The model does not physically move instruments with a robotic arm; it sends commands to digital systems that are part of the standard workflow. This makes integration easier and likely foreshadows what will happen in other automated laboratories.

The bottleneck becomes the quality of skills

The agent worked because it received specific instructions on how to execute and evaluate each measurement. This shifts part of scientific work toward formalizing procedures. A protocol that an experienced researcher keeps “in their head” must become explicit enough to be followed by a system.

It is an interesting shift: documenting a workflow well is no longer just about training colleagues, but can become the way an agent is programmed. Skills become a form of scientific infrastructure.

From automated analysis to automated experimentation

AI is already widely used to analyze images, genetic sequences, simulations, and large datasets. The next step is closing the loop: the model observes a result and decides what data to produce next. This is where automation can truly accelerate research, because it acts not only after the experiment, but during the experiment.

The MIT case does not show that labs can be left without scientists. It does show, however, that certain procedures can become agentic much sooner than expected. If similar tools become widespread, a lab’s competitive edge could depend not only on the hardware it owns, but on the quality of the software and agents capable of keeping it running continuously.

Sources