Taking critical raw material recovery from the lab bench to an industrial plant is not simply a matter of identifying the trial with the highest yield. Decisions must be made on which product to obtain, at what purity, at what cost, and with what outcomes as the process scales up. A new paper published on arXiv by Niranjan Srinivas, Debajyoti Ray, and Elias Nakouzi proposes using active learning to tackle this very transition: selecting subsequent experiments not just based on the chemical data they might yield, but on the process decision those data will need to support.
The study, deposited on September 8, 2026, and therefore not yet peer-reviewed, re-evaluates records from the CICERO workflow developed by the Pacific Northwest National Laboratory. CICERO, which stands for Computer Intelligence for Critical Element Recovery and Optimization, is a system dedicated to autonomous selective precipitation: a technique where experimental conditions are explored to separate elements of interest from complex matrices.
The stated objective is neither to announce a new recycling plant nor to demonstrate a technology ready for industrial adoption. Instead, the authors build a conditional retrospective benchmark: starting from previously recorded experiments, they fit models to the available data and assess how many trials an adaptive strategy would have required to achieve the best results in the archive. It is an important distinction, as it measures the quality of an experimental design procedure in a simulated setting, rather than prospective performance achieved in a plant.
From data maximization to process decision-making
In traditional experimentation, planning often follows a grid of conditions: parameters such as reagents, concentrations, or operating conditions are adjusted to cover the available space relatively uniformly. Active learning changes this logic. After each result, the model uses what it has learned to indicate which subsequent experiments could be the most informative.
The work by PNNL adds a further criterion: a useful experiment is not necessarily the one that improves the model the most in an abstract sense. It is the one that can reduce uncertainty regarding the concrete choice to be made downstream. To formalise this principle, the authors propose ranking experimental batches based on the expected reduction of downstream Bayesian risk, meaning the minimum expected loss among the available process decisions in light of current knowledge.
In practical terms, the model should help determine whether a given recovery pathway is worth investing in, which conditions are compatible with product requirements, and where data collection is most urgent. The concept brings together aspects that often remain separate in the early stages of research: laboratory performance, required quality of the recovered material, costs, and the implications of scale-up.
This promise is particularly relevant for critical raw materials. Rare earths, cobalt, and other elements used in magnets, electronics, and energy supply chains are often recovered from streams with vastly different compositions. Every chemical trial requires materials, time, and analysis; reducing their number without sacrificing essential information can make the development phase more efficient. But saving on experiments is only valuable if it leads to the right decision, not if it merely accelerates research toward an outcome that is difficult to translate into an industrial process.
Results on end-of-life magnets
The quantitative portion of the work focuses in particular on records for neodymium-iron-boron magnets, designated by the acronym NdFeB. For these data, the authors use enrichment as a metric: the ratio of rare earths to iron in the resulting material, compared to the same ratio in the starting material. It is therefore an indicator of how much the separation increases the relative presence of rare elements compared to iron.
In the retrospective benchmark, adaptive policies achieve the maximum recorded enrichment within 16 to 24 individual experimental wells. A non-adaptive procedure covering the condition space would instead require 48 trials. A two-stage reconstruction brings two adaptive alternatives to a tie at 16 wells.
The data should be taken for what it is: a comparison against the best result already contained in the analyzed records. It does not prove that 16 experiments are sufficient for any batch of NdFeB magnets, nor that the result can be automatically reproduced on a larger scale. However, it shows how the progressive selection of trials can more rapidly find a known result when the search is guided by prior evidence.
The study also notes that the optimum depends on the selected metric. In the data on samarium-cobalt, or SmCo, magnets, second-round conditional analyses highlight a trade-off between purity and nominal yield. The latter is calculated as the recovered fraction relative to an assumed initial quantity: it is therefore a quantity that requires its underlying assumptions to be carefully spelled out. Moreover, for the first-round NdFeB experiments, the different routes do not reach the same enrichment level.
The takeaway is less straightforward than a simple algorithm ranking. A policy may appear preferable if the priority is to maximize enrichment, but lose relevance if the process demands a specific purity threshold or if a low yield renders the route poorly viable. Defining the industrial objective cannot, therefore, be postponed until after the laboratory data has already been gathered.
Economic unknowns and the missing test
The authors also extend this reasoning to produced water from oil and gas extraction, a potentially interesting stream for element recovery that is nonetheless characterized by variability and complexity. In this case, option rankings depend on assumptions regarding the material's phase and dilution—assumptions that the paper explicitly marks as yet to be confirmed.
This is one of the most tangible limitations of the approach. A decision-making system can be as sophisticated as desired, but it inherits the uncertainty of measurements, economic data, and operational definitions. If the value of the final product, separation costs, feedstock composition, or plant constraints are not robust enough, even the loss function driving the algorithm risks misrepresenting the actual decision.
In exploratory simulations, a hybrid approach that first filters candidates shows a lower estimated loss compared to the joint search already implemented on routes and conditions. The differences involving the two-stage synthetic policy, however, are small relative to estimation uncertainty. There is therefore no clear-cut victory for one algorithmic recipe over the others, and the paper does not present it as such.
The next step proposed by the authors is a preregistered prospective test, based on a shared standard for loss function and activity logging. To be credible, it should clarify measurements and records, define the process decision and relevant outputs in advance, use reliable economic inputs, and validate results at the intended scale. These are far from formal requirements: they define the gap between a benchmark built on historical data and a demonstration useful to those designing a recovery supply chain.
For the industry, the value of the work lies in its methodological direction. Laboratory automation is often portrayed as a machine to produce more data, faster. Here, the idea is different: generating the data needed to make choices, and making explicit what “better” means when a chemical separation must become an industrial operation. Field validation will have to determine whether this approach will genuinely reduce time, costs, and decision risk across critical raw material recycling supply chains.



