The search for new antimicrobials increasingly relies on the automated reading of massive biological archives. In the laboratory led by César de la Fuente, a researcher known for his work on antimicrobial peptides, Codex and ChatGPT are being used to speed up the pathway leading from genetic sequence analysis to the identification of molecules potentially useful against drug-resistant infections.

The goal is not to entrust a language model with discovering a ready-to-use antibiotic. The work remains an experimental process made up of hypotheses, computational analysis, molecular synthesis, and laboratory testing. AI is positioned upstream in this pipeline: it helps the team write and edit code, organize analyses, query biological data, and reduce the time needed to screen sequences that warrant further verification.

The stakes are real. Antimicrobial resistance limits the effectiveness of medicines that have underpinned modern medicine for decades, from therapies for common infections and surgical procedures to cancer treatments. Finding compounds with novel or complementary mechanisms of action is therefore a research priority, but traditional discovery is time-consuming and carries a high failure rate.

From genomes to a candidate list

The laboratory's starting point is the idea that many useful molecular instructions may already be contained in the data available in nature. Genomes of modern organisms, as well as sequences reconstructed from ancient specimens, hold proteins and protein fragments that have not necessarily been studied as potential antimicrobials. Peptides, relatively short chains of amino acids, are one of the most promising areas: some can inhibit bacteria and other microorganisms through mechanisms distinct from those of conventional antibiotics.

The problem is scale. A genome is not an off-the-shelf catalog of drugs: it contains a vast amount of information, and even when focusing on specific sequence families, plausible candidates must be distinguished from those lacking the desired characteristics. This requires computational pipelines capable of extracting data, applying filters, comparing biological properties, and producing outputs that can be evaluated by expert researchers.

In this context, Codex is used to support software development. It can assist in creating, reviewing, and debugging the scripts required to handle complex datasets, as well as making technical steps more accessible that would otherwise consume a significant share of the workload. ChatGPT, on the other hand, is used as an interlocutor to explore ideas, structure problems, facilitate the interpretation of results, and support research and documentation tasks.

This does not mean the model replaces expertise in bioinformatics, microbiology, or chemistry. On the contrary, the quality of the output depends on the questions asked, the data provided, the chosen criteria, and the team's ability to identify errors or false leads. In a discipline where a minor variation in a peptide sequence can alter activity, toxicity, or stability, human oversight remains an integral part of the method.

Extinct genomes also become a source of hypotheses

One of the most distinctive aspects of the work described by de la Fuente is extending the research to the genomes of extinct organisms. Analyzing this material makes it possible to retrieve sequences that are absent, or no longer easily observable, in modern living organisms. It is an approach that broadens the scope of discovery: instead of searching solely within contemporary biodiversity, hypotheses can be formulated based on a biological archive built over the course of evolution.

However, these molecules are not automatically “brought back to life” as pharmaceuticals. A massive gap separates a digital sequence from a potential clinical application. A candidate identified computationally must first be synthesized and evaluated in appropriate experiments; it must demonstrate activity against the target microorganisms, avoid unacceptably damaging host cells, maintain efficacy under realistic conditions, and exhibit pharmacological properties compatible with future development.

Research into ancient sequences should therefore be seen as an exploratory strategy rather than the promise of an immediate cure. Its value lies in uncovering leads that methods based solely on chemical libraries or already known pathogens might overlook. It represents a shift in perspective: genetic heritage, both present and past, becomes a dataset to be queried using tools capable of finding patterns far faster than manual analysis.

Speed does not equal validation

The use of generative models in science is often framed around the theme of acceleration. In this case, that acceleration primarily affects the initial workflow cycle: turning a biological insight into a reproducible analysis, iterating on code, sorting results, and narrowing down a shortlist of molecules to test. This is a major advantage, as experiments cost time, materials, and specialized expertise.

Yet no coding assistant or chatbot can single-handedly solve the core challenge of drug discovery: proving that a molecule works, is safe, and can be manufactured and distributed sustainably. Models can generate inaccurate code, propose fragile interpretations, or reflect gaps present in the data and literature they are trained on. That is why every result must be verified, documented, and subjected to standard scientific research benchmarks.

Furthermore, in the case of antimicrobials, even a compound that proves active in the lab may fail to overcome the hurdles separating preliminary testing from an actual treatment. Factors such as target selectivity, toxicity, resistance that microorganisms might develop, delivery mechanisms, and clinical trials all come into play. AI can reduce the number of blind paths, but it cannot bypass these stages.

An operational use of AI in biology

The experience of the de la Fuente lab demonstrates a less flashy, yet more practical application of generative systems in research: not a machine announcing a discovery in place of scientists, but an added layer in the infrastructure of day-to-day work. Writing code to analyze sequences, modify a pipeline, compare hypotheses, and prepare an initial exploration of datasets are tasks where the time saved can be reinvested into experimental design and the validation of results.

For OpenAI, the case also represents a testing ground for Codex outside commercial software. Scientific programming is often fragmented, built around specialized tools, and carried out by teams that must balance biological and computational expertise. An assistant capable of working with code can streamline this transition, provided it is used in environments where reviews, version control, and checks are part of the process.

The next step for this line of research will not be a chat-generated answer, but the verdict of experiments. The identified sequences will need to be compared against measurable biological data and the constraints of drug development. If a portion of the candidates passes these filters, AI's contribution will have been at its most useful: not promising clinical shortcuts, but making the search for molecules that biology has already conceived broader and more systematic.

Sources