A clinical prediction model can appear convincing in a scientific paper and yet remain difficult to verify, reproduce, or adapt to a hospital other than the one where it was developed. Behind this gap often lies an element less visible than the data and final metrics: the analytical code used to prepare the information, train or configure the model, and measure its performance.

An analysis published in Nature Medicine captures just how uncommon code sharing still is in research on clinical prediction models. The research team examined 3,967 open access papers that cited the TRIPOD guidelines or their TRIPOD+AI update. Only 482 papers, or 12.2%, contained a code-sharing statement. This figure does not merely measure an editorial preference: it concerns the concrete ability to subject findings to independent scrutiny before a model intended for diagnosis or prognosis enters clinical workflows.

The review was carried out using a large language model-assisted pipeline designed to tackle a vast body of literature: identifying relevant studies, locating any repository links, and then examining the retrieved material. The authors evaluated the repositories against 14 predefined characteristics related to reproducibility. The outcome, beyond the scarce presence of code-sharing statements, is a pronounced inconsistency in the quality and documentation of the available software.

Code is part of the method, not an optional attachment

In clinical prediction models, code does not simply amount to the algorithm alone. It can encompass the rules used to select patients, handle missing data or outliers, transform variables, define outcomes, and split samples for development and validation. It also includes modeling choices, parameter tuning, and the procedures used to calculate accuracy, calibration, and other performance measures.

Two teams can therefore claim to have applied the same methodological approach and yet obtain different results if one of these steps remains implicit or cannot be reconstructed. In biomedical research, this is particularly relevant: clinical data is often protected and cannot always be freely shared, but code can still make the path leading from the dataset to the published result transparent. It does not eliminate data access hurdles, nor does it alone allow every analysis to be replicated. However, it does make the technical decisions that affected the outcome assessable.

Code availability also makes it possible to understand whether a model is genuinely transferable. A system built in one healthcare setting may encounter different nomenclature, inconsistent documentation practices, missing tests, or a population with different clinical characteristics elsewhere. To verify generalisability, knowing the value of a metric reported in a paper is not enough: repeatable, sufficiently documented procedures are required to perform external validation.

From repository presence to reusability

The study avoids a common oversimplification in the open science debate: uploading files to a repository does not automatically enable other researchers to use them. Indeed, an evaluation across 14 aspects reveals highly variable practices. The core issue lies in the information accompanying the code, the definition of software dependencies, and the organisation required to run the analysis.

A repository lacking clear instructions, without specifying required libraries or without a structure that clarifies the sequence of steps, may be formally accessible but scarcely operational. For an external team, reconstructing the working environment then becomes a lengthy and uncertain task; for a healthcare facility looking to test a model on local data, it can turn into a substantial roadblock.

This distinguishes documentary transparency from actual reproducibility. Journal and funder policies have helped make data and code availability statements more common, but the new analysis suggests that the clinical modeling field requires more specific expectations. Stating where a repository is located is helpful; explaining how to run it, with which dependencies, and following which execution workflow is what makes the material verifiable.

Variability across journals and countries

The rate of sharing is not uniform. The authors observe marked differences depending on the journal and the country. This snapshot does not justify turning the data into a ranking of national or editorial research quality: different disciplines, publication rules, infrastructures, institutional constraints, and the attitudes of individual teams all come into play. However, it signals that the adoption of reproducible practices is not progressing at the same pace across the sector.

For the digital health industry and those evaluating the procurement or adoption of clinical software, this fragmentation has practical implications. Predictive models are increasingly proposed as support for diagnostic or prognostic decision-making. If upstream research does not allow for a thorough examination of analytical procedures, it increases the work needed to verify the robustness, limitations, and behavior of the model outside the laboratory that created it.

Furthermore, the issue of reproducibility must be separated from that of clinical clearance. Well-documented code alone does not prove that a model is safe or appropriate for patient use. External validation, assessment of the operational context, clinical supervision, and, where applicable, regulatory pathways remain indispensable. But without transparency regarding the analytical process, even these assessments start from a less solid foundation.

The next step is making declared commitments verifiable

The review was designed to contribute to the development of TRIPOD-Code, a guideline dedicated to code availability and reproducibility in prediction model research. TRIPOD and TRIPOD+AI have already provided a benchmark for improving the completeness and clarity of reporting; extending this to software addresses a gap that has emerged alongside the spread of increasingly complex computational methods.

Future guidelines could make it more explicit what should accompany a repository: instructions for use, declared dependencies, an executable structure, and sufficient information to distinguish the code actually used in the analysis from partial or experimental scripts. This is not necessarily about mandating the indiscriminate publication of every single component. In healthcare, legitimate constraints exist around privacy, security, licensing, and intellectual property. Rather, the goal is to establish clear criteria explaining what is available, what cannot be, and to what extent the work can be scrutinized.

The 12.2% identified by the review therefore depicts a sector where code sharing remains the exception rather than the rule. For a field that aims to influence clinical decisions, the gap between a published result and one that others can examine is not merely an operational detail. It is one of the key steps on which trust in the model, its independent evaluation, and the ability to bring it into hospitals on more verifiable grounds depend.

Sources