The hardest question about artificial intelligence is not always “how smart is it?”. Sometimes it is much simpler: where does what it knows come from?

A controversy reported by The Verge has brought this question into one of the fields most sensitive to the authorship of ideas: mathematics. Several researchers have called on OpenAI for greater transparency regarding the potential use of non-public interactions and materials in model training. OpenAI has disputed the notion that it used specific private sessions in the manner suggested by critics.

Beyond this specific case, the issue is poised to become massive. If a model can contribute to a proof, a discovery or a new conjecture, data provenance is no longer just a copyright issue: it becomes part of the scientific method.

Science also works thanks to traceability

A scientific paper is not valuable solely because it reaches a correct conclusion. It must explain how it got there, what data it used, what prior work it consulted, and what limitations remain. This structure allows other researchers to verify, replicate, and challenge the result.

Generative models partially break this chain. They can produce a plausible idea without being able to reliably trace which fragments of their training contributed to that response.

The more powerful they become, the more problematic this opacity becomes.

A private conversation can contain research

For many people, chatting with an assistant is an everyday interaction. For a scientist, it can contain unpublished hypotheses, proof attempts, code snippets, or preliminary results. In other words, material with intellectual value well before it becomes a paper.

This changes the perception of consent. Stating that certain data can be used in de-identified form may be acceptable for improving a general-purpose product, but it becomes far more delicate when that information represents original scientific work.

The problem cannot be resolved with a disclaimer

Companies can improve disclosures, offer controls, and distinguish more clearly between data used and not used for training. But the real challenge is technical: being able to prove, when necessary, that an output does not stem from material that should not have contributed to it.

It is a form of provenance—that is, traceability of origin. In the world of AI-generated content, we demand it for images and video. In research, it could become even more important.

Scientific AI will need scientific rules

The point is not to prevent researchers from using advanced models. On the contrary, tools capable of exploring hypotheses and checking vast spaces of possibilities can accelerate human work. But the more AI participates in the production of knowledge, the closer it must come to the transparency standards of science.

A correct answer that is impossible to reconstruct may be useful. A scientific discovery, however, requires something more: knowing which assumptions, data, and ideas contributed to the result.

The next AI battle may therefore not only be about who builds the best model. It could be about who manages to build the model whose genealogy of ideas we can truly know.