The mathematical proofs announced by OpenAI are producing a far from marginal side effect: the academic community wants to understand precisely where the knowledge that led the models to those results came from. Following the objections raised in recent days by New York University professor Tristan Buckmaster, Andreas Thom is now demanding public and verifiable explanations regarding the use of interactions with ChatGPT.
Thom, a mathematician known for his work on non-sofic groups, argues that the responses received from OpenAI fail to clarify the point that truly matters to him. It is not enough for him to know whether his conversations with the chatbot can be directly retrieved by the system during reasoning: he wants to know whether that content, or that of other researchers, might have entered the data used to train or fine-tune the models. In statements published on Mastodon, he spoke of a lack of transparency and accused the company of conduct he considers unethical and dishonest.
The case concerns one of the ten mathematical results showcased prominently by OpenAI last month. One of these fell squarely within Thom’s area of research. The company acknowledged that the work was largely built on prior contributions by Thom and Gábor Kun, but the initial version of the documentation failed to adequately attribute that prior work. The omission had been criticized by other mathematicians, prompting OpenAI to later amend the published text.
The question is not just about the authorship of an idea
An incomplete bibliographic citation can be corrected; reconstructing the actual origin of a generative model’s capabilities is far more difficult. Thom explained that he began questioning his own conversations with ChatGPT following the dispute raised by Buckmaster over the potential use, by OpenAI models, of the work the professor had developed using Codex.
In Thom’s case, the perplexity also stems from the level of familiarity shown by OpenAI with specific methods from his line of research. According to the mathematician, the techniques employed were neither the most straightforward nor those considered, at the time, the most promising paths to reach a solution. While this alone does not prove an improper transfer of content, it makes the question of which sources the system might have learned those connections from highly relevant.
Non-sofic groups are infinite mathematical objects that, in very simplified terms, cannot be approximated by finite structures. It is a highly specialized field, far removed from the vast general-audience corpora that fuel the most visible side of generative AI. It is precisely this distance that makes the case significant: when a system produces advances in niche disciplines, with technical literature and results still circulating within academia, the provenance of the information cannot be treated as a mere operational detail.
Thom contacted Sébastien Bubeck, a researcher at OpenAI, and Mark Sellke, a statistician at Harvard, asking whether his interactions with ChatGPT could have been part of the training data or otherwise accessible to the reasoning process. The response received, according to the mathematician, reportedly addressed the second possibility—direct access to the conversations—without resolving the first. Hence his dissatisfaction and his call for greater clarity.
The fine line between using a product and contributing data
The issue is delicate because it touches on what has become a widespread practice. Researchers, programmers, students, and professionals use chatbots and coding assistants to discuss hypotheses, test ideas, rephrase notes, and verify steps. In a field like mathematics, a working session can contain unpublished insights, abandoned leads, draft proofs, or references to materials that have not yet passed through the peer-review process.
The terms and ways in which these interactions are handled therefore take on a new relevance. The issue is not limited to the risk of another user receiving a similar response: it also concerns the possibility that an as-yet confidential contribution could be absorbed, in any form, into the product improvement cycle. For researchers, this distinction affects scientific priority, credit attribution, and the ability to publish original results without having already exposed them to a commercial platform.
OpenAI must also confront an inherent challenge of large language models. These systems do not typically offer granular traceability regarding the origin of every generated output. They can cite sources if instructed to do so, but they are not designed to provide an evidentiary record capable of definitively separating what derives from public literature, licensed material, human demonstrations, or interactions that occurred while using the services.
This technical and organizational limitation does not constitute proof of the allegations raised by Thom. However, it makes it harder to provide convincing responses to specific requests for exclusion or verification. A general reassurance regarding the lack of direct accessibility to a conversation, for instance, does not automatically resolve the question of whether the content was used in other training, evaluation, or improvement processes.
A test for AI looking to enter research
OpenAI framed its mathematical achievements as evidence of the growing utility of models in frontier research. It is a significant promise: tools capable of exploring conjectures, identifying unexpected connections, or assisting in constructing proofs could transform work across many laboratories. But the closer these outputs come to original scientific contributions, the more essential it becomes to clarify who provided the ideas, what data was available, and how the sources were credited.
The discussion therefore does not end with the attribution concerning non-sofic groups. It concerns the relationship between companies operating AI infrastructure and scientific communities that produce knowledge in environments built on verification, publication, and author recognition. A model can accelerate research, but the acceptance of its findings will also depend on the ability to rule out that it is recycling, without proper credit, work already carried out by others.
For OpenAI, the case comes at a time when announcements of mathematical breakthroughs are heightening scrutiny of internal procedures. The attribution correction in the material concerning the result linked to Thom and Kun shows that community pressure can prompt rectifications. However, the current demand is broader: not a post-publication footnote, but clear guidelines regarding user conversation policies and the boundaries between service usage, training, and the generation of new conclusions.
Public evidence has not yet emerged showing that Thom's conversations contributed to the result announced by OpenAI. It is equally true, however, that the mathematician takes issue with the very lack of tools and responses that would make verification possible. In the coming weeks, the company's credibility in this arena will depend on the quality of the explanations it provides—not only to the two researchers involved in the controversy, but to anyone considering whether to entrust a chatbot with unpublished work.



