San Francisco — OpenAI recently announced a breakthrough by claiming its latest artificial intelligence technology had solved a million-dollar mathematical problem, specifically the Navier-Stokes existence and smoothness challenge that has perplexed researchers for decades. This development highlighted AI’s escalating capabilities in tackling complex scientific issues, but it also sparked concerns over the transparency and provenance of knowledge used in such achievements.
Tristan Buckmaster, a professor at New York University who has conducted research on related AI applications, raised questions about whether OpenAI’s system had incorporated elements of his work. OpenAI responded initially in a social media post, acknowledging the remote possibility that de-identified data from users’ interactions with their products, including potentially Buckmaster’s, might have indirectly contributed to improving their models. However, the company later issued a categorical denial, asserting it was impossible for Buckmaster’s recent prompts using Codex, an AI coding system, to have influenced their latest results.
This exchange underscores ongoing uncertainties about how AI models, which learn from vast datasets and billions of parameters, integrate individual users’ inputs. Experts note that when users enter information into AI platforms—whether public or private—these data can potentially be used to train future iterations, raising questions about intellectual ownership and data privacy.
Oren Etzioni, founding CEO of the Allen Institute for Artificial Intelligence and a University of Washington professor, emphasized the nuanced challenge of distinguishing an idea’s contribution within the immense pools of data processed by modern AI. “You cannot assume that the idea did not have an impact,” he said, noting that unique or rare insights—such as those involved in deep mathematical problems—might be especially susceptible to being absorbed by AI in ways difficult to trace.
Other researchers confirm that AI companies and organizations often have agreements to exclude sensitive academic or business data from training datasets. Aneesh Muppidi, a Stanford University AI researcher, explained that his institution uses a customized OpenAI platform version restricted to university affiliates, with contractual safeguards preventing data from being incorporated into broader AI training.
The intricate nature of these AI systems, which can deploy thousands of coordinated “agents” working over extended periods to solve problems, complicates efforts to interpret how specific inputs influence their outputs. Sanjeev Arora, a Princeton computer science professor, described the opacity of these processes, noting the difficulty in pinning down the exact pathways through which AI arrives at complex answers.
While many users’ interactions generally contribute to commonly known information that is widespread across data sources, rare scientific problems like Navier-Stokes pose a different challenge, increasing the likelihood that private research could indirectly affect AI training.
Despite these concerns, some experts suggest this issue may become less relevant as AI capabilities rapidly outpace human expertise. Arora predicted that within a year, AI systems will perform at levels rendering direct human competition obsolete in many scientific fields.
The broader debate about data use, attribution, and transparency in AI development persists, reflecting the complex interplay between innovation, intellectual property, and ethical considerations as AI technologies continue to evolve.
