Scientists and industry experts are grappling with emerging questions about authorship, attribution, and data ownership as artificial intelligence increasingly participates in scientific discovery, sometimes raising concerns over potential overlaps with unpublished research.

Mario Rodríguez Mestre, a computational biologist at the University of Copenhagen, recently encountered an unexpected situation when a major AI company, Anthropic, announced that its AI agent Claude had autonomously identified a new enzyme and related molecules—work closely aligned with Mestre’s own unpublished research. Mestre had been employing Claude to assist with coding and drafting a manuscript on a class of enzymes called reverse transcriptases, found in bacteria and used to detect viruses. Anthropic’s claims notably cited Mestre and colleagues in their references, yet the company described Claude’s discovery as independent.

Mestre expressed surprise that the source of the similar findings was not a competing academic lab but rather a prominent AI firm. He pointed out that the systems Claude highlighted included the same genes from the same bacteriophages that his team had identified. While Mestre acknowledged the overlap could be coincidental, he stressed he had no direct evidence that his unpublished work influenced Anthropic’s results and has not communicated with the company. He called for clearer standards regarding provenance and attribution in AI-driven research, emphasizing the need for transparency about the data and guidance an AI model uses when producing results. In response, Mestre is pivoting away from Claude, pursuing the development of other AI tools tailored for biological research.

Anthropic responded that it was unaware of any previously published studies describing the biological findings attributed to Claude and highlighted that the AI’s key contribution was recognizing that the reverse transcriptase was part of a larger system. An Anthropic spokesperson further stated that Claude had not been trained on any user transcripts and that their molecular biology team did not have access to such data.

This incident is part of wider ongoing tension surrounding AI’s role in scientific advancement. A similar case involved OpenAI’s work near a solution to the Navier-Stokes theorem, where mathematician Tristan Buckmaster noted significant parallels between the company’s output and his own team’s approaches, which had also utilized AI tools in their research.

The increasing use of AI in science poses challenges, especially when unpublished or confidential information—such as grant proposals or peer review reports—is uploaded to large language models (LLMs) during assessment processes. Major research funders like the Wellcome Trust in the UK have instituted bans on sharing sensitive materials with AI platforms, yet concerns remain that such data could still be exposed, potentially enabling AI companies to "snoop and scoop" novel ideas before researchers formally publish them.

At present, autonomous AI breakthroughs have not played a significant role in prize recognitions such as the Nobel Prizes. Officials from the Royal Swedish Academy of Sciences noted the awards typically honor human-led research prior to widespread AI integration, and their rules require recipients to be individual persons or groups of people, not machines.

As AI continues to be built on vast troves of human knowledge, experts emphasize that recognition and accountability must remain grounded in the contributions of human researchers. The ongoing debate highlights the need for comprehensive frameworks to manage attribution and data integrity in an era where AI increasingly assists scientific work.