When can we say AI made a scientific discovery?
2026-09-28 · MIT Technology Review
When Can We Say AI Made a Scientific Discovery?
Anthropic's Claim and the Backlash
Last Wednesday, Anthropic announced that it had launched a molecular biology lab earlier this year. In this setup, Claude agents read and conjecture about difficult biology problems, while human scientists run experiments based on their reports. Anthropic claimed that a system of 950 agents, after running for 21 hours, made its first discovery: flagging a repeating pattern surrounding a known enzyme that had not been catalogued before. The announcement described this pattern as "reminiscent" of the developments leading to CRISPR.
However, these claims have angered many biologists. A viral post by a biologist, later endorsed by the chair and CEO of Eli Lilly, argued that "finding a weird cluster of genes and repeats is often the easy part. The hard part, and where the real discoveries come from, is figuring out what the system actually does." According to critics, the AI merely assisted with laboratory grunt work, which does not constitute a true scientific discovery.
Prior Claims and Data Concerns
The controversy deepened when Mario Rodríguez Mestre, a biologist at the University of Copenhagen, stated that his team had already discovered this specific pattern, as reported by the New York Times. Mestre, who regularly used Claude in his work, wondered whether Anthropic's team had learned from his conversations with the AI. Anthropic denies this, but Mestre has stopped using Claude regardless.
The Narrative Conflict Between AI and Science
A core issue is that AI companies are not presenting their systems merely as tools for scientists, like microscopes or supercomputers. Instead, they insist that AI systems are making discoveries themselves. To many, this approach contradicts how science actually works, as new knowledge typically emerges from collaboration and an expanding arsenal of tools.
This framing also makes people more skeptical of genuine progress. Whittling 200,000 candidates down to a few worth exploring is a legitimate scientific feat. The fact that a general-purpose chatbot could do this is notable, even with human guidance. However, when the standard becomes whether Claude itself made a discovery, the nuance is lost, reducing the achievement to a binary debate: breakthrough or bust.
Shifting the Goalposts
Judging AI solely on whether it makes a discovery also tempts critics to shift the goalposts even when it achieves a win. Earlier this month, OpenAI announced its agents had solved a million-dollar math problem. But weeks later, AI skeptics widely shared an article questioning whether that specific math problem truly mattered. The piece didn't argue the solution was wrong, but rather that it wasn't the result mathematicians cared most about. Combined with accusations from a mathematician that the models used his work without credit, the public was left thinking OpenAI either cheated or that the solution was unimportant.
A Call for Higher Standards
Lucas Harrington, the biologist who critiqued Anthropic, suggested that AI companies should "set the bar high now, so that when an AI actually discovers a fundamentally new biological mechanism, everyone appreciates how big a deal it is." But as OpenAI's Sam Altman and Anthropic's Dario Amodei race to one-up each other, raising the bar for scientific breakthroughs might be the last thing on their minds.