In recent years, artificial intelligence (AI) has been increasingly embraced in healthcare settings with promises to enhance patient care and alleviate clinical workloads. However, experiences from frontline medical professionals and recent studies highlight significant challenges in validating the real-world effectiveness and safety of many AI tools once deployed in hospitals.

Kayla Secrest, a newly qualified physician at a hospital in Michigan, encountered these issues firsthand with an AI algorithm designed to detect sepsis by scanning patient records every 15 minutes and issuing alerts. Intended to expedite early identification of this potentially fatal condition, the system instead produced frequent false alarms, leading clinicians to become desensitized and eventually disregard the notifications. Secrest later learned that this widely implemented tool had not undergone extensive testing under actual hospital conditions before rollout.

Such instances underscore a broader concern within healthcare: AI products frequently enter clinical use with limited independent oversight or robust evidence demonstrating meaningful clinical benefit. Jess Morley, an associate research scientist at Yale Digital Ethics Centre and former AI advisor to the UK Department of Health and Social Care, notes that developers tend to prioritize statistical accuracy over proof of improved patient outcomes. “There is shockingly poor evidence that AI actually makes an impact on patient outcomes,” Morley said, emphasizing the distinction between laboratory validation and real-world clinical efficacy.

Despite these concerns, AI adoption in healthcare has accelerated faster than in other sectors, driven in part by administrative efficiencies, such as automating medical note-taking. Kimberly Powell, head of healthcare at Nvidia, highlighted how AI helped reduce documentation burdens. Additionally, companies like OpenAI report substantial decreases in time spent on administrative tasks and improvements in diagnostic accuracy in some settings, though independent validation remains sparse.

In areas like diagnostics and medical imaging, AI’s benefits are better documented. A 2024 study cited by Eric Topol of Scripps Research found that colonoscopies assisted by AI detected significantly more polyps than traditional procedures. Still, Topol warns that healthcare providers are often captivated by the excitement surrounding generative AI tools—such as those powered by ChatGPT—which can summarize large bodies of medical literature but lack conclusive evidence proving superior patient outcomes. He advocates for rigorous assessment of these tools in real-world clinical environments and notes that leading AI developers have shown limited interest in conducting such post-deployment evaluations.

Transparency issues further complicate efforts to assess AI effectiveness. Andrew Wong, director of clinical AI research at the University of Utah who worked on improving the Michigan sepsis algorithm, highlights that companies rarely disclose critical information about model training, data sources, or validation processes. Without access to this data, individual hospitals bear the burden of evaluating whether tools perform effectively across diverse patient populations and clinical settings.

Researchers also point to the prevalence of “spin” in AI research, with many studies recommending clinical adoption without external validation. Regulatory frameworks differ internationally but often rely on manufacturer-submitted reports lacking raw data, limiting independent scrutiny. In the United States, the Food and Drug Administration (FDA) authorized 258 AI medical devices in 2025, yet only about a quarter had undergone clinical trials similar to those required for pharmaceuticals.

Experts suggest that traditional regulatory methods may be ill-suited for AI technologies, which continuously evolve after deployment. Lawrence Tallon, chief executive of the UK’s Medicines and Healthcare products Regulatory Agency, advocates for conditional approvals paired with ongoing real-world monitoring to ensure safety and efficacy. Similar pilot programs are underway in the US to evaluate AI tools for chronic diseases based on evidence collected during routine use.

Nonetheless, concerns persist that regulatory advancement lags behind rapid AI development. John Paul Jeans, a UK anaesthetist and entrepreneur, warns that AI capabilities are doubling every four months, challenging the relevance of any regulation crafted for current tools. Additionally, resource constraints in smaller or rural hospitals may limit their ability to assess AI for potential biases or unintended consequences, risking exacerbation of existing health disparities.

Bias in AI systems has already been documented. Ziad Obermeyer of the University of California, Berkeley found that an AI-based risk stratification tool used for allocating healthcare resources perpetuated inequities by relying on past expenditure data, neglecting underserved groups who historically received lower care. Given the vast scale of AI adoption in patient care decisions, Obermeyer cautions that many stakeholders remain unaware of the profound influence these algorithms hold.

Gathering comprehensive data to monitor AI performance remains a technical and logistical challenge. Amy Abernethy, a former senior FDA official, emphasizes the need for advanced AI-powered data infrastructure to enable granular evaluation. Meanwhile, privacy protections and siloed hospital records hinder regulators’ ability to independently verify manufacturers’ claims. Initiatives such as the company Dandelion Health have emerged to compile large, diverse datasets to facilitate more reliable AI testing.

Academic collaborations are focusing on shifting attention from theoretical validation toward measuring actual clinical impact. Duke University researcher Sammy Choufani El Fassi stresses the importance of evaluating how AI tools affect patient outcomes in practice, moving beyond the question of whether they simply “work” in controlled environments.

As AI technology continues to integrate into healthcare, experts agree that establishing rigorous real-world evidence and transparent oversight is critical to balancing innovation with patient safety and equitable care delivery.