A new peer-reviewed analysis published on August 19 in PLOS Digital Health concludes that only three of 1,357 AI- and machine-learning-enabled medical devices cleared or approved by the U.S. Food and Drug Administration have been evaluated for patient-centered outcomes, exposing a wide gap between regulatory authorization and clinical validation.
Only 2.5% Are Linked To Registered Clinical Trials
The peer-reviewed study, led by Rawan Abulibdeh and colleagues at institutions including the University of Toronto and MIT Critical Data, systematically mapped every FDA-cleared AI/ML device through December 5, 2025 against ClinicalTrials.gov and PubMed. Fewer than 3% (n=34) of the 1,357 devices were linked to registered prospective trials, only 12 posted results, 12 had peer-reviewed publications, and just three (0.2%) evaluated patient-centered outcomes such as mortality, readmissions or quality of life. Sixty-two percent of the available studies were observational, and 68% of the registered trials were done exclusively in the United States.

Case Studies Highlight The Risk
The paper cites external-validation failures that have already drawn scrutiny. The Epic Sepsis Model missed 67% of sepsis cases while triggering alerts in 18% of hospitalizations, and IBM Watson for Oncology showed concordance with expert recommendations as low as 12% for gastric cancer in China and 33% in Denmark. The authors argue that these gaps are partly structural: 510(k) clearance requires only substantial equivalence to a predicate device, so vendors have no obligation to demonstrate patient-level benefit.
A Staged Framework To Close The Gap
To rebuild trust in medical AI, the authors propose a staged validation framework: retrospective validation on diverse datasets, then prospective workflow-embedded studies of at least 500 patients with pre-specified safety and usability endpoints, and finally multi-center trials of 2,000 or more participants that measure patient outcomes across relevant subgroups. The team also called for greater transparency around proprietary studies that never reach public registries.
The findings land as regulators, hospitals and vendors race to deploy diagnostic AI in radiology, cardiology and critical care. For related coverage, see Quantum X Labs' AI decoder work, Anthropic's Claude protein-binder results, and Cerebras' CS-4 inference launch.
Reporting based on coverage from PLOS Digital Health, News-Medical, Healio and MedicalXpress.
