Share:
Peer-Reviewed Publication
Front Artif Intell2025;81691499.January 1, 2025Journal Article

Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe.

Erin Palm1,2,3, Astrit Manikantan1, Herprit Mahal1,4,5, Srikanth Subramanya Belwadi1, Mark E Pepin1,6
1Suki AI, Redwood City, CA, United States.
2Santa Clara Valley Medical Center, San Jose, CA, United States.
3Division of General Surgery, Department of Surgery, Stanford University School of Medicine, Stanford, CA, United States.
4Hippocratic AI, Palo Alto, CA, United States.
5The Permanente Medical Group, Oakland, CA, United States.
6Stanford Cardiovascular Institute, Stanford University School of Medicine, Stanford, CA, United States.

Abstract

BACKGROUND: Generative artificial intelligence (AI) tools are increasingly being used as "ambient scribes" to generate drafts for clinical notes from patient encounters. Despite rapid adoption, few studies have systematically evaluated the quality of AI-generated documentation against physician standards using validated frameworks. OBJECTIVE: This study aimed to compare the quality of large langu…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.