Share:
Peer-Reviewed Publication
NPJ Digit Med2025;8(1):274.May 13, 2025Journal Article

A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation.

Elham Asgari1,2, Nina Montaña-Brown3, Magda Dubois3, Saleh Khalil3, Jasmine Balloch3, Joshua Au Yeung3, Dominic Pimenta3
1Tortus AI, London, UK. asgelham@gmail.com.
2Guy's and St Thomas NHS Trust, London, UK. asgelham@gmail.com.
3Tortus AI, London, UK.

Abstract

Integrating large language models (LLMs) into healthcare can enhance workflow efficiency and patient care by automating tasks such as summarising consultations. However, the fidelity between LLM outputs and ground truth information is vital to prevent miscommunication that could lead to compromise in patient safety. We propose a framework comprising (1) an error taxonomy for classifying LLM output…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.