Share:
Peer-Reviewed Publication
JMIR AI2025;4e72153.December 4, 2025Journal Article

Clinical Large Language Model Evaluation by Expert Review (CLEVER): Framework Development and Validation.

Veysel Kocaman1, Mustafa Aytuğ Kaya2, Andrei Marian Feier1, David Talby1
1John Snow Labs Inc, 16192 Coastal Highway, Lewes, DE, 19958, United States, +1 (302) 786-5227.
2Computational Sciences and Informatics (CSI), George Mason University, Fairfax, VA, United States.

Abstract

BACKGROUND: The proliferation of both general purpose and health care-specific large language models (LLMs) has intensified the challenge of effectively evaluating and comparing them. Data contamination plagues the validity of public benchmarks, self-preference distorts LLM-as-a-judge approaches, and there is a gap between the tasks used to test models and those used in clinical practice. OBJECTI…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.