Share:
Peer-Reviewed Publication
BMC Musculoskelet Disord2026;27(1)April 9, 2026Journal Article

Benchmarking of large language models to determine candidacy for spine surgery and comparison with conventional machine learning.

Christopher Rice1, Frederik Abel2, Alemu Sisay Nigru3, Andrej Rusakov4, Nikolai S Nikolaev5,6, Yurii Baiun7, Alessandra Falk4, Titenkov Yuriy5,6, Frolov Valerii5,6, Aleksey Kopeev5,6, Raphael Mourad8
1Plastic Surgery, Loma Linda University Health, 11234 Anderson St, Loma Linda, CA, 92354, USA.
2Department of Diagnostic and Interventional Radiology, University Hospital Zurich, University of Zurich, Zurich, 8091, Switzerland.
3Department of Information Engineering, University of Brescia, Brescia, 25127, Italy.
4Remedy Logic, 1177 Avenue of the Americas, 5Th Floor, New York, NY, 10036, USA.
5Orthopedics and Arthroplasty, Federal Center for Traumatology, Orthopedics and Arthroplasty, Federal State Budgetary Institution, Ministry of Health of the Russian Federation, Cheboksary, Russia.
6Federal State Budgetary Educational Institution of Higher Education, Chuvash State University Named After I.N. Ulyanov, Cheboksary, Russia.
7Center of Neurosurgery, Kyiv Regional Hospital, Kiev, Ukraine.
8University of Toulouse, CNRS, UPS, Toulouse, 31062, France. raphael.mourad@utoulouse.fr.

Abstract

BACKGROUND: Symptomatic lumbar spinal stenosis (LSS) is a disabling condition with a substantial economic impact. The determination of surgical candidacy for LSS relies on a subjective assessment of multiple clinical and imaging factors, leading to variability in recommendations. Artificial intelligence (AI), including traditional machine learning (ML) and emerging large language models (LLMs), ho…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.