Share:
Peer-Reviewed Publication
PLOS Digit Health2026;5(7):e0001506.July 1, 2026Journal Article

Vision-language models for human motion understanding: Lessons from stroke rehabilitation.

Victor Li1, Naveenraj Kamalakannan1,2, Avinash Parnandi3, Heidi Schambra4,5, Carlos Fernandez-Granda1,6
1Center for Data Science, New York University, New York, New York, United States of America.
2Tandon School of Engineering, New York University, Brooklyn, New York, United States of America.
3VitalConnect, San Jose, California, United States of America.
4Department of Neurology, NYU Grossman School of Medicine, New York, New York, United States of America.
5Department of Rehabilitation Medicine, NYU Grossman School of Medicine, New York, New York, United States of America.
6Courant Institute of Mathematical Sciences, New York University, New York, New York, United States of America.

Abstract

Vision-language models (VLMs) have demonstrated remarkable performance across a wide range of computer-vision tasks, sparking interest in their potential for digital health applications. Here, we apply VLMs to two fundamental challenges in data-driven stroke rehabilitation: automatic quantification of rehabilitation dose and impairment from videos. We formulate these problems as motion-identificat…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.