Share:
Peer-Reviewed Publication
Med Image Anal2023;90102939.December 1, 2023Journal Article

UNesT: Local spatial representation learning with hierarchical transformer for efficient medical segmentation.

Xin Yu1, Qi Yang1, Yinchi Zhou1, Leon Y Cai2, Riqiang Gao3, Ho Hin Lee1, Thomas Li2, Shunxing Bao4, Zhoubing Xu5, Thomas A Lasko6, Richard G Abramson7, Zizhao Zhang8, Yuankai Huo9, Bennett A Landman10, Yucheng Tang11
1Department of Computer Science, Vanderbilt University, Nashville TN, 37212, USA.
2Department of Biomedical Engineering, Vanderbilt University, Nashville, TN, 37212, USA.
3Department of Computer Science, Vanderbilt University, Nashville TN, 37212, USA; Digital Technology and Innovation, Siemens Healthineers, Princeton, NJ, 08540, USA.
4Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, 37212, USA.
5Digital Technology and Innovation, Siemens Healthineers, Princeton, NJ, 08540, USA.
6Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, 37212, USA.
7Department of Biomedical Engineering, Vanderbilt University, Nashville, TN, 37212, USA; Annalise-AI, Pty, Ltd, USA.
8Google Cloud AI, USA.
9Department of Computer Science, Vanderbilt University, Nashville TN, 37212, USA; Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, 37212, USA.
10Department of Computer Science, Vanderbilt University, Nashville TN, 37212, USA; Department of Biomedical Engineering, Vanderbilt University, Nashville, TN, 37212, USA; Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, 37212, USA; Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN, 37212, USA.
11Department of Electrical and Computer Engineering, Vanderbilt University, Nashville, TN, 37212, USA; Nvidia Corporation, USA. Electronic address: yuchengt@nvidia.com.

Abstract

Transformer-based models, capable of learning better global dependencies, have recently demonstrated exceptional representation learning capabilities in computer vision and medical image analysis. Transformer reformats the image into separate patches and realizes global communication via the self-attention mechanism. However, positional information between patches is hard to preserve in such 1D se…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.