Share:
Peer-Reviewed Publication
J Med Internet Res2020;22(7):e18055.July 15, 2020Journal Article

Exploring the Privacy-Preserving Properties of Word Embeddings: Algorithmic Validation Study.

Mohamed Abdalla1,2,3, Moustafa Abdalla4,5,6, Graeme Hirst1,2, Frank Rudzicz1,2,7,8
1Department of Computer Science, University of Toronto, Toronto, ON, Canada.
2The Vector Institute for Artificial Intelligence, Toronto, ON, Canada.
3Institute for Clinical Evaluative Sciences, Toronto, ON, Canada.
4Deptartment of Statistics, Computational Statistics & Machine Learning Group, University of Oxford, Oxford, United Kingdom.
5Wellcome Centre for Human Genetics, Nuffield Dept of Medicine, University of Oxford, Oxford, United Kingdom.
6Harvard Medical School, Boston, MA, United States.
7International Centre for Surgical Safety, Li Ka Shing Knowledge Institute, St Michael's Hospital, Toronto, ON, Canada.
8Surgical Safety Technologies Inc, Toronto, ON, Canada.

Abstract

BACKGROUND: Word embeddings are dense numeric vectors used to represent language in neural networks. Until recently, there had been no publicly released embeddings trained on clinical data. Our work is the first to study the privacy implications of releasing these models. OBJECTIVE: This paper aims to demonstrate that traditional word embeddings created on clinical corpora that have been deidenti…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.