Share:
Peer-Reviewed Publication
HGG Adv2021;2(3)July 1, 2021Journal Article

A data-driven architecture using natural language processing to improve phenotyping efficiency and accelerate genetic diagnoses of rare disorders.

Jignesh R Parikh1,2, Casie A Genetti3,2, Asli Aykanat3, Catherine A Brownstein3, Klaus Schmitz-Abe3, Morgan Danowski3, Andrew Quitadomo3,4, Jill A Madden3, Calum Yacoubian5, Richard Gain5, Tessa Williams5, Mary Meskell5, Andrew Brown5, Alison Frith5, Shira Rockowitz3,4, Piotr Sliz3,4, Pankaj B Agrawal3,6, Thomas Defay7, Paul McDonagh7,8, John Reynders7,9, Sebastien Lefebvre7, Alan H Beggs3
1J Square Labs, LLC, Natick, MA 01760, USA.
2These authors contributed equally to this work.
3The Manton Center for Orphan Disease Research, Division of Genetics and Genomics, Boston Children's Hospital, Harvard Medical School, Boston, MA 02115, USA.
4Computational Health Informatics Program, Boston Children's Hospital, Harvard Medical School, Boston, MA 02115, USA.
5Clinithink, Ltd., London N1 6DR, UK.
6Division of Newborn Medicine, Boston Children's Hospital, Harvard Medical School, Boston, MA 02115, USA.
7Alexion Pharmaceuticals, Inc., Boston, MA 02210, USA.
8Present address: Sema4, Stamford, CT 06902, USA.
9Present address: Latent Strategies, LLC, Newton, MA 02465, USA.

Abstract

Effective genetic diagnosis requires the correlation of genetic variant data with detailed phenotypic information. However, manual encoding of clinical data into machine-readable forms is laborious and subject to observer bias. Natural language processing (NLP) of electronic health records has great potential to enhance reproducibility at scale but suffers from idiosyncrasies in physician notes an…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.