Share:
Peer-Reviewed Publication
JMIR Form Res2025;9e78082.October 24, 2025Journal Article

Preprocessing Large-Scale Conversational Datasets: A Framework and Its Application to Behavioral Health Transcripts.

Paz Mor Naim1, Shiri Sadeh-Sharvit2,3, Samuel Jefroykin2, Eddie Silber2,3, Dennis P Morrison4, Ariel Goldstein1,5,6
1Department of Cognitive and Brain Sciences, Hebrew University of Jerusalem, Mount Scopus, Jerusalem, 9190500, Israel, 972 025882888.
2Eleos Health, Waltham, MA, United States.
3Palo Alto University, Palo Alto, CA, United States.
4Morrison Consulting, Bloomington, IN, United States.
5Business School, Hebrew University of Jerusalem, Jerusalem, Israel.
6Department of Psychology, Azrieli Israel Center for Addiction and Mental Health (Azrieli ICAMH), Hebrew University of Jerusalem, Jerusalem, Israel.

Abstract

BACKGROUND: The rise of artificial intelligence and accessible audio equipment has led to a proliferation of recorded conversation transcripts datasets across various fields. However, automatic mass recording and transcription often produce noisy, unstructured data that contain unintended recordings such as hallway conversations, media (eg, TV, radio), or transcription inaccuracies as speaker misa…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.