Speech Foundation Models for Parkinson’s Disease Detection: A Layer-Wise Comparison with Handcrafted Acoustics Across Two Cohorts
Speech Foundation Models for Parkinson’s Disease Detection: A Layer-Wise Comparison with Handcrafted Acoustics Across Two Cohorts
Where did the research take place?
The study site has not been established. Author addresses may differ from where the research occurred.
Publication status: preprint
A plain-language reading has not been prepared for this paper yet.
Original abstract
Abstract Foundation models have transformed speech processing, but their value for clinical voice analysis, particularly in Parkinson’s disease (PD), remains largely unexplored. In this study, we compare layer-wise representations from several foundation-model families (Whisper, Wav2Vec 2.0, OmniASR, NVIDIA Nemotron, and Qwen3-ASR) with a handcrafted acoustic baseline using sentence reading and sustained vowels from independent cohorts (PC-GITA and NeuroVoz). On sentence reading, the best foundation-model configurations achieved segment-level F1-scores of approximately 0.88 in both cohorts, above the handcrafted acoustic baseline. Their advantage on sustained vowels was smaller and more cohort-dependent, with overlapping uncertainty in participant-level comparisons. Soft-voting aggregation generally improved participant-level discrimination. Intermediate-to-late encoder layers most often yielded the strongest representations, although optimal depth varied by architecture and task. Stratified analyses revealed cohort disparities, F1-scores were higher for male speakers in PC-GITA and for senior speakers in NeuroVoz, so foundation-model representations alone do not remove demographic performance gaps. Overall, these results support speech foundation models as promising feature extractors for PD detection and motivate validation in larger clinical cohorts.