Longitudinal voice biomarker trajectory modelling for Parkinson's disease severity: domain-adaptive transfer learning on mPower real-world smartphone data.
Longitudinal voice biomarker trajectory modelling for Parkinson's disease severity: domain-adaptive transfer learning on mPower real-world smartphone data.
Where did the research take place?
The study site has not been established. Author addresses may differ from where the research occurred.
Chennai, IN · Author affiliation
School of Computer Science and Engineering, Vellore Institute of Technology, Chennai Campus, Chennai, India.Location evidence
A plain-language reading has not been prepared for this paper yet.
Original abstract
INTRODUCTION: Speech and voice changes affect up to 90% of people with Parkinson's disease (PD), a progressive neurodegenerative disorder affecting approximately 10 million people worldwide. Although continuous monitoring of disease severity is clinically important, most existing voice-based computational approaches focus on binary PD-versus-control classification and do not model longitudinal symptom progression. To address this gap, we propose a domain-adaptive transformer model, DAT-PD, for predicting continuous PD severity trajectories from real-world smartphone voice recordings. METHODS: DAT-PD was developed using the public mPower dataset, comprising 58,247 voice recordings from 5,800 participants. The proposed pipeline included noise-aware acoustic preprocessing, extraction of extended Geneva Minimalistic Acoustic Parameter Set (eGeMAPS) features, a domain-adaptive attention mechanism to reduce cross-device and cross-environment variability and a longitudinal trajectory decoder. The model was trained, inferred, and evaluated using continuous MDS-UPDRS Part-II scores as the sole prediction target. Confounder-aware domain adaptation was incorporated to address demographic imbalance, including the age gap between the PD cohort and healthy controls. Robustness was further evaluated under harsh acoustic conditions with signal-to-noise ratios as low as 0 dB. RESULTS: On the held-out test set, DAT-PD achieved a mean absolute error (MAE) of 2.74 MDS-UPDRS units (95% CI: 2.44-3.01), root mean squared error (RMSE) of 3.61 (95% CI: 3.18-4.04), and R² of 0.93 (95% CI: 0.91-0.95), outperforming six state-of-the-art baseline models. eGeMAPS features substantially outperformed MFCC-only representations, reducing MAE from 5.21 to 2.74. SHAP-based explainability identified MFCC-2, Shimmer (APQ5) and Jitter as the most influential longitudinal voice biomarkers. DISCUSSION: The superior performance of DAT-PD suggests that domain-adaptive longitudinal modeling can effectively capture clinically meaningful voice-based severity trajectories in PD. The advantage of eGeMAPS over MFCC-only features is likely due to its ability to represent phonatory and prosodic characteristics relevant to PD dysarthria, including F0 dynamics, loudness contour, shimmer, jitter and spectral flux. By maintaining robustness under noisy real-world acoustic conditions, DAT-PD supports unsupervised home-based monitoring using standard smartphones. These findings align with the Bridge2AI-Voice research agenda and position DAT-PD as a clinically implementable, non-invasive tool for continuous PD severity assessment.