Cross-sectional stratification of fall-related vulnerability in Parkinson's disease through multimodal fusion of speech, gait, facial expression, and upper-limb movement.
Cross-sectional stratification of fall-related vulnerability in Parkinson's disease through multimodal fusion of speech, gait, facial expression, and upper-limb movement.
Where did the research take place?
The study site has not been established. Author addresses may differ from where the research occurred.
GR · Author affiliation · country only
Electrical and Computer Engineering Department, Hellenic Mediterranean University, Heraklion, Greece.Location evidence
Bern, CH · Author affiliation
ARTORG Centre for Biomedical Engineering Research, University of Bern, Bern, Switzerland.Location evidence
Ioánnina, GR · Author affiliation
Unit of Medical Technology and Intelligent Information Systems, University of Ioannina, Ioannina, Greece.Location evidence
A plain-language reading has not been prepared for this paper yet.
Original abstract
Falls constitute a significant cause of morbidity in Parkinson's disease (PD), but current approaches to cross-sectional assessment of fall-related vulnerability frequently depend on unimodal mobility metrics with limited ability to characterize its multidimensional manifestations. This study presents a multimodal deep learning (DL) architecture that incorporates facial expressions, voice, gait dynamics, and upper-limb movements to classify cross-sectional fall-related vulnerability in PD using current postural-instability status and fall history. Data were gathered from 147 patients, and modality-specific models were systematically pruned to reduce model size, inference latency, and computational requirements while maintaining classification performance, providing an initial assessment of computational efficiency within the experimental environment. Various fusion strategies, including late, intermediate, attention-based, transformer-based, and large-language-model-assisted methods, were assessed using leave-one-subject-out cross-validation. The large-language-model-assisted analysis was exploratory and was based on a single recorded API run per participant-modality configuration. Pruning resulted in considerable efficiency improvements without a notable decrease in accuracy, with gait demonstrating the highest robustness. Fusion performance varied across methods and modality combinations. Attention-based fusion with face and gait achieved an observed macro-f1 of 91.67%. An exploratory large-language-model-assisted face-upper-limb configuration produced an observed macro-f1 of 93.75% in a single recorded run; because repeated-run stability and exact model-version reproducibility were not established, this result should not be interpreted as evidence of reproducible superiority. Transformer fusion produced an observed macro-f1 of 85.02% when all four modalities were included. Significantly, speech and particularly face expressivity provided supplementary signals to conventional gait and upper-limb assessments, underscoring the potential of underexplored axial characteristics as complementary digital markers associated with fall-related vulnerability. These findings support further study of multimodal AI for cross-sectional classification of fall-related vulnerability in PD. Prospective fall prediction, external generalizability, and performance on intended devices remain to be evaluated in independent future studies.