Quantifying Data Leakage in Multimodal Clinical Augmentation: A Decomposition Framework and a Leakage-Free Evaluation Protocol for Parkinson's Disease Classification
Quantifying Data Leakage in Multimodal Clinical Augmentation: A Decomposition Framework and a Leakage-Free Evaluation Protocol for Parkinson's Disease Classification
Publication status: preprint
A plain-language reading has not been prepared for this paper yet.
Original abstract
Abstract Generative augmentation is widely used to correct class imbalance in multimodal clinical machine learning, yet classification gains attributed to it are rarely checked for evaluation leakage, a gap with direct consequences for clinical trust in AI-assisted diagnosis. We ask whether the order in which heterogeneous modalities are embedded and rebalanced with synthetic samples affects downstream Parkinson's disease (PD) classification, and we introduce a decomposition framework that separates any apparent gain into a baseline term, a genuine augmentation effect, and a split-after-augmentation leakage term. Using a multimodal PD dataset combining wearable sensor time-series, clinical scales, and a non-motor symptom questionnaire, we compare Embed-Then-Balance (ETB) and Balance-Then-Embed (BTE) pipelines under a leakage-free 5-fold × 5-seed cross-validation protocol, and assess synthetic-sample fidelity with a classifier two-sample test, the Wasserstein distance, and the distance-to-closest-record. Classification performance is statistically indistinguishable across pipelines and no augmentation method improves on the unaugmented baseline (best AUC 0.838-0.842 across configurations). The two orderings diverge sharply in generative fidelity: raw-space balancing produces off-manifold samples trivially separable from real data (C2ST AUC = 1.000), whereas latent-space balancing produces samples statistically indistinguishable from real ones (C2ST AUC = 0.458). Gains commonly attributed to augmentation are, to a large extent, an artifact of evaluation leakage, contributing up to +0.17 AUC in multiclass settings. For clinical machine learning, this means that leakage-free evaluation is a prerequisite, not an optional refinement, before an augmentation strategy can be trusted for deployment in a diagnostic-support pipeline. The decomposition framework and protocol we release apply to any clinical machine learning study combining representation learning with synthetic-data augmentation.