About the journal
Browse
Collections
Multimedia collections
Authors & reviewers
MVSA: End-to-End Multi-View Sampling and Alignment for Multi-Modal Test-Time Adaptation
Xiao-Long Yin , De-Chuan Zhan , Yuan Jiang
In real-world environments, data distributions often shift due to changes in time, space, or modality, undermining the generalization and decision stability of machine learning models in open scenarios. Test-time adaptation (TTA) enables dynamic adaptation to unknown test distributions. As an important extension, multi-modal test-time adaptation (MM-TTA) introduces cross-modal collaboration and intra-modal variation, leading to higher complexity than single-modal settings. However, existing methods face two fundamental problems: (1) insufficient intra-modal information utilization during sampling due to random or fixed-interval sampling; (2) biased information from sampled views, where single-frame views are susceptible to noise and irrelevant factors. To address these issues, this paper proposes an end-to-end multi-view sampling and alignment (MVSA) framework. To address insufficient information utilization, we introduce an information density-guided sampling strategy that evaluates information density using normalized gradient energy and preferentially selects the most informative frames along with their temporally aligned audio segments. To address biased information interference, we introduce a cross-view alignment loss that enforces representational consistency between views via Jensen-Shannon divergence minimization. Experimental results on Kinetics50-C and VGGSound-C datasets demonstrate that MVSA achieves significant accuracy improvements over mainstream TTA methods such as Tent, MMT, and READ under modality reliability bias scenarios. Ablation studies and parameter sensitivity analysis further confirm the effectiveness of the proposed strategies in enhancing model robustness. MVSA enhances the adaptability of multi-modal systems in distributionally shifting environments requiring real-time responses, such as autonomous driving and remote healthcare.
Test-time adaptation / Multi-model Learning / Multi-view Learning
Higher Education Press 2026
/
| 〈 |
|
〉 |