MVSA: End-to-End Multi-View Sampling and Alignment for Multi-Modal Test-Time Adaptation

Xiao-Long Yin , De-Chuan Zhan , Yuan Jiang

Front. Comput. Sci. ››

PDF (850KB)
Front. Comput. Sci. ›› DOI: 10.1007/s11704-026-60998-9
REVIEW ARTICLE
MVSA: End-to-End Multi-View Sampling and Alignment for Multi-Modal Test-Time Adaptation
Author information +
History +
PDF (850KB)

Abstract

In real-world environments, data distributions often shift due to changes in time, space, or modality, undermining the generalization and decision stability of machine learning models in open scenarios. Test-time adaptation (TTA) enables dynamic adaptation to unknown test distributions. As an important extension, multi-modal test-time adaptation (MM-TTA) introduces cross-modal collaboration and intra-modal variation, leading to higher complexity than single-modal settings. However, existing methods face two fundamental problems: (1) insufficient intra-modal information utilization during sampling due to random or fixed-interval sampling; (2) biased information from sampled views, where single-frame views are susceptible to noise and irrelevant factors. To address these issues, this paper proposes an end-to-end multi-view sampling and alignment (MVSA) framework. To address insufficient information utilization, we introduce an information density-guided sampling strategy that evaluates information density using normalized gradient energy and preferentially selects the most informative frames along with their temporally aligned audio segments. To address biased information interference, we introduce a cross-view alignment loss that enforces representational consistency between views via Jensen-Shannon divergence minimization. Experimental results on Kinetics50-C and VGGSound-C datasets demonstrate that MVSA achieves significant accuracy improvements over mainstream TTA methods such as Tent, MMT, and READ under modality reliability bias scenarios. Ablation studies and parameter sensitivity analysis further confirm the effectiveness of the proposed strategies in enhancing model robustness. MVSA enhances the adaptability of multi-modal systems in distributionally shifting environments requiring real-time responses, such as autonomous driving and remote healthcare.

Keywords

Test-time adaptation / Multi-model Learning / Multi-view Learning

Cite this article

Download citation ▾
Xiao-Long Yin, De-Chuan Zhan, Yuan Jiang. MVSA: End-to-End Multi-View Sampling and Alignment for Multi-Modal Test-Time Adaptation. Front. Comput. Sci. DOI:10.1007/s11704-026-60998-9

登录浏览全文

4963

注册一个新账户 忘记密码

References

Rights & permissions

Higher Education Press 2026

PDF (850KB)

96

Accesses

0

Citation

Detail

Sections
Recommended

/