PDF
(2368KB)
Abstract
Recent advances in distributed computing enable privacy-preserving aggregation and artificial intelligence (AI)-driven analysis of multimodal medical data, empowering real-time distributed healthcare applications and large-scale disease detection systems. Inspired by its remarkable performance, we proposed to discover medical knowledge embedded in big data with a large multimodal model, offering intelligent disease analytics via this promising AI route. In this paper, we proposed a distributed multimodal representation learning structure for zero-shot medical image classification. Specifically, we simultaneously adopted the implicit knowledge extracted from a large multimodal model built on images, and the explicit knowledge extracted from the medical knowledge graph built on textual records. Facing the inconsistent alignment in latent space constructed by multimodal data, a cross-modal alignment strategy was proposed to adjust intra- and inter-modal representations for convinced learning. Experiments on several public datasets proved that the proposed framework could improve the accuracy of zero-shot medical image classification, achieving robust and accurate disease analytical results.
Keywords
distributed signal processing
/
contrastive vision-language model
/
zero-shot image classification
/
medical knowledge graph
Cite this article
Download citation ▾
Peng Lu, Xinfu Liu, Yuting Zhou.
A Distributed Cross-Modal Representation Structure for Zero-Shot Medical Image Classification.
Journal of Beijing Institute of Technology, 2026, 35 (4) : 411-426 DOI:10.15918/j.jbit1004-0579.2025.062