5 papers
Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models
Yiwen Liang, Hui Chen, Yizhe Xiong +7
Test-time adaptation (TTA) has been widely explored in single-label recognition, effectively mitigating distribution shifts, especially when combined with vision-language models. H…
A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation
Siyuan Ma, Liang He, Mengying Zhu +14
Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable co…
Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs
Chuangxin Zhao, Canran Xiao, Siyuan Ma +5
Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
Yiwen Liang, Hui Chen, Yizhe Xiong +7
Vision-language models (VLMs) exhibit remarkable zero-shot capabilities but struggle with distribution shifts in downstream tasks when labeled data is unavailable, which has motiva…
Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning
Mengyao Lyu, Yan Li, Huasong Zhong +5
The hypothesis that pretrained large language models (LLMs) necessitate only minimal supervision during the fine-tuning (SFT) stage (Zhou et al., 2024) has been substantiated by re…