activity
20232026
collaborators

5 papers

cs.CV2026

Towards Purified Multi-Label Test-Time Adaptation of Vision-Language Models

Yiwen Liang, Hui Chen, Yizhe Xiong +7

Test-time adaptation (TTA) has been widely explored in single-label recognition, effectively mitigating distribution shifts, especially when combined with vision-language models. H…

eess.IV2026

A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation

Siyuan Ma, Liang He, Mengying Zhu +14

Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable co…

cs.CV2026

Attention-Spectrum Regularization for Replay-Free Continual Multimodal LLMs

Chuangxin Zhao, Canran Xiao, Siyuan Ma +5

Multimodal large language models (MLLMs) are increasingly required to adapt to non-stationary streams of visual domains, question types, and user instructions, yet continual fine-t…

cs.CV2025

Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations

Yiwen Liang, Hui Chen, Yizhe Xiong +7

Vision-language models (VLMs) exhibit remarkable zero-shot capabilities but struggle with distribution shifts in downstream tasks when labeled data is unavailable, which has motiva…

cs.CV2025

Cream of the Crop: Harvesting Rich, Scalable and Transferable Multi-Modal Data for Instruction Fine-Tuning

Mengyao Lyu, Yan Li, Huasong Zhong +5

The hypothesis that pretrained large language models (LLMs) necessitate only minimal supervision during the fine-tuning (SFT) stage (Zhou et al., 2024) has been substantiated by re…