Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Seeing is Understanding: Unlocking Causal Attention into Modality-Mutual Attention for Multimodal LLMs
Wei-Yao Wang, Zhao Wang, Helen Suzuki +1
Recent Multimodal Large Language Models (MLLMs) have demonstrated significant progress in perceiving and reasoning over multimodal inquiries, ushering in a new research era for fou…
cs.CV2025
MSGCoOp: Multiple Semantic-Guided Context Optimization for Few-Shot Learning
Zhaolong Wang, Tongfeng Sun, Mingzheng Du +1
Vision-language pre-trained models (VLMs) such as CLIP have demonstrated remarkable zero-shot generalization, and prompt learning has emerged as an efficient alternative to full fi…