2 papers
cs.CV2026
SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs
Zi-Hao Bo, Yaqian Li, Anzhou Hou +6
Mixture-of-Experts (MoE) has become a prevalent backbone for large vision-language models (VLMs), yet how modality-specific signals should guide expert routing remains under-explor…
cs.CV2026
Step-CoT: Stepwise Visual Chain-of-Thought for Medical Visual Question Answering
Lin Fan, Yafei Ou, Zhipeng Deng +8
Chain-of-thought (CoT) reasoning has advanced medical visual question answering (VQA), yet most existing CoT rationales are free-form and fail to capture the structured reasoning p…