2 papers
cs.AI2026
DICA: Dual-Indicator Guided Contrastive Alignment in Multimodal Large Language Models
Hao Yang, Jin Wang, Xuejie Zhang
Human visual reasoning typically follows a coarse-to-fine attention process, starting from global scene understanding and gradually focusing on question-relevant regions. However,…
cs.CV2026
Visual Saliency Steering Distillation for Multimodal Chain-of-Thought Reasoning
Hao Yang, Jin Wang, Xuejie Zhang
Multimodal chain-of-thought (CoT) reasoning integrates visual and textual cues through step-by-step inference. In small models with limited token budgets, modality-interaction fusi…