Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
Xinpeng Dong, Min Zhang, Kairong Han +3
In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual informat…
cs.CV2026
iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Chang-Bin Zhang, Yujie Zhong, Qiang Zhang +1
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy duri…
cs.CV2026
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
Zeyu Chen, Jie Li, Kai Han
Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scar…