5 papers
Vision Inference Former: Sustaining Visual Consistency in Multimodal Large Language Models
Xinpeng Dong, Min Zhang, Kairong Han +3
In recent years, multimodal large language models (MLLMs) have achieved remarkable progress, primarily attributed to effective paradigms for integrating visual and textual informat…
iVGR: Internalizing Visually Grounded Reasoning for MLLMs with Reinforcement Learning
Chang-Bin Zhang, Yujie Zhong, Qiang Zhang +1
While visually grounded Chain-of-Thought (CoT) has emerged as a promising paradigm to enhance fine-grained perception in multimodal large language models (MLLMs), its efficacy duri…
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
Zeyu Chen, Jie Li, Kai Han
Multimodal representation alignment is pivotal for large language models and robotics. Traditional methods are often hindered by cross-modal information discrepancies and data scar…
Surgical Post-Training: Proximal On-Policy Distillation for Reasoning with Knowledge Retention
Wenye Lin, Kai Han
Injecting new reasoning knowledge into Large Language Models (LLMs) via post-training often induces catastrophic forgetting. Recent studies emphasize the importance of on-policy da…
VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse
Ying Nie, Kai Han, Hongguang Li +5
The rapid scaling of Large Language Models (LLMs) has achieved remarkable performance, but it also leads to prohibitive memory costs. Existing parameter-efficient approaches such a…