4 papers
Shot-Aware Frame Sampling for Video Understanding
Mengyu Zhao, Di Fu, Yongyu Xie +4
Video frame sampling is essential for efficient long-video understanding with Vision-Language Models (VLMs), since dense inputs are costly and often exceed context limits. Yet when…
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models
Jiaxin Zhang, Junjun Jiang, Haijie Li +3
Multimodal Large Language Models (MLLMs) demonstrate exceptional semantic reasoning but struggle with 3D spatial perception when restricted to pure RGB inputs. Despite leveraging i…
PVI: Plug-in Visual Injection for Vision-Language-Action Models
Zezhou Zhang, Songxin Zhang, Xiao Xiong +8
VLA architectures that pair a pretrained VLM with a flow-matching action expert have emerged as a strong paradigm for language-conditioned manipulation. Yet the VLM, optimized for…
Resolving Task Objective Conflicts in Unified Model via Task-Aware Mixture-of-Experts
Jiaxing Zhang, Hao Tang
Unified multimodal large language models (MLLMs) based on end-to-end autoregressive (AR) transformers effectively integrate both understanding and generation tasks within a single…