4 papers
LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies
Zheng Lu, Haoran Liao, Wanqi Zhong +9
Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic repr…
Token Predictors Are Not Planners: Building Physically Grounded Causal Reasoners
Zheng Lu, Mingqi Gao, Qinlei Xie +8
Current benchmarks for embodied vision-language planning often favor linguistic next-token prediction over physically grounded next-state reasoning. This rewards models that mimic…
Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts
Yunxin Li, Shenyuan Jiang, Baotian Hu +5
Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial comp…
A Comprehensive Evaluation of GPT-4V on Knowledge-Intensive Visual Question Answering
Yunxin Li, Longyue Wang, Baotian Hu +5
The emergence of multimodal large models (MLMs) has significantly advanced the field of visual understanding, offering remarkable capabilities in the realm of visual question answe…