4 papers
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
Jiaxing Liu, Zexi Zhang, Xiaoyan Li +3
Vision-Language Navigation (VLN) presents a unique challenge for Large Vision-Language Models (VLMs) due to their inherent architectural mismatch: VLMs are primarily pretrained on…
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
Zhenhao Zhang, Jiaxin Liu, Ye Shi +1
Planning physically feasible dexterous hand manipulation is a central challenge in robotic manipulation and Embodied AI. Prior work typically relies on object-centric cues or preci…
ViC-Bench: Benchmarking Visual-Interleaved Chain-of-Thought Capability in MLLMs with Free-Style Intermediate State Representations
Xuecheng Wu, Jiaxing Liu, Danlei Huang +8
Visual-Interleaved Chain-of-Thought (VI-CoT) enables Multi-modal Large Language Models (MLLMs) to continually update their understanding and decision space based on step-wise inter…
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
Chen Chen, ZeYang Hu, Fengjiao Chen +6
Multimodal Large Languages models have been progressing from uni-modal understanding toward unifying visual, audio and language modalities, collectively termed omni models. However…