16 papers
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
Mingyu Liu, Zheng Huang, Xiaoyi Lin +6
Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-…
ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
Muzhi Zhu, Hao Zhong, Canyu Zhao +9
Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…
Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
Hao Zhong, Muzhi Zhu, Shenyan Zeng +8
Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for…
Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?
Liyang Li, Muzhi Zhu, Zhiyue Zhao +5
Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passiv…
FLaG: Fine-Grained Latent Grouping for Hallucination Detection
Wentao Ye, Liyao Li, Zhiqing Xiao +6
Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global uncertainty score. In this wor…
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Zheng Huang, Mingyu Liu, Xiaoyi Lin +9
Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic fo…