collaborators

16 papers

cs.CV2026

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert

Mingyu Liu, Zheng Huang, Xiaoyi Lin +6

Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-…

cs.CV2026

ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Muzhi Zhu, Hao Zhong, Canyu Zhao +9

Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…

cs.CV2026

Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching

Hao Zhong, Muzhi Zhu, Shenyan Zeng +8

Wide-baseline matching (WBM) requires integrating geometric understanding, viewpoint changes, fine-grained perception, and occlusion reasoning, making it a challenging testbed for…

cs.CV2026

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration?

Liyang Li, Muzhi Zhu, Zhiyue Zhao +5

Humans can reproduce the viewpoint specified by a target image through active head and body motion, yet spatial intelligence in foundation models has largely been studied as passiv…

cs.LG2026

FLaG: Fine-Grained Latent Grouping for Hallucination Detection

Wentao Ye, Liyao Li, Zhiqing Xiao +6

Hallucinations in large language models (LLMs) arise from heterogeneous failure mechanisms, making reliable detection difficult for any single global uncertainty score. In this wor…

cs.RO2026

NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces

Zheng Huang, Mingyu Liu, Xiaoyi Lin +9

Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic fo…