From the 1 of 14 linked papers with an AI index.
14 papers
Perfect Demo Makes Poor Teacher: Learning Robust Alignment from Critical Motion Segments
Mingyu Liu, Zeju Li, Jiuhe Shu +4
The paper shows that smooth robot demonstrations can miss critical alignment moments, and proposes slowing down and resampling key motion segments, plus a spatio‑temporal feature c…
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
Mingyu Liu, Zheng Huang, Xiaoyi Lin +6
Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-…
ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
Muzhi Zhu, Hao Zhong, Canyu Zhao +9
Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…
NoTVLA: Semantics-Preserving Robot Adaptation via Narrative Action Interfaces
Zheng Huang, Mingyu Liu, Xiaoyi Lin +9
Vision-Language-Action (VLA) models represent a pivotal advance in embodied intelligence, yet they confront critical barriers to real-world deployment, most notably catastrophic fo…
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
Mingyu Liu, Jiuhe Shu, Hui Chen +6
A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and decision making. However, existing meth…
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
Yiduo Jia, Muzhi Zhu, Hao Zhong +7
To extend the reinforcement learning post-training paradigm to omni-modal models for concurrently bolstering video-audio understanding and collaborative reasoning, we propose OmniJ…