4 papers
RISE: Self-Improving Robot Policy with Compositional World Model
Jiazhi Yang, Kunyang Lin, Jinwei Li +10
Despite the sustained scaling on model capacity and data acquisition, Vision-Language-Action (VLA) models remain brittle in contact-rich and dynamic manipulation tasks, where minor…
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
Qinglong Yang, Haoming Li, Haotian Zhao +4
Mobile GUI agents are becoming critical tools to improve user experience on smart devices, with multimodal large language models (MLLMs) emerging as the dominant paradigms in this…
Agility Meets Stability: Versatile Humanoid Control with Heterogeneous Data
Yixuan Pan, Ruoyi Qiao, Li Chen +8
Humanoid robots are envisioned to perform a wide range of tasks in human-centered environments, requiring controllers that combine agility with robust balance. Recent advances in l…
Detect Anything 3D in the Wild
Hanxue Zhang, Haoran Jiang, Qingsong Yao +6
Despite the success of deep learning in close-set 3D object detection, existing approaches struggle with zero-shot generalization to novel objects and camera configurations. We int…