6 papers
Playful Agentic Robot Learning
Junyi Zhang, Jiaxin Ge, Hanjun Yoo +17
Current agentic robot systems can write executable Code-as-Policy programs, observe feedback, and revise behavior across multiple attempts, but they remain largely task-driven: reu…
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
Shaofeng Yin, Jiaxin Ge, Zora Zhiruo Wang +6
Vision-as-inverse-graphics, the concept of reconstructing images into editable programs, remains challenging for Vision-Language Models (VLMs), which inherently lack fine-grained s…
VisualMimic: Visual Humanoid Loco-Manipulation via Motion Tracking and Generation
Shaofeng Yin, Yanjie Ze, Hong-Xing Yu +2
Humanoid loco-manipulation in unstructured environments demands tight integration of egocentric perception and whole-body control. However, existing approaches either depend on ext…
RLVR-World: Training World Models with Reinforcement Learning
Jialong Wu, Shaofeng Yin, Ningya Feng +1
World models predict state transitions in response to actions and are increasingly developed across diverse modalities. However, standard training objectives such as maximum likeli…
Learning to Construct Knowledge through Sparse Reference Selection with Reinforcement Learning
Shao-An Yin
The rapid expansion of scientific literature makes it increasingly difficult to acquire new knowledge, particularly in specialized domains where reasoning is complex, full-text acc…
Trajectory World Models for Heterogeneous Environments
Shaofeng Yin, Jialong Wu, Siqiao Huang +4
Heterogeneity in sensors and actuators across environments poses a significant challenge to building large-scale pre-trained world models on top of this low-dimensional sensor info…