17 citations · 27 across the 13 of their papers we have counts for
6 papers · 1 filter
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
Shaofei Cai, Zhancun Mu, Anji Liu +1
We aim to develop a goal specification method that is semantically clear, spatially sensitive, domain-agnostic, and intuitive for human users to guide agent interactions in 3D envi…
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
Shaofei Cai, Bowei Zhang, Zihao Wang +4
Developing agents that can follow multimodal instructions remains a fundamental challenge in robotics and AI. Although large-scale pre-training on unlabeled datasets (no language i…
MineStudio: A Streamlined Package for Minecraft AI Agent Development
Shaofei Cai, Zhancun Mu, Kaichen He +4
Minecraft's complexity and diversity as an open world make it a perfect environment to test if agents can learn, adapt, and tackle a variety of unscripted tasks. However, the devel…
Preference Goal Tuning: Post-Training as Latent Control for Frozen Policies
Guangyu Zhao, Kewei Lian, Haoxuan Ru +8
Goal-conditioned policies enable decision-making models to execute diverse behaviors based on specified goals, yet their downstream performance is often highly sensitive to the cho…
JARVIS-1: Open-World Multi-task Agents with Memory-Augmented Multimodal Language Models
Zihao Wang, Shaofei Cai, Anji Liu +9
Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle cer…
GROOT: Learning to Follow Instructions by Watching Gameplay Videos
Shaofei Cai, Bowei Zhang, Zihao Wang +3
We study the problem of building a controller that can follow open-ended instructions in open-world environments. We propose to follow reference videos as instructions, which offer…