7 papers
FetchMan: Learning Visual Humanoid Loco-Manipulation Policies from Simulated Experiences
Omar Rayyan, Zhi Li, Max Argus +4
Visual loco-manipulation policies that can generalize to novel scenes and objects have long been a goal of robotics research. However, today's data-hungry algorithms make collectin…
MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction
Jianing Zhang, Chenhao Zheng, Yajun Yang +10
Motion forecasting is central to visual intelligence: agents must anticipate how objects will move in order to plan actions, reason about physical interactions, and synthesize real…
MolmoSpaces: A Large-Scale Open Ecosystem for Robot Navigation and Manipulation
Yejin Kim, Wilbert Pumacay, Omar Rayyan +23
Deploying robots at scale demands robustness to the long tail of everyday situations. The countless variations in scene layout, object geometry, and task specifications that charac…
cVLA: Towards Efficient Camera-Space VLAs
Max Argus, Jelena Bratulic, Houman Masnavi +4
Vision-Language-Action (VLA) models offer a compelling framework for tackling complex robotic manipulation tasks, but they are often expensive to train. In this paper, we propose a…
Efficient Learning of Object Placement with Intra-Category Transfer
Adrian Röfer, Russell Buchanan, Max Argus +2
Efficient learning from demonstration for long-horizon tasks remains an open challenge in robotics. While significant effort has been directed toward learning trajectories, a recen…
When and How Does CLIP Enable Domain and Compositional Generalization?
Elias Kempf, Simon Schrodi, Max Argus +1
The remarkable generalization performance of contrastive vision-language models like CLIP is often attributed to the diversity of their training distributions. However, key questio…