10 papers
Flex-: A Multi-Stream World-Action Model with Compute Flexibility
Ge Yan, Jinghao Liu, Yuzhi Fan +4
World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for t…
SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos
Jaehyeon Son, Junhyun Kim, Kyle Kam +7
Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an…
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Anthony Liang, Yigit Korkmaz, Jiahui Zhang +14
General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effecti…
Open-World Task and Motion Planning via Vision-Language Model Generated Constraints
Nishanth Kumar, William Shen, Fabio Ramos +4
Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve comp…
PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies
Jesse Zhang, Marius Memmel, Kevin Kim +6
Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-lev…
Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective
Xuning Yang, Clemens Eppner, Jonathan Tremblay +3
Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluati…