activity
20242026
collaborators

10 papers

cs.RO2026

Flex-: A Multi-Stream World-Action Model with Compute Flexibility

Ge Yan, Jinghao Liu, Yuzhi Fan +4

World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for t…

cs.RO2026

SeeTraceAct: Visibility-Aware Latent Planning from Cross-Embodiment Demonstration Videos

Jaehyeon Son, Junhyun Kim, Kyle Kam +7

Vision-language-action models (VLAs) are promising general-purpose robot policies, but adapting them to new tasks typically requires costly task-specific teleoperation data. As an…

cs.RO2026

Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons

Anthony Liang, Yigit Korkmaz, Jiahui Zhang +14

General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effecti…

cs.RO2026

Open-World Task and Motion Planning via Vision-Language Model Generated Constraints

Nishanth Kumar, William Shen, Fabio Ramos +4

Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve comp…

cs.RO2025

PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies

Jesse Zhang, Marius Memmel, Kevin Kim +6

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-lev…

cs.RO2025

Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective

Xuning Yang, Clemens Eppner, Jonathan Tremblay +3

Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluati…