most citedCosmos-Transfer1: Conditional World Generation with Adaptive Multimodal Control

2 citations · 2 across the 2 of their papers we have counts for

collaborators
Showing cs.ROShow all

6 papers · 1 filter

cs.RO2025

PEEK: Guiding and Minimal Image Representations for Zero-Shot Generalization of Robot Manipulation Policies

Jesse Zhang, Marius Memmel, Kevin Kim +6

Robotic manipulation policies often fail to generalize because they must simultaneously learn where to attend, what actions to take, and how to execute them. We argue that high-lev…

cs.RO2025

Robot Policy Evaluation for Sim-to-Real Transfer: A Benchmarking Perspective

Xuning Yang, Clemens Eppner, Jonathan Tremblay +3

Current vision-based robotics simulation benchmarks have significantly advanced robotic manipulation research. However, robotics is fundamentally a real-world problem, and evaluati…

cs.RO2025

GraspGen: A Diffusion-based Framework for 6-DOF Grasping with On-Generator Training

Adithyavairavan Murali, Balakumar Sundaralingam, Yu-Wei Chao +7

Grasping is a fundamental robot skill, yet despite significant research advancements, learning-based 6-DOF grasping approaches are still not turnkey and struggle to generalize acro…

cs.RO2025

HAMSTER: Hierarchical Action Models For Open-World Robot Manipulation

Yi Li, Yuquan Deng, Jesse Zhang +9

Large foundation models have shown strong open-world generalization to complex problems in vision and language, but similar levels of generalization have yet to be achieved in robo…

cs.RO2024

Aim My Robot: Precision Local Navigation to Any Object

Xiangyun Meng, Xuning Yang, Sanghun Jung +4

Existing navigation systems mostly consider "success" when the robot reaches within 1m radius to a goal. This precision is insufficient for emerging applications where the robot ne…

cs.RO2024

Open-World Task and Motion Planning via Vision-Language Model Generated Constraints

Nishanth Kumar, William Shen, Fabio Ramos +4

Foundation models like Vision-Language Models (VLMs) excel at common sense vision and language tasks such as visual question answering. However, they cannot yet directly solve comp…