4 papers
Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models
Li Wenjie, Yash Jangir, Ignacy Stepka +3
Action verbs describe not only the physical outcomes of actions, but also how those actions are performed. Yet action representations in vision-language-action models (VLAs) are ty…
MotionForesight: Re-purposing Video Models for Future 3D Scene-Flow Prediction
Homanga Bharadhwaj, Yash Jangir
Humans can infer how objects are likely to move from passive observation: a cup may be lifted, a drawer may slide, and a lid may rotate shut. Such predictions expose the physical c…
RobotArena : Scalable Robot Benchmarking via Real-to-Sim Translation
Yash Jangir, Yidi Zhang, Pang-Chi Lo +7
The pursuit of robot generalists, agents capable of performing diverse tasks across diverse environments, demands rigorous and scalable evaluation. Yet real-world testing of robot…
Dexterous Manipulation Policies from RGB Human Videos via 3D Hand-Object Trajectory Reconstruction
Hongyi Chen, Tony Dong, Tiancheng Wu +7
Multi-finger robotic hand manipulation and grasping are challenging due to the high-dimensional action space and the difficulty of acquiring large-scale training data. Existing app…