4 papers
Understanding and Mitigating the Video-Action Generalization Gap via Temporal Ratio
Utkarsh A. Mishra, Yongxin Chen, Danfei Xu +3
Generative video foundation models exhibit strong compositional priors, yet world-action models (WAMs) and video-action models (VAMs) often lose these priors after finetuning on ro…
ForceBand: Learning Forceful Manipulation with sEMG
Botao He, Zhi Wang, Linna Kuang +8
Human demonstrations are a scalable data source for learning robot manipulation policies. However, common sources of human demonstration data, such as motion-capture trajectories a…
Adversarial Game-Theoretic Algorithm for Dexterous Grasp Synthesis
Yu Chen, Botao He, Yuemin Mao +7
For many complex tasks, multi-finger robot hands are poised to revolutionize how we interact with the world, but reliably grasping objects remains a significant challenge. We focus…
Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models
Simeng Han, Howard Dai, Stephen Xia +7
Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based…