From the 1 of 23 linked papers with an AI index.
23 papers
: Reactive Real-time Flow Policies
Sungjae Park, Shubham Tulsiani
The paper introduces πR², a method that makes large pretrained manipulation policies reactive and real-time by separating fast proprioceptive inputs from slower vision-language inp…
RayRoPE: Projective Ray Positional Encoding for Multi-view Attention
Yu Wu, Minsik Jeon, Jen-Hao Rick Chang +2
We study positional encodings for multi-view transformers that process tokens from a set of posed input images, and seek a mechanism that encodes patches uniquely, allows SE(3)-inv…
Web2Grasp: Learning Functional Grasps from Web Images of Hand-Object Interactions
Hongyi Chen, Yunchao Yao, Yufei Ye +8
Functional grasping is essential for enabling dexterous multi-finger robot hands to manipulate objects effectively. Prior work largely focuses on power grasps, which only involve h…
Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild
Yehonathan Litman, Xiaoxuan Ma, Manan Shah +4
Reconstructing dynamic non-rigid objects from monocular video requires integrating visual cues from direct observations with data-driven priors over geometry and appearance. Prior…
DemoDiffusion: One-Shot Human Imitation using pre-trained Diffusion Policy
Sungjae Park, Homanga Bharadhwaj, Shubham Tulsiani
We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or p…
GHOST: Hierarchical Sub-Goal Policies for Generalizing Robot Manipulation
Sriram Krishna, Ben Eisner, Haotian Zhan +5
We present GHOST, a framework for learning visuomotor manipulation policies that generalize beyond the training distribution. GHOST factorizes control into (i) a high-level policy…