4 papers · 1 filter
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
Sammy Christen, Shreyas Hampali, Fadime Sener +5
Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furth…
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
Qihang Fang, Chengcheng Tang, Bugra Tekin +1
Recent advancements in models linking natural language with human motions have shown significant promise in motion generation and editing based on instructional text. Motivated by…
FoundPose: Unseen Object Pose Estimation with Foundation Features
Evin Pınar Ãrnek, Yann Labbé, Bugra Tekin +4
We propose FoundPose, a model-based method for 6D pose estimation of unseen objects from a single RGB image. The method can quickly onboard new objects using their 3D models withou…
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
Anna Kukleva, Fadime Sener, Edoardo Remelli +4
Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. Howe…