4 papers · 1 filter
MoCHA: Denoising Caption Supervision for Motion-Text Retrieval
Nikolai Warner, Cameron Ethan Taylor, Irfan Essa +1
Text-motion retrieval systems learn shared embedding spaces from motion-caption pairs via contrastive objectives. However, each caption is not a deterministic label but a sample fr…
AugLift: Depth-Aware Input Reparameterization Improves Domain Generalization in 2D-to-3D Pose Lifting
Nikolai Warner, Wenjin Zhang, Hamid Badiozamani +2
Lifting-based 3D human pose estimation infers 3D joints from 2D keypoints but generalizes poorly because coordinates alone are an ill-posed, sparse representation that disc…
Learning Complex Non-Rigid Image Edits from Multimodal Conditioning
Nikolai Warner, Jack Kolb, Meera Hahn +3
In this paper we focus on inserting a given human (specifically, a single image of a person) into a novel scene. Our method, which builds on top of Stable Diffusion, yields natural…
Text and Click inputs for unambiguous open vocabulary instance segmentation
Nikolai Warner, Meera Hahn, Jonathan Huang +2
Segmentation localizes objects in an image on a fine-grained per-pixel scale. Segmentation benefits by humans-in-the-loop to provide additional input of objects to segment using a…