2 papers
cs.CV2024
GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec Avatars
Keqiang Sun, Amin Jourabloo, Riddhish Bhalodia +9
Photo-realistic and controllable 3D avatars are crucial for various applications such as virtual and mixed reality (VR/MR), telepresence, gaming, and film production. Traditional m…
cs.CV2024
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
Wufei Ma, Kai Li, Zhongshi Jiang +5
Recent video-text foundation models have demonstrated strong performance on a wide variety of downstream video understanding tasks. Can these video-text models genuinely understand…