activity
20242026
collaborators

10 papers

cs.CV2026

Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation

Federico Nocentini, Kwanggyoon Seo, Qingju Liu +6

Speech-Driven Facial Animation (SDFA) has gained significant attention due to its applications in movies, video games, and virtual reality. However, most existing models are traine…

cs.CV2025

Zero-Shot Video Deraining with Video Diffusion Models

Tuomas Varanka, Juan Luis Gonzalez, Hyeongwoo Kim +2

Existing video deraining methods are often trained on paired datasets, either synthetic, which limits their ability to generalize to real-world rain, or captured by static cameras,…

cs.CV2025

Audio-Driven Universal Gaussian Head Avatars

Kartik Teotia, Helge Rhodin, Mohit Mendiratta +3

We introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP)…

eess.AS2025

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs

Eray Eren, Qingju Liu, Hyeongwoo Kim +2

Prosody conveys rich emotional and semantic information of the speech signal as well as individual idiosyncrasies. We propose a stand-alone model that maps text-to-prosodic feature…

cs.CV2025

Disentangling 3D from Large Vision-Language Models for Controlled Portrait Generation

Nick Yiwen Huang, Akin Caliskan, Berkay Kicanaoglu +2

We consider the problem of disentangling 3D from large vision-language models, which we show on generative 3D portraits. This allows free-form text control of appearance attributes…

eess.IV2025

VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing

Juan Luis Gonzalez Bello, Xu Yao, Alex Whelan +3

We present an implicit video representation for occlusions, appearance, and motion disentanglement from monocular videos, which we call Video SPatiotemporal Splines (VideoSPatS). U…