activity
20242026
collaborators

9 papers

cs.CV2026

Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation

Shengze Wang, Michael Stengel, Tianye Li +5

Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego…

cs.CV2026

Instant Expressive Gaussian Head Avatars at Over 100 FPS

Kaiwen Jiang, Xueting Li, Seonwook Park +3

Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods often compromise 3D consistency and…

cs.CV2026

DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction

Koki Nagano, Hongyu Liu, Seonwook Park +9

We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this f…

cs.CV2026

VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents

Amrita Mazumdar, Seonwook Park, Rajarshi Roy +6

Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smile…

cs.CV2026

COSY: Compositional 3DGS Synthesis for Disentangled Human Head Editing

Florian Barthel, Shalini De Mello, Koki Nagano +3

Recent 3D Gaussian Splatting (3DGS) GANs for human heads synthesize and render photorealistic 3D models in real-time and offer a vast variety in identity and appearance. However, c…

cs.GR2025

Dream, Lift, Animate: From Single Images to Animatable Gaussian Avatars

Marcel C. Bühler, Ye Yuan, Xueting Li +3

We introduce Dream, Lift, Animate (DLA), a novel framework that reconstructs animatable 3D human avatars from a single image. This is achieved by leveraging multi-view generation,…