13 papers
Grounded-Exo2Ego: Structured Semantic Grounding for Robust Exocentric-to-Egocentric Video Generation
Shengze Wang, Michael Stengel, Tianye Li +5
Generating egocentric video from a single exocentric video is an emerging and important topic for AR/VR and physical AI. Compared with conventional novel view synthesis, exo-to-ego…
World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video
Liyuan Zhu, Shengyu Huang, Amrita Mazumdar +6
We present World from Motion, a method for generating freely renderable dynamic 3D Gaussian representations from monocular videos. Our approach conditions a video model on dense, p…
Instant Expressive Gaussian Head Avatars at Over 100 FPS
Kaiwen Jiang, Xueting Li, Seonwook Park +3
Portrait animation has witnessed tremendous quality improvements thanks to recent advances in video diffusion models. However, these 2D methods often compromise 3D consistency and…
Real-time 3D Visualization of Radiance Fields on Light Field Displays
Jonghyun Kim, Cheng Sun, Michael Stengel +6
Radiance fields, including their recent efficient forms such as 3D Gaussian Splatting and Sparse Voxels, have revolutionized photorealistic 3D scene visualization by enabling high-…
DyaPlex: Full-Duplex Speech-Motion Model for Dyadic Interaction
Koki Nagano, Hongyu Liu, Seonwook Park +9
We present DyaPlex, a streaming, full-duplex speech-and-motion model designed for dyadic interaction. To capture the continuous and reciprocal nature of human communication, this f…
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
Amrita Mazumdar, Seonwook Park, Rajarshi Roy +6
Natural human conversation is full-duplex and audio-visual: people simultaneously speak and listen while continuously interpreting and producing nonverbal cues, such as nods, smile…