6 papers
S^2VG: 3D Stereoscopic and Spatial Video Generation via Denoising Frame Matrix
Peng Dai, Feitong Tan, Qiangeng Xu +6
While video generation models excel at producing high-quality monocular videos, generating 3D stereoscopic and spatial videos for immersive applications remains an underexplored ch…
Archon: A Unified Multimodal Model for Holistic Digital Human Generation
Chong Bao, Shichen Liu, Lijun Yu +9
Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open ch…
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
Mengyi Shan, Shouchieh Chang, Ziqian Bai +6
We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often prod…
CHOSEN: Contrastive Hypothesis Selection for Multi-View Depth Refinement
Di Qiu, Yinda Zhang, Thabo Beeler +5
We propose CHOSEN, a simple yet flexible, robust and effective multi-view depth refinement framework. It can be employed in any existing multi-view stereo pipeline, with straightfo…
IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos
Yuan Li, Ziqian Bai, Feitong Tan +3
We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expre…
Sandwiched Compression: Repurposing Standard Codecs with Neural Network Wrappers
Onur G. Guleryuz, Philip A. Chou, Berivan Isik +6
We propose sandwiching standard image and video codecs between pre- and post-processing neural networks. The networks are jointly trained through a differentiable codec proxy to mi…