4 papers
Archon: A Unified Multimodal Model for Holistic Digital Human Generation
Chong Bao, Shichen Liu, Lijun Yu +9
Digital humans are fundamental to immersive interaction, yet creating a unified model for holistic modalities, including text, audio, motion, and visual content, remains an open ch…
Talking Together: Synthesizing Co-Located 3D Conversations from Audio
Mengyi Shan, Shouchieh Chang, Ziqian Bai +6
We tackle the challenging task of generating complete 3D facial animations for two interacting, co-located participants from a mixed audio stream. While existing methods often prod…
LegacyAvatars: Volumetric Face Avatars For Traditional Graphics Pipelines
Safa C. Medin, Gengyan Li, Ziqian Bai +8
We introduce a novel representation for efficient classical rendering of photorealistic 3D face avatars. Leveraging recent advances in radiance fields anchored to parametric face m…
IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos
Yuan Li, Ziqian Bai, Feitong Tan +3
We propose a novel 3D-aware diffusion-based method for generating photorealistic talking head videos directly from a single identity image and explicit control signals (e.g., expre…