4 papers · 1 filter
VidTok: A Versatile and Open-Source Video Tokenizer
Anni Tang, Tianyu He, Junliang Guo +3
Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixe…
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
Anni Tang, Tianyu He, Xu Tan +2
Talking face generation aims at generating photo-realistic video portraits of a target person driven by input audio. Due to its nature of one-to-many mapping from the input audio t…
Efficient Dynamic-NeRF Based Volumetric Video Coding with Rate Distortion Optimization
Zhiyu Zhang, Guo Lu, Huanxiong Liang +3
Volumetric videos, benefiting from immersive 3D realism and interactivity, hold vast potential for various applications, while the tremendous data volume poses significant challeng…
Compositional 3D-aware Video Generation with LLM Director
Hanxin Zhu, Tianyu He, Anni Tang +3
Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain i…