Publications (13)
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…
Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait Generation
Weipeng Tan, Chuming Lin, Chengming Xu +6
Recent advances in Talking Head Generation (THG) have achieved impressive lip synchronization and visual quality through diffusion models; yet existing methods struggle to generate…
RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
Xiaozhong Ji, Chuming Lin, Zhonggan Ding +6
Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there…
NTIRE 2022 Challenge on Efficient Super-Resolution: Methods and Results
Yawei Li, Kai Zhang, Radu Timofte +108
This paper reviews the NTIRE 2022 challenge on efficient single image super-resolution with focus on the proposed solutions and results. The task of the challenge was to super-reso…
Learning Salient Boundary Feature for Anchor-free Temporal Action Localization
Chuming Lin, Chengming Xu, Donghao Luo +6
Temporal action localization is an important yet challenging task in video understanding. Typically, such a task aims at inferring both the action category and localization of the…
High-Resolution GAN Inversion for Degraded Images in Large Diverse Datasets
Yanbo Wang, Chuming Lin, Donghao Luo +3
The last decades are marked by massive and diverse image data, which shows increasingly high resolution and quality. However, some images we obtained may be corrupted, affecting th…
Fast Learning of Temporal Action Proposal via Dense Boundary Generator
Chuming Lin, Jian Li, Yabiao Wang +7
Generating temporal action proposals remains a very challenging problem, where the main issue lies in predicting precise temporal proposal boundaries and reliable action confidence…
SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion Model
Weipeng Tan, Chuming Lin, Chengming Xu +5
Talking Head Generation (THG), typically driven by audio, is an important and challenging task with broad application prospects in various fields such as digital humans, film produ…
JAVEDIT: Joint Audio-Visual Instruction-Guided Video Editing with Agentic Data Curation
Yinan Chen, Chuming Lin, Zhennan Chen +12
While instruction-based video editing has seen significant progress, joint audio-visual editing remains constrained by the absence of dedicated datasets and benchmarks. To bridge t…
Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
Xiaozhong Ji, Xiaobin Hu, Zhihong Xu +9
The study of talking face generation mainly explores the intricacies of synchronizing facial movements and crafting visually appealing, temporally-coherent animations. However, due…
Joint Learning Content and Degradation Aware Feature for Blind Super-Resolution
Yifeng Zhou, Chuming Lin, Donghao Luo +4
To achieve promising results on blind image super-resolution (SR), some attempts leveraged the low resolution (LR) images to predict the kernel and improve the SR performance. Howe…
Spectrum-to-Kernel Translation for Accurate Blind Image Super-Resolution
Guangpin Tao, Xiaozhong Ji, Wenzhuo Wang +6
Deep-learning based Super-Resolution (SR) methods have exhibited promising performance under non-blind setting where blur kernel is known. However, blur kernels of Low-Resolution (…
Frame and Feature-Context Video Super-Resolution
Bo Yan, Chuming Lin, Weimin Tan
For video super-resolution, current state-of-the-art approaches either process multiple low-resolution (LR) frames to produce each output high-resolution (HR) frame separately in a…