251 citations · 485 across the 38 of their papers we have counts for
62 papers · 1 filter
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
Fuchen Long, Zhaofan Qiu, Ting Yao +1
The recent innovations and breakthroughs in diffusion models have significantly expanded the possibilities of generating high-quality videos for the given prompts. Most existing wo…
Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
Peipei Li, Xing Cui, Yibo Hu +3
Point cloud analysis faces computational system overhead, limiting its application on mobile or edge devices. Directly employing small models may result in a significant drop in pe…
Selective Volume Mixup for Video Action Recognition
Yi Tan, Zhaofan Qiu, Yanbin Hao +2
The recent advances in Convolutional Neural Networks (CNNs) and Vision Transformers have convincingly demonstrated high learning capability for video action recognition on large da…
Learning and Evaluating Human Preferences for Conversational Head Generation
Mohan Zhou, Yalong Bai, Wei Zhang +3
A reliable and comprehensive evaluation metric that aligns with manual preference assessments is crucial for conversational head video synthesis methods development. Existing quant…
Deep Equilibrium Multimodal Fusion
Jinhong Ni, Yalong Bai, Wei Zhang +2
Multimodal fusion integrates the complementary information present in multiple modalities and has gained much attention recently. Most existing fusion approaches either learn a fix…
Transforming Radiance Field with Lipschitz Network for Photorealistic 3D Scene Stylization
Zicheng Zhang, Yinglu Liu, Congying Han +3
Recent advances in 3D scene representation and novel view synthesis have witnessed the rise of Neural Radiance Fields (NeRFs). Nevertheless, it is not trivial to exploit NeRF for t…