11 citations · 12 across the 6 of their papers we have counts for
6 papers
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
Jiazhi Guan, Zhiliang Xu, Hang Zhou +10
Lip-syncing videos with given audio is the foundation for various applications including the creation of virtual presenters or performers. While recent studies explore high-fidelit…
Accelerating Vision Transformers Based on Heterogeneous Attention Patterns
Deli Yu, Teng Xi, Jianwei Li +7
Recently, Vision Transformers (ViTs) have attracted a lot of attention in the field of computer vision. Generally, the powerful representative capacity of ViTs mainly benefits from…
HD-Fusion: Detailed Text-to-3D Generation Leveraging Multiple Noise Estimation
Jinbo Wu, Xiaobo Gao, Xing Liu +5
In this paper, we study Text-to-3D content generation leveraging 2D diffusion priors to enhance the quality and detail of the generated 3D models. Recent progress (Magic3D) in text…
StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based Generator
Jiazhi Guan, Zhanwang Zhang, Hang Zhou +8
Despite recent advances in syncing lip movements with any audio waves, current methods still struggle to balance generation quality and the model's generalization ability. Previous…
TRUST: An Accurate and End-to-End Table structure Recognizer Using Splitting-based Transformers
Zengyuan Guo, Yuechen Yu, Pengyuan Lv +6
Table structure recognition is a crucial part of document image analysis domain. Its difficulty lies in the need to parse the physical coordinates and logical indices of each cell…
UFO: Unified Feature Optimization
Teng Xi, Yifan Sun, Deli Yu +13
This paper proposes a novel Unified Feature Optimization (UFO) paradigm for training and deploying deep models under real-world and large-scale scenarios, which requires a collecti…