activity
20222024
most citedFastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis

9 citations · 11 across the 5 of their papers we have counts for

collaborators

5 papers

cs.CV2024

AUD-TGN: Advancing Action Unit Detection with Temporal Convolution and GPT-2 in Wild Audiovisual Contexts

Jun Yu, Zerui Zhang, Zhihong Wei +6

Leveraging the synergy of both audio data and visual data is essential for understanding human emotions and behaviors, especially in in-the-wild setting. Traditional methods for in…

cs.CV2024

Multimodal Fusion Method with Spatiotemporal Sequences and Relationship Learning for Valence-Arousal Estimation

Jun Yu, Gongpeng Zhao, Yongqi Wang +7

This paper presents our approach for the VA (Valence-Arousal) estimation task in the ABAW6 competition. We devised a comprehensive model by preprocessing video frames and audio seg…

cs.CV2024

Exploring Facial Expression Recognition through Semi-Supervised Pretraining and Temporal Modeling

Jun Yu, Zhihong Wei, Zhongpeng Cai +6

Facial Expression Recognition (FER) plays a crucial role in computer vision and finds extensive applications across various fields. This paper aims to present our approach for the…

eess.AS20232 cited

Make-A-Voice: Unified Voice Synthesis With Discrete Representation

Rongjie Huang, Chunlei Zhang, Yongqi Wang +7

Various applications of voice synthesis have been developed independently despite the fact that they generate "voice" as output in common. In addition, the majority of voice synthe…

cs.SD20229 cited

FastLTS: Non-Autoregressive End-to-End Unconstrained Lip-to-Speech Synthesis

Yongqi Wang, Zhou Zhao

Unconstrained lip-to-speech synthesis aims to generate corresponding speeches from silent videos of talking faces with no restriction on head poses or vocabulary. Current works mai…