39 citations · 45 across the 2 of their papers we have counts for
6 papers
UniCon: Unified Context Network for Robust Active Speaker Detection
Yuanhang Zhang, Susan Liang, Shuang Yang +4
We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candida…
Deformation Flow Based Two-Stream Network for Lip Reading
Jingyun Xiao, Shuang Yang, Yuanhang Zhang +2
Lip reading is the task of recognizing the speech content by analyzing movements in the lip region when people are speaking. Observing on the continuity in adjacent frames in the s…
Can We Read Speech Beyond the Lips? Rethinking RoI Selection for Deep Visual Speech Recognition
Yuanhang Zhang, Shuang Yang, Jingyun Xiao +2
Recent advances in deep learning have heightened interest among researchers in the field of visual speech recognition (VSR). Currently, most existing methods equate VSR with automa…
T: Multi-Modal Continuous Valence-Arousal Estimation in the Wild
Yuan-Hang Zhang, Rulin Huang, Jiabei Zeng +2
This report describes a multi-modal multi-task (T) approach underlying our submission to the valence-arousal estimation track of the Affective Behavior Analysis in-the-wild (A…
LRW-1000: A Naturally-Distributed Large-Scale Benchmark for Lip Reading in the Wild
Shuang Yang, Yuanhang Zhang, Dalu Feng +6
Large-scale datasets have successively proven their fundamental importance in several research fields, especially for early progress in some emerging topics. In this paper, we focu…
3D Feature Pyramid Attention Module for Robust Visual Speech Recognition
Jingyun Xiao
Visual speech recognition is the task to decode the speech content from a video based on visual information, especially the movements of lips. It is also referenced as lipreading.…