421 citations · 1.2k across the 37 of their papers we have counts for
54 papers · 1 filter
Cooperative Dual Attention for Audio-Visual Speech Enhancement with Facial Cues
Feixiang Wang, Shuang Yang, Shiguang Shan +1
In this work, we focus on leveraging facial cues beyond the lip region for robust Audio-Visual Speech Enhancement (AVSE). The facial region, encompassing the lip region, reflects a…
Clothes-Changing Person Re-identification with RGB Modality Only
Xinqian Gu, Hong Chang, Bingpeng Ma +3
The key to address clothes-changing person re-identification (re-id) is to extract clothes-irrelevant features, e.g., face, hairstyle, body shape, and gait. Most current works main…
SEGA: Semantic Guided Attention on Visual Prototype for Few-Shot Learning
Fengyuan Yang, Ruiping Wang, Xilin Chen
Teaching machines to recognize a new category based on few training samples especially only one remains challenging owing to the incomprehensive understanding of the novel category…
HRFormer: High-Resolution Transformer for Dense Prediction
Yuhui Yuan, Rao Fu, Lang Huang +4
We present a High-Resolution Transformer (HRFormer) that learns high-resolution representations for dense prediction tasks, in contrast to the original Vision Transformer that prod…
UniCon: Unified Context Network for Robust Active Speaker Detection
Yuanhang Zhang, Susan Liang, Shuang Yang +4
We introduce a new efficient framework, the Unified Context Network (UniCon), for robust active speaker detection (ASD). Traditional methods for ASD usually operate on each candida…
Gaze Estimation with an Ensemble of Four Architectures
Xin Cai, Boyu Chen, Jiabei Zeng +7
This paper presents a method for gaze estimation according to face images. We train several gaze estimators adopting four different network architectures, including an architecture…