234 citations · 871 across the 51 of their papers we have counts for
74 papers · 1 filter
Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in Transformers
Yasheng Sun, Hang Zhou, Kaisiyuan Wang +7
Previous studies have explored generating accurately lip-synced talking faces for arbitrary targets given audio conditions. However, most of them deform or generate the whole facia…
Cyclically Disentangled Feature Translation for Face Anti-spoofing
Haixiao Yue, Keyao Wang, Guosheng Zhang +4
Current domain adaptation methods for face anti-spoofing leverage labeled source domain data and unlabeled target domain data to obtain a promising generalizable decision boundary.…
CAE v2: Context Autoencoder with CLIP Target
Xinyu Zhang, Jiahui Chen, Junkun Yuan +10
Masked image modeling (MIM) learns visual representation by masking and reconstructing image patches. Applying the reconstruction supervision on the CLIP representation has been pr…
Group DETR v2: Strong Object Detector with Encoder-Decoder Pretraining
Qiang Chen, Jian Wang, Chuchu Han +12
We present a strong object detector with encoder-decoder pretraining and finetuning. Our method, called Group DETR v2, is built upon a vision transformer encoder ViT-Huge~\cite{dos…
U-HRNet: Delving into Improving Semantic Representation of High Resolution Network for Dense Prediction
Jian Wang, Xiang Long, Guowei Chen +3
High resolution and advanced semantic representation are both vital for dense prediction. Empirically, low-resolution feature maps often achieve stronger semantic representation, a…
RTFormer: Efficient Design for Real-Time Semantic Segmentation with Transformer
Jian Wang, Chenhui Gou, Qiman Wu +4
Recently, transformer-based networks have shown impressive results in semantic segmentation. Yet for real-time semantic segmentation, pure CNN-based approaches still dominate in th…