47 citations · 115 across the 5 of their papers we have counts for
6 papers
Audio2Gestures: Generating Diverse Gestures from Speech Audio with Conditional Variational Autoencoders
Jing Li, Di Kang, Wenjie Pei +4
Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one…
Animatable Neural Radiance Fields from Monocular RGB Videos
Jianchuan Chen, Ying Zhang, Di Kang +4
We present animatable neural radiance fields (animatable NeRF) for detailed human avatar creation from monocular videos. Our approach extends neural radiance fields (NeRF) to the d…
Model-based 3D Hand Reconstruction via Self-Supervised Learning
Yujin Chen, Zhigang Tu, Di Kang +5
Reconstructing a 3D hand from a single-view RGB image is challenging due to various hand configurations and depth ambiguity. To reliably reconstruct a 3D hand from a monocular imag…
Similarity Reasoning and Filtration for Image-Text Matching
Haiwen Diao, Ying Zhang, Lin Ma +1
Image-text matching plays a critical role in bridging the vision and language, and great progress has been made by exploiting the global alignment between image and sentence, or lo…
Consensus-Aware Visual-Semantic Embedding for Image-Text Matching
Haoran Wang, Ying Zhang, Zhong Ji +2
Image-text matching plays a central role in bridging vision and language. Most existing approaches only rely on the image-text instance pair to learn their representations, thereby…
Deep Mutual Learning
Ying Zhang, Tao Xiang, Timothy M. Hospedales +1
Model distillation is an effective and widely used technique to transfer knowledge from a teacher to a student network. The typical application is to transfer from a powerful large…