10 citations · 19 across the 5 of their papers we have counts for
5 papers
A comprehensive study on self-supervised distillation for speaker representation learning
Zhengyang Chen, Yao Qian, Bing Han +2
In real application scenarios, it is often challenging to obtain a large amount of labeled data for speaker representation learning due to speaker privacy concerns. Self-supervised…
The Microsoft System for VoxCeleb Speaker Recognition Challenge 2022
Gang Liu, Tianyan Zhou, Yong Zhao +4
In this report, we describe our submitted system for track 2 of the VoxCeleb Speaker Recognition Challenge 2022 (VoxSRC-22). We fuse a variety of good-performing models ranging fro…
i-Code: An Integrative and Composable Multimodal Learning Framework
Ziyi Yang, Yuwei Fang, Chenguang Zhu +17
Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to…
UniSpeech-SAT: Universal Speech Representation Learning with Speaker Aware Pre-Training
Sanyuan Chen, Yu Wu, Chengyi Wang +8
Self-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years witness…
To Trust, or Not to Trust? A Study of Human Bias in Automated Video Interview Assessments
Chee Wee Leong, Katrina Roohr, Vikram Ramanarayanan +6
Supervised systems require human labels for training. But, are humans themselves always impartial during the annotation process? We examine this question in the context of automate…