1 citations · 1 across the 3 of their papers we have counts for
3 papers
eess.AS2022
Augmenting Transformer-Transducer Based Speaker Change Detection With Token-Level Training Loss
Guanlong Zhao, Quan Wang, Han Lu +2
In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T base…
eess.AS2022
Attentive Temporal Pooling for Conformer-based Streaming Language Identification in Long-form Speech
Quan Wang, Yang Yu, Jason Pelecanos +2
In this paper, we introduce a novel language identification system based on conformer layers. We propose an attentive temporal pooling mechanism to allow the model to carry informa…
eess.AS2020★ 1 cited
Synth2Aug: Cross-domain speaker recognition with TTS synthesized speech
Yiling Huang, Yutian Chen, Jason Pelecanos +1
In recent years, Text-To-Speech (TTS) has been used as a data augmentation technique for speech recognition to help complement inadequacies in the training data. Correspondingly, w…