5 papers
Augmenting Transformer-Transducer Based Speaker Change Detection With Token-Level Training Loss
Guanlong Zhao, Quan Wang, Han Lu +2
In this work we propose a novel token-based training strategy that improves Transformer-Transducer (T-T) based speaker change detection (SCD) performance. The conventional T-T base…
Exploring Sequence-to-Sequence Transformer-Transducer Models for Keyword Spotting
Beltrán Labrador, Guanlong Zhao, Ignacio López Moreno +3
In this paper, we present a novel approach to adapt a sequence-to-sequence Transformer-Transducer ASR system to the keyword spotting (KWS) task. We achieve this by replacing the ke…
LSTM Acoustic Models Learn to Align and Pronounce with Graphemes
Arindrima Datta, Guanlong Zhao, Bhuvana Ramabhadran +1
Automated speech recognition coverage of the world's languages continues to expand. However, standard phoneme based systems require handcrafted lexicons that are difficult and expe…
Improved Techniques for Learning to Dehaze and Beyond: A Collective Study
Yu Liu, Guanlong Zhao, Boyuan Gong +8
Here we explore two related but important tasks based on the recently released REalistic Single Image DEhazing (RESIDE) benchmark dataset: (i) single image dehazing as a low-level…
PAD-Net: A Perception-Aided Single Image Dehazing Network
Yu Liu, Guanlong Zhao
In this work, we investigate the possibility of replacing the loss with perceptually derived loss functions (SSIM, MS-SSIM, etc.) in training an end-to-end dehazing neural…