46 citations · 91 across the 13 of their papers we have counts for
6 papers · 1 filter
Supervised Attention in Sequence-to-Sequence Models for Speech Recognition
Gene-Ping Yang, Hao Tang
Attention mechanism in sequence-to-sequence models is designed to model the alignments between acoustic features and output tokens in speech recognition. However, attention weights…
Vector-Quantized Autoregressive Predictive Coding
Yu-An Chung, Hao Tang, James Glass
Autoregressive Predictive Coding (APC), as a self-supervised objective, has enjoyed success in learning representations from large amounts of unlabeled data, and the learned repres…
Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT
Francois Grondin, Hao Tang, James Glass
This paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an…
VoiceID Loss: Speech Enhancement for Speaker Verification
Suwon Shon, Hao Tang, James Glass
In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly…
On The Inductive Bias of Words in Acoustics-to-Word Models
Hao Tang, James Glass
Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously diffic…
Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model
Suwon Shon, Hao Tang, James Glass
In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently…