activity
20162025
most citedAn Unsupervised Autoregressive Model for Speech Representation Learning

46 citations · 91 across the 13 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2022

Supervised Attention in Sequence-to-Sequence Models for Speech Recognition

Gene-Ping Yang, Hao Tang

Attention mechanism in sequence-to-sequence models is designed to model the alignments between acoustic features and output tokens in speech recognition. However, attention weights…

eess.AS20206 cited

Vector-Quantized Autoregressive Predictive Coding

Yu-An Chung, Hao Tang, James Glass

Autoregressive Predictive Coding (APC), as a self-supervised objective, has enjoyed success in learning representations from large amounts of unlabeled data, and the learned repres…

eess.AS2020

Audio-Visual Calibration with Polynomial Regression for 2-D Projection Using SVD-PHAT

Francois Grondin, Hao Tang, James Glass

This paper proposes a straightforward 2-D method to spatially calibrate the visual field of a camera with the auditory field of an array microphone by generating and overlaying an…

eess.AS2019

VoiceID Loss: Speech Enhancement for Speaker Verification

Suwon Shon, Hao Tang, James Glass

In this paper, we propose VoiceID loss, a novel loss function for training a speech enhancement model to improve the robustness of speaker verification. In contrast to the commonly…

eess.AS2018

On The Inductive Bias of Words in Acoustics-to-Word Models

Hao Tang, James Glass

Acoustics-to-word models are end-to-end speech recognizers that use words as targets without relying on pronunciation dictionaries or graphemes. These models are notoriously diffic…

eess.AS2018

Frame-level speaker embeddings for text-independent speaker recognition and analysis of end-to-end model

Suwon Shon, Hao Tang, James Glass

In this paper, we propose a Convolutional Neural Network (CNN) based speaker recognition model for extracting robust speaker embeddings. The embedding can be extracted efficiently…