activity
20192021
most citedAccent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

11 citations · 28 across the 10 of their papers we have counts for

collaborators

15 papers

eess.AS202111 cited

Accent-Robust Automatic Speech Recognition Using Supervised and Unsupervised Wav2vec Embeddings

Jialu Li, Vimal Manohar, Pooja Chitkara +5

Speech recognition models often obtain degraded performance when tested on speech with unseen accents. Domain-adversarial training (DAT) and multi-task learning (MTL) are two commo…

eess.AS20211 cited

On lattice-free boosted MMI training of HMM and CTC-based full-context ASR models

Xiaohui Zhang, Vimal Manohar, David Zhang +7

Hybrid automatic speech recognition (ASR) models are typically sequentially trained with CTC or LF-MMI criteria. However, they have vastly different legacies and are usually implem…

cs.CL20212 cited

Improved Language Identification Through Cross-Lingual Self-Supervised Learning

Andros Tjandra, Diptanu Gon Choudhury, Frank Zhang +6

Language identification greatly impacts the success of downstream tasks such as automatic speech recognition. Recently, self-supervised speech representations learned by wav2vec 2.…

eess.AS20211 cited

Kaizen: Continuously improving teacher using Exponential Moving Average for semi-supervised speech recognition

Vimal Manohar, Tatiana Likhomanenko, Qiantong Xu +5

In this paper, we introduce the Kaizen framework that uses a continuously improving teacher to generate pseudo-labels for semi-supervised speech recognition (ASR). The proposed app…

cs.CL2021

Contextualized Streaming End-to-End Speech Recognition with Trie-Based Deep Biasing and Shallow Fusion

Duc Le, Mahaveer Jain, Gil Keren +9

How to leverage dynamic contextual information in end-to-end speech recognition has remained an active research area. Previous solutions to this problem were either designed for sp…

cs.SD20215 cited

A Multi-View Approach To Audio-Visual Speaker Verification

Leda Sarı, Kritika Singh, Jiatong Zhou +3

Although speaker verification has conventionally been an audio-only task, some practical applications provide both audio and visual streams of input. In these cases, the visual str…