most citedSimple and Effective Zero-shot Cross-lingual Phoneme Recognition

10 citations · 18 across the 4 of their papers we have counts for

collaborators

5 papers

eess.AS2021

Word Order Does Not Matter For Speech Recognition

Vineel Pratap, Qiantong Xu, Tatiana Likhomanenko +2

In this paper, we study training of automatic speech recognition system in a weakly supervised setting where the order of words in transcript labels of the audio training data is n…

cs.CL202110 cited

Simple and Effective Zero-shot Cross-lingual Phoneme Recognition

Qiantong Xu, Alexei Baevski, Michael Auli

Recent progress in self-training, self-supervised pretraining and unsupervised learning enabled well performing speech recognition systems without any labeled data. However, in man…

eess.AS20211 cited

Kaizen: Continuously improving teacher using Exponential Moving Average for semi-supervised speech recognition

Vimal Manohar, Tatiana Likhomanenko, Qiantong Xu +5

In this paper, we introduce the Kaizen framework that uses a continuously improving teacher to generate pseudo-labels for semi-supervised speech recognition (ASR). The proposed app…

cs.LG20217 cited

CAPE: Encoding Relative Positions with Continuous Augmented Positional Embeddings

Tatiana Likhomanenko, Qiantong Xu, Gabriel Synnaeve +2

Without positional information, attention-based Transformer neural networks are permutation-invariant. Absolute or relative positional embeddings are the most popular ways to feed…

cs.SD2021

Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training

Wei-Ning Hsu, Anuroop Sriram, Alexei Baevski +8

Self-supervised learning of speech representations has been a very active research area but most work is focused on a single domain such as read audio books for which there exist l…