649 citations
- Carnegie Mellon UniversityUS23 papers
- Stanford UniversityUS20 papers
- Google (United States)US14 papers
- Georgia Institute of TechnologyUS13 papers
- Tel Aviv UniversityIL12 papers
- Cornell UniversityUS11 papers
- University of California, BerkeleyUS11 papers
- University College LondonGB10 papers
- Harvard University PressUS9 papers
- Johns Hopkins UniversityUS9 papers
- Massachusetts Institute of TechnologyUS9 papers
- The University of Texas at AustinUS9 papers
7 papers · 1 filter
NORESQA: A Framework for Speech Quality Assessment using Non-Matching References
Pranay Manocha, Buye Xu, Anurag Kumar
The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean refe…
Data Augmenting Contrastive Learning of Speech Representations in the Time Domain
Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve +4
Contrastive Predictive Coding (CPC), based on predicting future segments of speech based on past segments is emerging as a powerful algorithm for representation learning of speech…
Weak-Attention Suppression For Transformer Based Speech Recognition
Yangyang Shi, Yongqiang Wang, Chunyang Wu +5
Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acousti…
Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces
Frank Zhang, Yongqiang Wang, Xiaohui Zhang +3
In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state…
SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation
Arya D. McCarthy, Liezl Puzon, Juan Pino
We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio…
Phoneme Boundary Detection using Learnable Segmental Features
Felix Kreuk, Yaniv Sheena, Joseph Keshet +1
Phoneme boundary detection plays an essential first step for a variety of speech processing applications such as speaker diarization, speech science, keyword spotting, etc. In this…