output
20122024
most citedCaptum: A unified and generic model interpretability library for PyTorch

649 citations

Showing eess.ASShow all

7 papers · 1 filter

eess.AS202124 cited

NORESQA: A Framework for Speech Quality Assessment using Non-Matching References

Pranay Manocha, Buye Xu, Anurag Kumar

The perceptual task of speech quality assessment (SQA) is a challenging task for machines to do. Objective SQA methods that rely on the availability of the corresponding clean refe…

eess.AS202093 cited

Data Augmenting Contrastive Learning of Speech Representations in the Time Domain

Eugene Kharitonov, Morgane Rivière, Gabriel Synnaeve +4

Contrastive Predictive Coding (CPC), based on predicting future segments of speech based on past segments is emerging as a powerful algorithm for representation learning of speech…

eess.AS20205 cited

Weak-Attention Suppression For Transformer Based Speech Recognition

Yangyang Shi, Yongqiang Wang, Chunyang Wu +5

Transformers, originally proposed for natural language processing (NLP) tasks, have recently achieved great success in automatic speech recognition (ASR). However, adjacent acousti…

eess.AS20202 cited

Faster, Simpler and More Accurate Hybrid ASR Systems Using Wordpieces

Frank Zhang, Yongqiang Wang, Xiaohui Zhang +3

In this work, we first show that on the widely used LibriSpeech benchmark, our transformer-based context-dependent connectionist temporal classification (CTC) system produces state…

eess.AS20202 cited

SkinAugment: Auto-Encoding Speaker Conversions for Automatic Speech Translation

Arya D. McCarthy, Liezl Puzon, Juan Pino

We propose autoencoding speaker conversion for training data augmentation in automatic speech translation. This technique directly transforms an audio sequence, resulting in audio…

eess.AS2020

Phoneme Boundary Detection using Learnable Segmental Features

Felix Kreuk, Yaniv Sheena, Joseph Keshet +1

Phoneme boundary detection plays an essential first step for a variety of speech processing applications such as speaker diarization, speech science, keyword spotting, etc. In this…