133 citations · 226 across the 12 of their papers we have counts for
5 papers · 1 filter
Deep Contextualized Acoustic Representations For Semi-Supervised Speech Recognition
Shaoshi Ling, Yuzong Liu, Julian Salazar +1
We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we rec…
Masked Language Model Scoring
Julian Salazar, Davis Liang, Toan Q. Nguyen +1
Pretrained masked language models (MLMs) require finetuning for most NLP tasks. Instead, we evaluate MLMs out of the box via their pseudo-log-likelihood scores (PLLs), which are co…
Simple, Fast, Accurate Intent Classification and Slot Labeling for Goal-Oriented Dialogue Systems
Arshit Gupta, John Hewitt, Katrin Kirchhoff
With the advent of conversational assistants, like Amazon Alexa, Google Now, etc., dialogue systems are gaining a lot of traction, especially in industrial setting. These systems t…
Self-Attention Networks for Connectionist Temporal Classification in Speech Recognition
Julian Salazar, Katrin Kirchhoff, Zhiheng Huang
The success of self-attention in NLP has led to recent applications in end-to-end encoder-decoder architectures for speech recognition. Separately, connectionist temporal classific…
Multi-stream Network With Temporal Attention For Environmental Sound Classification
Xinyu Li, Venkata Chebiyyam, Katrin Kirchhoff
Environmental sound classification systems often do not perform robustly across different sound classification tasks and audio signals of varying temporal structures. We introduce…