activity
20162026
most citedBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

565 citations · 1.2k across the 52 of their papers we have counts for

collaborators
Showing eess.ASShow all

6 papers · 1 filter

eess.AS2023

AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement

Ju-Chieh Chou, Chung-Ming Chien, Karen Livescu

Speech enhancement systems are typically trained using pairs of clean and noisy speech. In audio-visual speech enhancement (AVSE), there is not as much ground-truth clean data avai…

eess.AS2022

Context-aware Fine-tuning of Self-supervised Speech Models

Suwon Shon, Felix Wu, Kwangyoun Kim +3

Self-supervised pre-trained transformers have improved the state of the art on a variety of speech tasks. Due to the quadratic time and space complexity of self-attention, they usu…

eess.AS2020★ 8 cited

A Correspondence Variational Autoencoder for Unsupervised Acoustic Word Embeddings

Puyuan Peng, Herman Kamper, Karen Livescu

We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis…

eess.AS2020

Whole-Word Segmental Speech Recognition with Acoustic Word Embeddings

Bowen Shi, Shane Settle, Karen Livescu

Segmental models are sequence prediction models in which scores of hypotheses are based on entire variable-length segments of frames. We consider segmental models for whole-word ("…

eess.AS2020

Unsupervised Pre-training of Bidirectional Speech Encoders via Masked Reconstruction

Weiran Wang, Qingming Tang, Karen Livescu

We propose an approach for pre-training speech representations via a masked reconstruction loss. Our pre-trained encoder networks are bidirectional and can therefore be used direct…

eess.AS2018

A Comparison of Techniques for Language Model Integration in Encoder-Decoder Speech Recognition

Shubham Toshniwal, Anjuli Kannan, Chung-Cheng Chiu +3

Attention-based recurrent neural encoder-decoder models present an elegant solution to the automatic speech recognition problem. This approach folds the acoustic model, pronunciati…