activity
20212026
most citedSelf-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

19 citations · 48 across the 44 of their papers we have counts for

collaborators
Showing 2022 · eess.ASShow all

6 papers · 2 filters

eess.AS2022★ 6 cited

Continual Learning for On-Device Speech Recognition using Disentangled Conformers

Anuj Diwan, Ching-Feng Yeh, Wei-Ning Hsu +4

Automatic speech recognition research focuses on training and evaluating on static datasets. Yet, as speech models are increasingly deployed on personal devices, such models encoun…

eess.AS2022

Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models

Reem Gody, David Harwath

Self-supervised learning (SSL) has been able to leverage unlabeled data to boost the performance of automatic speech recognition (ASR) models when we have access to only a small am…

eess.AS2022

Phoneme Segmentation Using Self-Supervised Speech Models

Luke Strgar, David Harwath

We apply transfer learning to the task of phoneme segmentation and demonstrate the utility of representations learned in self-supervised pre-training for the task. Our model extend…

eess.AS2022★ 2 cited

MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Alan Baade, Puyuan Peng, David Harwath

In this paper, we propose a simple yet powerful improvement over the recent Self-Supervised Audio Spectrogram Transformer (SSAST) model for speech and audio classification. Specifi…

eess.AS2022★ 19 cited

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Tran…

eess.AS2022

Word Discovery in Visually Grounded, Self-Supervised Speech Models

Puyuan Peng, David Harwath

We present a method for visually-grounded spoken term discovery. After training either a HuBERT or wav2vec2.0 model to associate spoken captions with natural images, we show that p…