activity
20222024
most citedSelf-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

19 citations · 25 across the 10 of their papers we have counts for

collaborators
Showing eess.ASShow all

5 papers · 1 filter

eess.AS20241 cited

Self-supervised Speech Models for Word-Level Stuttered Speech Detection

Yi-Jen Shih, Zoi Gkalitsiou, Alexandros G. Dimakis +1

Clinical diagnosis of stuttering requires an assessment by a licensed speech-language pathologist. However, this process is time-consuming and requires clinicians with training and…

eess.AS2022

Unsupervised Fine-Tuning Data Selection for ASR Using Self-Supervised Speech Models

Reem Gody, David Harwath

Self-supervised learning (SSL) has been able to leverage unlabeled data to boost the performance of automatic speech recognition (ASR) models when we have access to only a small am…

eess.AS2022

Phoneme Segmentation Using Self-Supervised Speech Models

Luke Strgar, David Harwath

We apply transfer learning to the task of phoneme segmentation and demonstrate the utility of representations learned in self-supervised pre-training for the task. Our model extend…

eess.AS20222 cited

MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Alan Baade, Puyuan Peng, David Harwath

In this paper, we propose a simple yet powerful improvement over the recent Self-Supervised Audio Spectrogram Transformer (SSAST) model for speech and audio classification. Specifi…

eess.AS202219 cited

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Tran…