activity
20202025
most citedSelf-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

19 citations · 30 across the 7 of their papers we have counts for

collaborators

5 papers

cs.CL2024

SyllableLM: Learning Coarse Semantic Units for Speech Language Models

Alan Baade, Puyuan Peng, David Harwath

Language models require tokenized inputs. However, tokenization strategies for continuous data like audio and vision are often based on simple heuristics such as fixed sized convol…

cs.CV20221 cited

Zero-shot Video Moment Retrieval With Off-the-Shelf Models

Anuj Diwan, Puyuan Peng, Raymond J. Mooney

For the majority of the machine learning community, the expensive nature of collecting high-quality human-annotated data and the inability to efficiently finetune very large state-…

eess.AS20222 cited

MAE-AST: Masked Autoencoding Audio Spectrogram Transformer

Alan Baade, Puyuan Peng, David Harwath

In this paper, we propose a simple yet powerful improvement over the recent Self-Supervised Audio Spectrogram Transformer (SSAST) model for speech and audio classification. Specifi…

eess.AS202219 cited

Self-Supervised Representation Learning for Speech Using Visual Grounding and Masked Language Modeling

Puyuan Peng, David Harwath

In this paper, we describe our submissions to the ZeroSpeech 2021 Challenge and SUPERB benchmark. Our submissions are based on the recently proposed FaST-VGS model, which is a Tran…

eess.AS20208 cited

A Correspondence Variational Autoencoder for Unsupervised Acoustic Word Embeddings

Puyuan Peng, Herman Kamper, Karen Livescu

We propose a new unsupervised model for mapping a variable-duration speech segment to a fixed-dimensional representation. The resulting acoustic word embeddings can form the basis…