activity
20162024
most citedFine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks

297 citations · 474 across the 15 of their papers we have counts for

collaborators

15 papers

cs.SD20242 cited

Masked Audio Generation using a Single Non-Autoregressive Transformer

Alon Ziv, Itai Gat, Gael Le Lan +6

We introduce MAGNeT, a masked generative sequence modeling method that operates directly over several streams of audio tokens. Unlike prior work, MAGNeT is comprised of a single-st…

cs.CL2023

Generative Spoken Language Model based on continuous word-sized audio tokens

Robin Algayres, Yossi Adi, Tu Anh Nguyen +4

In NLP, text language models based on words or subwords are known to outperform their character-based counterparts. Yet, in the speech community, the standard input of spoken LMs a…

cs.LG20232 cited

Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

Guy Yariv, Itai Gat, Sagie Benaim +3

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to b…

cs.CL2023

EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Tu Anh Nguyen, Wei-Ning Hsu, Antony D'Avirro +10

Recent work has shown that it is possible to resynthesize high-quality speech based, not on text, but on low bitrate discrete units that have been learned in a self-supervised fash…

cs.SD20233 cited

From Discrete Tokens to High-Fidelity Audio Using Multi-Band Diffusion

Robin San Roman, Yossi Adi, Antoine Deleforge +3

Deep generative models can generate high-fidelity audio conditioned on various types of representations (e.g., mel-spectrograms, Mel-frequency Cepstral Coefficients (MFCC)). Recent…

cs.CL2023116 cited

Scaling Speech Technology to 1,000+ Languages

Vineel Pratap, Andros Tjandra, Bowen Shi +13

Expanding the language coverage of speech technology has the potential to improve access to information for many more people. However, current speech technology is restricted to ab…