activity
20162021
most citedParallel WaveNet: Fast High-Fidelity Speech Synthesis

343 citations · 853 across the 13 of their papers we have counts for

collaborators

22 papers

cs.LG2021

Vector Quantized Models for Planning

Sherjil Ozair, Yazhe Li, Ali Razavi +3

Recent developments in the field of model-based RL have proven successful in a range of environments, especially ones where planning is essential. However, such successes have been…

cs.CV20211 cited

Divide and Contrast: Self-supervised Learning from Uncurated Data

Yonglong Tian, Olivier J. Henaff, Aaron van den Oord

Self-supervised learning holds promise in leveraging large amounts of unlabeled data, however much of its progress has thus far been limited to highly curated pre-training data suc…

cs.SD2021

Multimodal Self-Supervised Learning of General Audio Representations

Luyu Wang, Pauline Luc, Adria Recasens +2

We present a multimodal framework to learn general audio representations from videos. Existing contrastive audio representation learning methods mainly focus on using the audio mod…

cs.SD202130 cited

Multi-Format Contrastive Learning of Audio Representations

Luyu Wang, Aaron van den Oord

Recent advances suggest the advantage of multi-modal training in comparison with single-modal methods. In contrast to this view, in our work we find that similar gain can be obtain…

cs.CV202125 cited

Predicting Video with VQVAE

Jacob Walker, Ali Razavi, Aäron van den Oord

In recent years, the task of video prediction-forecasting future video given past video frames-has attracted attention in the research community. In this paper we propose a novel a…

cs.CV2021

Broaden Your Views for Self-Supervised Video Learning

Adrià Recasens, Pauline Luc, Jean-Baptiste Alayrac +11

Most successful self-supervised learning methods are trained to align the representations of two independent views from the data. State-of-the-art methods in video are inspired by…