activity
20172022
most citedData Augmenting Contrastive Learning of Speech Representations in the Time Domain

93 citations · 162 across the 11 of their papers we have counts for

collaborators

33 papers

cs.NE2022

Introducing topography in convolutional neural networks

Maxime Poli, Emmanuel Dupoux, Rachid Riad

Parts of the brain that carry sensory tasks are organized topographically: nearby neurons are responsive to the same properties of input signals. Thus, in this work, inspired by th…

cs.MA20223 cited

Emergent Communication: Generalization and Overfitting in Lewis Games

Mathieu Rita, Corentin Tallec, Paul Michel +4

Lewis signaling games are a class of simple communication games for simulating the emergence of language. In these games, two agents must agree on a communication protocol in order…

cs.MA20224 cited

On the role of population heterogeneity in emergent communication

Mathieu Rita, Florian Strub, Jean-Bastien Grill +2

Populations have often been perceived as a structuring component for language to emerge and evolve: the larger the population, the more structured the language. While this observat…

cs.SD2021

Speech Resynthesis from Discrete Disentangled Self-Supervised Representations

Adam Polyak, Yossi Adi, Jade Copet +5

We propose using self-supervised discrete representations for the task of speech resynthesis. To generate disentangled representation, we separately extract low-bitrate representat…

cs.SD2021

Learning spectro-temporal representations of complex sounds with parameterized neural networks

Rachid Riad, Julien Karadayi, Anne-Catherine Bachoud-Lévi +1

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often l…

cs.CL2021

Generative Spoken Language Modeling from Raw Audio

Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8

We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…