activity
20172022
most citedThe Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

42 citations · 136 across the 10 of their papers we have counts for

collaborators
Showing 2020 · cs.CLShow all

5 papers · 2 filters

cs.CL2020★ 42 cited

The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé +5

We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource S…

cs.CL2020

The Zero Resource Speech Challenge 2020: Discovering discrete subword and word units

Ewan Dunbar, Julien Karadayi, Mathieu Bernard +6

We present the Zero Resource Speech Challenge 2020, which aims at learning speech representations from raw audio signals without any labels. It combines the data sets and metrics f…

cs.CL2020

Perceptimatic: A human speech perception benchmark for unsupervised subword modelling

Juliette Millet, Ewan Dunbar

In this paper, we present a data set and methods to compare speech processing models and human behaviour on a phone discrimination task. We provide Perceptimatic, an open data set…

cs.CL2020

Analogies minus analogy test: measuring regularities in word embeddings

Louis Fournier, Emmanuel Dupoux, Ewan Dunbar

Vector space models of words have long been claimed to capture linguistic regularities as simple vector translations, but problems have been raised with this claim. We decompose an…

cs.CL2020★ 7 cited

The Perceptimatic English Benchmark for Speech Perception Models

Juliette Millet, Ewan Dunbar

We present the Perceptimatic English Benchmark, an open experimental benchmark for evaluating quantitative models of speech perception in English. The benchmark consists of ABX sti…