42 citations · 51 across the 4 of their papers we have counts for
6 papers
Are word boundaries useful for unsupervised language learning?
Tu Anh Nguyen, Maureen de Seyssel, Robin Algayres +3
Word or word-fragment based Language Models (LM) are typically preferred over character-based ones in many downstream applications. This may not be surprising as words seem more li…
textless-lib: a Library for Textless Spoken Language Processing
Eugene Kharitonov, Jade Copet, Kushal Lakhotia +8
Textless spoken language processing research aims to extend the applicability of standard NLP toolset onto spoken language and languages with few or no textual resources. In this p…
On the Feasibility of Predicting Questions being Forgotten in Stack Overflow
Thi Huyen Nguyen, Tu Nguyen, Tuan-Anh Hoang +1
For their attractiveness, comprehensiveness and dynamic coverage of relevant topics, community-based question answering sites such as Stack Overflow heavily rely on the engagement…
The Zero Resource Speech Challenge 2021: Spoken language modelling
Ewan Dunbar, Mathieu Bernard, Nicolas Hamilakis +6
We present the Zero Resource Speech Challenge 2021, which asks participants to learn a language model directly from audio, without any text or labels. The challenge is based on the…
Generative Spoken Language Modeling from Raw Audio
Kushal Lakhotia, Evgeny Kharitonov, Wei-Ning Hsu +8
We introduce Generative Spoken Language Modeling, the task of learning the acoustic and linguistic characteristics of a language from raw audio (no text, no labels), and a set of m…
The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling
Tu Anh Nguyen, Maureen de Seyssel, Patricia Rozé +5
We introduce a new unsupervised task, spoken language modeling: the learning of linguistic representations from raw audio signals without any labels, along with the Zero Resource S…