2 citations · 2 across the 5 of their papers we have counts for
5 papers
XLS-R fine-tuning on noisy word boundaries for unsupervised speech segmentation into words
Robin Algayres, Pablo Diego-Simon, Benoit Sagot +1
Due to the absence of explicit word boundaries in the speech stream, the task of segmenting spoken sentences into word units without text supervision is particularly challenging. I…
Generative Spoken Language Model based on continuous word-sized audio tokens
Robin Algayres, Yossi Adi, Tu Anh Nguyen +4
In NLP, text language models based on words or subwords are known to outperform their character-based counterparts. Yet, in the speech community, the standard input of spoken LMs a…
Big model only for hard audios: Sample dependent Whisper model selection for efficient inferences
Hugo Malard, Salah Zaiem, Robin Algayres
Recent progress in Automatic Speech Recognition (ASR) has been coupled with a substantial increase in the model sizes, which may now contain billions of parameters, leading to slow…
Fine-tuning Strategies for Faster Inference using Speech Self-Supervised Models: A Comparative Study
Salah Zaiem, Robin Algayres, Titouan Parcollet +2
Self-supervised learning (SSL) has allowed substantial progress in Automatic Speech Recognition (ASR) performance in low-resource settings. In this context, it has been demonstrate…
STOP: A dataset for Spoken Task Oriented Semantic Parsing
Paden Tomasello, Akshat Shrivastava, Daniel Lazar +12
End-to-end spoken language understanding (SLU) predicts intent directly from audio using a single model. It promises to improve the performance of assistant systems by leveraging a…