2 citations · 5 across the 12 of their papers we have counts for
4 papers · 1 filter
Streaming Speech-to-Text Translation with a SpeechLLM
Titouan Parcollet, Shucong Zhang, Xianrui Zheng +1
Normally, a system that translates speech into text consists of separate modules for speech recognition and text-to-text translation. Combining those tasks into a SpeechLLM promise…
Loquacious Set: 25,000 Hours of Transcribed and Diverse English Speech Recognition Data for Research and Commercial Use
Titouan Parcollet, Yuan Tseng, Shucong Zhang +1
Automatic speech recognition (ASR) research is driven by the availability of common datasets between industrial researchers and academics, encouraging comparisons and evaluations.…
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
Shucong Zhang, Titouan Parcollet, Rogier van Dalen +1
Self-attention relies on positional embeddings to encode input order. Relative Position (RelPos) embeddings are widely used in Automatic Speech Recognition (ASR). However, RelPos h…
Linear-Complexity Self-Supervised Learning for Speech Processing
Shucong Zhang, Titouan Parcollet, Rogier van Dalen +1
Self-supervised learning (SSL) models usually require weeks of pre-training with dozens of high-end GPUs. These models typically have a multi-headed self-attention (MHSA) context e…