4 papers
NAVER LABS Europe Submission to the Instruction-following 2026 Short Track
Marcely Zanon Boito, Hemant Yadav, Jean-Luc Meunier +1
In this paper, we describe NAVER LABS Europe's submission to the instruction-following speech processing short track at IWSLT 2026. We participate again in the constrained setting,…
Speech Representation Learning Revisited: The Necessity of Separate Learnable Parameters and Robust Data Augmentation
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah
Speech modeling methods learn one embedding for a fixed segment of speech, typically in between 10-25 ms. The information present in speech can be divided into two categories: "wha…
JOOCI: a Framework for Learning Comprehensive Speech Representations
Hemant Yadav, Rajiv Ratn Shah, Sunayana Sitaram
Information in speech can be categorized into two groups: Content (what is being said, such as linguistics) and Other (how it is expressed such as information about speaker and par…
MS-HuBERT: Mitigating Pre-training and Inference Mismatch in Masked Language Modelling methods for learning Speech Representations
Hemant Yadav, Sunayana Sitaram, Rajiv Ratn Shah
In recent years, self-supervised pre-training methods have gained significant traction in learning high-level information from raw speech. Among these methods, HuBERT has demonstra…