3 papers
cs.SD2026
VCNAC: A Variable-Channel Neural Audio Codec for Mono, Stereo, and Surround Sound
Florian Grötschla, Arunasish Sen, Alessandro Lombardi +2
We present VCNAC, a variable channel neural audio codec. Our approach features a single encoder and decoder parametrization that enables native inference for different channel setu…
eess.AS2025
SIFT-50M: A Large-Scale Multilingual Dataset for Speech Instruction Fine-Tuning
Prabhat Pandey, Rupak Vignesh Swaminathan, K V Vijay Girish +4
We introduce SIFT (Speech Instruction Fine-Tuning), a 50M-example dataset designed for instruction fine-tuning and pre-training of speech-text large language models (LLMs). SIFT-50…
cs.CL2024
Promptformer: Prompted Conformer Transducer for ASR
Sergio Duarte-Torres, Arunasish Sen, Aman Rana +5
Context cues carry information which can improve multi-turn interactions in automatic speech recognition (ASR) systems. In this paper, we introduce a novel mechanism inspired by hy…