3 papers
cs.SD2026
MambaVoiceCloning: Efficient and Expressive Text-to-Speech via State-Space Modeling and Diffusion Control
Sahil Kumar, Namrataben Patel, Honggang Wang +1
MambaVoiceCloning (MVC) asks whether the conditioning path of diffusion-based TTS can be made fully SSM-only at inference, removing all attention and explicit RNN-style recurrence…
cs.CL2024
KatzBot: Revolutionizing Academic Chatbot for Enhanced Communication
Sahil Kumar, Deepa Paikar, Kiran Sai Vutukuri +4
Effective communication within universities is crucial for addressing the diverse information needs of students, alumni, and external stakeholders. However, existing chatbot system…
cs.SD2024
Vision Transformer Segmentation for Visual Bird Sound Denoising
Sahil Kumar, Jialu Li, Youshan Zhang
Audio denoising, especially in the context of bird sounds, remains a challenging task due to persistent residual noise. Traditional and deep learning methods often struggle with ar…