collaborators

5 papers

eess.AS2026

Pixel-TTS: Image based Text Rendering for Robust Text-to-Speech

Adarsh Arigala, Arjun Gangwar, S Umesh +1

Recent advances in pixel-based text modeling show that representing text as images enables models to exploit visual cues for language understanding. Grounding text in its visual fo…

cs.SD2026

HybridCodec: Fast Dual-Stream, Semantically Enhanced Neural Audio Codec

Arjun Gangwar, S Umesh

The popularity of neural audio codecs as speech tokenizers has surged with the advent of Multimodal Large Language Models. New codec architectures with semantic and acoustic disent…

cs.SD2026

Multilingual Multi-Speaker Unit Vocoders: A Systematic Analysis of Discrete Speech Representations

Naman Kothari, Arjun Gangwar, Adarsh Arigala +1

Discrete speech units obtained via k-means clustering of self supervised embeddings entangle phonetic, speaker, and language information, causing speaker mixing and cross-lingual i…

cs.CL2025

Building Robust and Scalable Multilingual ASR for Indian Languages

Arjun Gangwar, Kaousheik Jayakumar, S. Umesh

This paper describes the systems developed by SPRING Lab, Indian Institute of Technology Madras, for the ASRU MADASR 2.0 challenge. The systems developed focuses on adapting ASR sy…

cs.SD2025

Improved Intelligibility of Dysarthric Speech using Conditional Flow Matching

Shoutrik Das, Nishant Singh, Arjun Gangwar +1

Dysarthria is a neurological disorder that significantly impairs speech intelligibility, often rendering affected individuals unable to communicate effectively. This necessitates t…