4 papers
DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis
Wasim Madha, Nityanand Mathur, Hamees Sayed +4
Current text-to-speech systems face a trade-off: autoregres- sive codec language models produce highly intelligible speech but require large-scale models and training data and deco…
How Do Instructions Shape Speech? Cross-Attention Attribution for Style-Captioned Text-to-Speech
Nityanand Mathur, Hamees Sayed, Wasim Madha +4
Style-captioned text-to-speech systems use natural language to control voice characteristics, but how individual words influence acoustic output remains unclear. Understanding this…
FlowLensing: Simulating Gravitational Lensing with Flow Matching
Hamees Sayed, Pranath Reddy, Michael W. Toomey +1
Gravitational lensing is one of the most powerful probes of dark matter, yet creating high-fidelity lensed images at scale remains a bottleneck. Existing tools rely on ray-tracing…
SPRING Lab IITM's submission to Low Resource Indic Language Translation Shared Task
Hamees Sayed, Advait Joglekar, Srinivasan Umesh
We develop a robust translation model for four low-resource Indic languages: Khasi, Mizo, Manipuri, and Assamese. Our approach includes a comprehensive pipeline from data collectio…